Technical
July / August 2026

Continuous Operational Control of PaaS in GxP Environments

Robin Wennemuth
Sebastian Koch
Florian Kekule
Robert McDowall, PhD
Continuous-Operational-Control-of-PaaS-in-GxP-Environments-750px.jpg

The transition from on-premise to cloud-based systems requires a shift in how IT infrastructure is managed in regulated environments. Rather than traditional static qualification, continuous operational control of Platform as a Service (PaaS) infrastructure supporting GxP applications enables life sciences companies to maintain compliance and accelerate innovation.

This article presents a practical, risk-based approach to establishing and maintaining confidence that critical PaaS services remain fit for intended use through continuous monitoring, automated testing, and real-time evidence generation, delivering both regulatory compliance and substantial business value.

Background

The transition from traditional on-premise IT infrastructures to cloud-based solutions marks a significant evolution in the life sciences industry. As companies strive for digital transformation, cloud computing offers unparalleled scalability, potential for innovation, and operational efficiency. However, ensuring compliance with GxP guidelines within these dynamic environments presents unique challenges. The need for cloud infrastructure qualification is well established in the industry. Nevertheless, current approaches and solutions fall short in accounting for the highly complex and dynamic nature of PaaS.

We explore the concept of continuous operational control of PaaS supporting GxP-regulated workloads to address this challenge. Ongoing assurance is also relevant for on-premise infrastructures, and the need becomes especially pronounced in cloud-based environments due to the high frequency of changes to cloud services and the shared responsibility between the cloud service provider (CSP) and the life sciences company. Hyperscale CSPs deploy code changes to their services at a higher rate than traditional on-premise IT, with public statements from major providers describing deployment rates in the millions per year.1 The life science company cannot control when and how these changes are deployed, but it is ultimately accountable for them.

This article presents a practical methodology for implementing continuous operational control that meets regulatory requirements, at the same time delivering strong business benefits.

The Move from On-Premise to Cloud Computing

Historically, life sciences companies managed their IT infrastructure on-premise, maintaining full control over their systems. Over time, they have transitioned to leveraging Infrastructure as a Service (IaaS), shifting parts of the responsibilities to IaaS providers but still retaining strong control. The shift to PaaS represents the next evolutionary step.

The National Institute of Standards and Technology (NIST) defines PaaS as “the capability provided to the consumer is to deploy onto the cloud infrastructure consumer-created or acquired applications created using programming languages, librar-ies, services, and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, including network, servers, operating systems, or storage, but has control over the deployed applications and possibly configuration settings for the application-hosting environment.”2

PaaS offers prebuilt, business-ready services that further enhance innovation and agility including but not limited to:

  • Highly abstracted services, close to using applications with comparatively high configurability; examples include managed databases, workflow management, document management, or portal services
  • Supporting and infrastructural services for authentication and authorization, application logging, or event automation.

This transition to PaaS not only accelerates development and streamlines operations but also necessitates a new approach to managing compliance and shared responsibilities within these complex environments.

Each layer of new abstraction by cloud providers shifts responsibility to the cloud provider, but the regulated user still remains accountable. Amazon Web Services, one of the most prominent cloud providers, states, “One of the concerns for regulated enterprise customers becomes how to qualify and demonstrate control over a system when so much of the responsibility is now shared with a supplier. The purpose of a Qualification Strategy is to answer this question.”3

Regulatory Framework for PaaS

Compliance with regulatory frameworks such as the European Union (EU) and The Pharmaceutical Inspection Convention and Pharmaceutical Inspection Co-operation Scheme (PIC/S) GMP Annex 114 and the US Food and Drug Administration (FDA) 21 CFR Parts 11 and 2115, 6 applies to regulated companies operating computerized systems used in GxP-relevant processes. The obligation rests on the regulated user, not on the CSP. Annex 11 states that “the application should be validated; IT infrastructure should be qualified.”4 In this context, the IT infrastructure supporting GxP applications—including cloud-based platforms—must be demonstrably controlled and fit for intended use by the regulated company.

ISPE GAMP® 5 Guide: A Risk-Based Approach to Compliant GxP Computerized Systems (Second Edition) (Appendix M11 on IT Infrastructure)7 classifies PaaS as IT infrastructure rather than as a GxP-regulated application. PaaS components are treated as Category 1 infrastructure software or as software tools (Appendix D9). This classification clarifies that PaaS itself is not subject to traditional application validation but still requires qualification as part of the underlying infrastructure. ISPE GAMP® 5 (Second Edition) recommends using “good IT practices” (e.g., Information Technology Infrastructure Library [ITIL]-based service man-agement, configuration management, operational monitoring, incident management, and continuous improvement) rather than classical installation qualification (IQ)/operational qualification (OQ)/performance qualification (PQ)-style protocols, which are not well suited to dynamic cloud environments. The goal is to achieve and maintain a controlled state of the infrastructure so that GxP applications operating on it remain in a validated state.

Regulatory expectations reinforce this approach. EU GMP Annex 11 requires that risks are managed throughout the system life cycle (Clause 1), that changes are controlled (Clause 10), and that systems are periodically evaluated to confirm they remain in a valid state (Clause 11).4 The US FDA’s guidance document FDA Computer Software Assurance for Production and Quality System Software8 acknowledges cloud-based deployment models and establishes a risk-based framework in which manufacturers “select an appropriate frequency for performing assurance activities based on their risk-based analysis.”

The guidance also notes that continuous performance monitoring can help ensure software maintains a validated state consistent with quality system obligations. FDA 21 CFR Part 11 further extends these expectations to infrastructure hosting electronic records, requiring that audit trails, security, and data integrity are ensured. In parallel, the International Council for Harmonization of Technical Requirements for Pharmaceuticals for Human Use (ICH) Q10 “Pharmaceutical Quality System” 9 requires organizations to establish and maintain conditions under which processes achieve intended results and are continuously improved, with suitable IT infrastructure being an integral part of this capability.

The ISPE GAMP® Good Practice Guide: IT Infrastructure Control and Compliance (Second Edition) provides practical guidance on meeting these expectations for PaaS environments.10 The “Qualification of Platforms—Building Block” concept supports a structured qualification approach by decomposing platforms into manageable components. Appendix 11 “Platform as a Service” provides specific guidance for qualifying PaaS, and further sections emphasize the importance of maintaining the qualified state during operation through ongoing monitoring, change control, and incident management. This continuous assurance of the operational environment is particularly critical for cloud-based infrastructures, in which frequent updates and configuration changes are the norm. Maintaining the qualified state during operation through risk-based, documented, and monitored controls is therefore central to demonstrating control and compliance in regulatory inspections and quality assurance audits.

The Challenge: Cloud Infrastructure Reality

Cloud providers deploy millions of updates annually. These updates occur outside regulated companies’ control. Traditional point-in-time infrastructure qualification is impractical in such environments. Public clouds are highly distributed, continuously evolving systems. Built on microservice architectures, they inherit the complexities of distributed computing—network partitions, cascading retries, partial failures, and emergent behavior that demand testing approaches beyond traditional service-level checks.

Recent incidents11 illustrate this point: hardware anomalies, network disruptions, and seemingly harmless configuration changes at a single service can propagate into multiservice outages. These ripple effects underscore how hidden dependencies and constant platform changes can erode confidence in system reliability.

As a consequence, point-in-time preproduction testing alone cannot provide sustained assurance that a cloud service remains fit for intended use at a risk level acceptable for life sciences applications. Modern distributed systems exhibit emergent failure modes that arise from interactions between continuously changing components and that are not exhaustively enumerable in advance; the partial failure, network partition, and configuration propagation observed in recent industry incidents illustrate this directly. A defensible assurance approach therefore combines rigorous preproduction qualification with ongoing, production-safe testing and monitoring of the platform during operation.

The life science application is embedded into the cloud environment. Application changes are typically infrequent, controlled, and explicitly validated at each release, whereas the underlying cloud platform changes continuously and largely unnoticed to the regulated user. Platform change frequency therefore exceeds application change frequency by orders of magnitude. The validated state of the overall system is a function of both layers, so whereas application validation can remain event-driven (triggered on release), platform assurance must be ongoing in order to detect changes that could erode the qualified state of the infrastructure on which the application depends. The system consists of the components controlled by the life science company and the cloud components outside its control, as shown in Figure 1.

Figure 1: Different levels and responsibilities of a PaaS infrastructure (adapted from ISPE GAMP® 5 (Second Edition).7

Image
McDowall_Fig-1_NEW-md.jpg

This is particularly evident in PaaS scenarios. PaaS continuously evolves, for example, by delivering security patches or improving the quality of service. The release cycles of such platform updates and the life science applications using them usually do not align. Reviewing every individual change is practically infeasible due to short lead times, limited transparency of internal changes, and the effort required for impact analysis. In addition, PaaS providers offer tenant-level separation (e.g., subaccounts, subscriptions, or projects) that allows customers to isolate their development, quality assurance (QA)/validation, and production environments. However, the underlying managed service implementation—the control plane and the multi-tenant run time—is shared across all customer environments. Service-level changes and defects therefore propagate to every environment simultaneously, regardless of how customers structure their own tenancy. Customer applications in different life-cycle stages consume the same service versions and inherit the same platform-side changes.

The qualification strategy needs to shift toward a platform-qualification approach, focusing on managing the platform in a controlled state and fitness for use, independent of specific applications. This involves qualifying the platform provider and qualifying standardized services using the Building Block concept from the ISPE GAMP® Good Practice Guide: IT Infrastructure Control and Compliance (Second Edition).10 Various building blocks (e.g., components such as managed databases) possess inherently different risk profiles. Leveraging qualification data can provide valuable insights into these risks, and regular risk reassessments should be embedded within a continuous quality improvement process.

Given the continuous change on the platform side, an agile, risk-based, and electronically documented change management procedure is required to assess the impact of cloud service updates on GxP applications and to determine requalification needs (see Figure 2). Additionally, establishing company-wide standards for use of cloud services is essential. The goal is to develop an internal “Cloud Service Catalog” that defines a set of approved services within the company, thereby creating a “Regulated Landing Zone.”3 Within this controlled environment, life sciences companies can operate GxP-relevant applications on public cloud platforms in a risk-managed manner.

Figure 2: Continuous risk-based monitoring of changes and disruptions of cloud services.

Image
McDowall_Fig-2_NEW-md.jpg

Why Platform-Level Continuous Control Is GAMP-Aligned

ISPE GAMP® 5 (Second Edition) treats PaaS components as Category 1 infrastructure or software tools (Appendix D9) and recommends ITIL-style good IT practice rather than application-style IQ/OQ/PQ. The following approach details that recommendation: the deliverables are IT service management (ITSM) artifacts and monitoring evidence, and the activity is operational control of the platform, not validation of a regulated application. The question is whether the three usual controls—application validation, supplier qualification of the CSP, and infrastructure-level incident management—are enough on their own for GxP workloads on PaaS. They cover most of the risk envelope, but they miss a specific set of failures that none of them are well placed to catch. Platform-level continuous control is the direct response to that gap in the GAMP® framework of severity × probability × detectability.

The following failures listed share two features: they can have real GxP impact, and the three usual controls are poorly placed to spot them. The limitation is one of vantage point and timing, not of effort.

Drift in a Stable API

The API version is unchanged, but a bug fix or implementation change shifts response meaning, error codes, latency, or ordering. Release-time application tests are not re-run, and supplier audits cover process, not per-service behavior.

Tail-Latency or Tail-Error Regression

The mean is unchanged but p99 or p999 moves; tail-sensitive batch jobs miss their window. Provider service level agreements (SLAs) report coarse availability, not workload-specific tails.

Failover-Only Failure

The failover region has quietly drifted to a different version, configuration, or replication topology. The mismatch shows up only during an actual failover, when response time is shortest.

Side Effects of Background Platform Work

Schedule backup, replication, key rotation, certificate renewal, or patching changes quality of service (QoS) at specific times. Application tests are driven by the application and never exercise platform-initiated work.

Interference from Other Tenants

Resource contention from co-tenants degrades QoS within SLA bounds; the effect is bursty, load-correlated, not reproducible on demand, and not covered by most provider SLAs.

Deprecation Overrun

A deprecated feature still works past its announced removal date, and the application still depends on it. Application tests check current behavior, not deprecation calendars.

Failure in Service Composition

Each managed service the application uses works correctly in isolation, but the way they compose has changed. The application’s direct-call tests do not exercise the transitive dependency graph.

Provider Signals the Application Does Not See

The platform emits a new error class or telemetry signal that the application does not subscribe to but that indicates degraded fitness for intended use.

The list is not exhaustive, but it is enough to show that the residual set is real, can have material impact, and is not covered by the three usual controls. Severity is workload-dependent and often material; probability per change is low to moderate but cumulative over operating life; detectability under the usual controls is poor. The product is a residual risk that, by GAMP® principles, calls for a complementary detection control.

That control is continuous monitoring of platform behavior, run by the regulated user or a delegated vendor and not by the platform provider. It combines independent probes against the provider’s published application programming interfaces, intake of the provider’s own health, deprecation, and change-notification feeds, and statistical analysis of those signals over time against a baseline. What sets it apart from application-level testing is not who runs it, but what it sees. This includes signals from inside the platform and probe surfaces the application itself does not exercise. The provider is not asked to do anything beyond keeping the observability surfaces it already runs for its own site reliability work.

The control is additive. It does not replace application validation, supplier qualification, or ITIL incident and problem management; it covers the detectability gap those three leave open. The deliverables are ITSM artefacts: a per-service behavior baseline, the probes and feeds tied to it, an alerting policy that routes anomalies into the existing incident process, and an evidence log of operation. The methodology in the next section applies these principles in ITSM-aligned terms.

Continuous Cloud Service Operational Control and Compliance

Achieving and maintaining a state of control requires mitigating or, where possible, eliminating risks. Though, in PaaS, risk can never be fully eliminated unless you stop using the platform service altogether. Regulated companies and their PaaS providers should aim for a controlled environment by implementing operational controls that enhance the detectability of issues through continuous qualification, thereby reducing risks to patient safety, product quality, and data integrity.

Operational control of the platform layer is credible only if the CSP itself is qualified as a supplier. Because the regulated user has no direct visibility into the corporate service provider’s internal change, release, and quality processes, supplier qualification must rely on documented evidence and contractual instruments. In practice, this comprises the following:

  • Review of third-party attestations such as SOC 2 Type II12 and ISO/IEC 27001/27017/2701813 reports, mapped against the regulated user’s GxP control requirements
  • Review of provider-issued GxP white papers and shared-responsibility matrices, recognizing that these are vendor statements and not substitutes for the user’s own risk assessment
  • A quality (or technical) agreement covering change-notification obligations, incident communication, audit rights, data residency, and subprocessor management
  • A documented procedure for triggering supplier reassessment on material events (loss of certification, jurisdictional changes, security incidents, or substantive changes to the shared-responsibility boundary)

The continuous operational controls described next complement, rather than replace, this supplier qualification activity.

The following methodology in Table 1 outlines the process for establishing and maintaining continuous qualification and is shown in Figure 3, cloud service qualification assessment in a cloud application life cycle.7 This framework is designed with GxP requirements in mind but is versatile enough to be applied to both GxP and nonGxP applications. Although the framework is built on GxP-specific guidance (e.g., Annex 11, FDA CSA, ISPE GAMP® 5, ICH Q10), its underlying principles of risk-based scoping, continuous evidence generation, and platform-level controls are transferable to other regulated industries that operate workloads on managed cloud services.

Using historical service data on failures can be crucial for evaluating service risks and the potential failure space. Beyond mitigating the effects of changes by increased detectability, more information on the failure space (e.g., virtual machine errors) can help determine corrective actions and failure traps as part of the application design. Upon detecting and classifying a known failure scenario, corrective actions can be taken automatically such as provisioning more compute resources or establishing self-correction capabilities. Furthermore, external monitoring information such as the official CSP availability can help to quickly identify, classify, and mitigate changes. Proactive change impact analysis may also serve as an effective control mechanism using historical information, as well as information from the CSP on scheduled changes such as deprecation of features.

Table 1. Phases and activities for continuous qualification of PaaS infrastructure.
PhaseActivities
1. Assessment of Cloud Services Used• Identify critical solution components: Identify the critical solution components in software applications and document the cloud services used
• Analyze impact: Analyze the impact of cloud services on the application
• Intended use definition: Clearly define how each cloud service will be used within the GxP application
• Assess basic requirements: Access control, audit trail, monitoring, data confidentiality (data encryption at both interface level [data transfer] and repository level [data storage]), backup and restore, disaster recovery, business continuity, and, as appropriate, archiving and retrieval
• Functional service assessment: Assess specific features and functions that are used from each cloud service and define specific requirements based on its intended use
2. Perform Risk Assessment• Hazard identification: Identify potential hazards associated with the usage of each service, focusing on the aspects of service functionality, data integrity, security, and availability
• Risk analysis: Assess the severity, probability, and detectability of potential failures
• Risk evaluation: Prioritize risks based on their impact on GxP compliance and overall business operations
3. Define Continuous Monitoring Plan• Define test strategy: Develop a comprehensive test strategy that includes scope, extent of testing for a service, and frequency of tests
• Automated testing implementation: Implement automated tests that continuously verify the compliance of cloud services
• Mitigation strategies: Define strategies to mitigate identified risks, including alert triggers for immediate investigation in case of deviations, fallback procedures, and corrective actions
• Continuous monitoring plan: Outline a plan for automated and continuous infrastructure qualification that enables traceability from cloud service to test case; ensure review and approval in the change management process
4. Release Gate and Run Phase: Initial Qualification, Continuous Testing, and Deviation Management• Service admission baseline: Establish the documented baseline of expected service behavior, monitoring probes, telemetry subscriptions, and alerting rules for the service, and submit it for quality review and approval. Releasing the service admission baseline admits the service to the cloud service catalog and marks the transition to operational monitoring
• Test execution: Regularly execute probes and ingest telemetry feeds according to the continuous monitoring plan; record results against the service admission baseline
• Incident and problem management: Route monitoring anomalies and detected deviations from baseline into the regulated user’s existing ITIL incident management process for triage, root cause analysis, and corrective action (including application-level mitigations such as end-user notification where appropriate); address and manage provider-attributable incidents through change-notification and support channels with the PaaS provider, with documentation, resolution, and communication captured against the originating incident record
5. Reporting• Real-time monitoring: Provide access to a monitoring dashboard with real-time reports that provide insights into the compliance status of cloud services
• Continuous monitoring evidence log: Maintain an evidence log demonstrating end-to-end traceability from monitoring probe and telemetry signal to detected anomaly, incident record, and corrective action; the log constitutes the documented evidence that operational controls operated as designed during the reporting period
• Continuous test improvement/adaptation: Monitor announced changes of cloud services and adapt automated test case implementations based on updates
• Review process: Perform a regular review of the continuous infrastructure qualification process

Case Study: SAP Business Technology Platform (BTP)

SAP Business Technology Platform (BTP) is a comprehensive PaaS environment that provides tools and services for developing, integrating, and managing enterprise applications. Life sciences companies leverage SAP BTP services to build and maintain software for GxP processes. Examples of practical use cases include:

  • Cockpit for batch releases: Leveraging a cloud-based solution to streamline, simplify, and support the batch release process at the same time maintaining compliance with regulatory standards
  • Cell and gene therapy process orchestration: Using SAP BTP for orchestrating complex cell and gene therapy treatment processes, helping ensure that the right patients get the right treatments without errors
  • Temperature management solution: Deploying a solution for monitoring and managing the temperature of pharmaceutical products throughout the supply chain

Life sciences companies leverage continuous qualification on SAP BTP to demonstrate control over the cloud services used. An integrated infrastructure monitoring tool was developed, which is used by multiple companies for this purpose. It covers the entire process of cloud service qualification presented in the previous section, including running automated tests delivered in a predefined cloud service catalog. Continuous qualification integrates with the operational life cycle of GxP applications on SAP BTP, and life sciences companies gain continuous monitoring, visibility over change impact, and deviation management.

In practice, this keeps a plan-and-report approach viable despite continuous platform change. The monitoring plan for each service is established once and thereafter maintained under change control. It is adjusted only when the application’s use of a BTP service changes, and only after human review and approval; the platform’s ongoing changes are caught by the monitoring itself rather than by revising the plan. The supporting evidence is generated continuously by the tool rather than compiled manually, so the maintenance effort tracks the application’s controlled release cycle rather than the platform’s much faster pace of change. The result is a structured, automation-driven way to manage qualification and compliance that stays operationally sustainable and produces inspection-ready evidence as a byproduct of monitoring.

Figure 3: Cloud service qualification assessment in a cloud application life cycle [adapted from.7

Image
McDowall_Fig-3_NEW-md.jpg

Figure 4: Example of a dashboard demonstrating real-time control over cloud services.

Image
figure4-md.jpg

Figure 5: Example of qualification report documenting cloud service qualification.

Image
figure5-md.jpg

To maintain and visualize control over cloud services, a comprehensive dashboard and over-time reporting capabilities are in place. The dashboard provides real-time insights into service performance and compliance status, enabling proactive management of any issues (see Figure 4).

Over-time reporting allows companies to track trends, identify recurring issues, and assess the effectiveness of corrective actions. Figure 5 illustrates a system extract of an over-time service qualification report, including the full end-to-end traceability in one document.

Conclusion

Continuous operational control of PaaS infrastructure represents the practical implementation of regulatory requirements for infrastructure qualification in cloud environments. By combining risk-based assessment, continuous testing, automated monitoring, and real-time evidence generation, life sciences companies can maintain compliance, accelerate innovation, and reduce operational costs. This approach is grounded in fundamental regulatory principles (Annex 11, FDA CSA, ISPE GAMP® 5 (Second Edition) , ICH Q10). The future of infrastructure control in GxP environments is continuous, automated, and intelligent.

Not a Member Yet?

To continue reading this article and to take advantage of full access to Pharmaceutical Engineering magazine articles, technical reports, white papers and exclusive content on the latest pharmaceutical engineering news, join ISPE today. In addition to exclusive access to all of the content in Pharmaceutical Engineering magazine, you will get online access to 24 ISPE Good Practice Guides, exclusive networking events, regulatory resources, Communities of Practice, and more.

Learn more about the valuable benefits you'll receive with an ISPE membership.

Join Today


About Pharmaceutical Engineering

ISPE members receive an annual subscription to ISPE’s award-winning Pharmaceutical Engineering magazine as part of their membership benefits. Published six times yearly, each issue features contributions from expert authors and technical articles highlighting the latest industry trends and innovations.

Learn more

Join ISPE Today

Becoming a member of ISPE offers numerous benefits, including access to a vast network of professionals, exclusive training events, and valuable resources. As a member, you'll join more than 22,000 of your professional peers from over 120 countries in advancing solutions that lead to improved patient health. Membership provides access to 20+ complimentary ISPE Good Practice Guides, a robust library of on-demand training and e-learning resources, and much more. Learn more and consider joining today.

Become an ISPE member 

ISPE members: Get more involved by volunteering.