Showing Posts From

Platform engineering

The ungoverned cloud: Why cloud strategies fail at execution.

The ungoverned cloud: Why cloud strategies fail at execution.

In boardrooms across Europe, cloud strategy is undergoing a harsh reality check. For years, the narrative was centered on speed and migration. Today, executive teams face a very different set of challenges: unpredictable cloud expenditure, strict regulatory mandates (NIS2, BIO2, EU AI Act), and diffuse operational accountability. When external audits reveal that two-thirds of cloud environments lack proper control, the executive reflex is predictable: install a heavy Governance Board, write 80-page policy manuals, and require manual sign-offs for every change. This approach fails every time. It creates shadow IT, paralyzes delivery teams, and fails to eliminate actual risk. Personally, I view cloud governance not as a bureaucratic brake, but as an operational operating system. True governance provides clear guardrails, automated compliance, and organizational clarity—allowing engineering teams to move fast safely. To achieve this, organizations must move away from theoretical policies and implement a functional Cloud Center of Excellence (CCoE).The 5 pillars of cloud governance Before structuring your team, you must define what cloud governance actually encompasses. Mature cloud governance covers five distinct operational domain pillars:Pillar 1: Financial Management (FinOps)Shifting from static annual IT budgets to dynamic unit economics, continuous cost allocation, and real-time optimization.Pillar 2: Security & Regulatory Compliance (NIS2 / BIO2)Enforcing baseline controls aligned with NIS2, BIO2, ISO 27001, and GDPR across all cloud landing zones.Pillar 3: Automation & Platform EngineeringEliminating manual infrastructure configuration through Infrastructure as Code (IaC) and automated developer platforms.Pillar 4: Identity & Data Control (Zero Trust)Implementing Zero Trust architecture, strict least-privilege principles, and explicit data boundaries.Pillar 5: AI & Emerging Tech Governance (ISO/IEC 42001)Setting parameters for responsible AI use under the EU AI Act and ISO/IEC 42001, preventing unmanaged shadow-AI implementations.What is expected of C-level leadership? Cloud governance cannot be delegated away to IT or a compliance team. Real governance requires active C-level involvement, clear sponsorship, and strategic alignment. Here is what is explicitly expected of executive leadership across each core domain:C-Level Role Core Executive Expectation & Operational ResponsibilityChief Executive Officer (CEO) & Board Treat Cloud Governance as Risk Management: Recognize that cloud failure, data breaches, and non-compliance carry direct board liability under NIS2. Establish risk appetite boundaries and mandate cross-functional governance across the company.Chief Operating Officer (COO) Align the Operating Model & CCoE Mandate: Provide the CCoE with formal authority to set organizational standards. Break down functional silos between IT, Security, and Business units, ensuring that delivery speed never bypasses compliance.Chief Financial Officer (CFO) Enforce Financial Accountability (FinOps): Shift financial oversight from traditional CapEx IT depreciation to dynamic OpEx management. Demand unit-cost transparency and require Business/Product Owners to account for cloud consumption within their P&L.Chief Information / Technology Officer (CIO/CTO) Drive Modern Architecture & Enablement: Transition engineering teams away from manual ticketing towards self-service platforms (IDPs). Enforce "Policy as Code" and ensure cloud infrastructure aligns with architecture goals.Chief Information Security Officer (CISO) Automate Guardrails Over Gatekeeping: Move from reactive security reviews to proactive, automated policy enforcement. Integrate NIS2, ISO 27001, and AI compliance directly into deployment pipelines.Key Leadership Takeaway: Executive leadership is not expected to manage cloud settings or review code. Leadership is expected to set parameters, grant mandate, enforce accountability, and model the culture required for operational discipline.Enablement, not control The central execution engine of cloud governance is the Cloud Center of Excellence (CCoE). Too many companies misinterpret the CCoE as an architectural approval committee that meets every Thursday to review tickets. That is the quickest way to kill organizational momentum. Gatekeeper vs. EnablementAnti-Pattern: The Gatekeeper CCoE Modern Pattern: The Enablement CCoEManually reviews and approves architectural change requests. Builds automated guardrails and self-service templates.Writes static policy PDFs that engineers rarely read. Embeds policy directly into deployment pipelines (Policy as Code).Acts as a centralized bottleneck for cloud adoption. Functions as an internal product team serving delivery teams.Measures success by policy compliance and audit logs. Measures success by engineering velocity, security, and cost efficiency.Structure & core roles A successful CCoE is a lean, cross-functional team that brings together key domains. It does not replace engineering teams; it empowers them.Executive Sponsor (COO / VP Operations): Secures budget, aligns governance with corporate P&L goals, and resolves organizational friction between business units. Cloud Lead / Architect: Defines overall multi-cloud strategy, Landing Zone standards, and reference architectures. Cloud Security & Risk Specialist: Translates regulatory requirements (NIS2, ISO 27001, EU AI Act) into actionable security policies and automated checks. Platform Lead / Software Architect: Drives Platform Engineering, building Internal Developer Platforms (IDPs) and self-service "Golden Paths". FinOps Practitioner: Analyzes cloud consumption data, establishes unit-cost metrics, and works directly with product owners on cost accountability.Practical implementation: A 4-phase roadmap Implementing cloud governance across an organization requires a phased, practical approach. Phase 1: Establish the charter & landing zone architectureDefine the CCoE Charter: Formally declare the team's purpose, scope, and mandate across the business. Build Landing Zones: Create standard multi-account cloud structures (e.g., AWS Organizations or Azure Management Groups). Isolate workloads by environment (Dev, Test, Prod) and business unit. Implement Centralized Logging: Ensure audit trails, identity logs, and network traffic are automatically ingested into a central SIEM system from day one.Phase 2: Automate guardrails (Policy-as-Code)Define Preventive & Detective Controls: Use native cloud policies (e.g., Azure Policy, AWS Service Control Policies) to enforce mandatory constraints: Preventative: Block public S3 buckets or unencrypted storage volumes from ever being created. Detective: Automatically flag and alert security teams when a resource drifts from baseline configuration.Tagging Strategy Enforcement: Mandate metadata tags (Owner, CostCenter, Environment, DataClassification) at deployment time. If a resource lacks tags, auto-remediate or reject the build.Phase 3: Platform engineering & self-service (Golden Paths)Build the Internal Developer Platform (IDP): Provide engineering teams with a self-service portal (e.g., Backstage) to provision compliant infrastructure in minutes. Publish Golden Paths: Pre-package approved architectures (e.g., secure microservice deployment, compliant SQL cluster) that include security, monitoring, and backups by default. Community of Practice: Establish cloud guilds to train product teams, share best practices, and accelerate internal skills development.Phase 4: FinOps maturity & Responsible AI governanceShift-Left Cost Management: Integrate cost-estimation tools into CI/CD pipelines so developers see the estimated monthly bill before merging code. Establish AI Guardrails: Deploy private API endpoints for Generative AI. Ensure corporate data is isolated and protected under strict tenant boundaries. Continuous Executive Dashboards: Provide board-level visibility into compliance posture, operational risks, and cloud cost efficiency.What the board needs to see To ensure your CCoE is delivering real value, track concrete operational metrics rather than subjective milestones:Metric Target / Good Practice Executive FocusLanding Zone Coverage > 95% of workloads in governed Landing Zones Risk & ComplianceUntagged Cloud Resources < 2% of total cloud assets Financial AccountabilityPolicy Drift MTTR < 4 hours to remediate non-compliant resources NIS2 / Security PostureGolden Path Adoption > 80% of new microservices deployed via IDP Velocity & StandardizationCloud Unit Cost Decreasing cost per business transaction P&L & ScalabilityClosing thoughts Solving cloud governance is not a technical problem; it is an organizational design challenge. Relying on manual audits, reactive firefighting, and bureaucratic approvals inevitably leads to higher costs and increased business risk. True operational leadership means building a system where compliance, security, and cost control are automated and frictionless. By establishing a modern Cloud Center of Excellence, embedding Policy as Code, and adopting Platform Engineering, executive teams can bridge the gap between high-level strategy and ground-level execution. When governance is built directly into your operating model, compliance stops being a burden—it becomes a competitive advantage that enables rapid, resilient, and profitable growth.

The future of managed services is letting go of control

The future of managed services is letting go of control

For decades, managed services have been built around a simple idea: The provider builds. The customer consumes. We standardized desktops. We standardized servers. We standardized networks. We defined what users were allowed to do, locked everything else down, and called it governance. It made perfect sense. Technology was complex. Expertise was scarce. Standardization created stability. But if I look at the direction our industry has taken over the past fifteen years, I don't see a story about better infrastructure. I see a story about increasing autonomy. And I don't think we've fully realized what that means for the future of managed services. This didn't start with AI AI is getting all the attention. But the shift started long before large language models. Think about what we've introduced over the last decade. Infrastructure as Code allowed engineers to describe infrastructure instead of manually configuring it. Cloud platforms removed the need to provision hardware. The modern workplace allowed users to work from anywhere, on almost any device. Power Platform enabled business users to automate processes without waiting for IT. Platform engineering is giving development teams self-service platforms instead of ticket queues. These aren't isolated innovations. They all move in exactly the same direction. Every generation of technology removes another dependency on central IT. Every generation gives more capability directly to the people creating value. AI simply accelerates that trend. Customers don't want fewer capabilities They want fewer dependencies. That's an important difference. Organizations don't want to submit tickets to deploy an application. They want to deploy it themselves. They don't want to wait three weeks for an environment. They want it in three minutes. They don't want IT departments approving every workflow. They want to automate their own. For years, many managed service providers viewed this as a threat. I think it's exactly the opposite. Because customers aren't trying to eliminate the MSP. They're trying to eliminate unnecessary friction. The MSP is no longer the builder Imagine a product team in three years. A product owner describes a new customer portal. An AI engineering team generates the application. Another agent provisions infrastructure. Security agents validate policies. Test agents perform functional and performance testing. Deployment agents roll everything into production. None of that feels unrealistic anymore. The interesting question isn't whether this will happen. It's what role the MSP still plays. I don't believe the answer is "building the platform." Because increasingly, customers will do that themselves. Or rather, their AI agents will. The foundation becomes the product If customers can build, deploy and operate faster than ever before, then the value of the MSP shifts underneath the visible work. The platform becomes the product. Not the portal. Not the virtual machine. Not the Kubernetes cluster. The invisible foundation beneath all of it. The landing zones. Identity. Networking. Compliance. Policies. Guardrails. Observability. Knowledge. Recovery. Customers won't ask an MSP to deploy an application. They'll expect an environment where deploying applications is safe by default. That's a fundamentally different business. Governance stops saying "no" Many organizations still think governance means restricting users. Removing permissions. Blocking installations. Limiting change. That approach worked when IT was responsible for every change. It breaks down completely when hundreds of developers, business users and AI agents are continuously creating new workloads. The answer cannot be to review every deployment. It cannot be to manually approve every prompt. And it certainly cannot be to lock everything down. Governance has to evolve from permission to policy. Instead of deciding who may build, we decide the conditions under which anything may be built. Instead of reviewing every change, we continuously validate every outcome. Instead of configuring environments manually, we enforce compliance automatically. Control doesn't disappear. It simply moves to a different layer. The MSP becomes an enabler of autonomy This may be the biggest mindset shift our industry has ever faced. For years, success was measured by how much work the provider performed. Tomorrow, success may be measured by how little intervention is required. The best managed service providers won't be the ones operating every workload. They'll be the ones enabling thousands of safe deployments that never required them in the first place. Their customers will move faster. Developers will have more freedom. Business teams will automate more processes. AI agents will continuously improve solutions. And underneath all of it, the MSP quietly ensures that security, compliance and operational resilience remain intact. Invisible when everything works. Essential when it doesn't. Expertise doesn't disappear Some people interpret AI as the end of expertise. History suggests otherwise. Every abstraction has increased demand for people who understand the layer beneath it. Cloud didn't eliminate infrastructure expertise. Infrastructure as Code didn't eliminate architects. Platform engineering didn't eliminate operations. It simply changed where expertise creates value. AI will do exactly the same. The future MSP won't spend its days deploying resources. It will design the ecosystems in which autonomous systems can safely deploy themselves. Closing thought I don't believe the future of managed services is about doing more work for customers. I think it's about making customers capable of doing more themselves. Not because the MSP becomes less relevant. But because relevance is moving. From operating technology... ...to enabling autonomy. The organizations that understand this will stop asking how AI fits into managed services. They'll realize managed services are being redefined by the same force that is reshaping every other part of IT: giving more control to the people closest to the problem, while ensuring the platform beneath them remains secure, compliant and resilient. That, to me, is what the next generation of managed services looks like.