Case Analysis | Practice of Operation & Maintenance Management Platform for a World-leading Central SOE in Chemical Engineering Construction
PART 01 Project Background
01 Client Profile
The client of this case is a large central enterprise directly supervised by the State-owned Assets Supervision and Administration Commission of the State Council (SASAC). It stands as a leader and core backbone in China’s chemical engineering construction sector, with business operations spanning more than 80 countries and regions worldwide.
02 Pain Point Analysis
Driven by the in-depth advancement of the Digital China strategy, the enterprise has witnessed rapid expansion of its IT infrastructure and a remarkable rise in the complexity of its application systems. Meanwhile, propelled by the national Information Technology Application Innovation (ITAI) strategy, the group faces dual challenges of digital transformation and ITAI renovation. Specific operational and maintenance pain points are outlined as follows:
Insufficient Monitoring & Alert Capabilities
- Fragmented monitoring tools: Original monitoring tools are scattered across separate systems. O&M staff must switch between multiple platforms, making it impossible to build a unified O&M view.
- Severe alert storms: Lack of effective alert convergence and suppression mechanisms hinders rapid identification of critical faults, resulting in low efficiency of response and troubleshooting.
Weak Visualization & Reporting Functions
- Absence of a unified O&M dashboard: Senior management cannot intuitively grasp the overall health status of the group’s IT assets. The correlation between business operations and IT resources remains ambiguous, preventing quick assessment of fault impact scope on businesses.
- Dispersed data hinders decision-making: O&M data stored on isolated platforms lacks automated statistical and analytical tools. Manual data compilation fails to deliver reliable support for O&M decision-making.
Disordered Asset Ledger Management
- Manual-dependent asset administration: Assets are tracked purely via Excel spreadsheets, leading to delayed information updates and high error rates. Disconnection between asset records and monitoring data means resource changes cannot be automatically synchronized, generating massive redundant manual work.
- Opaque asset correlations: Without a dedicated asset management platform, automatic discovery of inter-asset dependencies is unavailable. Resource chains underpinning business systems remain unclear, making it hard to evaluate the impact scope of configuration changes.
Low Level of Automated Network O&M
- Low-efficiency manual inspections: Technicians log into devices to conduct scheduled manual checks, which are inefficient, prone to missed alarms and delayed fault detection.
- Difficult mass job execution: No capability for automated batch operations requires repetitive single-device configuration, consuming extensive working hours and carrying high risks of misoperation.
Complex Multi-cloud Governance
Multiple private and public clouds (Huawei Cloud, Tencent Cloud, China Telecom Cloud, ZStack, China Unicom Cloud, Intelligent Computing Cloud) operate under independent control. Cloud management platforms cannot exchange data with the existing O&M system, forming information silos with no unified dashboard for cloud resource oversight.
Heavy Pressure of ITAI Compatibility Mandates
Mandatory domestic IT substitution policies impose strict compliance requirements. Existing open-source tools (e.g., Prometheus, Zabbix) feature inadequate compatibility with domestic hardware and software, exposing the enterprise to compliance risks. Full-stack compatibility with domestic ITAI environments is therefore imperative.
PART 02 Lerwee Solution
Based on the client’s actual business conditions, Lerwee delivers a customized end-to-end informatized O&M management platform solution built around the core concepts of Unified Monitoring, Intelligent Alerting, Refined Management, Automated O&M. Adopting a phased implementation roadmap, we completed deployment covering production environments, disaster recovery environments and 15 collection agents, establishing an O&M management framework covering all IT infrastructure across the group.
- Achieve unified access and centralized monitoring for operating systems, databases, middleware, network devices, virtualization platforms and cloud platforms.
- Deep integration and interconnection with the group’s existing ITSM, master data system, IAM unified identity authentication system, cloud portal/cloud management platform and WeCom.
- Functional optimization tailored to real O&M scenarios, including optimized alert convergence, automatic business topology discovery, 3D computer room visualization and a dedicated automated script library.
- Full ITAI compatibility adaptation for Kylin V10 OS and Kingbase V9 database.

01 Monitoring & Alerting
Unified monitoring coverage includes 701 operating systems, 6 database brands (Oracle, MySQL, DB2, DM, Kingbase, GaussDB), over 200 network devices, 2 virtualization platforms, 3 cloud platforms and a wide range of middleware products.
Alert convergence & classification: More than 200 alert rules are configured to realize centralized aggregation, convergence, suppression and hierarchical classification of alerts. Alerts are divided into five tiers: Critical, Major, Minor, Warning and Information, with matched notification channels (SMS, Email, WeCom). Critical alerts trigger immediate phone calls; Major alerts automatically generate ITSM work orders.

02 Visualization & Reporting
- O&M Control Center: A group-wide large-screen O&M cockpit displays an overview of total group IT resources, alert statistics, health scores, resource distribution and the operational status of core business systems.

- 3D Computer Room: Digital twin 3D modeling for 118 cabinets across 3 computer rooms, supporting real-time display of cabinet-level device status, U-space management, temperature distribution and flashing alert reminders.
- Business Topology: Automatically identifies business resources, dependencies, application processes and service ports based on business IP addresses, generating business topology diagrams for over 200 business systems and supporting business SLO health scoring.


03 CMDB Asset Management
Asset Discovery & Onboarding: Powered by the Perseus collector, resource data is captured via monitoring templates, with automatic mapping rules linking monitoring metrics to CMDB model attributes. Over 6,000 assets are automatically imported into CMDB, covering five asset categories (compute, network, storage, computer room, software) across more than 80 data models.

Full-lifecycle management: End-to-end lifecycle tracking for devices covering registration, utilization, modification, decommissioning and recovery. Asset QR codes support mobile on-site inspection, while spare parts management enables traceable inquiries.
04 Network Management
- Network device monitoring: Monitors over 200 network devices from vendors including Huawei, H3C, Cisco and Ruijie, tracking key metrics such as interface status, traffic, packet errors, CPU and memory utilization.
- Configuration backup & audit: Stores more than 5,000 configuration backup records, supporting baseline comparison and change auditing. The system automatically sends notifications upon backup failures or detected configuration discrepancies.
- IP Address Management: Governs over 50 address segments and more than 5,000 IP allocation records, enabling IP status checking, MAC address tracing, and pinpointing of access devices and ports.

05 Expansion & Implementation of Automated O&M Functions
- Script Library Development: A repository of over 200 O&M scripts (Shell/Python/PowerShell) covering daily inspection, log cleanup, service restart, backup verification and other common scenarios, equipped with version control capabilities.
- Job Management Platform: Supports scheduled, cyclic and immediate job execution to enable scheduled batch task distribution and operation. The platform stores over 100 defined jobs and more than 10,000 historical execution records.
- Fault Self-healing Configuration: Automated self-recovery scripts are deployed for common faults such as insufficient disk space, suspended services and log accumulation to cut manual intervention costs.

06 System Integration & Interconnection Implementation
- ITSM Work Order Linkage: Major and higher-tier alerts automatically generate ITSM fault tickets with auto-mapped and populated ticket fields. Ticket status updates are synchronized back to the monitoring platform, forming a closed-loop alert-ticket management workflow.
- Master Data Synchronization: Interconnected with the group’s master data system to realize automatic synchronization of organizational structures, positions, staff profiles, business systems and supplier information with a 100% synchronization success rate.
- IAM Unified Authentication: Implemented Single Sign-On (SSO) integration with the group’s unified identity authentication system, delivering accurate user permission mapping.
- Cloud Platform Interconnection: Cloud API integration with Huawei private cloud, Tencent public cloud and ZStack private cloud for automatic cloud resource onboarding and synchronized monitoring data.
- WeCom Notifications: Alert reminders, approval notifications and daily reports are pushed in real time via WeCom, compatible with privatized deployment interconnection.
PART 03 Project Outcomes
- Comprehensive unified monitoring covers products from over 20 vendors, more than 300 device models and over 50,000 refined monitoring metrics. With over 200 alert rules deployed, the alert convergence rate exceeds 85%, thoroughly resolving alert storm issues.
- Over 6,000 assets are automatically brought under centralized management, with complete visualization of dependency chains across more than 200 business systems, lifting asset management efficiency by over 70%.
- Rapid location and tracing of more than 5,000 IP addresses accelerates troubleshooting and asset inventory workflows. Refined IP resource control ensures long-term stable operation of large-scale multi-terminal networks.
- Daily inspection efficiency rises by over 300%: inspection tasks that previously took hours are completed within 30 minutes. Large-scale batch configuration rollouts that once required days now finish within one hour.
- 100% success rate for ITSM ticket-alert linkage, boosting average alert processing efficiency by 30%. Master data synchronization maintains a 100% success rate, while IAM SSO enables “one login, full-network access” with drastically improved user experience.
- The platform fully complies with central SOE domestic IT substitution policies through complete ITAI adaptation for operating systems and databases. It comprehensively safeguards the stability, security and reliability of the entire system within ITAI environments, successfully fulfilling compliance targets for informatization localization renovation.
Conclusion
Moving forward, alongside continuous advancement in digital technologies and sustained business innovation of enterprises, Lerwee will keep exploring more intelligent O&M management application scenarios. We will stand guard over enterprises’ digital transformation journeys and jointly forge a brighter future.

- Case Interpretation: O&M System Construction for Listed Mfg. Enterprise
- Case Study: Monitoring & Network Management for a Listed Electronic Circuit Substrate Enterprise
- O&M Practice | Lerwee Monitoring Helps Stable Operation of Medical Business
- Case Study | Lerwee Monitoring Helps a Large Cigarette Factory Build an Efficient O&M Monitoring System
- 案例解读 | 专业化赋能,乐维助力某大型信息技术企业数字化转型升级
- 案例解读 | 某大型国际证券企业智能运维平台建设实践