Case Study: Operation and Monitoring Implementation Practice of a Large-Scale Comprehensive Science and Innovation Industrial Park
Project Background
As an integrated industrial park integrating scientific research experiments, corporate offices and supporting services, the science park is equipped with large-scale Huawei IT infrastructure, including firewalls, vulnerability scanners, switches, routers, wireless AC controllers and massive AP access points. These facilities support core businesses such as park scientific research network services, office wireless network and border security protection. The stability of network and security infrastructure is directly critical to the normal operation of scientific research work within the park.
With the continuous expansion of the park’s informatization scale and the growing number of network, security and wireless devices, the drawbacks of the original decentralized operation and maintenance model have become prominent. Multiple independent device management systems lack a unified monitoring foundation, and operation and maintenance work relies heavily on manual operations. Delayed fault detection and low troubleshooting efficiency can no longer meet the park’s requirements for high-reliability business operation. Therefore, it is urgent to build an intelligent monitoring platform to realize centralized monitoring and operation of the park’s IT infrastructure.
Core User Pain Points
The park features a wide variety and large quantity of security, network and wireless devices. The original operation and maintenance system has multiple practical problems that seriously threaten business continuity:
01. Isolated Management of Multi-Type Devices and Fragmented O&M Perspectives
The park’s security, network and wireless devices are managed through separate entrances. Firewalls, switches, routers, wireless AC/AP devices and vulnerability scanning devices operate on independent management systems. O&M personnel have to frequently switch between multiple platforms to check device statuses, making it impossible to obtain a unified overview of the overall operation status of IT infrastructure.
02. Inconsistent Unified Standards for Device Monitoring Metrics and High O&M Difficulty
The native monitoring metric standards of Huawei security, network and wireless devices in the park are inconsistent, with redundant metrics on some devices and missing key operating metrics. O&M personnel need to memorize different viewing logic for various devices, making it difficult to horizontally compare the operating status of similar devices.
03. Lack of Unified Visual Early Warning and Passive Fault Handling
Without a centralized visual monitoring platform, the operating status of infrastructure cannot be displayed intuitively. Hidden risks such as device performance bottlenecks, link anomalies and AP offline failures cannot be warned in advance. Operation and maintenance teams can only detect faults after users report network lagging and abnormal business access.
04. Heavy Reliance on Manual Inspection, Time-Consuming Fault Localization
Daily O&M relies heavily on manual inspection by logging into each device individually. This brings heavy workload and high risk of missed checks. In case of network outages or large-scale wireless failures, there is no unified alerting or data support. Engineers have to troubleshoot device by device, resulting in hour-long fault localization, which poses substantial risks to scientific research and office services.
05. Poor Native Alert Quality; Alert Storms Obscure Real Faults
Built-in alert rules vary widely across different devices. Some overly sensitive thresholds generate massive redundant, invalid alerts. Others have misconfigured alert severity levels, tagging minor incidents as critical risks. Genuine severe faults get buried under floods of alerts, making it difficult for O&M staff to identify valid hazards.
Actual Business Requirements
Combined with the above pain points, the park’s O&M team put forward requirements for building integrated monitoring:
- Centralized onboarding of all devices: Uniformly connect all Huawei security, network and wireless devices within the park to break silos between separate systems. Complete device monitoring and management on a single platform.
- Real-time collection and visualization of core metrics: Gather key metrics including device CPU, memory, port traffic, link status, AP online status, number of wireless connected users and signal quality for intuitive data presentation.
- Standardized alerting and noise reduction: Implement unified alert management with alert grading and filtering to suppress invalid alerts. Push alert notifications to the corresponding O&M owners in a timely manner.
- Dedicated wireless O&M capabilities: Provide exclusive monitoring views for large-scale AP deployments to quickly identify wireless issues such as offline APs, poor signal and abnormal connections, and ensure stable wireless network operation across the park.
Solution Implementation
Targeting the park’s business scenarios, the intelligent monitoring platform addresses all pain points via protocol adaptation, alert governance and dedicated view construction. The core implementation capabilities are as follows:
In-depth Multi-device Protocol Adaptation for Rapid Full Device Onboarding
Based on SNMP and API protocols, the platform delivers deep monitoring adaptation for Huawei firewalls, vulnerability scanners, switches, routers, wireless AC controllers and APs. Devices can be quickly onboarded to the monitoring platform without developing complex custom scripts.
It provides a unified device asset list, centrally displaying the IP address, business affiliation and operating status of all monitored devices. O&M staff can overview the online status of all devices on one single page, eliminating frequent switching between multiple systems.
Value: Build a unified asset inventory, gain global visibility of device availability and end fragmented management.

Alert Grading & Noise Reduction, Standardized Alert Governance to Filter Invalid Alerts
The platform supports full alert rule configuration, allowing custom alert severity, descriptions and filter conditions. To resolve overly sensitive native alerts and misclassified severity levels on Huawei devices, O&M staff can adjust alert severity and configure filtering rules to block duplicate and low-risk nuisance alerts.
Differentiated alert policies can be configured by device groups (routers, switches, firewalls, security appliances), with separate notification strategies for different device types. Notification channels such as email can be triggered both on alert generation and alert recovery, delivering targeted notifications to assigned O&M administrators.
Value: Eliminate alert storms and retain only valid risk alerts, ensuring O&M staff can promptly detect real faults.

Custom Dedicated Wireless Monitoring View to Stabilize Wireless Networks
A dedicated wireless monitoring dashboard is built for the park’s large number of wireless AC and AP devices. It collects real-time core data including AP monitoring status, online status, IP address, number of connected terminals, MAC address and uptime, and visually presents AP online rate, connected user volume and wireless signal quality.
O&M staff can quickly spot offline APs and abnormal access points without logging into each AC controller, rapidly locating wireless fault points and maintaining good wireless experience across office and research zones in the park.

Project Outcomes
- Improved O&M Efficiency: Centralized unified monitoring for security, network and wireless devices. No more switching between multiple systems. One platform handles full-scope device inspections.
- Reduced Management Costs: All device status data is aggregated and displayed in one place. It cuts the learning and operation overhead for O&M staff working across multiple systems and delivers full visibility of IT infrastructure.
- Strengthened Wireless Network Assurance: Batch-monitor AP operating status via dedicated wireless views to quickly locate wireless faults and maintain stable, available wireless networks throughout the park.
- Faster Fault Response: Proactive risk early warning replaces the fire-fighting, reactive O&M model. The average fault localization time is reduced from hours to minutes, minimizing the impact of failures on research and office services.

- Case: Unified Monitoring & Alerting Platform in “Double First-Class” Uni
- 案例解读 | 某三甲医院运维监控体系升级实例
- Case Analysis | Practice of Operation & Maintenance Management Platform for a World-leading Central SOE in Chemical Engineering Construction
- 案例解读 | 乐维助力某期货企业综合运维平台建设实践
- Case Study: Operation and Monitoring Implementation Practice of a Large-Scale Comprehensive Science and Innovation Industrial Park
- Case Interpretation: O&M System Construction for Listed Mfg. Enterprise