Multi-mode intelligent linkage 3D visual data center operation and maintenance system and method

Through the 3D visual data center operation and maintenance system of multimodal data fusion and three-dimensional particle modeling, the problem of multi-source information separation in traditional systems is solved, the intelligence and real-time improvement of the data center is achieved, and the operation and maintenance efficiency and fault prevention and control capabilities are enhanced.

CN120451381AActive Publication Date: 2025-08-08NANJING ARSENIC ELECTRONIC TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510501887.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-08
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

Traditional data center operation and maintenance monitoring systems are difficult to achieve rapid perception, spatial positioning and trend prediction of multi-source information, and lack unified modeling and semantic linkage of multi-modal data, resulting in insufficient intelligent analysis and prediction capabilities, affecting system stability and operation and maintenance efficiency.

Method used

Build a multimodal intelligent linkage 3D visual data center operation and maintenance system, and realize three-dimensional visual interactive display through multimodal data fusion, three-dimensional particle modeling, semantic scoring analysis and structured prediction inference, combining knowledge graphs to perform causal relationship analysis and control instructions optimization, so as to realize three-dimensional visual interactive display.

Benefits of technology

It enhances the intelligence, real-time and visualization of data center operation and maintenance, supports unified modeling and linkage control of multimodal data, improves the capabilities of fault prevention and control and operation and maintenance planning, and improves the spatial perception and operation and maintenance efficiency of operation and maintenance personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451381A_ABST
    Figure CN120451381A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode intelligent linkage 3D visual data center operation and maintenance system and method, and relates to the technical field of data center operation and maintenance. The system comprises a multi-modal data fusion module used for mapping physical sensor data and video monitoring and operation and maintenance logs to a unified three-dimensional coordinate system and constructing a multi-modal fusion feature tensor; the three-dimensional particle modeling module is used for constructing a particle space structure based on a Voronoi diagram and Delaunay triangulation and adaptively adjusting the particle resolution; the multi-modal score analysis module is used for calculating a particle multi-dimensional score and generating a semantic heat map to recognize an abnormal region; the structured noise prediction module is used for simulating future responses under different instructions based on a diffusion model; the instruction generation and regulation module is used for realizing causal analysis and instruction optimization in combination with a knowledge graph; and the three-dimensional visual interaction module supports real-time rendering and interaction operation at a client. According to the method, the intelligence, the real-time performance and the visualization level of data center operation and maintenance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data center operation and maintenance technology, and in particular to a multi-modal intelligent linked 3D visualization data center operation and maintenance system and method. Background Art

[0002] As data centers continue to expand and their operating environments become more complex, traditional operations and maintenance monitoring methods struggle to meet the requirements for rapid multi-source information perception, spatial positioning, and trend prediction. Current common data center monitoring systems primarily rely on two-dimensional dashboards, tabular data, or partial video displays. These systems suffer from significant deficiencies such as information fragmentation, delayed response to exceptions, and a lack of causal support for control operations. Especially when faced with complex events like sudden temperature rises and the spread of equipment failures, the lack of unified modeling and semantic linkage mechanisms for multimodal data leads to insufficient intelligent analysis and prediction capabilities, impacting system stability and operational efficiency.

[0003] Some studies have proposed combining sensor data and image information for condition monitoring, such as combining camera images with thermal sensor data for anomaly detection. However, these approaches are often limited to local regions, single-modal features, or static analysis, lacking multimodal fusion, interactive visualization, and control decision support based on a unified spatial model. In the field of image processing, although multimodal vision models such as CLIP and ViT have been widely used for semantic understanding of images and text, there is still a lack of mature solutions for collaborating with heterogeneous information such as physical data and operation and maintenance logs to build a system-level analysis and control framework.

[0004] Therefore, there is an urgent need for a data center operation and maintenance system that can integrate multimodal heterogeneous data, support semantic analysis and three-dimensional visualization, and have predictive simulation and linkage control capabilities, so as to solve the problems of perception delay, analysis fragmentation and regulation lag in existing technologies, and improve the system's intelligence, real-time performance and visual interaction level. Summary of the Invention

[0005] The present invention proposes a 3D visualization data center operation and maintenance system and method with multimodal intelligent linkage. By introducing an integrated architecture including multimodal feature tensor construction, three-dimensional particle modeling, semantic scoring analysis, structured predictive reasoning, and visualization linkage control, it aims to solve the problems existing in the existing technology, such as multi-source data fragmentation, lack of real-time spatial correlation and predictive control capabilities.

[0006] The present invention achieves the above-mentioned purpose through the following technical solutions:

[0007] A multi-modal intelligently linked 3D visualization data center operation and maintenance system, comprising:

[0008] The multimodal data fusion module is used to collect physical sensor data, video surveillance data, and operation and maintenance log data from the data center. Through time synchronization and spatial alignment, it maps the different modal data to the same three-dimensional spatial coordinate system and constructs a multimodal fusion feature tensor.

[0009] A three-dimensional particle modeling module is used to construct the particle space structure of the data center based on the Voronoi diagram and Delaunay triangulation method, establish a mapping relationship between particles and data items in the multimodal fusion feature tensor, dynamically adjust the visualization resolution level of particles according to the distance between the particles and the current viewpoint, and output the current particle state set;

[0010] The multimodal scoring analysis module is used to calculate the local concept score, global concept score, local visual score, and global visual score of each particle in the current particle state set, generate a semantic heat map based on the weighted scoring results, and identify abnormal hot spots that meet the abnormality prediction conditions;

[0011] The structured noise prediction module is used to construct a sketch of the future particle state for prediction. It generates predicted particle states based on the structured noise inverse initialization and diffusion inference method to simulate the system response that different control instructions may cause in the future.

[0012] An instruction generation and control module is used to combine the abnormal hotspot area, predicted particle state and knowledge graph rules to perform causal relationship analysis and control instruction optimization, and output a device control instruction set;

[0013] The three-dimensional visualization interaction module is used to realize three-dimensional visualization interactive display on the client through GPU acceleration and WebGL rendering technology, including at least real-time particle distribution rendering, abnormal hot spot area highlighting, trend evolution trajectory animation and device control instruction interactive operation.

[0014] A further improvement of the present invention is that the multimodal data fusion module includes:

[0015] A sensor data unit is used to obtain physical sensor data from various devices in the data center, including at least temperature, humidity, and power consumption, synchronize the physical sensor data according to timestamps, and map the data to a three-dimensional spatial coordinate system based on the spatial installation location of each sensor to form physical features;

[0016] The video surveillance data unit is used to obtain the video surveillance data of the data center and perform image frame-level decoding, extract the visual features of the target area in the image, and map the visual features into three-dimensional space by combining the internal and external parameter information of the corresponding camera;

[0017] The operation and maintenance log data unit is used to perform semantic analysis on the operation and maintenance log text and structure it into event triples, which are then mapped to a three-dimensional spatial coordinate system according to preset spatial semantic mapping rules.

[0018] Data fusion unit, used to fuse physical features, visual features and semantic features in a unified three-dimensional spatial coordinate system and time dimension to generate a multimodal fusion feature tensor Among them, x, y, z represent three-dimensional space coordinates, and t represents time.

[0019] A further improvement of the present invention is that, in the three-dimensional particle modeling module, the method for constructing the particle spatial structure includes:

[0020] Taking the physical sensor nodes deployed in the data center space as spatial seed points, a weighted Voronoi diagram is used to divide the three-dimensional space into regions. The weight value of each physical sensor node is related to the sensor's sensing coverage, the frequency of historical abnormal records, or the importance of the region.

[0021] Extracting a centroid set within the weighted Voronoi region and performing Delaunay triangulation on the centroid set to construct a three-dimensional particle network structure;

[0022] After the construction is completed, a one-to-one mapping relationship is established based on the tensor value in the multimodal fusion feature tensor in the area where each particle is located;

[0023] The three-dimensional particle modeling module monitors the local change gradient of the multimodal fusion feature tensor during system operation, and dynamically adjusts the particle distribution structure based on the preset reconstruction threshold, performs particle refinement in high-variability areas, and performs particle aggregation reconstruction in stable areas.

[0024] A further improvement of the present invention is that the three-dimensional particle modeling module determines the visualization resolution level of each particle based on the spatial distance between each particle and the current viewing angle center point and the temperature gradient of the area where the particle is located. The resolution level scheduling mechanism specifically includes:

[0025] For each particle p i , calculate the Euclidean distance d between the particle and the center of the viewing angle i , and calculate the temperature gradient amplitude of the local area where the particle is located

[0026] Construct particle rendering weight function:

[0027]

[0028] Where W i is the particle rendering weight; σ is the Sigmoid function; β and γ are the system preset adjustment coefficients;

[0029] According to the particle rendering weight W i The value in the range of [0,1] corresponds to different visualization resolution levels and determines whether to retain, aggregate or refine the rendering granularity of the particle in the current frame;

[0030] During system operation, the visual resolution level of each particle can be dynamically updated to support visually continuous transition and smooth presentation, ensuring that particles in abnormal hot spots are rendered delicately while areas far from the viewing angle and changing slowly are dynamically simplified.

[0031] A further improvement of the present invention is that the current particle state set includes at least:

[0032] The three-dimensional coordinate position of the particle (x i ,y i ,z i ), used for spatial positioning;

[0033] Real-time monitored temperature value T i , humidity value H i And power consumption value P i , used to characterize the physical environment state of the spatial region corresponding to the particle;

[0034] Fusion feature vector v i , the fusion feature vector is composed of the visual features and semantic features extracted from the multimodal fusion feature tensor of the region where the particle is located;

[0035] Visualization level label L i , used to indicate the visualization resolution level of particles at the current viewing angle;

[0036] State entropy It is used to measure the normalized uncertainty of the particle's temperature, humidity, and power consumption changes in the last several consecutive moments. The calculation formula is:

[0037]

[0038] Where p k is the distribution probability of any physical property k of the particle within a time window;

[0039] Life cycle counter C i , used to record the number of state updates of the particle since its generation, and trigger the particle resampling process when the upper limit of the example update life cycle set by the system is exceeded;

[0040] Abnormal cluster label A i , used to identify whether the particle is clustered and belongs to any abnormal hotspot area or potential fault domain in abnormal pattern recognition, and is used to support the priority calculation and abnormal tracking of equipment control instructions.

[0041] A further improvement of the present invention is that, in the multimodal scoring analysis module, the local concept score LC is based on a weighted average method that integrates local semantic heat and mask area ratio to measure the semantic concept significance of the local area where the particle is located;

[0042] The global concept score GC is calculated through the image-text embedding similarity of the Alpha-CLIP model to evaluate the global semantic consistency between the particle image region and the target semantic text;

[0043] The local visual score LV is obtained by calculating the feature similarity between the support image and the current image in the local area of the particle. The calculation process includes the local average aggregation of the spatial matching matrix.

[0044] The global visual score GV is the overall embedding distribution between the support image extracted by the ViT or DINO model and the current image, and the EMD is used to measure the similarity of the particle region feature distribution between different images.

[0045] A further improvement of the present invention is that the contextual semantic consistency factor γ(p i ), the multimodal large model scores the consistency between the particle semantics and the global description of the current operation and maintenance scenario. The value range is [0,1], which is used to adjust the local concept score and the global concept score to form the modified scoring item:

[0046] LC′(p i )=LC(p i )·γ(p i ), GC′(p i )=GC(p i )·γ(p i );

[0047] Where p i represents the i-th particle, LC(p i ) is the particle p i The local concept score, GC(p i ) is the particle p i The global concept score of LC′(p i ) is the corrected particle p i The local concept score, GC′(p i ) is the corrected particle p i Global concept score of

[0048] Introducing the semantic-behavior association vector a(p i ), which is obtained by collaborative reasoning of text and video modalities, is used to predict the potential causal relationship between the particle area and historical operation and maintenance events, and calculate the behavior enhancement weight term R(pi ):

[0049] R(p i )= <a(p i ),Emb event >

[0050] Where, Emb event is the semantic embedding vector of historical known fault events; <·,·> represents the vector dot product;

[0051] According to the multimodal score distribution of the local area where the particle is located, the particle information entropy is calculated:

[0052]

[0053] Where H local (p i ) is the particle p i Information entropy of the surrounding local score distribution; j is the score type index, j∈{1,2,3,4}; S j (p i ) is the particle p i The j-th score value of H global is the average entropy of the score distribution of all particles; mean i It means taking the average of all particles numbered i;

[0054] Define the local-global entropy difference △H(p i )=H local (p i )-H global , and is used to adjust the weights of each weighted item as follows:

[0055] β′ j (p i )=β j ·(1+λ·△H(p i ));

[0056] In the formula, △H(p i ) is the particle p i The local-global entropy difference; λ is the entropy difference adjustment coefficient; β j represents the basic weight of the j-th score; β′ j (p i ) is for particle p i The j-th weighting coefficient after dynamic adjustment;

[0057] Calculate the final weighted score S of the particle:

[0058] S(p i )=β′1(p i )·LC′(p i )+β′2(pi )·GC′(p i )+β′3(p i )·LV(p i )+β′4(p i )·GV(p i );

[0059] Where, S(p i ) is the particle p i The final weighted score of LV(p i ) is the particle p i Local visual score; GV(p i ) is the particle p i Global visual score of

[0060] Determine whether it constitutes an abnormal hotspot area. The abnormal prediction conditions are as follows:

[0061]

[0062] Where τ1 is the abnormality score threshold; For particle p i The temperature change gradient at the position; θ1 is the temperature change gradient threshold; M heat (p i ) is the historical fault mask; θ2 is the temperature mask heat intensity threshold; ρ is the behavior correlation judgment threshold.

[0063] A further improvement of the present invention is that the structured noise prediction module includes:

[0064] The particle state sketch generator is used to construct a future particle state sketch within the prediction period based on the current particle state set, abnormal hotspot areas, and semantic heat maps. The sketch includes particle positions, attribute trends, and semantic label distributions, and supports the generation of particle state evolution profiles at multiple time granularities.

[0065] Structured noise initializer, used to perform structured noise inverse initialization on the future particle state sketch encoding to generate the initial state The calculation formula is:

[0066]

[0067] Where z0 is the latent variable representation of the future particle state sketch; α t is the noise scheduling coefficient of the diffusion time step t; ∈ is the standard Gaussian structured noise;

[0068] The anti-diffusion reasoner is used to As the starting point, the reverse diffusion process is performed to generate the predicted particle state sequence P (t+△t), and simulate the evolution trajectory of system response under different control instructions respectively;

[0069] Anomaly enhancement mechanism is used to add anomaly weight coefficients to particles located in abnormal hot spots during particle prediction, so that the prediction process can enhance the focus in high temperature and high entropy areas;

[0070] The closed-loop feedback update unit is used to compare the actual feedback state of the device with the predicted particle state and construct a prediction error function:

[0071]

[0072] Where E(t) is the prediction error function value at time t; N is the total number of particles; is the predicted state vector of the i-th particle; p i real is the actual feedback state vector of the i-th particle;

[0073] When the prediction error function value E(t) exceeds the set threshold, the structured noise generation parameters and the future particle state sketch are automatically updated.

[0074] A further improvement of the present invention is that the instruction generation and control module includes:

[0075] A causal path modeler is used to combine the weighted scoring results of the abnormal hotspot areas, the semantic heat map text and the device behavior rules defined in the knowledge graph to construct a causal path diagram of events, causal relationships and controllable variables;

[0076] A control instruction candidate generator is used to identify controllable variables from the causal path diagram and call candidate device control instructions corresponding to the controllable variables from the historical operation and maintenance instruction knowledge base. i , and combined with the predicted particle state to simulate the system response of each candidate device control instruction at different time granularities;

[0077] The semantic response evaluation module is used to call the multimodal language model to generate a natural language summary of the device response in the simulated prediction state for each candidate control instruction, and compare the semantic consistency with the predicted particle scene to calculate the response credibility score R (a i );

[0078] The instruction conflict detection module is used to detect the combinations of mutually exclusive operations, resource occupation conflicts or device status restrictions in the candidate device control instruction set, and eliminate the instruction groups that cannot be executed in parallel;

[0079] The execution risk assessment module is used to construct the instruction execution risk score ρ(a i), as the penalty factor of the control instruction utility function;

[0080] Fusion optimization module, used to integrate the final weighted score S and response credibility score R (a i ) and execution risk score ρ(a i ), calculate the control instruction utility function U(a i ):

[0081] U(a i )=λ1·S+λ2·R(a i )-λ3·ρ(a i );

[0082] Where λ1, λ2, and λ3 are the weighting coefficients set by the system;

[0083] Instruction output selector, used to control the utility function U(a i ) priority, sort the candidate device control instruction sets, and output the optimal device control instruction set that can be executed by the current system to the device control end.

[0084] A multi-modal intelligent linkage 3D visualization data center operation and maintenance method is based on the multi-modal intelligent linkage 3D visualization data center operation and maintenance system described above, and the method includes:

[0085] Collect physical sensor data, video surveillance data, and operation and maintenance log data from the data center. Through time synchronization and spatial alignment, the different modal data are uniformly mapped to the same three-dimensional spatial coordinate system to construct a multimodal fusion feature tensor.

[0086] The particle space structure of the data center is constructed based on the Voronoi diagram and Delaunay triangulation method, a mapping relationship is established between the particles and the data items in the multimodal fusion feature tensor, and the visualization resolution level of the particles is dynamically adjusted according to the distance between the particles and the current viewing angle, and the current particle state set is output;

[0087] Calculate the local concept score, global concept score, local visual score, and global visual score of each particle in the current particle state set, generate a semantic heat map based on the weighted scoring results, and identify abnormal hot spots that meet the abnormality prediction conditions;

[0088] Construct a sketch of future particle states for prediction, and generate predicted particle states based on structured noise inverse initialization and diffusion inference methods to simulate the system responses that may be caused by different control instructions in the future;

[0089] Combining the abnormal hotspot area, predicted particle state and knowledge graph rules, causal relationship analysis and control instruction optimization are performed to output a device control instruction set;

[0090] 3D visual interactive display is achieved on the client through GPU acceleration and WebGL rendering technology.

[0091] The beneficial effects of the present invention are as follows: by constructing a three-dimensional multimodal fusion feature tensor, the representation of different types of data is unified, and information integrability and spatiotemporal traceability are enhanced; the particle space structure is constructed using the Voronoi diagram and Delaunay triangulation method, and the particle density and visualization resolution are adaptively adjusted in combination with local gradient and viewing angle factors, thereby improving the information presentation accuracy and global rendering efficiency of key areas; multimodal scores are calculated based on local and global concept scores, visual similarity analysis, and semantic causal enhancement to form a semantic heat map, which supports the rapid identification and spatial focusing of potential fault areas; based on structured noise modeling and diffusion inference mechanism, the future particle state change trajectory is predicted, and the system evolution process under the action of different control instructions can be simulated, thereby enhancing fault prevention and control and operation and maintenance planning capabilities; causal paths are constructed by combining predicted states, knowledge graphs, and language models, and the semantic consistency and risk scores of candidate control instructions are generated and evaluated, realizing intelligent optimization and execution conflict avoidance of equipment control instructions; dynamic display and operation feedback of particle states, abnormal areas, and control processes are realized based on GPU acceleration and WebGL rendering, enhancing the spatial perception and operation decision-making efficiency of operation and maintenance personnel. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0093] in:

[0094] Figure 1 It is a system structure block diagram of the present invention;

[0095] Figure 2 is a structural block diagram of a multimodal data fusion module in an embodiment of the present invention;

[0096] Figure 3 3D particle modeling module implementation flow chart of the embodiment of the present invention;

[0097] Figure 4 is a flowchart of an implementation of a structured noise prediction module in an embodiment of the present invention;

[0098] Figure 5 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0099] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of the present invention.

[0100] like Figure 1 As shown, an embodiment of the present invention provides a multimodal intelligent linkage 3D visualization data center operation and maintenance system, which at least includes a multimodal data fusion module, a three-dimensional particle modeling module, a multimodal scoring analysis module, a structured noise prediction module, an instruction generation and control module, and a three-dimensional visualization interaction module.

[0101] (1) Multimodal data fusion module

[0102] The multimodal data fusion module is used to collect physical sensor data, video surveillance data, and operation and maintenance log data from the data center. Through time synchronization and spatial alignment processing, different modal data are uniformly mapped to the same three-dimensional spatial coordinate system to construct a multimodal fusion feature tensor.

[0103] like Figure 2 As shown, in one embodiment, the multimodal data fusion module includes:

[0104] The sensor data unit is used to obtain physical sensor data from various devices in the data center, including at least temperature, humidity, and power consumption. The physical sensor data is synchronized according to timestamps (such as timestamp alignment based on NTP) and mapped to a three-dimensional spatial coordinate system based on the spatial installation location of each sensor to form physical features.

[0105] The video surveillance data unit is used to obtain the video surveillance data of the data center and perform image frame-level decoding, extract the visual features of the target area in the image, and map the visual features into three-dimensional space by combining the internal and external parameter information of the corresponding camera;

[0106] The operation and maintenance log data unit is used to perform semantic analysis on the operation and maintenance log text (such as named entity recognition and event extraction technology) and structure it into event triples (for example, "server X", "temperature abnormality", "2024-10-01 14:22"), and map it to a three-dimensional spatial coordinate system according to preset spatial semantic mapping rules (such as a mapping table between equipment number and location, and a matching rule between fault events and facility areas);

[0107] Data fusion unit, used to fuse physical features, visual features and semantic features in a unified three-dimensional spatial coordinate system and time dimension to generate a multimodal fusion feature tensor Among them, x, y, z represent three-dimensional space coordinates, and t represents time.

[0108] (2) 3D particle modeling module

[0109] The three-dimensional particle modeling module is used to construct the particle space structure of the data center based on the Voronoi diagram and Delaunay triangulation method, establish a mapping relationship between particles and data items in the multimodal fusion feature tensor, dynamically adjust the particle visualization resolution level according to the distance between the particle and the current viewing angle, and output the current particle state set.

[0110] Specifically, if Figure 3 As shown, the method for constructing the particle space structure includes:

[0111] Using physical sensor nodes deployed within the data center (such as temperature and humidity sensors, power consumption meters, and smoke detectors) as spatial seed points, a weighted Voronoi diagram is used to partition the three-dimensional space. The weight of each physical sensor node is related to the sensor's coverage, the frequency of historical anomaly records, or the importance of the area (for example, the core server area has a greater weight than the peripheral area). Based on these weighted values, the three-dimensional spatial sub-area dominated by each sensor node is delineated, ensuring that the spatial partitioning is both reasonable and fault-sensitive.

[0112] A set of centroids (i.e., the geometric center or characteristic points of the region) is extracted within the weighted Voronoi region and Delaunay triangulation is performed on the centroids to construct a three-dimensional particle network structure. This network has good spatial coverage and topological resolvability, which is conducive to subsequent information dissemination, reasoning, and visualization operations.

[0113] After the construction is completed, a one-to-one mapping relationship is established based on the tensor values in the multimodal fusion feature tensor in the area where each particle is located, so that each particle can carry the fusion perception information of its area; the mapping data includes but is not limited to: current values of physical sensors (temperature, humidity, power consumption), video visual features, semantic event labels, etc.; this mapping relationship gives particles multi-source attributes, making particles the carrier unit of spatial data.

[0114] During system operation, the three-dimensional particle modeling module monitors the local change gradient of the multimodal fusion feature tensor and dynamically adjusts the particle distribution structure based on a preset reconstruction threshold. It refines particles in high-variability areas (such as drastic temperature changes) and increases particle density to improve monitoring accuracy. It performs particle aggregation reconstruction in stable areas (with smooth changes and low risks) and reduces particle density to reduce the computational burden.

[0115] Through the three-dimensional particle modeling method described in this embodiment, the system can build an efficient and adjustable spatial particle model based on the actual deployment environment, significantly improving the visualization and spatial analysis capabilities of multimodal data.

[0116] Furthermore, the 3D particle modeling module determines the visualization resolution level of each particle based on the spatial distance between each particle and the current viewing center point and the temperature gradient of the area where the particle is located. The resolution level scheduling mechanism specifically includes:

[0117] For each particle p i , calculate the Euclidean distance d between the particle and the center of the viewing angle i , and calculate the temperature gradient amplitude of the local area where the particle is located

[0118] d i =‖X i -C view ‖;

[0119]

[0120] Where, X i For particle p i The corresponding three-dimensional coordinates; C view The 3D coordinates of the center point of the current user's perspective obtained by the system in real time during the client rendering process; For particle p i The set of spatial neighborhood particles of Any particle of T i 、T b Respectively represent the real-time temperature value of the particles;

[0121] Construct particle rendering weight function:

[0122]

[0123] Where W i is the particle rendering weight; σ is the Sigmoid function; β and γ are the system preset adjustment coefficients;

[0124] According to the particle rendering weight W i The value in the range of [0,1] corresponds to different visualization resolution levels and determines whether to retain, aggregate or refine the rendering granularity of the particle in the current frame;

[0125] During system operation, the visual resolution level of each particle can be dynamically updated to support visually continuous transition and smooth presentation, ensuring that particles in abnormal hot spots are rendered delicately while areas far from the viewing angle and changing slowly are dynamically simplified.

[0126] In one preferred embodiment, the current particle state set includes at least:

[0127] The three-dimensional coordinate position of the particle (x i ,y i ,z i ), used for spatial positioning;

[0128] Real-time monitored temperature value T i , humidity value H i And power consumption value P i , used to characterize the physical environment state of the spatial region corresponding to the particle;

[0129] Fusion feature vector v i , which is composed of the visual features and semantic features extracted from the multimodal fusion feature tensor of the region where the particle is located;

[0130] Visualization level label L i , used to indicate the visualization resolution level of particles at the current viewing angle;

[0131] State entropy It is used to measure the normalized uncertainty of the particle's temperature, humidity, and power consumption changes in the last several consecutive moments. The calculation formula is:

[0132]

[0133] Where p k is the distribution probability of any physical property k of the particle within a time window;

[0134] Life cycle counter C i , used to record the number of state updates of the particle since its generation, and trigger the particle resampling process when the upper limit of the example update life cycle set by the system is exceeded;

[0135] Abnormal cluster label A i , used to identify whether the particle is clustered and belongs to any abnormal hotspot area or potential fault domain in abnormal pattern recognition, and is used to support the priority calculation and abnormal tracking of equipment control instructions.

[0136] (3) Multimodal rating analysis module

[0137] The multimodal scoring analysis module is used to receive the current particle state set and the multimodal fusion feature tensor, calculate the local concept score, global concept score, local visual score and global visual score of each particle in the current particle state set, generate a semantic heat map based on the weighted scoring results, and identify abnormal hot spots that meet the abnormality prediction conditions.

[0138] Specifically, the local concept score LC is based on a weighted average method that combines local semantic heat and mask area ratio to measure the semantic concept significance of the local area where the particle is located.

[0139] The calculation formula of the local concept score LC (Local Conceptual Score) is:

[0140]

[0141] Where p i represents the i-th particle, LC(p i ) is the particle p i The local concept score; α is the weighting factor of local significance and heat, 0<α<1; m i For particle p i Corresponding semantic mask area; RTA(x,y) represents the semantic heat map intensity (Refined Text Alignment) at the two-dimensional coordinate (x,y); Area(m i ) is the mask m i The area of the corresponding region; Area(M) is the total area of all masks.

[0142] The global concept score GC is calculated through the image-text embedding similarity of the Alpha-CLIP model to evaluate the global semantic consistency between the particle image region and the target semantic text;

[0143] The global concept score GC is calculated by the Alpha-CLIP model as follows:

[0144]

[0145] Where, GC(p i ) is the particle p i Global concept score of is the normalized semantic embedding vector corresponding to the target text (such as alarm description); For particle p i The normalized visual embedding vector corresponding to the image region.

[0146] The local visual score (LV) is obtained by calculating the feature similarity between the supporting image and the current image in the local area of the particle. The calculation process includes the local average aggregation of the spatial matching matrix.

[0147] The local visual score is used to evaluate the visual consistency between the support image and the current image in the area where the current particle is located. This score measures the degree of visual semantic similarity between the known equipment or status areas in the support sample and the corresponding areas in the current operation and maintenance scenario. It is one of the important indicators for achieving spatial anomaly recognition and anomaly tracking.

[0148] In the specific implementation process, the calculation of the local vision score LV includes the following steps:

[0149] Support image I s and the current image I q Also input a pre-trained visual encoder (such as DINO v2, CLIP or Vision Transformer);

[0150] For each particle p i The three-dimensional space coordinates of the image are located in the projection area R in the image plane. i , extract the corresponding set of visual embedding vectors:

[0151] F s ={f s (x,y)},F q ={f q (x′,y′)},(x,y),(x′,y′)∈R i ;

[0152] The support image and the current image in the particle local area R i The embedded features in the are subjected to pairwise cosine similarity calculation to obtain the spatial matching matrix:

[0153]

[0154] Based on the matching matrix M, the maximum matching value is locally averaged and aggregated to calculate the local visual score of the particle. Based on this, the maximum matching value is locally averaged and aggregated to calculate the local visual score of the particle:

[0155]

[0156] This score reflects whether the local area of the current image has a similar semantic area in the support image, which can be used to determine whether the device status is consistent, whether the structure has changed, or whether anomalies exist.

[0157] The visual encoder used in the present invention can use the current mainstream image representation learning network, including but not limited to:

[0158] CLIP (Contrastive Language–Image Pretraining) model;

[0159] DINO or DINO v2;

[0160] LoFTR (Detector-Free Local Feature Matching);

[0161] Lightweight self-attention structures such as Swin Transformer and ViT-B.

[0162] The global visual score GV is the overall embedding distribution between the support image extracted by the ViT or DINO model and the current image, and EMD (Earth Mover's Distance, i.e., distribution alignment distance metric) is used to measure the similarity of particle region feature distribution between different images.

[0163] The calculation formula of the global vision score GV is as follows:

[0164] GV(p i )=1-EMD(F s (p i ),F q (p i ));

[0165] In the formula, GV(p i ) is the particle p i The global visual score of F is closer to 1. s (p i ) and F q (p i ) are the supporting image and the current image on particle p i The characteristic distribution of the region.

[0166] Furthermore, the contextual semantic consistency factor γ(p i ), a modeling structure that gives the model the ability to understand context. A multimodal large model (such as LLaVA or GPT-4o) scores the consistency between the particle semantics and the global description of the current operation and maintenance scenario. The value range is [0,1], which is used to adjust the local concept score and the global concept score to form a modified scoring item:

[0167] LC′(p i )=LC(p i )·γ(p i ), GC′(p i )=GC(p i )·γ(p i );

[0168] Where, LC′(p i ) is the corrected particle p iThe local concept score of GC′(p i ) is the corrected particle p i Global concept score of

[0169] Introducing the semantic-behavior association vector a(p i ), which is obtained by collaborative reasoning of text and video modalities, is used to predict the potential causal relationship between the particle area and historical operation and maintenance events, and calculate the behavior enhancement weight term R(p i ), used to assess the degree of match between particles and expected risk events:

[0170] R(p i )= <a(p i ),Emb event >

[0171] Where, Emb event is the semantic embedding vector of historical known fault events; <·,·> represents the vector dot product, representing the potential causal relationship between particles and fault semantics;

[0172] According to the multimodal score distribution of the local area where the particle is located, the particle information entropy is calculated:

[0173]

[0174] Where H local (p i ) is the particle p i The information entropy of the surrounding local score distribution; j is the score type index, j∈{1,2,3,4}, corresponding to the local concept score, global concept score, local visual score and global visual score respectively; S j (p i ) is the particle p i The j-th score value of H global is the average entropy of the score distribution of all particles; mean i It means taking the average of all particles numbered i;

[0175] Define the local-global entropy difference △H(p i )=H local (p i )-H global , and is used to adjust the weights of each weighted item as follows:

[0176] β′ j (p i )=β j ·(1+λ·△H(p i ));

[0177] In the formula, △H(p i ) is the particle p iThe local-global entropy difference is used to perceive the abnormal score distribution deviation; λ is the entropy difference adjustment coefficient, which reflects the amplification effect of score difference on weight; β j represents the basic weight of the j-th score; β′ j (p i ) is for particle p i The j-th weighting coefficient after dynamic adjustment;

[0178] Calculate the final weighted score S of the particle:

[0179] S(p i )=β′1(p i )·LC′(p i )+β′2(p i )·GC′(p i )+β′3(p i )·LV(p i )+β′4(p i )·GV(p i );

[0180] Where, S(p i ) is the particle p i The final weighted score of LV(p i ) is the particle p i Local visual score;

[0181] Determine whether it constitutes an abnormal hotspot area. The abnormal prediction conditions are as follows:

[0182]

[0183] Where τ1 is the abnormal score threshold, which is used to filter the areas with significant scores; For particle p i The temperature change gradient at the location indicates the thermal change trend; θ1 is the temperature change gradient threshold, reflecting the risk boundary of the rapidly changing area; M heat (p i ) is the historical fault mask; θ2 is the temperature mask heat intensity threshold, which is used to filter out high-heat areas and eliminate areas with insufficient heat changes to improve recognition accuracy; ρ is the behavior correlation judgment threshold, which is used to filter out areas most relevant to known fault types to prevent false alarms.

[0184] Introducing a contextual semantic consistency factor to enhance the adaptability of the scoring to the context of operation and maintenance events and prevent misjudgment of local anomalies; introducing a behavior latent vector a(p) obtained by log / video cross-modeling i), and the dot product with the scoring result is used as an adjustment coefficient to improve the prediction ability of the scoring heat map for possible future fault trends; the fusion weight is dynamically adjusted according to the difference in score distribution between the particle surrounding area and the overall space, and the scoring information entropy difference is introduced and mapped into a coefficient adjustment function. By adaptively strengthening the weight expression of the abnormal concentration area, the balance between the model sensitivity and robustness is improved.

[0185] (4) Structured noise prediction module

[0186] The structured noise prediction module is used to receive the current particle state set and abnormal hotspot areas, construct a sketch of the future particle state for prediction, and generate predicted particle states based on the structured noise inverse initialization and diffusion inference method to simulate the system responses that may be caused by different control instructions in the future.

[0187] like Figure 4 As shown, in one embodiment, the structured noise prediction module includes:

[0188] The particle state sketch generator is used to construct a future particle state sketch within the prediction period based on the current particle state set, abnormal hotspot areas, and semantic heat maps. The sketch includes particle positions, attribute trends, semantic label distribution, etc. It supports the generation of particle state evolution profiles at multiple time granularities (such as 5 minutes, 30 minutes, and 2 hours).

[0189] Structured noise initializer, used to perform structured noise inverse initialization on the future particle state sketch encoding to generate the initial state The calculation formula is:

[0190]

[0191] Where z0 is the latent variable representation of the future particle state sketch; α t is the noise scheduling coefficient of the diffusion time step t; ∈ is the standard Gaussian structured noise;

[0192] The anti-diffusion reasoner is used to As the starting point, the reverse diffusion process is performed to generate the predicted particle state sequence P (t+△t) , and simulate the evolution trajectory of system response under different control instructions respectively;

[0193] Anomaly enhancement mechanism is used to add anomaly weight coefficients to particles located in anomaly hotspots during particle prediction, so that the prediction process is more focused in high-temperature and high-entropy areas, thereby improving the sensitivity and accuracy of anomaly trend responses;

[0194] The closed-loop feedback update unit is used to compare the actual feedback state of the device with the predicted particle state and construct a prediction error function:

[0195]

[0196] Where E(t) is the prediction error function value at time t, which represents the total deviation between the predicted state of all particles and the actual state at time point t by the structured noise prediction module; N is the total number of particles, that is, the number of particles participating in the prediction evaluation in the current frame or the current prediction cycle;

[0197] is the predicted state vector of the i-th particle, which represents the particle attribute state predicted by the system based on the structured noise model and the diffusion back-inference algorithm. It is generally in vector form, such as: Contains spatial coordinates (x i ,y i ,z i ) and the predicted values of key physical indicators (temperature T, humidity H, power consumption P);

[0198] is the actual feedback state vector of the i-th particle, which represents the real state corresponding to the position of the i-th particle collected by the system through physical sensors in real time. It is also in vector form and contains the measurement data corresponding to the predicted state vector;

[0199] When the prediction error function value E(t) exceeds the set threshold, it indicates that the prediction model has failed and needs to trigger model update or parameter relearning. At this time, the structured noise generation parameters and the future particle state sketch are automatically updated to achieve adaptive adjustment of the prediction model, forming a closed-loop control chain of prediction-execution-feedback-correction.

[0200] (5) Instruction generation and control module

[0201] The instruction generation and control module is used to combine abnormal hot spots, predicted particle states and knowledge graph rules to perform causal relationship analysis and control instruction optimization, and output device control instruction sets;

[0202] In a preferred embodiment, the instruction generation and control module includes:

[0203] A causal path modeler is used to combine the weighted scoring results of abnormal hotspot areas, semantic heat map text, and device behavior rules defined in the knowledge graph to build a causal path diagram of events, causal relationships, and controllable variables;

[0204] The control instruction candidate generator is used to identify the controllable variables from the causal path diagram and call the candidate device control instructions corresponding to the controllable variables from the historical operation and maintenance instruction knowledge base. i , and combined with the predicted particle state to simulate the system response of each candidate device control instruction at different time granularities;

[0205] The semantic response evaluation module is used to call the multimodal language model to generate a natural language summary of the device response in the simulated prediction state for each candidate control instruction, and compare the semantic consistency with the predicted particle scene to calculate the response credibility score R (a i );

[0206] The instruction conflict detection module is used to detect the combinations of mutually exclusive operations, resource occupation conflicts or device status restrictions in the candidate device control instruction set, and eliminate the instruction groups that cannot be executed in parallel;

[0207] The execution risk assessment module is used to construct the instruction execution risk score ρ(a i ), as the penalty factor of the control instruction utility function;

[0208] Fusion optimization module, used to integrate the final weighted score S and response credibility score R (a i ) and execution risk score ρ(a i ), calculate the control instruction utility function U(a i ):

[0209] U(a i )=λ1·S+λ2·R(a i )-λ3·ρ(a i );

[0210] Where λ1, λ2, and λ3 are the weighting coefficients set by the system;

[0211] Instruction output selector, used to control the utility function U(a i ) priority, sort the candidate device control instruction sets, and output the optimal device control instruction set that can be executed by the current system to the device control end.

[0212] (6) 3D visualization interaction module

[0213] The three-dimensional visualization interaction module is used to receive the current particle state set, semantic heat map and predicted particle state, and realize three-dimensional visualization interactive display on the client through GPU acceleration and WebGL rendering technology, which at least includes real-time particle distribution rendering, abnormal hot spot area highlighting, trend evolution trajectory animation and device control command interactive operation.

[0214] like Figure 5 FIG. 1 is another embodiment of the present invention, which provides a multi-modal intelligent linkage 3D visualization data center operation and maintenance method, including the following contents:

[0215] Collect physical sensor data, video surveillance data, and operation and maintenance log data from the data center. Through time synchronization and spatial alignment, the different modal data are uniformly mapped to the same three-dimensional spatial coordinate system to construct a multimodal fusion feature tensor.

[0216] The particle space structure of the data center is constructed based on the Voronoi diagram and Delaunay triangulation method. A mapping relationship is established between the particles and the data items in the multimodal fusion feature tensor. The visualization resolution level of the particles is dynamically adjusted according to the distance between the particles and the current viewpoint, and the current particle state set is output.

[0217] Calculate the local concept score, global concept score, local visual score, and global visual score of each particle in the current particle state set, generate a semantic heat map based on the weighted scoring results, and identify abnormal hot spots that meet the abnormality prediction conditions;

[0218] Construct a sketch of future particle states for prediction, and generate predicted particle states based on structured noise inverse initialization and diffusion inference methods to simulate the system responses that may be caused by different control instructions in the future;

[0219] Combine abnormal hot spots, predicted particle states, and knowledge graph rules to perform causal relationship analysis and control instruction optimization, and output device control instruction sets;

[0220] 3D visual interactive display is achieved on the client through GPU acceleration and WebGL rendering technology.

[0221] In summary, the present invention has the integrated capabilities of multimodal intelligent fusion, three-dimensional structural modeling, predictive simulation and deduction, and linkage control decision-making, which significantly improves the intelligence, real-time and interactivity of the data center operation and maintenance system, and has broad application prospects and promotion value.

[0222] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any other combination. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product, which includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0223] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the aforementioned integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0224] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A multi-modal intelligent linkage 3D visualization data center operation and maintenance system, characterized by: The system comprises: The multimodal data fusion module is used to collect physical sensor data, video surveillance data, and operation and maintenance log data from the data center. Through time synchronization and spatial alignment, it maps the different modal data to the same three-dimensional spatial coordinate system and constructs a multimodal fusion feature tensor. A three-dimensional particle modeling module is used to construct the particle space structure of the data center based on the Voronoi diagram and Delaunay triangulation method, establish a mapping relationship between particles and data items in the multimodal fusion feature tensor, dynamically adjust the visualization resolution level of particles according to the distance between the particles and the current viewpoint, and output the current particle state set; The multimodal scoring analysis module is used to calculate the local concept score, global concept score, local visual score, and global visual score of each particle in the current particle state set, generate a semantic heat map based on the weighted scoring results, and identify abnormal hot spots that meet the abnormality prediction conditions; The structured noise prediction module is used to construct a sketch of the future particle state for prediction. It generates predicted particle states based on the structured noise inverse initialization and diffusion inference method to simulate the system response that different control instructions may cause in the future. An instruction generation and control module is used to combine the abnormal hotspot area, predicted particle state and knowledge graph rules to perform causal relationship analysis and control instruction optimization, and output a device control instruction set; The three-dimensional visualization interaction module is used to realize three-dimensional visualization interactive display on the client through GPU acceleration and WebGL rendering technology, including at least real-time particle distribution rendering, abnormal hot spot area highlighting, trend evolution trajectory animation and device control instruction interactive operation.

2. The multi-modal intelligent linkage 3D visualization data center operation and maintenance system according to claim 1 is characterized in that: The multimodal data fusion module includes: A sensor data unit is used to obtain physical sensor data from various devices in the data center, including at least temperature, humidity, and power consumption, synchronize the physical sensor data according to timestamps, and map the data to a three-dimensional spatial coordinate system based on the spatial installation location of each sensor to form physical features; The video surveillance data unit is used to obtain the video surveillance data of the data center and perform image frame-level decoding, extract the visual features of the target area in the image, and map the visual features into three-dimensional space by combining the internal and external parameter information of the corresponding camera; The operation and maintenance log data unit is used to perform semantic analysis on the operation and maintenance log text and structure it into event triples, which are then mapped to a three-dimensional spatial coordinate system according to preset spatial semantic mapping rules. Data fusion unit, used to fuse physical features, visual features and semantic features in a unified three-dimensional spatial coordinate system and time dimension to generate a multimodal fusion feature tensor Among them, x, y, z represent three-dimensional space coordinates, and t represents time.

3. The multi-modal intelligent linkage 3D visualization data center operation and maintenance system according to claim 1 is characterized in that: In the three-dimensional particle modeling module, the method for constructing the particle space structure includes: Taking the physical sensor nodes deployed in the data center space as spatial seed points, a weighted Voronoi diagram is used to divide the three-dimensional space into regions. The weight value of each physical sensor node is related to the sensor's sensing coverage, the frequency of historical abnormal records, or the importance of the region. Extracting a centroid set within the weighted Voronoi region and performing Delaunay triangulation on the centroid set to construct a three-dimensional particle network structure; After the construction is completed, a one-to-one mapping relationship is established according to the tensor value in the multimodal fusion feature tensor in the area where each particle is located; the three-dimensional particle modeling module monitors the local change gradient of the multimodal fusion feature tensor during the operation of the system, and dynamically adjusts the particle distribution structure based on the preset reconstruction threshold, performs particle refinement in high-variability areas, and performs particle aggregation reconstruction in stable areas.

4. The multi-modal intelligent linkage 3D visualization data center operation and maintenance system according to claim 3 is characterized in that: The three-dimensional particle modeling module determines the visualization resolution level of each particle based on the spatial distance between each particle and the current viewing center point and the temperature gradient of the area where the particle is located. The resolution level scheduling mechanism specifically includes: For each particle p i , calculate the Euclidean distance d between the particle and the center of the viewing angle i , and calculate the temperature gradient amplitude of the local area where the particle is located Construct particle rendering weight function: Where W i is the particle rendering weight; σ is the Sigmoid function; β and γ are the system preset adjustment coefficients; According to the particle rendering weight W i The value in the range of [0,1] corresponds to different visualization resolution levels and determines whether to retain, aggregate or refine the rendering granularity of the particle in the current frame; During system operation, the visual resolution level of each particle can be dynamically updated to support visually continuous transition and smooth presentation, ensuring that particles in abnormal hot spots are rendered delicately while areas far from the viewing angle and changing slowly are dynamically simplified.

5. The multi-modal intelligent linkage 3D visualization data center operation and maintenance system according to claim 4 is characterized in that: The current particle state set includes at least: The three-dimensional coordinate position of the particle (x i ,y i ,z i ), used for spatial positioning; Real-time monitored temperature value T i , humidity value H i And power consumption value P i , used to characterize the physical environment state of the particle corresponding to the spatial region; fusion feature vector v i , the fusion feature vector is composed of the visual features and semantic features extracted from the multimodal fusion feature tensor of the region where the particle is located; Visualization level label L i , used to indicate the visualization resolution level of particles at the current viewing angle; State entropy It is used to measure the normalized uncertainty of the particle's temperature, humidity, and power consumption changes in the last several consecutive moments. The calculation formula is: Where p k is the distribution probability of any physical property k of the particle within a time window; Life cycle counter C i , used to record the number of state updates of the particle since its generation, and trigger the particle resampling process when the upper limit of the example update life cycle set by the system is exceeded; Abnormal cluster label A i , used to identify whether the particle is clustered and belongs to any abnormal hotspot area or potential fault domain in abnormal pattern recognition, and is used to support the priority calculation and abnormal tracking of equipment control instructions.

6. The multi-modal intelligent linkage 3D visualization data center operation and maintenance system according to claim 1 is characterized in that: In the multimodal scoring analysis module, the local concept score LC is based on a weighted average method that integrates local semantic heat and mask area ratio to measure the semantic concept significance of the local area where the particle is located; The global concept score GC is calculated through the image-text embedding similarity of the Alpha-CLIP model to evaluate the global semantic consistency between the particle image region and the target semantic text; The local visual score LV is obtained by calculating the feature similarity between the support image and the current image in the local area of the particle. The calculation process includes the local average aggregation of the spatial matching matrix. The global visual score GV is the overall embedding distribution between the support image extracted by the ViT or DINO model and the current image, and the EMD is used to measure the similarity of the particle region feature distribution between different images.

7. The multi-modal intelligent linkage 3D visualization data center operation and maintenance system according to claim 6 is characterized in that: In the multimodal scoring analysis module, the contextual semantic consistency factor γ(p i ), the multimodal large model scores the consistency between the particle semantics and the global description of the current operation and maintenance scenario. The value range is [0,1], which is used to adjust the local concept score and the global concept score to form the modified scoring item: LC′(p i )=LC(p i )·γ(p i ),GC′(p i )=GC(p i )·γ(p i ); Where p i represents the i-th particle, LC(p i ) is the particle p i The local concept score, GC(p i ) is the particle p i The global concept score of LC′(p i ) is the corrected particle p i The local concept score, GC′(p i ) is the corrected particle p i The global concept score of the semantic-behavior association vector a(p i ), which is obtained by collaborative reasoning of text and video modalities, is used to predict the potential causal relationship between the particle area and historical operation and maintenance events, and calculate the behavior enhancement weight term R(p i ): R(p i )=<a(p i ),Emb event >; Where, Emb event is the semantic embedding vector of historical known fault events; <·,·> represents the vector dot product; According to the multimodal score distribution of the local area where the particle is located, the particle information entropy is calculated: H global =mean i (H local (p i )); Where H local (p i ) is the particle p i Information entropy of the surrounding local score distribution; j is the score type index, j∈{1,2,3,4}; S j (p i ) is the particle p i The j-th score value of H global is the average entropy of the score distribution of all particles; mean i It means taking the average of all particles numbered i; Define the local-global entropy difference △H(p i )=H local (p i )-H global , and is used to adjust the weights of each weighted item as follows: b′ j (p i )=β j ·(1+λ·△H(p i )); In the formula, △H(p i ) is the particle p i The local-global entropy difference; λ is the entropy difference adjustment coefficient; β j represents the basic weight of the j-th score; β′ j (p i ) is for particle p i The j-th weighting coefficient after dynamic adjustment; Calculate the final weighted score S of the particle: S(p i )(β′1(p i )·LC′(p i )+β′2(p i )·GC′(p i )+β′3(p i )·LV(p i )+β′4(p i )·GV(p i )4 Where, S(p i ) is the particle p i The final weighted score of LV(p i ) is the particle p i Local visual score; GV(p i ) is the particle p i Global visual score of Determine whether it constitutes an abnormal hotspot area. The abnormal prediction conditions are as follows: Where τ1 is the abnormality score threshold; For particle p i The temperature change gradient at the position; θ1 is the temperature change gradient threshold; M heat (p i ) is the historical fault mask; θ2 is the temperature mask heat intensity threshold; ρ is the behavior correlation judgment threshold.

8. The multi-modal intelligent linkage 3D visualization data center operation and maintenance system according to claim 1 is characterized in that: The structured noise prediction module includes: The particle state sketch generator is used to construct a future particle state sketch within the prediction period based on the current particle state set, abnormal hotspot areas, and semantic heat maps. The sketch includes particle positions, attribute trends, and semantic label distributions, and supports the generation of particle state evolution profiles at multiple time granularities. Structured noise initializer, used to perform structured noise inverse initialization on the future particle state sketch encoding to generate the initial state The calculation formula is: Where z0 is the latent variable representation of the future particle state sketch; α t is the noise scheduling coefficient of the diffusion time step t; ∈ is the standard Gaussian structured noise; The anti-diffusion reasoner is used to As the starting point, the reverse diffusion process is performed to generate the predicted particle state sequence P (t +△t) , and simulate the evolution trajectory of system response under different control instructions respectively; Anomaly enhancement mechanism is used to add anomaly weight coefficients to particles located in abnormal hot spots during particle prediction, so that the prediction process can enhance the focus in high temperature and high entropy areas; The closed-loop feedback update unit is used to compare the actual feedback state of the device with the predicted particle state and construct a prediction error function: Where E(t) is the prediction error function value at time t; N is the total number of particles; is the predicted state vector of the i-th particle; is the actual feedback state vector of the i-th particle; When the prediction error function value E(t) exceeds the set threshold, the structured noise generation parameters and the future particle state sketch are automatically updated.

9. The multi-modal intelligent linkage 3D visualization data center operation and maintenance system according to claim 7, characterized in that: The instruction generation and control module includes: A causal path modeler is used to combine the weighted scoring results of the abnormal hotspot areas, the semantic heat map text and the device behavior rules defined in the knowledge graph to construct a causal path diagram of events, causal relationships and controllable variables; A control instruction candidate generator is used to identify controllable variables from the causal path diagram and call candidate device control instructions corresponding to the controllable variables from the historical operation and maintenance instruction knowledge base. i , and combined with the predicted particle state to simulate the system response of each candidate device control instruction at different time granularities; The semantic response evaluation module is used to call the multimodal language model to generate a natural language summary of the device response in the simulated prediction state for each candidate control instruction, and compare the semantic consistency with the predicted particle scene to calculate the response credibility score R (a i ); an instruction conflict detection module for detecting combinations of mutually exclusive operations, resource occupancy conflicts, or device state restrictions in a candidate device control instruction set, and eliminating instruction groups that cannot be executed in parallel; The execution risk assessment module is used to construct the instruction execution risk score ρ(a i ), as the penalty factor of the control instruction utility function; Fusion optimization module, used to integrate the final weighted score S and response credibility score R (a i ) and execution risk score ρ(a i ), calculate the control instruction utility function U(a i ): U(a i )=λ1·S+λ2·R(a i )-λ3·ρ(a i ); Where λ1, λ2, and λ3 are the weighting coefficients set by the system; Instruction output selector, used to control the utility function U(a i ) priority, sort the candidate device control instruction sets, and output the optimal device control instruction set that can be executed by the current system to the device control end.

10. A multi-modal intelligent linkage 3D visualization data center operation and maintenance method, based on the multi-modal intelligent linkage 3D visualization data center operation and maintenance system according to any one of claims 1 to 9, characterized in that: The method comprises: Collect physical sensor data, video surveillance data, and operation and maintenance log data from the data center. Through time synchronization and spatial alignment, the different modal data are uniformly mapped to the same three-dimensional spatial coordinate system to construct a multimodal fusion feature tensor. The particle space structure of the data center is constructed based on the Voronoi diagram and Delaunay triangulation method, a mapping relationship is established between the particles and the data items in the multimodal fusion feature tensor, and the visualization resolution level of the particles is dynamically adjusted according to the distance between the particles and the current viewing angle, and the current particle state set is output; Calculate the local concept score, global concept score, local visual score, and global visual score of each particle in the current particle state set, generate a semantic heat map based on the weighted scoring results, and identify abnormal hot spots that meet the abnormality prediction conditions; Construct a sketch of future particle states for prediction, and generate predicted particle states based on structured noise inverse initialization and diffusion inference methods to simulate the system responses that may be caused by different control instructions in the future; Combining the abnormal hotspot area, predicted particle state and knowledge graph rules, causal relationship analysis and control instruction optimization are performed to output a device control instruction set; 3D visual interactive display is achieved on the client through GPU acceleration and WebGL rendering technology.

Citation Information

Patent Citations

  • Cloud GIS two-dimensional and three-dimensional integrated visualization system and two-dimensional and three-dimensional integrated visualization method

    CN116662435A

  • Real estate surveying and mapping data processing method and system based on intelligent data

    CN119646483A

  • Real-time system for multi-modal 3D geospatial mapping, object recognition, scene annotation and analytics

    US20150269438A1

Cited By

  • AI talent skill portrait generation method and system based on multi-modal image recognition

    CN121121272A

  • Underground pipeline three-dimensional visualization and real-time sensing system based on AR-GIS fusion

    CN121236255A

  • AI human body motion capture sensor and automatic control method thereof

    CN121256249A

  • Information visualization method and system based on industrial Internet of Things

    CN121478867A

  • Information visualization method and system based on industrial internet of things

    CN121478867B