Industrial scene automatic monitoring method, system and product based on multi-modal data
Through multimodal data collection and processing technology, combined with visual recognition, knowledge graphs and language models, automated and intelligent monitoring of industrial scenarios is achieved, solving the problems of low monitoring efficiency and misjudgment caused by scattered equipment, and improving the accuracy and timeliness of monitoring.
Patent Information
- Application Number
- CN202510868893.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-17
AI Technical Summary
In modern industrial scenarios, equipment is dispersed and data is dense, resulting in low monitoring efficiency, easy misjudgments and information omissions, making it difficult to respond to emergencies in a timely manner.
A multimodal data acquisition method is adopted to collect data using visible light cameras, infrared thermal imaging cameras and acoustic imaging cameras. Anomaly detection and parameter association are performed by combining visual recognition models, knowledge graphs and graph neural networks, and natural language monitoring results are output through a multimodal language model.
It realizes automated and intelligent monitoring of industrial scenarios, improves monitoring efficiency and accuracy, reduces manual intervention, detects equipment anomalies in a timely manner, and improves the interpretability and reliability of monitoring.
Smart Images

Figure CN120805847A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of industrial scene monitoring, and in particular to an industrial scene automatic monitoring method and system based on multi-modal data. BACKGROUND
[0002] Modern industrial production processes are increasingly complex, involving many links, equipment and processes. Currently, industrial scene monitoring mainly uses equipment data monitoring and environmental video monitoring, and then relies on workers to monitor and judge the running status of equipment in the industrial scene according to the equipment data monitoring and environmental video monitoring.
[0003] However, due to the factors such as the large number of equipment and their dispersion in the industrial scene, the dense display of equipment data and the excessive number of monitoring screens, when checking the monitoring information, not only will it cause a large viewing burden to the workers, but also will be inefficient, and it is easy to make a mistake due to the omission of monitoring information, so that when a sudden event occurs, the evaluation and decision-making will take a long time, it is difficult to form a response processing scheme in time, leading to the occurrence and expansion of the accident. SUMMARY
[0004] In view of the above problems, the present application provides an industrial scene automatic monitoring method and system based on multi-modal data, which is used to solve the problems of low efficiency and easy omission caused by relying on workers to check monitoring information in the current industrial scene monitoring.
[0005] In a first aspect, the present application provides an industrial scene automatic monitoring method based on multi-modal data, which comprises: Obtaining environmental multi-modal data in an industrial scene, performing anomaly detection on each device in the industrial scene based on the environmental multi-modal data, and outputting an environmental perception result in natural language form; Obtaining running data of each device based on the environmental perception result, obtaining real-time correlation coefficients between different running parameters of each device when running, and outputting a parameter correlation result of the device when running in natural language form; Aligning the time information of the environmental multi-modal data and the running data, describing the state of the industrial scene based on the environmental perception result and the parameter correlation result, and outputting an industrial scene monitoring result including the environmental state, the device state and the overall trend in natural language form.
[0006] Further, the environmental multi-modal data is multi-source images of visible light, temperature field and acoustic field in the industrial scene collected by a visible light camera, an infrared thermal imaging camera and an acoustic imaging camera respectively.
[0007] Furthermore, the multimodal environmental data is input into a visual recognition model to perform anomaly detection on various devices in the industrial scene; wherein, the visual recognition model is used to detect anomalies in the environment or equipment, associate the anomalies with the source, and output the detection and recognition results in natural language form.
[0008] Furthermore, the real-time correlation coefficient is obtained by using a knowledge graph to describe the correlation between various devices, establishing a graph neural network based on the knowledge graph, obtaining a real-time correlation coefficient matrix between different operating parameters of each device during operation, and then obtaining the real-time correlation coefficient.
[0009] Furthermore, the real-time relevance matrix The formula is: ; in, Indicates the device A real-time relevance matrix for physical centers; represents the normalization function; Represents the query matrix, which is the number of devices except Real-time operation data of each device other than Represents the key matrix, reflecting the device Real-time operation data; represents transpose; Indicates the dimension of the key matrix, corresponding to the number of types of running parameters.
[0010] Furthermore, the method for describing the industrial scene status based on the environmental perception results and the parameter association results is: establishing a multimodal language model with an attention mechanism; wherein the input interface of the multimodal language model is used to receive the natural language text of the environmental perception results, and the attention coefficient input interface of the multimodal language model is used to receive the real-time correlation coefficient in the parameter association results; with the current device as the physical center, the natural language text of the received environmental perception results and the real-time correlation coefficient in the parameter association results are input, and the industrial scene monitoring results are output.
[0011] Furthermore, the multimodal language model includes an encoder and a decoder; the encoder is used to encode the natural language text of the input environmental perception results and parameter association results, and the decoder is used to generate natural language text of the industrial scene monitoring results including environmental status, equipment status and overall trends.
[0012] Furthermore, the attention coefficient is: ; in, Indicates the Moment a weight coefficient of an input feature; a feature value of the i th input feature.
[0013] In a second aspect, the present application provides an industrial scene automatic monitoring system based on multi-modal data, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the above methods.
[0014] In a third aspect, the present application provides a computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of any of the above methods.
[0015] Overall, the present application provides an industrial scene automatic monitoring method, system and product based on multi-modal data, which can achieve the following beneficial effects compared with the prior art by the technical solutions conceived by the present application: (1) The present application addresses the problems of low efficiency, heavy burden and easy misjudgment caused by manual review of environmental monitoring data in the industrial scene with massive and scattered equipment, utilizes different types of information to effectively express the key features of industrial environment and industrial equipment in natural language, simultaneously aligns and associates multi-modal data, describes the industrial scene state according to the natural language form of the environmental perception result and the natural language form of the parameter association result, and outputs the industrial scene monitoring result including the environmental state, the equipment state and the overall trend in natural language form; forms a multi-modal information fusion link for the automatic and intelligent detection, extraction and evaluation of the complex state of the industrial scene, and solves the problems of low efficiency and easy omission caused by relying on workers to check monitoring information in the current industrial scene monitoring.
[0016] (2) The present application can comprehensively obtain visible light, temperature field and acoustic field data in the environmental state by fusing multi-source visual data of various sensors such as visible light cameras, infrared thermal imaging cameras and acoustic imaging cameras, and utilizes a visual large model to detect and identify the environmental multi-modal data, thereby improving the comprehensiveness and accuracy of environmental perception.
[0017] (3) The present application can analyze the correlation between equipment operating parameters in real time through a knowledge graph and a graph neural network, thereby obtaining real-time correlation coefficients between operating parameters, simultaneously aligns the time information of environmental multi-modal data and operating data to ensure the time synchronization of all information, which helps to timely discover abnormal situations in equipment operation, makes information processing more efficient and accurate, improves monitoring efficiency and reliability, and provides important technical support for industrial automation.
[0018] (4) The industrial scene automatic monitoring method provided by the application can automatically process environment and device state information by using an attention mechanism language model, convert multi-modal data and parameter association results into natural language descriptions, and output industrial scene monitoring results in the form of natural language texts, so that the intelligent understanding and description capability improves the interpretability and practicality of the monitoring results, reduces the need for manual intervention, and improves the automation and intelligent level of monitoring. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0020] Figure 1 is a method step schematic diagram of an industrial scene automatic monitoring method, system and product provided by the application based on multi-modal data; Figure 2 is an environment perception result schematic diagram of an industrial scene automatic monitoring method, system and product provided by the application based on multi-modal data; Figure 3 is a parameter association result schematic diagram of an industrial scene automatic monitoring method, system and product provided by the application based on multi-modal data; Figure 4 is a multi-modal language model schematic diagram of an industrial scene automatic monitoring method, system and product provided by the application based on multi-modal data. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the application more clear, the technical solutions in the application will be described clearly and completely in the following with reference to the drawings and embodiments in the application. Obviously, the described embodiments are some embodiments of the application, but not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.
[0022] It should be noted that in the description of the embodiments of the application, the terms “include”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the method, step or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such method, step or device. Without more limitation, the element defined by the sentence “including a…” does not exclude the presence of another same element in the method, step or device including the element.
[0023] The application provides an industrial scene automatic monitoring method and system based on multi-modal data and a product, which solves the problems of low efficiency and easy omission caused by relying on workers to check monitoring information in the current industrial scene monitoring.
[0024] Specifically, as shown in the method comprises: Figure 1 S101: Obtain environment multi-modal data in an industrial scene, perform anomaly detection on each device in the industrial scene based on the environment multi-modal data, and output an environment perception result in natural language form.
[0025] As an embodiment, the environment multi-modal data is multi-source images of visible light, temperature field, and acoustic field in the industrial scene collected by a visible light camera, an infrared thermal imaging camera, and an acoustic imaging camera.
[0026] For example, as shown in the boiler steam industrial scene, the boiler steam industrial scene includes a boiler, a regulating valve, a steam turbine, a condenser, a circulating water pump, a condensate pump, a feed water pump, and other devices. Visible light images of the boiler steam system are obtained by a visible light camera, temperature field data of each device are obtained by an infrared thermal imaging camera, and acoustic field data are obtained by an acoustic imaging camera. Figure 2 As an embodiment, the environment multi-modal data is input into a visual recognition model to perform anomaly detection on each device in the industrial scene and output an environment perception result.
[0027] The input information of the visual recognition model is the environment multi-modal data collected by the visible light camera, the infrared thermal imaging camera, and the acoustic imaging camera; and the output information of the visual recognition model is the environment perception result. The environment perception result is the detection and recognition result of the visual recognition model, which can be an environment without abnormal state or a device with abnormal state, but no matter how, it is output in natural language form.
[0028] It should be noted that the visual recognition model is used to detect anomalies in the environment or devices, associate the anomalies with the source, and output the detection and recognition result in natural language form.
[0029] In other words, the environment state, target device type, temperature field, and acoustic field are detected and recognized, the recognition result of the target device type is aligned with the detection result of the temperature and acoustic field, and the source of the temperature anomaly and acoustic anomaly is determined.
[0030] The visual recognition model can be YOLO, CLIP, U-Net, VGG, etc., which is a relatively mature technical means in the prior art and will not be described here.
[0031]
[0032] Taking the boiler steam industrial scene as an example, the visual recognition model is used to detect and identify the environmental state such as smoke concentration of the boiler steam system, the target equipment type, the temperature field distribution, and the acoustic field distribution; the detection results of the target equipment type, temperature, and acoustics are aligned to determine the source of the temperature anomaly (such as local overheating) or acoustic anomaly (such as leakage) of the equipment, and finally the environmental perception result is output in natural language form.
[0033] For example: the infrared thermal imaging camera detects temperature anomaly in a certain area, and the visible light detects that the equipment type corresponding to the area is the top of the boiler; the acoustic imaging camera detects that there is a leakage phenomenon in a certain area, and the visible light detects that the equipment type corresponding to the area is a steam pipeline; the areas with temperature anomaly and acoustic anomaly are matched with the recognized equipment type to obtain the abnormal result, and then output “temperature detection finds that the temperature of the top of the boiler is abnormally high, and acoustic detection finds that there is a leakage phenomenon in the steam pipeline”.
[0034] In addition, in order to improve the overall recognition effect, the visual recognition model can further include a fusion coefficient for fusing multiple source images.
[0035] After the fusion of multiple source images, the fusion coefficient can guide the model to complement and filter the features of different modal images, form more representative comprehensive feature vectors, not only can avoid the waste of time and computing resources caused by processing a large amount of redundant single modal features, but also can reduce the feature dimension to a certain extent, and improve the inference speed and recognition efficiency of the model.
[0036] Further, the fusion coefficient is: ; Wherein, represents the fusion coefficient between the m-th modal data and the n-th modal data; represents a nonlinear activation function such as sigmoid; represents the weight coefficient of the m-th input feature; represents the m-th input feature of the m-th modal data; represents the m-th input feature of the m-th modal data.
[0037] S102: Obtain running data corresponding to each device based on the environmental perception result, obtain real-time correlation coefficients between different running parameters of each device when running, and output the parameter correlation result of the device when running in natural language form.
[0038] The running data is the data value of pressure, flow, rotating speed, torque and the like collected by the sensors installed on each device in real time for the running state of the device.
[0039] As an embodiment, as shown in Figure 3 , the real-time correlation coefficient is obtained by using the knowledge graph to describe the correlation between devices, establishing a graph neural network based on the knowledge graph, obtaining a real-time correlation coefficient matrix between different running parameters of each device in running, and then obtaining the real-time correlation coefficient.
[0040] The correlation between devices is described by using the knowledge graph, that is, taking the device type as the entity: boiler V1, regulating valve V2, condenser V3, circulating water pump V4, condensate pump V5, regulating valve V6, and feed water pump V7, and taking the production process as the correlation relationship: e1, e2, e3. The entity is taken as the node in the knowledge graph, and the correlation relationship is taken as the edge in the knowledge graph to construct the graph structure of the knowledge graph.
[0041] The knowledge graph describes the dynamic correlation between devices in the physical or process flow, and the graph neural network learns the features and dynamic relationships of nodes (devices) on this knowledge graph structure.
[0042] Common graph neural network models include graph convolution network (GCN) and graph attention network (GAT). Based on these graph neural network models, the knowledge graph data (including node features, edge information and corresponding labels, such as the actual running state of the device) is input into the graph neural network model. In the forward propagation process, the model aggregates and propagates information according to the graph structure and the features of nodes and edges, and calculates the output result. The error between the predicted result and the true label is calculated by using the loss function, and the parameters of the model are updated by using the back propagation algorithm and the optimizer.
[0043] This process is repeated continuously, and after multiple training iterations, the model can learn the patterns contained in the device nodes and their correlation relationships in the knowledge graph, so as to be used for predicting the real-time correlation between different running parameters of the device in running.
[0044] The input information of the graph neural network model is the real-time running data of each device based on the knowledge graph, and the output information of the graph neural network model is the real-time correlation coefficient between different running parameters of each device in running, that is, the attention vector.
[0045] As an embodiment, the formula of the real-time correlation matrix is as follows: ; wherein, represents the device A real-time relevance matrix for physical centers; represents the normalization function; Represents the query matrix, which is the number of devices except Real-time operation data of each device other than Represents the key matrix, reflecting the device Real-time operation data; represents transpose; Indicates the dimension of the key matrix, corresponding to the number of types of running parameters.
[0046] Taking the boiler and steam industry scenario as an example, a knowledge graph is used to describe the relationships between various devices in the boiler and steam industry, such as boilers, steam pipes, and valves. A graph neural network is established based on the knowledge graph to calculate the real-time correlation coefficients between device operating parameters. For example, the correlation coefficient between pressure and flow is 0.85, and the correlation coefficient between pressure and temperature is 0.92.
[0047] S103: Align the time information of the environmental multimodal data and the operating data, describe the industrial scene status based on the environmental perception results and parameter association results, and output the industrial scene monitoring results including the environmental status, equipment status and overall trend in the form of natural language.
[0048] Aligning time information means taking each device as the physical center and aligning the time scales of environmental status information and device status information to ensure that all information is synchronized. This helps to promptly detect abnormal conditions in device operation, making information processing more efficient and accurate, and improving monitoring efficiency and reliability.
[0049] For example, taking the boiler as the physical center, align the time scales of environmental status information such as temperature field and acoustic field with equipment status information such as pressure and flow.
[0050] As an example, Figure 4 As shown in Figure 2, the method for describing the industrial scene status based on the environmental perception results and parameter association results is: A multimodal language model with an attention mechanism is established; the input interface of the multimodal language model is used to receive the natural language text of the environmental perception results, and the attention coefficient input interface of the multimodal language model is used to receive the real-time correlation coefficient in the parameter association results; with the current device as the physical center, the natural language text of the environmental perception results and the real-time correlation coefficient in the parameter association results are input, and the industrial scene monitoring results are output.
[0051] The attention coefficient is: ; in, Indicates the Moment a weight coefficient of an input feature; representing the feature value of the input feature.
[0052] The multi-modal language model is used to process the state information of the environment and the device, and the input information is: the natural language text of the environment perception result and the correlation coefficient of the graph neural network; and the output information is: the industrial scene monitoring result. The industrial scene monitoring result includes the environment state, the device state and the overall trend.
[0053] Further, the multi-modal language model includes an encoder and a decoder; the encoder is used to encode the input natural language text of the environment perception result and the correlation coefficient, and the decoder is used to generate the natural language text of the industrial scene monitoring result including the environment state, the device state and the overall trend.
[0054] The multi-modal language model can be a Transformer model, a BERT model, a CLIP model, etc. The models are relatively mature, and the training process is not described again.
[0055] For example, the natural language text of the environment perception result is input into the multi-modal language model, and the natural language text is: "The temperature detection finds that the temperature at the top of the boiler abnormally rises, and the acoustic detection finds that there is a leakage phenomenon in the steam pipeline"; and the correlation coefficient of the graph neural network is also input into the multi-modal language model, and the correlation coefficient is: "The correlation coefficient between pressure and temperature is 0.92". Then, taking the device focused on by the current attention as the center, the industrial scene monitoring result is described as: "The temperature at the top of the boiler abnormally rises, which may be caused by the pressure change caused by the steam pipeline leakage, and it is suggested to check the sealing performance of the steam pipeline." In a second aspect, the present application provides an industrial scene automatic monitoring system based on multi-modal data, which includes a memory, a processor and a computer program stored in the memory, and the processor executes the computer program to realize the steps of the method in any of the above aspects.
[0056] In a third aspect, the present application provides a computer program product, which includes computer programs / instructions, and the computer programs / instructions are executed by the processor to realize the steps of the method in any of the above aspects.
[0057] The technical features of the system and the product are consistent, and will not be described again.
[0058] In summary, the present application aims at the problems of massive and scattered equipment in industrial scenes, low efficiency, heavy burden, and easy misjudgment of manual review of environmental monitoring data, uses different information categories to effectively express the key features of industrial environment and industrial equipment in natural language, aligns and associates multi-modal data, describes the industrial scene state according to the natural language form of the environmental perception result and the natural language form of the parameter association result, and outputs the industrial scene monitoring result including the environment state, the equipment state and the overall trend in the form of natural language; forms a multi-modal information fusion link of automatic and intelligent detection, extraction and evaluation of the complex state of the industrial scene, and solves the problems of low efficiency and easy omission caused by the dependence of workers on monitoring information in the current industrial scene monitoring.
[0059] It should be noted that, for the foregoing various embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.
[0060] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0061] In several embodiments provided in the present application, it should be understood that the disclosed method or system can be implemented by other ways. For example, the above-described embodiments are only illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division way, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0062] The above-described are only exemplary embodiments of the present disclosure, and cannot limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. Those skilled in the art will easily think of the embodiments of the present disclosure after considering the specification and practicing the disclosure herein. The present application aims to cover any variations, uses or adaptive changes of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and examples are only considered as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
[0063] Any technical features in the above embodiments can be combined, and for the sake of brevity, not all possible combinations are described above, and it is understood that any combination of the technical features is within the scope of the present disclosure.
[0064] Those skilled in the art easily understand that the above description is only the preferred embodiments of the present application, and is not used to limit the present application, and any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. An automated monitoring method for industrial scenes based on multimodal data, characterized in that: The method comprises: Acquire multimodal environmental data in the industrial scene, perform anomaly detection on each device in the industrial scene based on the multimodal environmental data, and output environmental perception results in natural language form; Based on the environmental perception results, operating data corresponding to each device is obtained, a real-time correlation coefficient between different operating parameters of each device during operation is obtained, and parameter correlation results of the device during operation are output in a natural language form; Align the time information of the environmental multimodal data and the operating data, describe the industrial scene status based on the environmental perception results and the parameter association results, and output the industrial scene monitoring results including the environmental status, equipment status and overall trend in the form of natural language.
2. The method for automated monitoring of industrial scenes based on multimodal data according to claim 1, characterized in that: The environmental multimodal data is: multi-source images of visible light, temperature field, and acoustic field in industrial scenes are collected by using a visible light camera, an infrared thermal imaging camera, and an acoustic imaging camera.
3. The method for automated monitoring of industrial scenes based on multimodal data according to claim 2, characterized in that: The multimodal environmental data is input into a visual recognition model to perform anomaly detection on various devices in the industrial scene; wherein the visual recognition model is used to detect anomalies in the environment or equipment, associate the anomalies with the source, and output the detection and recognition results in natural language form.
4. The method for automated monitoring of industrial scenes based on multimodal data according to claim 1, characterized in that: The real-time correlation coefficient is obtained by using a knowledge graph to describe the correlation between various devices, establishing a graph neural network based on the knowledge graph, obtaining a real-time correlation coefficient matrix between different operating parameters of various devices during operation, and then obtaining the real-time correlation coefficient.
5. The method for automated monitoring of industrial scenes based on multimodal data according to claim 4, characterized in that: The real-time relevance matrix The formula is: ; in, Indicates the device A real-time relevance matrix for physical centers; represents the normalization function; Represents the query matrix, which is the number of devices except Real-time operation data of each device other than Represents the key matrix, reflecting the device Real-time operation data; represents transpose; Indicates the dimension of the key matrix, corresponding to the number of types of running parameters.
6. The method for automated monitoring of industrial scenes based on multimodal data according to claim 1, characterized in that: The method for describing the industrial scene status based on the environmental perception results and the parameter association results is as follows: establishing a multimodal language model with an attention mechanism; wherein the input interface of the multimodal language model is used to receive the natural language text of the environmental perception results, and the attention coefficient input interface of the multimodal language model is used to receive the real-time correlation coefficient in the parameter association results; with the current device as the physical center, the natural language text of the received environmental perception results and the real-time correlation coefficient in the parameter association results are input, and the industrial scene monitoring results are output.
7. The method for automated monitoring of industrial scenes based on multimodal data according to claim 6, characterized in that: The multimodal language model includes an encoder and a decoder; the encoder is used to encode the natural language text of the input environmental perception results and parameter association results, and the decoder is used to generate natural language text of the industrial scene monitoring results including environmental status, equipment status and overall trend.
8. The method for automated monitoring of industrial scenes based on multimodal data according to claim 7, characterized in that: The attention coefficient is: ; in, Indicates the Moment The weight coefficient of the input features; Indicates the The eigenvalues of the input features.
9. A mathematics test system based on a large language model, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Multi-mode industrial data acquisition method and system based on multi-core processor
CN121705026A
A multi-core processor-based multi-mode industrial data acquisition method and system
CN121705026B