Machine learning based multi-modal information fusion target recognition method and system
By fusing multimodal data features from the factory using machine learning models and convolutional neural networks, the problem of locating abnormal areas in multimodal information fusion was solved, improving factory safety and identification efficiency.
Patent Information
- Application Number
- CN202411770964.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-12-04
AI Technical Summary
Existing multimodal information fusion technology has difficulty quickly locating abnormal areas in factory target identification, leading to deviations and time redundancy in the identification process, which affects factory safety.
By monitoring multimodal data from the factory, machine learning models are used to extract audio mutation features and image anomaly features. These features are then fused using convolutional neural networks to determine the anomaly location area. Finally, anomaly impact index is used for layer-by-layer analysis to identify abnormal situations.
It enables rapid location of abnormal areas in the factory, improving safety and identification efficiency during factory operation and reducing identification time.
Smart Images

Figure CN119807994B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information fusion, more particularly, the present application relates to a multi-modal information fusion target recognition method and system based on machine learning. BACKGROUND
[0002] Information fusion refers to the process of collecting data from multiple sources, formats or sensors and processing it comprehensively to obtain more accurate, comprehensive and reliable information or conclusions. The goal of information fusion is to improve the quality and efficiency of decision-making by integrating information from different sources to reduce uncertainty and improve the credibility of results.
[0003] Multi-modal information fusion target recognition based on machine learning refers to the use of machine learning techniques to integrate data from multiple modalities (such as images, text, sound, etc.) to improve the accuracy and robustness of target recognition. By fully utilizing the complementary information between different modalities and combining machine learning models to extract and fuse features, more effective target recognition is achieved. In the existing multi-modal information fusion target recognition process, features are extracted from each piece of information in the multi-modal information of the factory, then all the extracted features are fused, and the abnormal situation in the factory is identified. Since the collected multi-modal information includes information other than abnormal information, directly extracting features from multi-modal information will result in features containing other interference information in addition to target features, which will cause deviations in the identification of the factory and redundant time in the identification process. Therefore, how to quickly locate the abnormal area in the factory through multi-modal information to improve the safety of the factory during operation has become a problem in the industry. SUMMARY
[0004] The present application provides a multi-modal information fusion target recognition method and system based on machine learning, which can quickly locate the abnormal area in the factory through multi-modal information, thereby improving the safety of the factory during operation.
[0005] In a first aspect, the present application provides a multi-modal information fusion target recognition method based on machine learning, comprising the following steps:
[0006] Monitoring the multi-modal data of the current target factory during operation, the multi-modal data including picture data and audio data of the target factory;
[0007] Extracting audio mutation features and picture abnormal features of the current target factory from the multi-modal data based on the trained machine learning model;
[0008] The picture abnormality features are decoupled to obtain a plurality of abnormal segmentation maps for abnormality identification of the current target plant, the abnormal boundary when the sound abnormality occurs in the current target plant is determined according to the audio mutation feature and the spatial layout of the current target plant, the convolutional neural network based on machine learning is used to perform feature fusion on all the abnormal segmentation maps and the abnormal boundary, and then the abnormal positioning area of the current target plant is obtained;
[0009] The abnormal positioning area is used as a constraint condition for abnormality identification of the current target plant, and the abnormal influence index of the current target plant is obtained by layer-by-layer analysis of the area of the current target plant in combination with a machine learning algorithm;
[0010] The abnormality of the current target plant in operation is identified by using the abnormal influence index.
[0011] In some embodiments, the audio mutation feature and the picture abnormality feature of the current target plant are extracted from the multi-modal data based on a trained machine learning model, and specifically include:
[0012] A trained machine learning model is obtained;
[0013] The multi-modal data is used as an initialization parameter of the machine learning model;
[0014] The audio mutation feature and the picture abnormality feature of the current target plant are determined according to the machine learning model.
[0015] In some embodiments, the picture abnormality features are decoupled to obtain a plurality of abnormal segmentation maps for abnormality identification of the current target plant, and specifically include:
[0016] An image segmentation threshold of the current target plant is determined;
[0017] An abnormal picture of the picture abnormality features is selected as a selected abnormal picture, and the selected abnormal picture is segmented by using the image segmentation threshold to obtain a plurality of segmented pictures of the selected abnormal picture;
[0018] An abnormal segmentation map corresponding to the selected abnormal picture for abnormality identification of the current target plant is determined according to all the segmented pictures;
[0019] The abnormal segmentation maps corresponding to the remaining abnormal pictures for abnormality identification of the current target plant are continuously determined.
[0020] In some embodiments, the abnormal boundary when the sound abnormality occurs in the current target plant is determined according to the audio mutation feature and the spatial layout of the current target plant, and specifically includes:
[0021] The spatial layout of the current target plant is obtained;
[0022] determining a plurality of audio mutation regions of the current target factory in operation according to the audio mutation feature and the spatial layout;
[0023] determining an abnormal boundary in the current target factory when the sound is abnormal through all the audio mutation regions.
[0024] In some embodiments, a convolutional neural network based on machine learning fuses features of all the abnormal segmentation maps and the abnormal boundary, and further obtains an abnormal positioning region of the current target factory, which specifically includes:
[0025] The convolutional neural network based on machine learning fuses features of all the abnormal segmentation maps and the abnormal boundary, and further obtains abnormal fusion information of the current target factory.
[0026] determining a plurality of abnormal fusion regions of the current target factory according to the abnormal fusion information;
[0027] determining the abnormal positioning region of the current target factory through all the abnormal fusion regions.
[0028] In some embodiments, the abnormal positioning region is used as a constraint condition for abnormal identification of the current target factory, and a machine learning algorithm is combined to analyze the region of the current target factory layer by layer, and further obtain an abnormal influence index of the current target factory in operation, which specifically includes:
[0029] the abnormal positioning region is used as a constraint condition for abnormal identification of the current target factory;
[0030] determining a plurality of layer-by-layer abnormal values in the current target factory according to the constraint condition, the spatial layout of the current target factory and the machine learning algorithm;
[0031] determining the abnormal influence index of the current target factory in operation through all the layer-by-layer abnormal values.
[0032] In some embodiments, the abnormal situation of the current target factory in operation is identified through the abnormal influence index, which specifically includes:
[0033] determining an abnormal influence threshold range of the abnormal situation of the current target factory in operation;
[0034] if the abnormal influence index is less than the lower limit value of the abnormal influence threshold range, the abnormal situation of the current target factory in operation is a safe state;
[0035] if the abnormal influence index is in the abnormal influence threshold range, the abnormal situation of the current target factory in operation is a to-be-handled state;
[0036] If the abnormality influence index is greater than the upper limit of the abnormality influence threshold range, the abnormal condition of the current target factory in operation is an emergency state.
[0037] In a second aspect, the present application provides a multi-modal information fusion target recognition system based on machine learning, comprising:
[0038] A monitoring module is configured to monitor multi-modal data of the current target factory in operation, wherein the multi-modal data comprises picture data and audio data of the target factory.
[0039] A processing module is configured to extract audio mutation features and picture abnormal features of the current target factory from the multi-modal data based on a trained machine learning model.
[0040] The processing module is further configured to perform feature decoupling on the picture abnormal features to obtain a plurality of abnormal segmentation maps for abnormality recognition of the current target factory, determine an abnormal boundary for sound abnormality in the current target factory according to the audio mutation features and a spatial layout of the current target factory, and perform feature fusion on all the abnormal segmentation maps and the abnormal boundary based on a convolutional neural network of machine learning, so as to obtain an abnormal positioning area of the current target factory.
[0041] The processing module is further configured to take the abnormal positioning area as a constraint condition for abnormality recognition of the current target factory, and perform layer-by-layer analysis on a region of the current target factory in combination with a machine learning algorithm, so as to obtain an abnormality influence index of the current target factory in operation.
[0042] An execution module is configured to identify an abnormal condition of the current target factory in operation through the abnormality influence index.
[0043] In a third aspect, the present application provides a computer device, comprising a memory and a processor, wherein the memory stores a code, and the processor is configured to acquire the code and execute the multi-modal information fusion target recognition method based on machine learning.
[0044] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the multi-modal information fusion target recognition method based on machine learning.
[0045] The technical scheme provided by the embodiments of the present application has the following beneficial effects:
[0046] The machine learning-based multi-modal information fusion target identification method and system provided in the application first monitors multi-modal data of a current target factory in operation, wherein the multi-modal data includes picture data and audio data of the target factory; extracts audio mutation features and picture abnormal features of the current target factory from the multi-modal data based on a trained machine learning model; decouples the picture abnormal features to obtain multiple abnormal segmentation maps for abnormal identification of the current target factory, determines an abnormal boundary when a sound is abnormal in the current target factory according to the audio mutation features and a spatial layout of the current target factory, and performs feature fusion on all the abnormal segmentation maps and the abnormal boundary based on a convolutional neural network of machine learning, thereby obtaining an abnormal positioning area of the current target factory; takes the abnormal positioning area as a constraint condition for abnormal identification of the current target factory, and performs layer-by-layer analysis on the area of the current target factory in combination with a machine learning algorithm, thereby obtaining an abnormal influence index of the current target factory in operation; and identifies an abnormal condition of the current target factory in operation through the abnormal influence index.
[0047] It can be seen that, in the multi-modal information fusion target identification process, first, based on machine learning, audio mutation features and picture abnormal features are extracted from the multi-modal data in the current target factory. The audio mutation features are used to identify the abnormal sound situation of the current target factory. The picture abnormal features are used to identify the abnormal objects in the current target factory, so as to timely locate the abnormal objects in the current target factory. The segmentation map of the current target factory in the abnormal identification is extracted from the picture abnormal features, and a plurality of abnormal segmentation maps are obtained. The abnormal segmentation map is used to quickly locate the abnormal area in the current target factory. The boundary area with sound mutation in the current target factory is extracted from the audio mutation features, and an abnormal boundary is obtained. The abnormal boundary can locate the abnormal situation in the current target factory, and then judge the abnormal situation. Secondly, based on the convolutional neural network of machine learning, all the abnormal segmentation maps and the abnormal boundary are fused, and then the area of the current target factory for locating the abnormal area when the abnormal situation occurs is analyzed from the fused features. The abnormal positioning area of the current target factory is determined. The abnormal positioning area can quickly locate the abnormal area in the current target factory, and reduce the time of identifying the abnormality in the current target factory. Therefore, the abnormal positioning area is used as a constraint condition for identifying the abnormality in the current target factory, and the area of the current target factory is analyzed layer by layer combined with the machine learning algorithm, and then the abnormal influence index of the current target factory in operation is obtained. The abnormal influence index represents the parameter value of the influence degree of the abnormal situation of the current target factory in operation, and is used to divide the abnormal situation in the current target factory. Finally, the abnormal situation of the current target factory in operation is identified by the abnormal influence index. The above scheme can quickly locate the abnormal area in the factory through multi-modal information, thereby improving the safety of the factory in operation. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 is an example flowchart of a multi-modal information fusion target identification method based on machine learning according to some embodiments of the present application;
[0049] Figure 2 is a schematic diagram for realizing abnormal picture average segmentation according to some embodiments of the present application;
[0050] Figure 3 is an example flowchart of determining an abnormal influence index according to some embodiments of the present application;
[0051] Figure 4 is a structural schematic diagram of a multi-modal information fusion target identification system based on machine learning according to some embodiments of the present application;
[0052] Figure 5is a structural schematic diagram of a computer device for implementing a machine learning-based multi-modal information fusion target recognition method according to some embodiments of the present application. DETAILED DESCRIPTION
[0053] In order to better understand the technical solutions of the present application, the technical solutions of the present application will be described in detail below in combination with the accompanying drawings and specific embodiments.
[0054] Reference Figure 1 The figure is an exemplary flowchart of a machine learning-based multi-modal information fusion target recognition method 100 according to some embodiments of the present application, which mainly includes the following steps:
[0055] In step 101, the multi-modal data of the current target factory in operation is monitored, and the multi-modal data includes picture data and audio data of the target factory.
[0056] In implementation, the multi-modal data of the current target factory in operation is monitored by a multi-modal acquisition device (such as an edge computing terminal, a multi-modal bracelet) in the prior art, and in the present embodiment, the multi-modal data includes picture data and audio data of the target factory, the picture data represents a set of pictures of each area in the current target factory, and the audio data represents audio generated in the current target factory. In other embodiments, other ways of monitoring can also be used, which are not limited here.
[0057] In step 102, the audio mutation features and picture abnormal features of the current target factory are extracted from the multi-modal data based on a trained machine learning model.
[0058] In some embodiments, the audio mutation features and picture abnormal features of the current target factory are extracted from the multi-modal data based on a trained machine learning model, which is implemented by the following steps:
[0059] Obtain a trained machine learning model;
[0060] Use the multi-modal data as the initialization parameters of the machine learning model;
[0061] Determine the audio mutation features and picture abnormal features of the current target factory according to the machine learning model.
[0062] In a specific implementation, the trained machine learning model can be obtained in the following manner: obtaining the trained machine learning model from the database of the current target factory, wherein the machine learning model includes an audio feature extraction algorithm and an image feature extraction algorithm, and the machine learning model can extract the features of the image data and the features of the audio data in the multi-modal data. The machine learning model is trained based on historical multi-modal data in the current target factory. As a preferred embodiment, the machine learning model can be trained by combining the historical multi-modal data with the deep learning method in the prior art. In other embodiments, other methods can also be used to determine the machine learning model, which is not limited herein.
[0063] It should be noted that the audio mutation feature in the present application represents the mutation of the audio data in the current target factory. The audio mutation feature includes a set of all mutated audios, which is used to identify the sound situation of the current target factory and facilitate the finding of abnormal situations in the current target factory. The picture abnormal feature represents the abnormal situation of the image data in the current target factory. The picture abnormal feature includes a set of all abnormal images. The image content in the picture abnormal feature is the image of the abnormal area in the current target factory, which is used to identify the abnormal objects in the current target factory and facilitate the timely positioning of the abnormal objects in the current target factory.
[0064] In step 103, the picture abnormal feature is decoupled to obtain a plurality of abnormal segmentation maps for identifying the abnormality of the current target factory. The abnormal boundary of the sound abnormality in the current target factory is determined based on the audio mutation feature and the spatial layout of the current target factory. The convolutional neural network based on machine learning is used to fuse the features of all abnormal segmentation maps and the abnormal boundary, and then the abnormal positioning area of the current target factory is obtained.
[0065] In some embodiments, the picture abnormal feature can be decoupled to obtain a plurality of abnormal segmentation maps for identifying the abnormality of the current target factory in the following steps:
[0066] Determining the image segmentation threshold of the current target factory;
[0067] Selecting an abnormal picture of the picture abnormal feature as a selected abnormal picture, and segmenting the selected abnormal picture by the image segmentation threshold to obtain a plurality of segmented pictures of the selected abnormal picture;
[0068] Determining the abnormal segmentation map corresponding to the selected abnormal picture for identifying the abnormality of the current target factory according to all the segmented pictures;
[0069] Continuing to determine the abnormal segmentation map corresponding to each remaining abnormal picture for identifying the abnormality of the current target factory.
[0070] In a specific implementation, the image segmentation threshold of the current target factory building can be determined in the following manner: the image segmentation threshold of the current target factory building is determined by combining the image data in the multi-modal data through an iterative threshold segmentation method, where the image segmentation threshold is a parameter value representing the segmentation degree of the image in the current target factory building, and is used to segment the image in the current target factory building; and the selected abnormal picture is segmented by using the image segmentation threshold to obtain a plurality of segmented pictures of the selected abnormal picture, which can be achieved in the following manner: the selected abnormal picture is evenly segmented by using the image segmentation threshold, and each picture obtained by the even segmentation is taken as a segmented picture of the selected abnormal picture, for example, if the image segmentation threshold is 6, the selected abnormal picture is evenly segmented into 6 pictures, and each of the 6 pictures is taken as a segmented picture, as shown in the schematic diagram of the even segmentation of the abnormal picture in some embodiments. Figure 2 The abnormal picture even segmentation shown in some embodiments can also be determined in other manners in other embodiments, which are not limited herein.
[0071] In a specific implementation, the abnormal segmentation picture corresponding to the selected abnormal picture when the current target factory building is subjected to abnormal identification can be determined in the following manner: a segmented picture is selected as a selected segmented picture, an abnormal contour is extracted from the selected segmented picture by using an edge detection method (such as a Canny edge detector or a Sobel operator) in the prior art, the remaining abnormal contours are determined, and all the abnormal contours are merged by using a morphological operation technique (such as an inflation operation or a connected region analysis) in the prior art, and the merged picture is taken as the abnormal segmentation picture corresponding to the selected abnormal picture when the current target factory building is subjected to abnormal identification; the abnormal segmentation picture can also be determined in other manners in other embodiments, which are not limited herein.
[0072] It should be noted that the abnormal segmentation picture in the present application is a segmentation picture when the current target factory building is subjected to abnormal identification, which is used to judge the abnormal situation in the current target factory building, so as to facilitate timely adjustment of the current target factory building.
[0073] In some embodiments, the abnormal boundary when the sound is abnormal in the current target factory building can be determined in the following steps:
[0074] The spatial layout of the current target factory building is obtained;
[0075] A plurality of audio mutation regions in the current target factory building during operation are determined according to the audio mutation feature and the spatial layout;
[0076] The abnormal boundary when the sound is abnormal in the current target factory building is determined by using all the audio mutation regions.
[0077] In a specific implementation, the spatial layout of the current target factory building can be obtained in the following manner: the spatial layout of the current target factory building is obtained from the drawings of the current target factory building, wherein the spatial layout represents the arrangement of objects, facilities, and functional areas in the current target factory building, and the spatial layout includes the specific positions of objects, facilities, and functional areas, which are used to determine the positions of the current target factory building; the multiple audio mutation areas of the current target factory building during operation can be determined in the following manner: the distances and directions from the sound sources of each audio in the audio mutation feature to the multi-modal acquisition device in the current target factory building are obtained from the multi-modal database of the current target factory building, the positions of each audio in the current target factory building are located by combining the distances and directions of each audio, the spatial layout, and the audio positioning system in the prior art, the adjacent positions among all positions are connected, and each region obtained by the connection is regarded as the multiple audio mutation areas of the current target factory building during operation, i.e., since there are multiple abnormal areas in the current target factory building, not all positions are adjacent, and multiple regions exist after connection, wherein the audio mutation area represents a region where a mutation sound exists in the current target factory building, i.e., a mutation sound may exist in an abnormal area, which is used to determine the abnormal conditions of the current target factory building, and other ways can also be used to determine in other embodiments, which are not limited here.
[0078] In a specific implementation, the abnormal boundary of the current target factory building when the sound is abnormal can be determined in the following manner: the farthest distance and the nearest distance from each audio mutation area to the multi-modal acquisition device are calculated by the Euclidean distance in the prior art, the points corresponding to each farthest distance are connected to obtain the farthest envelope line, the points corresponding to each nearest distance are connected to obtain the nearest envelope line, and the region between the farthest distance and the nearest distance is regarded as the abnormal boundary of the current target factory building when the sound is abnormal, and other ways can also be used to determine in other embodiments, which are not limited here.
[0079] It should be noted that the abnormal boundary in the present application is a boundary region reflecting the existence of a sound mutation in the current target factory building, which is used to locate the abnormal conditions in the current target factory building, and then determine the abnormal conditions, so that the current target factory building can make appropriate responses.
[0080] In some embodiments, since the convolutional neural network is based on machine learning, which is a branch of machine learning, the convolutional neural network based on machine learning in the present application can perform feature fusion on all the abnormal segmentation maps and the abnormal boundary, and then obtain the abnormal positioning area of the current target factory building in the following steps:
[0081] The convolutional neural network based on machine learning fuses features of all the anomaly segmentation maps and the anomaly boundary to obtain anomaly fusion information of the current target plant;
[0082] A plurality of anomaly fusion regions of the current target plant are determined according to the anomaly fusion information.
[0083] An anomaly positioning region of the current target plant is determined through all the anomaly fusion regions.
[0084] In a specific implementation, the convolutional neural network based on machine learning fuses features of all the anomaly segmentation maps and the anomaly boundary to obtain anomaly fusion information of the current target plant, which can be implemented in the following manner: the convolutional neural network based on machine learning fuses features of all the anomaly segmentation maps and the anomaly boundary, and information obtained through the fusion is taken as the anomaly fusion information of the current target plant, where the anomaly fusion information represents fusion information of anomaly region features in the current target plant, and the anomaly fusion information includes anomaly region information and boundary information; a plurality of anomaly fusion regions of the current target plant are determined from the anomaly fusion information through a fusion feature map segmentation method in the prior art, where the anomaly fusion region represents an anomaly region identified by combining audio and picture recognition in the current target plant, and is used to position the anomaly region in the current target plant; in other embodiments, the plurality of anomaly fusion regions of the current target plant can also be determined in other manners, which are not limited here.
[0085] In a specific implementation, the determination of the abnormal positioning area of the current target factory by all abnormal fusion areas can be implemented in the following manner: historical abnormal area data in the current target factory is obtained from a database of the current target factory, wherein the historical abnormal area data is a set of all historical abnormal areas, a maximum historical abnormal area in the historical abnormal area data is divided by a minimum historical abnormal area, a logarithm operation with a base of 10 is performed on a value obtained by the division, a value obtained by the logarithm operation is multiplied by an average area of all historical abnormal areas in the historical abnormal area data, a value obtained by the multiplication is taken as an abnormal area threshold of the current target factory, wherein the abnormal area threshold represents a judgment value of an abnormal area that has no impact on the current target factory, and is used for judging the abnormal area in the current target factory, one abnormal fusion area is selected as a selected abnormal fusion area, an area of the selected abnormal fusion area is compared with the abnormal area threshold, if the area of the selected abnormal fusion area is greater than or equal to the abnormal area threshold, the selected abnormal fusion area is taken as an abnormal positioning sub-area of the current target factory, if the area of the selected abnormal fusion area is less than the abnormal area threshold, the selected abnormal fusion area is removed, and the remaining abnormal fusion areas are continuously judged to obtain multiple abnormal positioning sub-areas, and a set of all the abnormal positioning sub-areas is taken as the abnormal positioning area. In other embodiments, other manners can also be used for the determination, which is not limited here.
[0086] It should be noted that the abnormal positioning area in the present application is an area for positioning an abnormal area when the current target factory appears abnormal, and is used for analyzing the abnormal situation in the current target factory, so as to quickly position the abnormal area in the current target factory and reduce the time for identifying the abnormality in the current target factory.
[0087] In step 104, the abnormal positioning area is taken as a constraint condition for identifying the abnormality in the current target factory, and a machine learning algorithm is combined to analyze the area of the current target factory layer by layer, so as to obtain the abnormal influence index of the current target factory in operation.
[0088] In some embodiments, referring to Figure 3 The figure is a flowchart for determining the abnormal influence index in some embodiments of the present application. In the present embodiment, the abnormal positioning area is taken as a constraint condition for identifying the abnormality in the current target factory, and a machine learning algorithm is combined to analyze the area of the current target factory layer by layer, so as to obtain the abnormal influence index of the current target factory in operation, which can be implemented in the following steps:
[0089] First, in step 1041, the abnormal positioning area is taken as a constraint condition for identifying the abnormality in the current target factory.
[0090] Secondly, in step 1042, a plurality of layer-by-layer abnormal values in the current target factory are determined according to the constraint condition, the spatial layout of the current target factory and the machine learning algorithm.
[0091] Finally, in step 1043, an abnormal influence index of the current target factory in operation is determined by all the layer-by-layer abnormal values.
[0092] In specific implementation, the plurality of layer-by-layer abnormal values in the current target factory can be determined according to the constraint condition, the spatial layout of the current target factory and the machine learning algorithm in the following manner, that is, a layer-by-layer abnormal value model is established by using a machine learning algorithm (such as a regression algorithm, a neural network, etc.), the constraint condition is taken as a constraint parameter of the layer-by-layer abnormal value model, the spatial layout of the current target factory is taken as an initialization parameter of the layer-by-layer abnormal value model, and the plurality of layer-by-layer abnormal values in the current target factory are output by the layer-by-layer abnormal value model, wherein the layer-by-layer abnormal value represents a parameter value of an abnormal situation of a corresponding layer of the current target factory, and is used for analyzing the abnormal situation of each layer in the current target factory; the layer-by-layer abnormal value model is, for example, a plurality of layer-by-layer abnormal values = constraint condition * A + spatial layout of current target factory * B, wherein A and B are weight coefficients, and A and B can be determined according to a large number of layer-by-layer abnormal values, and in other embodiments, other manners can also be used for determination, which is not limited here.
[0093] In specific implementation, the abnormal influence index of the current target factory in operation can be determined by all the layer-by-layer abnormal values in the following manner, that is, the maximum layer-by-layer abnormal value is divided by the minimum layer-by-layer abnormal value, a logarithmic operation with base 2 is performed on the value obtained by the division, the value obtained by the logarithmic operation is multiplied by the average value of all the layer-by-layer abnormal values, and the value obtained by the multiplication is taken as the abnormal influence index of the current target factory in operation, and in other embodiments, other manners can also be used for implementation, which is not limited here.
[0094] It should be noted that the abnormal influence index in the present application represents a parameter value of an influence degree of an abnormal situation of the current target factory in operation, and is used for dividing the abnormal situation in the current target factory, so as to facilitate the target factory to make an existing coping strategy.
[0095] In step 105, the abnormal situation of the current target factory in operation is identified by the abnormal influence index.
[0096] In some embodiments, the abnormal situation of the current target factory in operation can be identified by the abnormal influence index in the following steps:
[0097] An abnormal influence threshold range of the abnormal situation of the current target factory in operation is determined.
[0098] if the abnormal influence index is less than the lower limit value of the abnormal influence threshold range, the abnormal condition of the current target factory in operation is a safe state;
[0099] if the abnormal influence index is in the abnormal influence threshold range, the abnormal condition of the current target factory in operation is a to-be-handled state;
[0100] if the abnormal influence index is greater than the upper limit value of the abnormal influence threshold range, the abnormal condition of the current target factory in operation is an emergency state.
[0101] In specific implementation, when the abnormal condition of the current target factory in operation is a safe state, no processing is performed on the current target factory, when the abnormal condition of the current target factory in operation is a to-be-handled state, a coping strategy needs to be made for the current target factory in time, and when the abnormal condition of the current target factory in operation is an emergency state, the operation of the current target factory is stopped and a coping strategy is made. In other embodiments, other ways can also be used for implementation, which are not limited here.
[0102] It should be noted that in the present application, the abnormal influence threshold range can be set according to the specific needs of the current target factory, for example, if the product processed in the current target factory has a high safety factor, the abnormal influence threshold range can be set in a high range, and if the product processed in the current target factory is dangerous, the abnormal influence threshold range can be set in a low range. In other embodiments, for example, if chemical products are processed in the current target factory, the abnormal influence threshold range can be set in a high range, thereby improving the safety of the current target factory in operation.
[0103] In addition, another aspect of the present application, in some embodiments, the present application provides a multi-modal information fusion target recognition system based on machine learning, referring to Figure 4 The figure is a structural schematic diagram of a multi-modal information fusion target recognition system based on machine learning according to some embodiments of the present application. The multi-modal information fusion target recognition system 400 includes a monitoring module 401, a processing module 402 and an execution module 403, which are described as follows:
[0104] The monitoring module 401 is mainly used for monitoring the multi-modal data of the current target factory in operation in the present application. The multi-modal data includes picture data and audio data of the target factory.
[0105] The processing module 402 is used for extracting the audio mutation features and picture abnormal features of the current target factory from the multi-modal data based on the trained machine learning model in the present application.
[0106] It should be noted that the processing module 402 is further configured to decouple the picture abnormal features to obtain a plurality of abnormal segmentation maps for identifying the current target factory, determine an abnormal boundary in the current target factory when the sound is abnormal according to the audio mutation feature and the spatial layout of the current target factory, and perform feature fusion on all the abnormal segmentation maps and the abnormal boundary based on a convolutional neural network of machine learning, so as to obtain an abnormal positioning area of the current target factory.
[0107] In addition, it should be noted that the processing module 402 is further configured to take the abnormal positioning area as a constraint condition for identifying the current target factory, and perform layer-by-layer analysis on the area of the current target factory in combination with a machine learning algorithm, so as to obtain an abnormal influence index of the current target factory when running.
[0108] The execution module 403 is mainly configured to identify the abnormal condition of the current target factory when running by using the abnormal influence index.
[0109] In addition, the present application further provides a computer device, which comprises a memory and a processor, the memory stores a code, and the processor is configured to acquire the code and execute the above-mentioned machine learning-based multi-modal information fusion target identification method.
[0110] In some embodiments, with reference to Figure 5 The figure is a structural schematic diagram of a computer device for implementing the machine learning-based multi-modal information fusion target identification method according to some embodiments of the present application. The machine learning-based multi-modal information fusion target identification method in the above-mentioned embodiments can be implemented by the computer device shown in the figure, which comprises at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504. Figure 5
[0111] The processor 501 can be a general central processing unit (CPU) or an application specific integrated circuit (ASIC).
[0112] The communication bus 502 can be used to transmit information between the above-mentioned components.
[0113] The memory 503 can be a read only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read only memory (EEPROM), a compact disc read only memory (CD ROM) or other optical disk storage, a magnetic disk or other magnetic storage device, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to. The memory 503 can exist independently, and is connected to the processor 501 through the communication bus 502. The memory 503 can also be integrated with the processor 501.
[0114] The memory 503 is configured to store program codes for implementing the solutions of the present application, and the processor 501 is configured to control the execution of the program codes. The processor 501 is configured to execute the program codes stored in the memory 503. The program codes can include one or more software modules. The methods used in the above embodiments can be implemented by the processor 501 and one or more software modules in the program codes in the memory 503.
[0115] The communication interface 504 is configured to communicate with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc., using any transceiver-like mechanism.
[0116] In specific implementations, as an example, the computer device can include multiple processors, each of which can be a single CPU processor or a multi-CPU processor. The processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0117] The computer device described above can be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device can be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of the computer device.
[0118] In addition, the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the machine learning based multi-modal information fusion target recognition method described above.
[0119] Although the preferred embodiments of the present application have been described, those skilled in the art who are familiar with the basic inventive concept can make additional changes and modifications to the embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0120] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A multimodal information fusion target recognition method based on machine learning, characterized in that, Includes the following steps: Monitor the multimodal data of the target factory building during its operation, including image data and audio data of the target factory building; The trained machine learning model extracts audio mutation features and image anomaly features of the current target factory from the multimodal data; The abnormal features of the image are decoupled to obtain multiple abnormal segmentation maps when identifying anomalies in the current target factory. The abnormal boundary when the sound is abnormal in the current target factory is determined based on the audio mutation features and the spatial layout of the current target factory. The convolutional neural network based on machine learning fuses all the abnormal segmentation maps and the abnormal boundary to obtain the abnormal location area of the current target factory. The abnormal location area is used as a constraint condition for anomaly identification of the current target factory building, and machine learning algorithms are combined to analyze the area of the current target factory building layer by layer, thereby obtaining the abnormal impact index of the current target factory building during operation. The abnormal conditions in the operation of the current target plant are identified by the aforementioned abnormal impact index; Specifically, determining the abnormal boundary when there is sound abnormality in the current target factory based on the audio mutation characteristics and the spatial layout of the current target factory includes: Obtain the spatial layout of the current target factory building; Based on the audio mutation characteristics and the spatial layout, multiple audio mutation areas are determined during the operation of the current target factory building; Determine the abnormal boundaries when there is sound abnormality in the current target factory by identifying all audio mutation regions; Specifically, determining the abnormal boundary when sound anomalies occur in the current target factory by using all audio mutation regions includes: calculating the farthest and nearest distances from each audio mutation region to the multimodal acquisition device using Euclidean distance; connecting the points corresponding to each farthest distance to obtain the farthest envelope; connecting the points corresponding to each nearest distance to obtain the nearest envelope; and using the region between the farthest and nearest distances as the abnormal boundary when sound anomalies occur in the current target factory.
2. The method as described in claim 1, characterized in that, The machine learning model based on training extracts audio mutation features and image anomaly features of the current target factory from the multimodal data, specifically including: Obtain the trained machine learning model; The multimodal data is used as the initialization parameters of the machine learning model; The machine learning model is used to determine the audio mutation features and image anomaly features of the current target factory.
3. The method as described in claim 1, characterized in that, By decoupling the abnormal features of the image, multiple abnormal segmentation maps are obtained for anomaly identification of the current target factory building. These specifically include: Determine the image segmentation threshold for the current target factory building; Select an abnormal image with the aforementioned abnormal features as the selected abnormal image, and segment the selected abnormal image using the aforementioned image segmentation threshold to obtain multiple segmented images of the selected abnormal image; When identifying anomalies in the current target factory building based on all segmented images, select the anomaly segmentation image corresponding to the anomaly image. Continue to determine the anomaly segmentation images corresponding to each remaining anomaly image when performing anomaly identification on the current target factory.
4. The method as described in claim 1, characterized in that, A machine learning-based convolutional neural network fuses features from all anomaly segmentation maps and anomaly boundaries to obtain the anomaly location area of the current target factory building, specifically including: A convolutional neural network based on machine learning fuses the features of all the anomaly segmentation maps and the anomaly boundaries to obtain the anomaly fusion information of the current target factory building. Based on the abnormal fusion information, multiple abnormal fusion areas of the current target factory building are determined; The abnormal location area of the current target plant is determined by identifying all abnormal fusion areas.
5. The method as described in claim 1, characterized in that, Using the aforementioned abnormal location area as a constraint condition for anomaly identification of the current target factory building, and combining machine learning algorithms to perform layer-by-layer analysis of the current target factory building's area, the specific abnormal impact index of the current target factory building during operation is obtained, including: The abnormal location area is used as a constraint condition when identifying anomalies in the current target factory building. Based on the constraints, the spatial layout of the current target factory, and machine learning algorithms, multiple layer-by-layer outliers in the current target factory are determined. The abnormal impact index of the current target plant during operation is determined by identifying all the outliers at each level.
6. The method as described in claim 1, characterized in that, The identification of abnormal conditions in the operation of the current target plant through the aforementioned abnormal impact index specifically includes: Determine the threshold range of abnormal impacts of the current target plant's abnormal operating conditions; If the abnormal impact index is less than the lower limit of the abnormal impact threshold range, then the abnormal condition of the current target plant in operation is a safe state. If the abnormal impact index is within the abnormal impact threshold range, then the abnormal condition of the current target plant in operation is in a state of pending processing. If the abnormal impact index is greater than the upper limit of the abnormal impact threshold range, then the abnormal situation of the current target plant in operation is an emergency state.
7. A multimodal information fusion target recognition system based on machine learning, which uses the method described in any one of claims 1 to 6 for multimodal information fusion target recognition, characterized in that, The system includes: The monitoring module is used to monitor the multimodal data of the target factory building during its operation. The multimodal data includes image data and audio data of the target factory building. The processing module is used to extract audio mutation features and image anomaly features of the current target factory from the multimodal data based on the trained machine learning model; The processing module is also used to decouple the abnormal features of the image to obtain multiple abnormal segmentation maps when identifying abnormalities in the current target factory building. Based on the audio mutation features and the spatial layout of the current target factory building, the abnormal boundary when the sound is abnormal in the current target factory building is determined. Based on the convolutional neural network of machine learning, all the abnormal segmentation maps and the abnormal boundary are fused to obtain the abnormal location area of the current target factory building. The processing module is also used to use the abnormal location area as a constraint condition when identifying anomalies in the current target factory building, and to combine machine learning algorithms to perform layer-by-layer analysis of the area of the current target factory building, thereby obtaining the abnormal impact index of the current target factory building during operation. The execution module is used to identify abnormal conditions in the operation of the current target plant through the abnormal impact index.
8. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing code, and the processor being configured to retrieve the code and execute the machine learning-based multimodal information fusion target recognition method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multimodal information fusion target recognition method based on machine learning as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Leakage detection method and system based on sound image fusion
CN117739289A
IDC machine room operation and maintenance method and system, electronic equipment and storage medium
CN118229648A