Intelligent monitoring system and method based on vision and language fusion

By integrating visual and language data into an intelligent monitoring method, the problem of poor monitoring effect of a single data source in existing technologies is solved, deep perception and intelligent analysis of target scenes are achieved, the accuracy and timeliness of monitoring are improved, and anomalies can be identified in a timely manner and safety analysis reports can be generated.

CN120635828APending Publication Date: 2025-09-12LIANYUNGANG DIGITAL IND INVESTMENT DEVELOPMENT CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202511102230.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The existing public safety monitoring system relies on a single data source, making it difficult to comprehensively and accurately capture and analyze safety information in the target scenario. Especially in large-scale open environments, a single data source cannot provide sufficient information depth and breadth, resulting in poor monitoring results.

Method used

Integrate visual and language data, collect spatiotemporal sequence images and text data synchronously through visual sensors, perform multi-dimensional spatial analysis and contextual association analysis, establish a multi-level interactive channel for hierarchical feature information and semantic coding information, perform local feature matching and trend consistency calibration, generate joint representation data, and dynamically correct the evaluation boundary parameters in combination with environmental feedback parameters to determine the public safety status.

Benefits of technology

It achieves deep perception and intelligent analysis of target scenarios, improves the accuracy and timeliness of monitoring, and can identify anomalies in a timely manner and generate safety analysis reports to ensure public safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635828A_ABST
    Figure CN120635828A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of public safety monitoring, and particularly relates to an intelligent monitoring system and method based on vision and language fusion. According to the method, multi-dimensional spatial analysis is carried out on a space-time sequence image, hierarchical feature information can be extracted, context association analysis is carried out on text data, semantic coding information is generated, then image features and text entity description are in butt joint through a multi-stage interaction channel of the hierarchical feature information and the semantic coding information, and the text entity description is extracted. And performing trend consistency calibration, eliminating phase deviation on a time sequence, predicting a future trend of hierarchical features and semantic coding, generating joint characterization data, dynamically correcting evaluation boundary parameters by combining environment feedback parameters on the basis, and finally evaluating the current joint characterization data according to the corrected evaluation boundary parameters. And determining the public safety state level in the target scene, thereby realizing the intelligent monitoring process of the public safety information in the target scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of public safety monitoring, and in particular relates to an intelligent monitoring system and method based on vision and language fusion. Background Art

[0002] With the rapid development of information technology, the field of public safety monitoring is facing increasingly complex and changing security challenges. Traditional monitoring methods often rely on single sensor data or manual monitoring, which makes it difficult to comprehensively and accurately capture and analyze security information in the target scenario. Especially in large-scale, open public environments such as urban blocks and transportation hubs, a single data source often cannot provide sufficient information depth and breadth, resulting in poor monitoring results. Therefore, how to integrate multiple data sources to achieve comprehensive perception and intelligent analysis of information has become an urgent problem to be solved in the field of public safety monitoring.

[0003] In the existing technology, there have been some attempts to improve monitoring effects by fusing different types of sensor data. However, most of these methods are limited to the simple superposition of image data and often ignore the important value of text data in public safety monitoring, especially social media texts and monitoring logs. They can not only provide key semantic clues and event context for monitoring, but also complement image data to enhance the accuracy and timeliness of monitoring. Ignoring text data will obviously limit the performance improvement and application scope of intelligent monitoring systems. Based on this, this solution provides an intelligent monitoring method based on the fusion of vision and language to solve the above problems. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent monitoring system and method based on vision and language fusion, which can integrate visual and language data in the target scene to achieve deep perception and intelligent analysis of public safety information in the target scene.

[0005] The technical solutions adopted by the present invention are as follows: An intelligent monitoring method based on vision and language fusion, comprising: The visual sensor collects spatiotemporal sequence images of the target scene and simultaneously collects text data associated with the target scene; Perform multi-dimensional spatial analysis on spatiotemporal sequence images to extract hierarchical feature information including temporal variation patterns; Perform context-sensitive analysis on text data to generate semantically coherent semantically encoded information; Establish a multi-level interactive channel between hierarchical feature information and semantic encoding information, perform local feature matching in the spatial dimension, and perform trend consistency calibration in the temporal dimension to generate joint representation data; Based on the historical joint characterization data within a preset time period, the evaluation boundary parameters are dynamically corrected in combination with the environmental feedback parameters, and the current joint characterization data is evaluated according to the corrected evaluation boundary parameters to determine the public safety status in the target scenario.

[0006] In a preferred embodiment, the acquisition of spatiotemporal sequence images and text data includes the following steps: The pre-deployed multispectral sensor array synchronously acquires three-dimensional spatiotemporal sequence image data of visible light band, infrared band and depth information; Establish a text collection thread pool to capture monitoring logs, social media text, and speech transcription data in real time through a natural language interface, and attach timestamp tags that match the spatiotemporal sequence images; The spatiotemporal sequence images and text data are spatiotemporally calibrated, and the buffers of the spatiotemporal sequence images and text data are aligned using a sliding window mechanism, and the spatiotemporal deviation is eliminated through feature point matching.

[0007] In a preferred embodiment, the step of performing multi-dimensional spatial analysis on the spatiotemporal sequence images to extract hierarchical feature information including temporal variation patterns includes: Perform multi-scale spatial decomposition on spatiotemporal sequence images, extract feature sets including object contours, surface textures, and motion trajectories at different resolution levels, and form a multi-level feature structure; Construct a spatiotemporal correlation matrix of multi-level feature structures, and perform cross-level fusion processing on object contours, surface textures, and motion trajectories to generate a spatiotemporal feature map; Perform time series analysis on multi-level feature structures to identify periodic change features, trend change features, and sudden abnormal features; Based on the periodic change characteristics and trend change characteristics, combined with the spatiotemporal feature map, topological extension is performed to generate a regular hierarchical feature information database; Based on the sudden abnormality features, similar abnormal patterns are searched in the historical feature library, and adjacent feature fragments of the sudden abnormality features under similar abnormality patterns are collected and summarized into a hierarchical feature information library for sudden abnormality reference.

[0008] In a preferred embodiment, the step of performing context-related parsing on text data to generate semantically coherent semantically encoded information includes: Based on the preset encoding mechanism of the text data context, the text content in the text data is converted into a semantic vector sequence with time attributes; The corresponding entity objects in the semantic vector sequence are associated and mapped with the entities with the same name in the spatiotemporal sequence images to build a bidirectional matching relationship between entities and semantics; Collect multiple semantic vector sequence fragments adjacent to the current timestamp, build a semantic context network through time series analysis, and then combine them with events in the historical semantic library for semantic enhancement processing to form a dynamically updated semantic coding information library.

[0009] In a preferred embodiment, the step of establishing a multi-level interactive channel of hierarchical feature information and semantic coding information and performing local feature matching in the spatial dimension includes: Construct a hierarchical structure of multi-level interactive channels and map hierarchical feature information and semantic encoding information layer by layer, including: At the low-level semantic layer, it is the entity mapping from object contours to text semantics; At the middle semantic layer, there is the semantic association from surface texture and motion trajectory to event description; At the high-level semantic layer, it is the fusion of spatiotemporal feature maps to scene semantics; Calculate the matching degree between local features of the image and text entity descriptions, and achieve the connection between multi-level features and semantic information through spatial coordinate mapping and semantic vector distance; Among them, when the matching degree between the local features of the image and the text entity description is lower than the preset matching threshold, the corresponding local features are marked as abnormal areas, and the abnormal warning mechanism is triggered simultaneously. Then, a secondary verification is performed in combination with the contextual relationship in the semantic coding information library. In the secondary verification, if the matching degree is still lower than the matching threshold, the abnormal area is marked as high risk and an abnormality report is generated.

[0010] In a preferred embodiment, the step of performing trend consistency calibration in the time dimension to generate joint characterization data includes: Extract the periodic and trend change patterns in the hierarchical feature information, and perform dynamic time warping with the timestamp event sequence in the semantic coding information to eliminate the phase deviation in the timing; Through time series analysis, the future trends of hierarchical features and semantic encoding are predicted and recorded as prediction trajectories and semantic evolution paths respectively; The consistency between the predicted trajectory and the semantic evolution path is verified through cross-validation, and the verification deviation is output simultaneously; When the verification deviation exceeds the preset deviation threshold, it indicates that there is inconsistency between the hierarchical features and the semantic encoding, triggering the abnormal warning mechanism; When the verification deviation does not exceed the preset deviation threshold, it indicates that the hierarchical features and semantic coding are consistent, and the hierarchical feature information and semantic coding information are fused to generate joint representation data.

[0011] In a preferred embodiment, the step of dynamically revising the evaluation boundary parameters based on the historical joint characterization data within a preset time period in combination with the environmental feedback parameters includes: Convert environmental feedback parameters and historical joint representation data into the same dimension and align them along the time axis; Calculate the initial value of the evaluation boundary parameter based on the distribution characteristics of the historical joint representation data and the changing trend of the environmental feedback parameters; Collect the dynamic changes of the current environmental feedback parameters in real time and record them as correction condition parameters; When the correction condition parameters are within the allowable fluctuation range, the original assessment boundary parameters are maintained unchanged, and the dynamic changes of the environmental feedback parameters continue to be monitored; If it exceeds the allowable fluctuation range, the correction condition parameters are substituted into the boundary adjustment function, the new evaluation boundary parameters are calculated, and the original evaluation boundary parameters are replaced.

[0012] In a preferred embodiment, the step of evaluating the current joint representation data based on the modified evaluation boundary parameters to determine the public safety status in the target scenario includes: Setting safety level thresholds based on assessment boundary parameters; Compare the current joint representation data with the security level threshold and output the current public safety status level, where the public safety status level is divided into three levels: safe, risky, and emergency; When the public safety status level is risk or emergency, the corresponding early warning mechanism is triggered and a safety analysis report is generated. The early warning level in the emergency state is higher than that in the risk state. When the public security status level is safe, continuous monitoring is maintained.

[0013] The present invention also provides an intelligent monitoring system based on vision and language fusion, using the above-mentioned intelligent monitoring method based on vision and language fusion, comprising: A data acquisition module is used to collect spatiotemporal sequence images of the target scene through a visual sensor, and simultaneously collect text data associated with the target scene; Image feature extraction module, used to perform multi-dimensional spatial analysis on spatiotemporal sequence images and extract hierarchical feature information including temporal variation patterns; The text feature extraction module is used to perform context-related analysis on text data and generate semantically coherent semantic coding information; The feature fusion module is used to establish a multi-level interactive channel between hierarchical feature information and semantic encoding information, perform local feature matching in the spatial dimension, and perform trend consistency calibration in the temporal dimension to generate joint representation data; The status assessment module is used to dynamically correct the assessment boundary parameters based on the historical joint characterization data within a preset time period and the environmental feedback parameters, and to evaluate the current joint characterization data based on the corrected assessment boundary parameters to determine the public safety status in the target scenario.

[0014] And, an electronic device, comprising: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned intelligent monitoring method based on vision and language fusion.

[0015] The technical effects achieved by the present invention are: The present invention constructs an intelligent monitoring method by integrating visual and language data in the target scene. First, the spatiotemporal sequence images and related text data in the target scene are collected, multi-dimensional spatial analysis is performed on the spatiotemporal sequence images, hierarchical feature information is extracted, and context association analysis is performed on the text data to generate semantic coding information. Through a multi-level interactive channel of hierarchical feature information and semantic coding information, image features are connected with text entity descriptions, and trend consistency calibration is performed to eliminate phase deviation in time sequence, and future trends of hierarchical features and semantic coding are predicted, and joint representation data is generated. Based on this, the evaluation boundary parameters are dynamically corrected in combination with environmental feedback parameters. Finally, the current joint representation data is evaluated according to the corrected evaluation boundary parameters to determine the public safety status level in the target scene, thereby realizing an intelligent monitoring process for public safety information in the target scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic flow chart of the method of the present invention; Figure 2 It is a schematic diagram of the system modules of the present invention; Figure 3 It is a schematic structural diagram of an electronic device of the present invention. DETAILED DESCRIPTION

[0017] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0018] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0019] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in a preferred embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it constitute a separate or selective embodiment that is mutually exclusive of other embodiments.

[0020] See also Figure 1 As shown, the present invention provides an intelligent monitoring method based on vision and language fusion, comprising: S1, collecting spatiotemporal sequence images of the target scene through a visual sensor, and synchronously collecting text data associated with the target scene; In step S1, when it is necessary to monitor the target scene, the visual sensor is first used to collect spatiotemporal sequence images of the target scene. The spatiotemporal sequence images can capture information about the scene changing over time, including the movement of objects, changes in appearance, etc. At the same time, text data associated with the target scene, such as monitoring logs and social media texts, is also collected synchronously. The text data contains rich semantic information and helps to fully understand the security status of the target scene. The collection of spatiotemporal sequence images and text data includes the following steps: The pre-deployed multispectral sensor array synchronously acquires three-dimensional spatiotemporal sequence image data of visible light band, infrared band and depth information; Establish a text collection thread pool to capture monitoring logs, social media text, and speech transcription data in real time through a natural language interface, and attach timestamp tags that match the spatiotemporal sequence images; Perform spatiotemporal calibration on spatiotemporal sequence images and text data, use a sliding window mechanism to align the buffers of spatiotemporal sequence images and text data, and eliminate spatiotemporal deviations through feature point matching; Specifically, when collecting spatiotemporal sequence images and text data in the target scene, a multispectral sensor array is first pre-deployed in the target scene. The multispectral sensor array can synchronously obtain three-dimensional spatiotemporal sequence image data of visible light band, infrared band and depth information. The visible light band image provides intuitive visual information of the scene, and the infrared band image can capture the thermal radiation information of the object, which is used to identify hidden heat sources or abnormal temperature areas. The depth information provides the three-dimensional spatial position relationship of the objects in the scene, thereby realizing multi-level monitoring of the target scene. At the same time, in order to synchronously collect text data associated with the target scene, a text collection system is established. Thread pool, the text collection thread pool captures monitoring logs, social media texts and voice transcription data in real time through the natural language interface, and attaches timestamp labels that match the spatiotemporal sequence images. The text data contains semantic information such as event descriptions and character behaviors, so as to facilitate a detailed understanding of the security status of the target scene. After collecting the spatiotemporal sequence images and text data, the spatiotemporal sequence images and text data need to be calibrated in time and space. Specifically, the sliding window mechanism can be used to align the buffers of the spatiotemporal sequence images and text data, and the spatiotemporal deviation can be eliminated through feature point matching to ensure the consistency of the spatiotemporal sequence images and text data in time and space.

[0021] S2. Perform multi-dimensional spatial analysis on the spatiotemporal sequence images to extract hierarchical feature information including temporal variation patterns; In step S2, when analyzing the spatiotemporal sequence images, hierarchical feature information including temporal variation patterns is extracted from the acquired spatiotemporal sequence images to reveal key information such as the motion trajectory and behavior pattern of objects in the scene. The step of performing multi-dimensional spatial analysis on the spatiotemporal sequence images to extract hierarchical feature information including temporal variation patterns includes: Perform multi-scale spatial decomposition on spatiotemporal sequence images, extract feature sets including object contours, surface textures, and motion trajectories at different resolution levels, and form a multi-level feature structure; Construct a spatiotemporal correlation matrix of multi-level feature structures, and perform cross-level fusion processing on object contours, surface textures, and motion trajectories to generate a spatiotemporal feature map; Perform time series analysis on multi-level feature structures to identify periodic change features, trend change features, and sudden abnormal features; Based on the periodic change characteristics and trend change characteristics, combined with the spatiotemporal feature map, topological extension is performed to generate a regular hierarchical feature information database; Based on the sudden abnormality characteristics, similar abnormal patterns are searched in the historical feature library, and adjacent feature fragments of sudden abnormality characteristics under similar abnormality patterns are collected and summarized into a hierarchical feature information library for sudden abnormality reference; Specifically, when performing spatial analysis of spatiotemporal sequence images, the spatiotemporal sequence images are first subjected to multi-scale spatial decomposition so that the spatiotemporal sequence images can be observed at different scales, thereby extracting a multi-level feature set including object contours, surface textures, and motion trajectories. The multi-level feature set can reveal different details and layers of objects in the scene at different resolution levels, providing a basis for subsequent spatiotemporal correlation analysis. In order to effectively fuse features at different levels, this embodiment constructs a spatiotemporal correlation matrix of a multi-level feature structure. The spatiotemporal correlation matrix can reflect the cross-level connection between object contours, surface textures, and motion trajectories. Through fusion processing, a spatiotemporal feature map of multi-level features can be generated. The spatiotemporal feature map not only contains the static features of the object, but also incorporates dynamic information that changes over time, providing corresponding support for subsequent anomaly detection and trend prediction. Then, a time series analysis is performed on the multi-level feature structure to identify the periodic change characteristics, trend change characteristics, and sudden abnormal characteristics of objects in the scene. Time series analysis can be based on autoregressive integral sliding. Time series prediction algorithms such as the Automatic Regression Average (ARIMA) model and the Long Short-Term Memory (LSTM) network are used for implementation. Periodic change features reveal the regular changes of objects in the scene over time, such as periodic fluctuations in traffic flow. Trend change features reflect the long-term change trends of objects in the scene over time, such as urban expansion and population migration. Sudden anomaly features indicate sudden abnormal events or behaviors in the scene, such as traffic accidents and sudden gatherings. This provides key information for the subsequent construction of a hierarchical feature information library. Based on periodic and trend change features, topological extension is performed in combination with the spatiotemporal feature map to generate a regular hierarchical feature information library. The hierarchical feature information library not only contains the static features of objects in the scene, but also incorporates the dynamic laws of change over time, providing corresponding support for subsequent trend prediction and anomaly detection. In addition, for sudden anomaly features, similar anomaly patterns are searched in the historical feature library. Through comparison and analysis, adjacent feature fragments of sudden anomaly features under similar anomaly patterns are collected and summarized into a hierarchical feature information library for sudden reference.

[0022] S3, performing context-related analysis on text data to generate semantically coherent semantically encoded information; In step S3, for the collected text data, context association analysis is performed on the text data, and natural language processing technology is used to generate semantically coherent semantic coding information to facilitate understanding of the meaning of the text data, thereby complementing the spatiotemporal sequence image. The step of performing context association analysis on the text data to generate semantically coherent semantic coding information includes: Based on the preset encoding mechanism of the text data context, the text content in the text data is converted into a semantic vector sequence with time attributes; The corresponding entity objects in the semantic vector sequence are associated and mapped with the entities with the same name in the spatiotemporal sequence images to build a bidirectional matching relationship between entities and semantics; Collect multiple semantic vector sequence fragments adjacent to the current timestamp, build a semantic context network through time series analysis, and then perform semantic enhancement processing based on events in the historical semantic library to form a dynamically updated semantic coding information library; Specifically, when analyzing text data, we first use a preset encoding mechanism based on the contextual relationship of the text data. The encoding mechanism can be implemented through pre-trained language models such as Word2Vec and BERT, so that each word or sentence in the text data is converted into a corresponding semantic vector. The semantic vector has a time attribute, which can reflect the changes in the text content in time sequence, thereby constructing a semantic vector sequence. In order to associate the text data with the spatiotemporal sequence image, the corresponding entity objects in the semantic vector sequence, such as people, places, events, etc., will be associated and mapped with the entities of the same name in the spatiotemporal sequence image. This mapping relationship realizes The connection between text description and image features. In addition, in order to enrich the connotation of semantic information, multiple semantic vector sequence fragments adjacent to the current timestamp will be collected to jointly constitute the context environment of the text data. Through time series analysis, a semantic context network can be constructed. The context network can reveal the temporal correlation and evolution trend of text data. Finally, semantic enhancement processing is performed in combination with events in the historical semantic library. The historical semantic library stores historical events and their related semantic information. Through comparison and analysis, the accuracy and richness of the semantic coding information can be further improved, and finally a dynamically updated semantic coding information library is formed.

[0023] S4. Establish a multi-level interactive channel between hierarchical feature information and semantic encoding information, perform local feature matching in the spatial dimension, and perform trend consistency calibration in the temporal dimension to generate joint representation data; In step S4, a multi-level interactive channel between hierarchical feature information and semantic coding information is established. Local feature matching is performed in the spatial dimension to identify corresponding parts in the image and text data; trend consistency calibration is performed in the temporal dimension to ensure that the correspondence between the image and text data in the time series is accurate. In this way, joint representation data is generated, which integrates the multi-dimensional information of the image and text. The multi-level interactive channel between hierarchical feature information and semantic coding information is established, and the step of performing local feature matching in the spatial dimension includes: Construct a hierarchical structure of multi-level interactive channels and map hierarchical feature information and semantic encoding information layer by layer, including: At the low-level semantic layer, it is the entity mapping from object contours to text semantics; At the middle semantic layer, there is the semantic association from surface texture and motion trajectory to event description; At the high-level semantic layer, it is the fusion of spatiotemporal feature maps to scene semantics; Calculate the matching degree between local features of the image and text entity descriptions, and achieve the connection between multi-level features and semantic information through spatial coordinate mapping and semantic vector distance; When the degree of match between the local features of the image and the text entity description is lower than the preset matching threshold, the corresponding local features are marked as abnormal areas, and the abnormal warning mechanism is triggered simultaneously. A secondary verification is then performed in combination with the contextual relationship in the semantic coding information library. If the matching degree is still lower than the matching threshold in the secondary verification, the abnormal area is marked as high risk and an abnormality report is generated. Specifically, when performing feature matching on hierarchical feature information and semantic coding information, a multi-level interactive channel for information interaction is first constructed between the hierarchical feature information and the semantic coding information, and layer-by-layer mapping is performed in the hierarchical structure of the multi-level interactive channel. At the low-level semantic layer, entity mapping from object contours to text semantics is mainly realized. The mapping relationship of the low-level semantic layer is relatively direct. By comparing the object contours in the spatiotemporal sequence images with the entity descriptions in the text data, the association between the image and the text can be preliminarily established. For example, the vehicle contour identified in the spatiotemporal sequence image is compared with the "vehicle intrusion" event mentioned in the text data. Row association, in the middle semantic layer, it mainly realizes the semantic association from surface texture and motion trajectory to event description. The mapping relationship of the middle semantic layer is more complex than that of the low semantic layer. It needs to combine the surface texture and motion trajectory in the spatiotemporal sequence image and the event description in the text data. For example, the slippery road texture shown in the spatiotemporal sequence image is associated with the "slippery road in rainy days, please pay attention to safety" event mentioned in the text data, so as to infer that the slippery road surface may cause vehicle driving safety problems. In the high semantic layer, it mainly realizes the fusion of spatiotemporal feature map to scene semantics. The mapping relationship of the high semantic layer is the most complex, and the spatiotemporal feature map needs to be combined with the spatiotemporal feature map. The multi-level features in the spectrum are deeply integrated with the scene semantics in the text data. Through comprehensive analysis, the complex relationships hidden in the scene are revealed. For example, the periodic changes in traffic flow shown in the spatiotemporal feature map are associated with the "traffic congestion during peak hours in the morning and evening" mentioned in the text data, thereby inferring the causes and laws of traffic congestion. After constructing a multi-level interactive channel, the matching degree between the local features of the image and the text entity description will be calculated. The matching degree calculation can be achieved through spatial coordinate mapping and semantic vector distance. Spatial coordinate mapping can ensure the consistency of image features and text descriptions in spatial position, while semantic vector distance can It can reflect the semantic similarity between image features and text descriptions. When the matching degree between the local features of the image and the text entity description is lower than the preset matching threshold, the corresponding local features will be judged to be abnormal. At this time, the corresponding local features will be marked as abnormal areas, and the abnormal warning mechanism will be triggered simultaneously to remind relevant personnel to pay attention. To ensure the accuracy of abnormal judgment, a secondary verification will be performed in combination with the contextual relationship in the semantic coding information library. In the secondary verification, if the matching degree is still lower than the matching threshold, the abnormal area will be marked as high risk and an abnormal report will be automatically generated so that relevant personnel can take timely response measures.

[0024] Secondly, trend consistency calibration is performed in the time dimension to generate joint representation data, including the following steps: Extract the periodic and trend change patterns in the hierarchical feature information, and perform dynamic time warping with the timestamp event sequence in the semantic coding information to eliminate the phase deviation in the timing; Through time series analysis, the future trends of hierarchical features and semantic encoding are predicted and recorded as prediction trajectories and semantic evolution paths respectively; The consistency between the predicted trajectory and the semantic evolution path is verified through cross-validation, and the verification deviation is output simultaneously; When the verification deviation exceeds the preset deviation threshold, it indicates that there is inconsistency between the hierarchical features and the semantic encoding, triggering the abnormal warning mechanism; When the verification deviation does not exceed the preset deviation threshold, it indicates that the hierarchical features and semantic encoding are consistent, and the hierarchical feature information and semantic encoding information are fused to generate joint representation data; Specifically, after the local feature matching is completed, the trend consistency calibration of the successfully matched hierarchical feature information and semantic coding information will be performed in the time dimension. First, the periodic change law and trend change law will be extracted from the hierarchical feature information. The periodic change law and trend change law reflect the dynamic change characteristics of the objects in the target scene. Then, the periodic change law and trend change law are dynamically time-warped with the timestamp event sequence in the semantic coding information. Dynamic time warping is used to measure the similarity between two time series. Specifically, the regularization method can be implemented by a dynamic programming algorithm to eliminate the phase deviation in the time series, so that the hierarchical feature information and the semantic coding information are aligned in time. Then, the future trends of the hierarchical features and semantic coding are predicted respectively through time series analysis, and the prediction results are recorded as prediction trajectories and semantic evolution paths respectively. The prediction trajectory can reflect the dynamic change law that objects in the target scene may present over time, and the semantic evolution path reflects the possible development trend of events or behaviors in the text data over time. In order to verify the accuracy of the prediction results, Through cross-validation methods, such as K-fold cross-validation, the consistency of the predicted trajectory and the semantic evolution path is evaluated, and the verification deviation is output simultaneously. The verification deviation reflects the degree of difference between the predicted trajectory and the semantic evolution path. When the verification deviation exceeds the preset deviation threshold, it indicates that there is inconsistency in the trend of hierarchical features and semantic encoding in the time dimension, which may be caused by data noise, mismatching or other external interference factors. At this time, the abnormal warning mechanism will be triggered to remind relevant personnel to conduct further inspection and analysis of the inconsistent areas to ensure the accuracy and reliability of the data. When the verification deviation does not exceed the preset deviation threshold, it indicates that the trend of hierarchical features and semantic encoding in the time dimension is consistent. At this time, the hierarchical feature information and semantic encoding information will be fused to generate joint representation data. The joint representation data integrates the multi-dimensional information of images and texts, including visual features such as object contours, textures, and motion trajectories, and incorporates semantic information in text data, which can provide corresponding data support for subsequent anomaly detection, trend prediction and event recognition.

[0025] S5. Dynamically modify the evaluation boundary parameters based on historical joint characterization data within a preset time period and in combination with environmental feedback parameters, and evaluate the current joint characterization data based on the modified evaluation boundary parameters to determine the public safety status in the target scenario; In step S5, after the joint characterization data is determined, the evaluation boundary parameters are dynamically corrected based on the historical joint characterization data within a preset time period in combination with the environmental feedback parameters. The evaluation boundary parameters are used to define the boundary between normal behavior and abnormal behavior. The current joint characterization data is evaluated based on the corrected evaluation boundary parameters to determine the public safety status in the target scenario. When abnormal behavior or potential security threats are discovered, an alarm signal can be issued in a timely manner to facilitate the adoption of corresponding preventive or response measures. The step of dynamically correcting the evaluation boundary parameters based on the historical joint characterization data within a preset time period in combination with the environmental feedback parameters includes: Convert environmental feedback parameters and historical joint representation data into the same dimension and align them along the time axis; Calculate the initial value of the evaluation boundary parameter based on the distribution characteristics of the historical joint representation data and the changing trend of the environmental feedback parameters; Collect the dynamic changes of the current environmental feedback parameters in real time and record them as correction condition parameters; When the correction condition parameters are within the allowable fluctuation range, the original assessment boundary parameters are maintained unchanged, and the dynamic changes of the environmental feedback parameters continue to be monitored; If it exceeds the allowable fluctuation range, the correction condition parameters are substituted into the boundary adjustment function, the new evaluation boundary parameters are calculated, and the original evaluation boundary parameters are replaced.

[0026] Specifically, when evaluating the joint characterization data, the environmental feedback parameters and the historical joint characterization data will first be preprocessed to ensure that the environmental feedback parameters and the historical joint characterization data are in the same dimension and aligned along the time axis. The environmental feedback parameters include temperature, humidity, light intensity, etc. Changes in environmental feedback parameters may directly affect objects and events in the target scene. For example, in a high temperature environment, monitoring equipment may cause image quality degradation due to overheating. Therefore, it is necessary to align the environmental feedback parameters with the historical joint characterization data in order to accurately evaluate the public safety status in the target scene. Specifically, the initial value of the evaluation boundary parameter will be calculated based on the distribution characteristics of the historical joint characterization data and the changing trend of the environmental feedback parameters. The initial value = the mean of the historical joint characterization data + the environmental parameter variation coefficient × the environmental parameter variation. Among them, the historical data mean reflects the average level of the historical joint characterization data within the preset time period. After obtaining the initial value, it will be collected in real time. The dynamic change value of the current environmental feedback parameter is recorded as the correction condition parameter. Specifically, the dynamic change value can be determined by the sliding window mechanism, that is, the monitoring period is divided into multiple time windows, and the mean of the historical environmental feedback parameter is calculated in each time window. The mean value closest to the current node is selected and the difference is processed with the current environmental feedback parameter to output the correction condition parameter. When the correction condition parameter is within the allowable fluctuation range, it means that the change of the current environment has not exceeded expectations. At this time, the original evaluation boundary parameter will be maintained unchanged, and the dynamic changes of the environmental feedback parameter will continue to be monitored to ensure the accuracy and stability of the evaluation. However, if the correction condition parameter exceeds the allowable fluctuation range, it means that the change of the current environment has exceeded expectations. At this time, the correction condition parameter will be substituted into the boundary adjustment function to calculate the new evaluation boundary parameter and replace the original evaluation boundary parameter to adapt to the change of the environment and ensure the accuracy and effectiveness of the evaluation. The expression of the boundary adjustment function is: , where represents the new evaluation boundary parameter, represents the original evaluation boundary parameter, represents the control coefficient, which is the ratio of the standard deviation of the historical joint representation data to the mean. Indicates the correction condition parameter, It represents the mean of the historical environmental feedback parameters, and finally evaluates the current joint representation data according to the revised evaluation boundary parameters to determine the public safety status under the target scenario. When abnormal behavior or potential security threats are found, an alarm signal can be issued in time to facilitate the adoption of corresponding preventive or response measures to ensure public safety.

[0027] In a preferred embodiment, the step of evaluating the current joint representation data based on the modified evaluation boundary parameters to determine the public safety status in the target scenario includes: Setting safety level thresholds based on assessment boundary parameters; Compare the current joint representation data with the security level threshold and output the current public safety status level, where the public safety status level is divided into three levels: safe, risky, and emergency; When the public safety status level is risk or emergency, the corresponding early warning mechanism is triggered and a safety analysis report is generated. The early warning level in the emergency state is higher than that in the risk state. When the public security status level is safe, continuous monitoring is maintained.

[0028] In this embodiment, after the evaluation boundary parameters are corrected, the current joint characterization data will be divided into security levels. First, a security level threshold will be set based on the evaluation boundary parameters. The security level threshold is used to define the standards for different public safety states. By comparing the current joint characterization data with the security level threshold, the current public safety state level can be output. The public safety state level is divided into three levels, including safety, risk, and emergency, to reflect the different levels of public safety in the target scenario. When the public safety state level is risk or emergency, it indicates that there may be abnormal behavior or potential security threats in the target scenario. At this time, the corresponding early warning mechanism will be triggered and a safety analysis report will be generated, including the time, location, type, etc. of the occurrence, so that relevant personnel can take preventive or response measures in a timely manner. Among them, the early warning level in the emergency state is higher than the risk state. The emergency state indicates that there may be serious security threats in the target scene and immediate action is needed to avoid accidents. When the public safety state level is safe, it means that the public safety state in the target scene is at a normal level. At this time, continuous monitoring will be maintained to ensure the continued stability of public safety.

[0029] See also Figure 2 , an intelligent monitoring system based on vision and language fusion, using the above-mentioned intelligent monitoring method based on vision and language fusion, including: A data acquisition module is used to collect spatiotemporal sequence images of the target scene through a visual sensor, and simultaneously collect text data associated with the target scene; Image feature extraction module, used to perform multi-dimensional spatial analysis on spatiotemporal sequence images and extract hierarchical feature information including temporal variation patterns; The text feature extraction module is used to perform context-related analysis on text data and generate semantically coherent semantic coding information; The feature fusion module is used to establish a multi-level interactive channel between hierarchical feature information and semantic encoding information, perform local feature matching in the spatial dimension, and perform trend consistency calibration in the temporal dimension to generate joint representation data; The status assessment module is used to dynamically correct the assessment boundary parameters based on the historical joint characterization data within a preset time period and the environmental feedback parameters, and to evaluate the current joint characterization data based on the corrected assessment boundary parameters to determine the public safety status in the target scenario.

[0030] In the above, the data acquisition module is responsible for capturing the dynamic changes of the target scene in real time through the visual sensor. At the same time, the text data acquisition unit synchronously collects text data related to the target scene to provide a basis for subsequent feature extraction and fusion. The image feature extraction module performs multi-dimensional spatial analysis on the collected spatiotemporal sequence images, and extracts hierarchical feature information including time change rules. The hierarchical feature information reflects the visual features such as the contour, texture, and motion trajectory of the objects in the target scene. The text feature extraction module uses natural language processing technology to perform context-related analysis on the text data and generate semantic coding information with semantic coherence. The semantic coding information reveals the semantic connotation of events or behaviors in the text data. The feature fusion module acts as a negative It is responsible for establishing a multi-level interactive channel for hierarchical feature information and semantic coding information. First, local feature matching is performed in the spatial dimension. Through spatial coordinate mapping and semantic vector distance calculation, accurate matching of local image features and text entity descriptions is achieved. Then, trend consistency calibration is performed in the time dimension. Through time series analysis and cross-validation, the consistency of hierarchical features and semantic coding in future trends is ensured. Finally, joint representation data that integrates multi-dimensional information of images and texts is generated. The status assessment module dynamically corrects the assessment boundary parameters based on the historical joint representation data within a preset time period and the real-time collected environmental feedback parameters, and evaluates the current joint representation data according to the corrected assessment boundary parameters, thereby determining the public safety status in the target scenario.

[0031] See also Figure 3 , an electronic device, the electronic device comprising: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by at least one processor so that the at least one processor can execute the above-mentioned intelligent monitoring method based on vision and language fusion.

[0032] The processor of the above-mentioned electronic device can be a central processing unit (CPU), a graphics processing unit (GPU) or a digital signal processor (DSP), etc., and the memory can be a non-volatile memory, such as a solid-state drive (SSD) or flash memory, which is used to store an operating system, application programs and the above-mentioned computer program of the intelligent monitoring method based on vision and language fusion. When the computer program is executed by the processor, it is used to implement the above-mentioned intelligent monitoring method. In addition, the electronic device can also include an operator, an input device and an output device, etc. The operator can be an arithmetic logic unit (ALU) for performing various arithmetic and logical operations, an input device such as a keyboard, a mouse or a touch screen for receiving user operation instructions, and an output device such as a display or a printer for displaying processing results or printing output.

[0033] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.

[0034] The foregoing is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained herein shall, unless otherwise specified or limited, be implemented in accordance with conventional means in the art.

Claims

1. An intelligent monitoring method based on vision and language fusion, characterized by: include: The visual sensor collects spatiotemporal sequence images of the target scene and simultaneously collects text data associated with the target scene; Perform multi-dimensional spatial analysis on spatiotemporal sequence images to extract hierarchical feature information including temporal variation patterns; Perform context-sensitive analysis on text data to generate semantically coherent semantically encoded information; Establish a multi-level interactive channel between hierarchical feature information and semantic encoding information, perform local feature matching in the spatial dimension, and perform trend consistency calibration in the temporal dimension to generate joint representation data; Based on the historical joint characterization data within a preset time period, the evaluation boundary parameters are dynamically corrected in combination with the environmental feedback parameters, and the current joint characterization data is evaluated according to the corrected evaluation boundary parameters to determine the public safety status in the target scenario.

2. The intelligent monitoring method based on vision and language fusion according to claim 1, characterized in that: When collecting spatiotemporal sequence images and text data, the following steps are included: The pre-deployed multispectral sensor array synchronously acquires three-dimensional spatiotemporal sequence image data of visible light band, infrared band and depth information; Establish a text collection thread pool to capture monitoring logs, social media text, and speech transcription data in real time through a natural language interface, and attach timestamp tags that match the spatiotemporal sequence images; The spatiotemporal sequence images and text data are spatiotemporally calibrated, and the buffers of the spatiotemporal sequence images and text data are aligned using a sliding window mechanism, and the spatiotemporal deviation is eliminated through feature point matching.

3. The intelligent monitoring method based on vision and language fusion according to claim 1, characterized in that: The step of performing multi-dimensional spatial analysis on the spatiotemporal sequence images to extract hierarchical feature information including temporal variation patterns includes: Perform multi-scale spatial decomposition on spatiotemporal sequence images, extract feature sets including object contours, surface textures, and motion trajectories at different resolution levels, and form a multi-level feature structure; Construct a spatiotemporal correlation matrix of multi-level feature structures, and perform cross-level fusion processing on object contours, surface textures, and motion trajectories to generate a spatiotemporal feature map; Perform time series analysis on multi-level feature structures to identify periodic change features, trend change features, and sudden abnormal features; Based on the periodic change characteristics and trend change characteristics, combined with the spatiotemporal feature map, topological extension is performed to generate a regular hierarchical feature information database; Based on the sudden abnormality features, similar abnormal patterns are searched in the historical feature library, and adjacent feature fragments of the sudden abnormality features under similar abnormality patterns are collected and summarized into a hierarchical feature information library for sudden abnormality reference.

4. The intelligent monitoring method based on vision and language fusion according to claim 1, characterized in that: The step of performing context-related parsing on text data to generate semantically coherent semantically encoded information includes: Based on the preset encoding mechanism of the text data context, the text content in the text data is converted into a semantic vector sequence with time attributes; The corresponding entity objects in the semantic vector sequence are associated and mapped with the entities with the same name in the spatiotemporal sequence images to build a bidirectional matching relationship between entities and semantics; Collect multiple semantic vector sequence fragments adjacent to the current timestamp, build a semantic context network through time series analysis, and then combine them with events in the historical semantic library for semantic enhancement processing to form a dynamically updated semantic coding information library.

5. The intelligent monitoring method based on vision and language fusion according to claim 1, characterized in that: The step of establishing a multi-level interactive channel of hierarchical feature information and semantic coding information and performing local feature matching in the spatial dimension includes: Construct a hierarchical structure of multi-level interactive channels and map hierarchical feature information and semantic encoding information layer by layer, including: At the low-level semantic layer, it is the entity mapping from object contours to text semantics; At the middle semantic layer, there is the semantic association from surface texture and motion trajectory to event description; At the high-level semantic layer, it is the fusion of spatiotemporal feature maps to scene semantics; Calculate the matching degree between local features of the image and text entity descriptions, and achieve the connection between multi-level features and semantic information through spatial coordinate mapping and semantic vector distance; Among them, when the matching degree between the local features of the image and the text entity description is lower than the preset matching threshold, the corresponding local features are marked as abnormal areas, and the abnormal warning mechanism is triggered simultaneously. Then, a secondary verification is performed in combination with the contextual relationship in the semantic coding information library. In the secondary verification, if the matching degree is still lower than the matching threshold, the abnormal area is marked as high risk and an abnormality report is generated.

6. The intelligent monitoring method based on vision and language fusion according to claim 1, characterized in that: The step of performing trend consistency calibration in the time dimension to generate joint representation data includes: Extract the periodic and trend change patterns in the hierarchical feature information, and perform dynamic time warping with the timestamp event sequence in the semantic coding information to eliminate the phase deviation in the timing; Through time series analysis, the future trends of hierarchical features and semantic encoding are predicted and recorded as prediction trajectories and semantic evolution paths respectively; The consistency between the predicted trajectory and the semantic evolution path is verified through cross-validation, and the verification deviation is output simultaneously; When the verification deviation exceeds the preset deviation threshold, it indicates that there is inconsistency between the hierarchical features and the semantic encoding, triggering the abnormal warning mechanism; When the verification deviation does not exceed the preset deviation threshold, it indicates that the hierarchical features and semantic coding are consistent, and the hierarchical feature information and semantic coding information are fused to generate joint representation data.

7. The intelligent monitoring method based on vision and language fusion according to claim 1, characterized in that: The step of dynamically correcting the evaluation boundary parameters based on the historical joint characterization data within a preset time period in combination with the environmental feedback parameters includes: Convert environmental feedback parameters and historical joint representation data into the same dimension and align them along the time axis; Calculate the initial value of the evaluation boundary parameter based on the distribution characteristics of the historical joint representation data and the changing trend of the environmental feedback parameters; Collect the dynamic changes of the current environmental feedback parameters in real time and record them as correction condition parameters; When the correction condition parameters are within the allowable fluctuation range, the original assessment boundary parameters are maintained unchanged, and the dynamic changes of the environmental feedback parameters continue to be monitored; If it exceeds the allowable fluctuation range, the correction condition parameters are substituted into the boundary adjustment function, the new evaluation boundary parameters are calculated, and the original evaluation boundary parameters are replaced.

8. The intelligent monitoring method based on vision and language fusion according to claim 1, characterized in that: The step of evaluating the current joint representation data according to the revised evaluation boundary parameters to determine the public safety status in the target scenario includes: Setting safety level thresholds based on assessment boundary parameters; Compare the current joint representation data with the security level threshold and output the current public safety status level, where the public safety status level is divided into three levels: safe, risky, and emergency; When the public safety status level is risk or emergency, the corresponding early warning mechanism is triggered and a safety analysis report is generated. The early warning level in the emergency state is higher than that in the risk state. When the public security status level is safe, continuous monitoring is maintained.

9. An intelligent monitoring system based on the fusion of vision and language, characterized by: The intelligent monitoring method based on vision and language fusion according to any one of claims 1 to 8 comprises: A data acquisition module is used to collect spatiotemporal sequence images of the target scene through a visual sensor, and simultaneously collect text data associated with the target scene; Image feature extraction module, used to perform multi-dimensional spatial analysis on spatiotemporal sequence images and extract hierarchical feature information including temporal variation patterns; The text feature extraction module is used to perform context-related analysis on text data and generate semantically coherent semantic coding information; The feature fusion module is used to establish a multi-level interactive channel between hierarchical feature information and semantic encoding information, perform local feature matching in the spatial dimension, and perform trend consistency calibration in the temporal dimension to generate joint representation data; The status assessment module is used to dynamically correct the assessment boundary parameters based on the historical joint characterization data within a preset time period and the environmental feedback parameters, and to evaluate the current joint characterization data based on the corrected assessment boundary parameters to determine the public safety status in the target scenario.

10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the intelligent monitoring method based on vision and language fusion as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Trusted industrial agent task distribution method and system

    CN120893794A

  • Digital twin system for machine room equipment state perception and prediction

    CN121116699A

  • Identification information content intelligent detection method based on visual ergonomics

    CN121281035A

  • Visible light communication system

    CN121356685A

  • Intelligent fire-fighting emergency management method and system

    CN121660395A