Safety production detection data processing method and system based on digitization

CN119989244AActive Publication Date: 2025-05-13GUIZHOU NEW THINKING TECH CO LTD

Patent Information

Application Number
CN202510473145.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

Smart Images

  • Figure CN119989244A_ABST
    Figure CN119989244A_ABST
Patent Text Reader

Abstract

The invention provides a safety production detection data processing method and system based on digitization, and the method comprises the steps: obtaining safety detection data in a target scene, and receiving a detection instruction for the safety detection data, the safety detection data comprising a continuous time sequence data stream collected by a multi-dimensional sensor; performing semantic analysis on the detection instruction to generate a detection instruction feature containing a task type and an entity distribution relationship; performing cross-modal fusion processing on the security detection data through a multi-modal feature extraction model of joint training, and generating a security feature vector of each time sequence node; based on a dynamic matching result of the detection instruction feature and the security feature vector of each time sequence node, screening out a key data segment from the time sequence data stream; and calling a pre-trained security decision model to carry out abnormal mode identification and risk assessment on the key data segment, and generating a structured security report corresponding to the detection instruction. According to the invention, the accuracy and reliability of safety detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a method and system for processing safe production detection data based on digitization. Background Art

[0002] In modern safety production in various fields, there are many urgent problems to be solved in the processing of safety production detection data. With the widespread application of advanced sensors and monitoring equipment, the sources of safety detection data are complex and diverse. The multi-dimensional data formats, sampling frequencies and meanings collected by different sensors are different, and data integration and processing are difficult. Due to the large amount of safety detection data and a large amount of redundant information, the existing data screening methods lack effective use of instruction intent and data timing characteristics, and the accuracy and reliability of the screening results are low. In addition, for the key data segments screened out, the existing abnormal pattern recognition and risk assessment methods mostly rely on a single feature or model, which makes it difficult to fully and accurately capture complex information, and easily causes potential safety hazards to be missed or misjudged. In summary, the detection accuracy of existing safety production detection data processing technologies needs to be improved. Summary of the invention

[0003] The purpose of the present invention is to provide a digital-based method for processing production safety detection data.

[0004] The embodiment of the present invention is implemented as follows: In the first aspect, an embodiment of the present invention provides a method for processing safety production detection data based on digitization, comprising: acquiring safety detection data in a target scenario, and receiving detection instructions for the safety detection data, wherein the safety detection data includes a continuous time series data stream collected by a multi-dimensional sensor; performing semantic parsing on the detection instructions to generate detection instruction features including a task type and an entity distribution relationship; performing cross-modal fusion processing on the safety detection data through a jointly trained multimodal feature extraction model to generate a safety feature vector for each time series node; based on the dynamic matching results of the detection instruction features and the safety feature vectors of each time series node, screening out key data segments from the time series data stream; calling a pre-trained safety decision model to perform abnormal pattern recognition and risk assessment on the key data segments, and generating a structured safety report corresponding to the detection instruction.

[0005] In a second aspect, the present invention provides a computer system comprising: one or more processors; a memory; one or more computer programs; wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method as described above is implemented.

[0006] The technical effects of the present invention include at least: the present invention obtains safety detection data from the target scene, which contains rich information related to safe production. Then, the detection instruction is semantically parsed and converted into detection instruction features containing the relationship between task type and entity distribution, providing clear goals and directions for subsequent data screening and analysis. Then, the safety detection data is cross-modally fused using a jointly trained multimodal feature extraction model to mine the association and complementary information between different modal data, and generate a safety feature vector for each time series node to comprehensively and accurately represent the safety production status. Afterwards, based on the dynamic matching results of the detection instruction features and the safety feature vectors, key data segments are screened out from the time series data stream, focusing on important data related to the detection instructions. Finally, the pre-trained safety decision model is called to perform abnormal pattern recognition and risk assessment on key data segments, generate a structured safety report, and provide a scientific basis for safety production management.

[0007] Through the multimodal feature extraction model, the safety detection data is cross-modally fused, and the visual modal data, acoustic modal data and equipment status modal data are integrated to fully explore the potential information between different modal data. Through cross-modal fusion processing, a more comprehensive and accurate safety feature vector can be generated, which solves the problem of insufficient information in a single modal data and improves the perception of the safety production status. Through semantic analysis of the detection instructions, the detection instruction features containing the relationship between the task type and the entity distribution are generated. The specific steps include syntax tree parsing, semantic expansion, intent classification and feature fusion, which can deeply mine the semantic information of the detection instructions, determine the task type priority, and generate an optimized detection instruction feature vector. Based on the dynamic matching results of the detection instruction features and the security feature vectors of each time series node, the key data segments are screened out from the time series data stream. Through the steps of calculating similarity scores, time window sliding analysis, length weight adjustment, context dependency analysis and density clustering, the task type priority of the detection instruction and the time series characteristics of the data can be comprehensively considered to screen out the most relevant key data segments. The pre-trained safety decision model is called to perform abnormal pattern recognition and risk assessment on the key data segments to generate a structured safety report.

[0008] The present invention effectively solves key problems such as the complexity and diversity of data processing, the understanding and execution of detection instructions, the accuracy and efficiency of data screening, and abnormal pattern recognition and risk assessment. It improves the accuracy and efficiency of production safety detection, timely discovers potential safety hazards, provides strong support for production safety management, and ensures the safe and stable operation of the production process. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for describing the embodiments of the present invention are briefly introduced below.

[0010] Figure 1 The present invention provides a flowchart of a method for processing production safety detection data based on digitization.

[0011] Figure 2 It is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0012] The execution subject of the digital production safety detection data processing method in the embodiment of the present invention is a computer system, including but not limited to a server, a personal computer, a laptop, a tablet computer, a smart phone, etc. Among them, the computer system can be operated alone to implement the present invention, and can also be connected to a network and implement the present invention through interactive operations with other computer systems in the network. Among them, the network where the computer system is located includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, etc.

[0013] The embodiment of the present invention provides a method for processing safety production detection data based on digitization, such as Figure 1 As shown, the method includes: Step 100: Acquire safety detection data in a target scenario and receive detection instructions for the safety detection data, where the safety detection data includes a continuous time series data stream collected by a multi-dimensional sensor.

[0014] In actual operation, different target scenes may be faced, such as industrial production workshops, intelligent building environments, transportation hubs, etc., and the present invention does not limit the target scenes. The embodiments of the present invention can collect safety detection data with the help of various multi-dimensional sensors, which are distributed in various key positions of the workshop to obtain relevant information comprehensively and in real time. For example, in terms of vision, a high-definition camera can be used to continuously monitor the operation of personnel, the operation status of equipment, and the stacking of materials in the workshop, and the collected video data is the safety detection data of the visual modality. In terms of acoustics, a microphone sensor can be used to collect sound information in the workshop. Under normal circumstances, the operation of the machine in the workshop will have a specific sound frequency and intensity range. If there is an abnormal sharp noise or irregular sound fluctuation, it may be that the equipment is faulty. In terms of equipment status, various sensors can be used to monitor the temperature, pressure, current, voltage and other parameters of the equipment in real time. These data collected by the multi-dimensional sensors are continuous time-series data streams, that is, the data is continuously generated over time, and each data point has a corresponding timestamp to reflect the order in which the data is generated. The data collected by the sensor can be transmitted to the data processing center through a wired or wireless network for storage and further analysis.

[0015] Step 200: semantically parse the detection instruction to generate detection instruction features including the relationship between task type and entity distribution.

[0016] For the received detection instructions, it is necessary to clarify the type of task required by the instructions and the scope of detection, so as to process the safety detection data in a targeted manner in the future. Taking an industrial production scenario as an example, assuming that the received detection instruction is "check whether the temperature and pressure of all production equipment in workshop A in the past 24 hours are abnormal", it is necessary to conduct in-depth semantic analysis of this instruction. Specifically, natural language processing technology can be used to parse the detection instructions. At the grammatical level, analyze the various words in the instructions and the grammatical relationships between them to identify key entity objects and operation predicates. In the above example, the entity objects are "workshop A" and "production equipment", and the operation predicates are "check" and "whether abnormal". Through this grammatical analysis, the basic structure and key elements of the instructions can be preliminarily grasped. Then, the entity objects are semantically expanded. For example, with the help of a pre-trained domain knowledge graph, which contains various knowledge and information in the field of industrial production, entity objects such as "workshop A" and "production equipment" are associated and expanded with relevant knowledge in the graph. For example, the specific layout of workshop A, the models and specifications of production equipment, and other information can be obtained from the knowledge graph to form a more comprehensive entity association network, which helps the system understand the entity distribution relationship more accurately.

[0017] For the operation predicate, intent classification can be performed to determine the task type priority of the detection instruction. In this example, "check" indicates the core intention of the operation. According to the preset rules and models, it is judged that this is a safety detection task, and the urgency and importance of the task in the entire safety production process are further analyzed. Through intent classification, different task types can be assigned corresponding priorities to ensure that high-priority tasks can be processed first when processing data.

[0018] After completing the semantic extension of entity objects and the intent classification of operation predicates, relevant information can be fused to generate multi-dimensional detection instruction features. These features include entity distribution vectors and task weight matrices. The entity distribution vector reflects the distribution of entity objects within the entire entity distribution relationship, and the task weight matrix reflects the priority and importance of different task types. For example, in this example, the entity distribution vector may show the distribution of production equipment in different areas of workshop A, and the task weight matrix will highlight the high priority of the "temperature and pressure anomaly detection" task. Finally, the multi-dimensional detection instruction features can be dynamically weighted through the attention mechanism. The attention mechanism automatically adjusts the weights of different features according to the context and specific requirements of the instructions, so that the system pays more attention to important information.

[0019] Step 300: Perform cross-modal fusion processing on the security detection data through the jointly trained multi-modal feature extraction model to generate a security feature vector for each time series node.

[0020] Safety detection data is collected by multi-dimensional sensors and has multiple modalities, such as vision, acoustics, and equipment status. These different modal data are fused to extract more comprehensive and accurate safety features. First, the safety detection data is modally split into visual modal data, acoustic modal data, and equipment status modal data. For visual modal data, a special visual encoder can be called for processing. The visual encoder can extract spatial features from the video screen. For acoustic modal data, an acoustic encoder can be used for spectral decomposition and abnormal waveform detection. The acoustic encoder can analyze the frequency, intensity, and waveform of the sound to determine whether there is an abnormal sound signal. The state encoder can determine whether the operating status of the device is normal based on parameters such as the temperature, pressure, and current of the device. For example, for a motor, the state encoder can determine whether the motor is overloaded or has faults such as short circuits based on the change in current, and generate a device operating state vector. After completing the feature extraction of different modal data, these features are aligned and fused through a cross-modal attention network. The cross-modal attention network can automatically adjust the weights of different modal features so that the system pays more attention to important information. For example, in some cases, visual features may be more able to reflect the appearance of damage to the device, while acoustic features may be more able to detect internal faults in the device. The cross-modal attention network can assign appropriate weights to features of different modalities according to the specific situation, fuse visual features, acoustic energy distribution features, and device operation state vectors, and generate a safety feature vector for each time series node. It integrates information from multiple modalities and can more comprehensively and accurately reflect the safety status of the target scene at each time point, providing strong support for subsequent data analysis and decision-making. Deep learning algorithms can be used to jointly train multi-modal feature extraction models, and the performance and accuracy of the models can be improved by continuously adjusting the parameters of the models.

[0021] Step 400: based on the dynamic matching result between the detection instruction feature and the security feature vector of each timing node, the key data segment is screened out from the timing data stream.

[0022] Specifically, the detection instruction feature vector can be dynamically matched with the security feature vector of each timing node, and the similarity score between them can be calculated. The higher the similarity score, the stronger the correlation between the data of the timing node and the detection instruction. Then, the initial matching sequence can be analyzed by sliding the time window. The time window is a time period of fixed length. The window is slid on the timing data stream to check the similarity scores of multiple consecutive timing nodes in the window. If the similarity scores in a certain window exceed the dynamic threshold, it means that the data in this time period is highly correlated with the detection instruction, and the data in this time period will be identified as a candidate data segment. Then, the length weight of the candidate data segment can be adjusted according to the task type priority in the detection instruction feature. If the task that is focused on in the detection instruction is of high priority, the length weight of the corresponding candidate data segment will increase; conversely, the length weight of the candidate data segment corresponding to the low priority task will decrease. For example, if the high priority task in the detection instruction is to detect sudden failures of the equipment, the length weight of the candidate data segment related to the task will be increased, which can ensure that key data is not missed.

[0023] After that, the optimized candidate data segment set can be input into the pre-trained time series association model to analyze the contextual dependencies of each candidate data segment in the time series data stream. For those isolated candidate data segments whose correlation with the previous and next time series nodes is less than the correlation threshold, they can be eliminated because these isolated data may be interference information. At the same time, for candidate data segments with time overlap, they can be merged to make the data more compact and meaningful. Finally, density clustering can be performed according to the distribution density of each data segment in the time series data stream in the association optimization data segment set. If the distribution density of a data segment is higher than the preset clustering threshold, it is screened out as the final key data segment set. The key data segment contains the preset minimum number of continuous time series nodes and covers the entity object set specified in the detection instruction feature, which can provide accurate and effective data support for subsequent abnormal pattern recognition and risk assessment. Machine learning algorithms can be used to implement operations such as similarity calculation, time window sliding analysis, time series association analysis, and density clustering to ensure the accuracy and efficiency of key data segment screening.

[0024] Step 500: Call the pre-trained security decision model to perform abnormal pattern recognition and risk assessment on the key data segment, and generate a structured security report corresponding to the detection instruction.

[0025] In step 500, the pre-trained safety decision model is called to perform abnormal pattern recognition and risk assessment on the key data segment, and a structured safety report corresponding to the detection instruction is generated. In the safety production detection scenario, after the key data segments are screened out in the previous steps, the safety information hidden behind these data is further mined to ensure the safety of production activities.

[0026] For example, multi-granularity slicing is performed on key data segments, dividing them into data blocks of different levels, each of which contains global state features and local abnormal features. Next, the hierarchical data blocks are input into the convolutional recursive network of the safety decision model for spatiotemporal feature extraction. Then, the pattern matching layer in the safety decision model will compare the extracted spatiotemporal features with the preset abnormal pattern library for similarity. The abnormal pattern library stores various known safety abnormalities and their characteristics. By comparing, it can be determined whether the current data has similarities with the known abnormal patterns. For example, if the temperature and pressure change patterns of the reactor are similar to the "reaction out of control" pattern in the abnormal pattern library, then there may be a risk of reaction out of control.

[0027] Based on the comparison results, risk assessment indicators including risk level, abnormality type and impact range can be generated. Risk level can be divided into high, medium and low levels, abnormality type can be specific to equipment failure, operating error, etc., and the impact range can be clearly defined as whether it only affects the reactor itself or affects surrounding equipment or personnel.

[0028] Finally, according to the risk assessment indicators and the task types in the detection instruction characteristics, a structured safety report containing visual charts and disposal suggestions is generated. Visual charts can intuitively display the risk situation, such as using line charts to show the changing trends of temperature and pressure, and using bar charts to compare the risk levels in different time periods. Disposal suggestions give specific response measures based on the risk assessment results. For example, for the risk of reaction runaway, it is recommended to immediately stop feeding and start the emergency cooling system. Deep learning algorithms can be used to train safety decision-making models, image processing technology can be used to generate visual charts, and disposal suggestions can be formulated in combination with expert knowledge in the field of safe production, thereby providing a scientific and effective decision-making basis for safe production management.

[0029] As an implementation mode, step 200, semantically parsing the detection instruction to generate a detection instruction feature including a task type and an entity distribution relationship, includes: Step 210: parsing the syntax tree of the detection instruction to extract the entity object set and the operation predicate set in the instruction; Step 220: Input the entity object set into the pre-trained domain knowledge graph for semantic expansion to generate an expanded entity association network; Step 230: classify the intent of the operation predicate set and determine the task type priority of the detection instruction; Step 240: Fusion of the topological features of the entity association network and the task type priority to generate a multi-dimensional detection instruction feature, wherein the detection instruction feature includes an entity distribution vector and a task weight matrix; Step 250: Dynamically weight the multi-dimensional detection instruction features through the attention mechanism to generate an optimized detection instruction feature vector.

[0030] In step 210, the syntax tree of the detection instruction is parsed to extract the entity object set and the operation predicate set in the instruction. Taking the detection instruction of an industrial production scenario "Check whether the temperature and pressure of all production equipment in workshop A are normal" as an example, the sentence can be grammatically analyzed to identify the entity objects and operation predicates therein. The entity object refers to the specific thing or concept involved in the instruction. In this example, "workshop A" and "production equipment" are entity objects; the operation predicate indicates the action required by the instruction. Here, "check" and "whether normal" reflect the core intention of the operation. Natural language processing toolkits such as NLTK (Natural Language Toolkit) or SpaCy can be used to extract entity objects and operation predicates.

[0031] In step 220, the entity object set is input into the pre-trained domain knowledge graph for semantic expansion to generate an expanded entity association network. In the field of industrial production, the knowledge graph may contain information such as equipment models, specifications, performance parameters, workshop layout, safety regulations, etc. A graph database, such as Neo4j, can be used to store and manage the domain knowledge graph, and the entity information can be retrieved and expanded through a graph query language.

[0032] In step 230, the set of operation predicates is classified for intent, and the priority of the task type of the detection instruction is determined. The operation predicate reflects the core intent of the detection instruction, but different operation predicates may have different priorities and importance. These operation predicates can be classified and prioritized so that high-priority tasks can be prioritized in subsequent data processing. Taking the instruction "check whether the temperature and pressure of all production equipment in workshop A are normal" as an example, the intent of the operation predicate "check" is to perform status detection. It is necessary to judge the urgency and importance of this detection task in the entire safe production process according to the preset rules and models. If the production equipment in workshop A is performing critical production tasks, then the abnormal temperature and pressure may cause serious production accidents, and the priority of this detection task will be higher. Machine learning algorithms, such as support vector machines (SVM) or deep learning models, can be used to classify the intent of operation predicates and set priority rules for task types based on domain knowledge and experience.

[0033] In step 240, the topological features of the entity association network and the task type priority are integrated to generate a multi-dimensional detection instruction feature, wherein the detection instruction feature includes an entity distribution vector and a task weight matrix. The topological features of the entity association network reflect the relationship and structure between entity objects, while the task type priority reflects the importance of different tasks. The two are combined to generate a multi-dimensional vector and matrix that can fully reflect the detection instruction feature. In the above example, the entity distribution vector can represent the distribution of production equipment in workshop A, such as which equipment is located in the key position of the workshop and which equipment is closely related to other equipment. The task weight matrix can assign corresponding weights to different detection tasks according to the task type priority. If the priority of the temperature detection task is higher, then in the task weight matrix, the weight corresponding to the temperature detection will be larger. The vector space model and matrix operation can be used to realize the generation of the entity distribution vector and the task weight matrix, and this information is integrated into the detection instruction feature by calculating the similarity between entities and the importance of tasks.

[0034] In step 250, the multi-dimensional detection instruction features are dynamically weighted through the attention mechanism to generate an optimized detection instruction feature vector. The attention mechanism is a technology that can automatically adjust the weights of different features. It can make the system pay more attention to important information according to the context and specific requirements of the instructions. In production safety detection, different detection tasks and entity objects may have different importance in different situations. For example, in some cases, the temperature of the equipment may be more critical than the pressure. At this time, the attention mechanism can automatically give higher weights to temperature-related features. By dynamically weighting the multi-dimensional detection instruction features, a more optimized detection instruction feature vector can be generated. This vector can more accurately reflect the core requirements of the detection instruction and provide more precise guidance for subsequent data screening and analysis. A deep learning-based attention model, such as the multi-head attention mechanism in Transformer, can be used to achieve dynamic weighting of detection instruction features.

[0035] Steps 210-250 provide a complete set of methods and processes for semantic analysis of detection instructions. Through syntax analysis, entity expansion, intent classification, feature fusion and dynamic weighting of instructions, the meaning of detection instructions can be accurately understood, and detection instruction features containing the relationship between task type and entity distribution can be generated, providing strong support for data processing and analysis of safety production detection.

[0036] As an implementation method, step 230, classifying the operation predicate set to determine the task type priority of the detection instruction, includes: Step 231: Standardize the operation predicate set, remove redundant stop words and fuzzy expressions, and generate a standardized operation predicate set; Step 232: Call the pre-trained intent recognition model to perform multi-granular semantic analysis on the standardized operation predicate set, extract the logical dependency relationship and action triggering conditions between the operation predicates, and generate an initial intent label set; Step 233: semantically verify the initial intention label set based on the preset safety production domain knowledge base, verify the matching degree of each label in the initial intention label set with the standard operation scenario in the domain knowledge base, and generate the confidence of the verified intention label; Step 234: dynamically assign weights to the initial intent tag set according to the verified intent tag confidences, wherein intent tags that match the standard operation scenario with a degree higher than a preset verification threshold are assigned high confidence weights, and a weighted intent tag set is generated; Step 235: input the weighted intention label set into a predefined task priority rule base, and map each intention label to a corresponding priority weight according to the mapping relationship between the security response level and the operation urgency defined in the rule base; Step 236: Combine the context features of the detection instruction, extract the entity object set and operation timeliness constraint conditions involved in the detection instruction, and dynamically modify the priority weight; Step 237: Sort the weighted intent label set based on the revised priority weights to generate a task type priority queue, where the arrangement order of each element in the queue reflects the execution urgency of its corresponding operation predicate in production safety detection, and the head element of the queue is marked as the core task type of the detection instruction.

[0037] In step 231, the set of operation predicates is standardized, redundant stop words and fuzzy expressions are removed, and a standardized set of operation predicates is generated. In natural language expression, the detection instruction may contain some words that do not affect the core semantics, such as modal particles, auxiliary words, etc., which will interfere with the accurate understanding of the operation predicate. For the instruction "Please carefully check whether the temperature and pressure of all production equipment in workshop A are normal", words such as "please", "carefully", and "a moment" are redundant stop words and can be removed. At the same time, some fuzzy expressions, such as "probably" and "maybe", will also be processed. Through this standardized processing, a more concise and clear set of operation predicates can be obtained, for example, the above instruction is processed as "check whether the temperature and pressure of production equipment in workshop A are normal". Stop word lists and regular expressions can be used to remove redundant stop words and fuzzy expressions. The stop word list contains common words without actual semantics, and regular expressions can be used to match and replace fuzzy expressions.

[0038] In step 232, the pre-trained intention recognition model is called to perform multi-granular semantic analysis on the standardized set of operation predicates, extract the logical dependencies and action triggering conditions between the operation predicates, and generate an initial intention label set. The pre-trained intention recognition model is trained based on a large amount of text data and can understand the semantics of the operation predicates and the relationship between them. For the standardized operation predicate "check whether the temperature and pressure of the production equipment in workshop A are normal", the intention recognition model will analyze that "check" is the core action, "production equipment in workshop A" is the operation object, "temperature and pressure" are the specific parameters of the detection, and "whether normal" is the judgment standard. At the same time, the model will identify the logical dependencies between the operation predicates, that is, first determine the operation object, then check the specific parameters of the object, and finally judge according to the standard. The action triggering condition may be to check at a certain time interval, or to trigger the check when certain abnormal signs appear. Based on these analyses, the model will generate an initial intention label set, such as "equipment status detection", "temperature and pressure detection", etc. Deep learning models such as long short-term memory networks (LSTM) or gated recurrent units (GRU) can be used as intent recognition models. These models have good performance in processing sequence data and semantic understanding.

[0039] In step 233, semantic verification is performed on the initial intention tag set based on the preset safety production domain knowledge base, and the matching degree of each tag in the initial intention tag set is verified with the standard operation scenario in the domain knowledge base, and the confidence of the verified intention tag is generated. The safety production domain knowledge base contains various standard operation scenarios and corresponding intention tags, and the initial intention tags can be compared with the tags in the knowledge base. For the two initial intention tags of "equipment status detection" and "temperature and pressure detection", the standard scenarios that match them can be found in the knowledge base. If the knowledge base clearly stipulates that temperature and pressure detection of production equipment in a workshop environment is a normal equipment status detection operation, then the matching degree of these two tags with the standard scenario will be higher. Confidence is generated for each tag based on the matching degree, and the higher the matching degree, the higher the confidence. Among them, methods such as cosine similarity can be used to calculate the matching degree between tags.

[0040] In step 234, the initial intent tag set is dynamically weighted according to the verified intent tag confidence, wherein the intent tags whose matching degree with the standard operation scenario is higher than the preset verification threshold are assigned high confidence weights to generate a weighted intent tag set. The preset verification threshold is a pre-set standard used to determine whether the matching degree between the tag and the standard scenario is high enough. If the confidence of an intent tag is higher than the threshold, it means that it has a high matching degree with the standard operation scenario and can be assigned a high confidence weight. For the two tags "device status detection" and "temperature and pressure detection", if their confidences are both higher than the preset threshold, they will be assigned higher weights in the weighted intent tag set. Through this dynamic weight allocation, tags with high matching degree with the standard operation scenario can be highlighted, providing a more accurate basis for subsequent priority determination.

[0041] In step 235, the weighted intention tag set is input into the predefined task priority rule base, and each intention tag is mapped to the corresponding priority weight according to the mapping relationship between the security response level and the operation urgency defined in the rule base. The predefined task priority rule base is formulated based on the actual needs and experience of safe production, and it stipulates the urgency of various operations under different security response levels. For the two tags "equipment status detection" and "temperature and pressure detection", the rule base may stipulate that under high security response levels, the temperature and pressure detection of the equipment are very urgent tasks and need to be executed immediately, so these two tags are given a higher priority weight. At low security response levels, the urgency of these tasks may be relatively low, and the priority weight will be reduced accordingly. By querying the rule base, each tag in the weighted intention tag set is mapped to the corresponding priority weight, so as to preliminarily determine the priority of each task.

[0042] In step 236, the contextual features of the detection instruction are combined to extract the entity object set and the operation timeliness constraints involved in the detection instruction, and the priority weight is dynamically corrected. The contextual features of the detection instruction include information such as the time and place of the instruction issuance, the current production status, etc. For the instruction "check whether the temperature and pressure of all production equipment in workshop A are normal", if workshop A is currently carrying out critical production tasks, then the temperature and pressure detection of the equipment becomes more urgent, and the priority weight of the relevant tasks can be increased. At the same time, the operation timeliness constraints will also affect the priority weight. For example, if the instruction requires the inspection to be completed within 1 hour, the priority of the task will be increased accordingly. Taking these factors into consideration, the previously determined priority weight is dynamically corrected to make the determination of the priority more in line with the actual situation.

[0043] In step 237, the weighted intent tag set is sorted based on the revised priority weights to generate a task type priority queue, in which the order of each element in the queue reflects the execution urgency of its corresponding operation predicate in production safety detection, and the head element of the queue is marked as the core task type of the detection instruction. The weighted intent tag set is sorted in descending order according to the revised priority weights, and the tags with higher priority weights are ranked higher. For the two tags "equipment status detection" and "temperature and pressure detection", if the priority weight of "temperature and pressure detection" is higher after correction, it will be ranked at the front of the queue. The generated task type priority queue can clearly show the execution urgency of each task, and the subsequent data processing and analysis work can be arranged according to this queue, giving priority to the core task types at the head of the queue.

[0044] As an implementation method, step 300 performs cross-modal fusion processing on the security detection data through a jointly trained multi-modal feature extraction model to generate a security feature vector for each time series node, including: Step 310: split the safety detection data into visual modal data, acoustic modal data and device status modal data; Step 320: respectively calling the visual encoder to extract spatial features from the visual modality data, and generating visual features including regional thermal distribution and motion trajectory; Step 330: calling an acoustic encoder to perform spectrum decomposition and abnormal waveform detection on the acoustic modal data to generate acoustic energy distribution features; Step 340: input the equipment state modal data into the state encoder for working mode recognition to generate an equipment operation state vector; Step 350: Align and fuse the visual features, acoustic energy distribution features, and device operation state vectors through a cross-modal attention network to generate a security feature vector for each time series node.

[0045] In step 310, the safety detection data is split into visual modal data, acoustic modal data, and equipment status modal data. Data of different modalities reflect different aspects of the production environment and equipment operation. Data collected by different sensors can be classified into corresponding modalities according to their sources and characteristics.

[0046] In step 320, the visual encoder is a model specifically used to process image and video data, which can extract valuable spatial features from visual modal data. In the example of workshop A, for the video images captured by the camera, the visual encoder can analyze the color, brightness and other information of different areas in the picture to generate regional thermal distribution characteristics. At the same time, the visual encoder can also track the movement trajectory of people and objects, and determine their movement direction and speed by analyzing the position changes of objects in consecutive frames. For example, it can track the walking path of workers in the workshop and the movement trajectory of materials on the conveyor belt. A convolutional neural network (CNN) can be used as a visual encoder. CNN has a strong feature extraction capability in image and video processing. Through multi-layer convolution and pooling operations, spatial features at different levels can be extracted.

[0047] In step 330, the acoustic encoder is called to perform spectral decomposition and abnormal waveform detection on the acoustic modal data to generate acoustic energy distribution characteristics. The acoustic encoder can convert the acoustic signal into a frequency domain representation and analyze the frequency components and energy distribution of the sound. In workshop A, the sound of the equipment running under normal circumstances has a specific spectral characteristic. When the equipment fails, the spectrum changes. The acoustic encoder performs a fast Fourier transform (FFT) on the collected acoustic data, converts the time domain signal into a frequency domain signal, and obtains a spectrogram of the sound. By analyzing the spectrogram, the energy distribution of different frequency components can be identified. Deep learning-based acoustic models, such as variants of recurrent neural networks (RNNs) or convolutional neural networks (CNNs), can be used to achieve spectral decomposition and abnormal waveform detection of acoustic data.

[0048] In step 340, the equipment state modal data is input into the state encoder for working mode recognition to generate an equipment operation state vector. The state encoder can determine the current working mode of the equipment based on various operating parameters of the equipment. In workshop A, the equipment state modal data includes multiple parameters such as temperature, pressure, and current, and the state encoder will conduct a comprehensive analysis of these parameters. For example, for a motor, the state encoder can determine whether the motor is in a startup, normal operation, or overload state based on the magnitude and change trend of the current. At the same time, the overall operating state of the equipment is further determined by combining parameters such as temperature and pressure. The state encoder converts this information into an equipment operation state vector, and each element in the vector represents a specific working condition feature. For example, the first element of the vector may represent the temperature state of the equipment, the second element represents the pressure state, and so on. Machine learning algorithms, such as support vector machines (SVM) or decision trees, can be used to classify and model the equipment state modal data to achieve working mode recognition.

[0049] In step 350, the visual features, acoustic energy distribution features and equipment operation state vector are feature aligned and fused through a cross-modal attention network to generate a safety feature vector for each time series node. The cross-modal attention network can automatically adjust the weights of different modal features so that the system pays more attention to important information. In the scene of workshop A, features of different modalities may reflect the safety status of the equipment to different degrees in different situations. For example, in some cases, visual features may be more able to reflect the appearance damage of the equipment, while acoustic features may be more able to detect internal faults of the equipment. The cross-modal attention network will assign appropriate weights to features of different modalities according to the current task and data situation. For example, when detecting whether the equipment has obvious appearance damage, the weight of visual features can be increased; when detecting whether there is a fault inside the equipment, the weight of acoustic features can be increased. Through feature alignment and fusion, the features of different modalities are integrated into a unified safety feature vector, which can comprehensively and accurately reflect the safety status of the equipment at each time series node. A deep learning model based on the attention mechanism, such as a variant of the Transformer architecture, can be used to achieve cross-modal feature alignment and fusion.

[0050] As an implementation method, step 400, based on the dynamic matching result of the detection instruction feature and the security feature vector of each timing node, filters out the key data segment from the timing data stream, including: Step 410: Based on the dynamic matching result between the detection instruction feature vector and the security feature vector of each timing node, the similarity score between each timing node and the detection instruction feature vector is calculated to generate an initial matching degree sequence; Step 420: Perform a time window sliding analysis on the initial matching degree sequence, identify candidate data segments in which the similarity scores of multiple consecutive time series nodes in the time series data stream exceed the dynamic threshold, and generate a set of candidate data segments; Step 430: Based on the task type priority in the detection instruction feature, adjust the length weight of each candidate data segment in the candidate data segment set, wherein the length weight of the candidate data segment corresponding to the high priority task is increased, and the length weight of the candidate data segment corresponding to the low priority task is reduced, to generate an optimized candidate data segment set; Step 440: input the optimized candidate data segment set into the pre-trained time series association model, analyze the context dependency of each candidate data segment in the time series data stream, remove isolated candidate data segments whose correlation with the previous and next time series nodes is less than the correlation threshold, and merge the candidate data segments with time overlap to generate a set of associated optimized data segments; Step 450: Based on the distribution density of each data segment in the associated optimized data segment set in the timing data stream, density cluster the associated optimized data segment set, and select the data segments with distribution density higher than the preset clustering threshold as the final key data segment set, wherein each key data segment in the key data segment set contains a preset minimum number of consecutive timing nodes and covers the entity object set specified in the detection instruction feature.

[0051] In step 410, based on the dynamic matching result between the detection instruction feature vector and the security feature vector of each timing node, the similarity score between each timing node and the detection instruction feature vector is calculated to generate an initial matching degree sequence. First, the detection instruction feature vector and the security feature vector of each timing node are normalized respectively so that they are in the same numerical range to avoid affecting the accuracy of similarity calculation due to different numerical scales. For example, the value of an element in the detection instruction feature vector may be between 0-100, while the value of the corresponding element in the security feature vector may be between 0-1. This difference can be eliminated after normalization. Then, based on the semantic coverage of the entity distribution vector in the detection instruction feature vector and the dynamic distribution of the task weight matrix, the weight of the normalized detection instruction feature vector can be dynamically adjusted to generate a dynamically weighted detection instruction feature vector. This is because the importance of different entities and tasks in the detection instruction may be different, and it is necessary to adjust according to the actual situation. For example, in the instruction to detect whether the temperature and pressure of the production equipment in workshop A are normal, if the priority of the temperature detection task is higher, then in the dynamically weighted detection instruction feature vector, the weight of the temperature-related elements will increase.

[0052] Then, the dynamically weighted detection instruction feature vector is compared with the normalized safety feature vector of each timing node for multi-dimensional similarity, including three independent calculation paths: semantic coverage overlap, entity distribution direction consistency, and task weight sensitivity. For the semantic coverage overlap calculation path, the entity association network topology features in the dynamically weighted detection instruction feature vector can be extracted, and the regional overlap analysis can be performed with the entity spatial distribution features in the normalized safety feature vector to generate a semantic coverage overlap score. For example, if the detection instruction focuses on the equipment in several specific areas of workshop A, and the safety feature vector of a certain timing node shows that there is operating data of related equipment in these areas, then the semantic coverage overlap score will be higher. For the entity distribution direction consistency calculation path, the entity distribution vector in the dynamically weighted detection instruction feature vector and the entity distribution vector in the normalized safety feature vector can be directional projected to generate an entity distribution direction consistency score. If the directions of the two vectors are relatively consistent, it means that the entity distribution is similar and the score will be high. For the task weight sensitivity calculation path, the task weight matrix in the dynamically weighted detection instruction feature vector can be element-by-element multiplied with the risk sensitivity matrix in the normalized security feature vector to generate a task weight sensitivity score. If the weight of a task in the task weight matrix is ​​large and the risk sensitivity of the task in the security feature vector is also high, then the task weight sensitivity score will be high.

[0053] The semantic coverage overlap score, entity distribution direction consistency score and task weight sensitivity score are input into the pre-trained fusion decision model, and the three scores are dynamically weighted and fused according to the task type priority in the detection instruction feature to generate a comprehensive similarity score for each timing node. For example, if the temperature detection task in the detection instruction has a high priority, the weight of the temperature-related part of the semantic coverage overlap score will be increased during fusion. After that, the timing dependency analysis model is called to perform context-related correction on the comprehensive similarity score. The model performs smooth correction on the comprehensive similarity score of the current timing node based on the dynamic matching relationship between the change trend of the security feature vector of the adjacent timing nodes and the detection instruction feature vector. Finally, the corrected comprehensive similarity score is arranged in time sequence to generate an initial matching sequence. The value of each element in the sequence reflects the dynamic matching degree between the corresponding timing node and the detection instruction feature vector, and the value change of the adjacent elements follows the context-related constraint output by the timing dependency analysis model.

[0054] In step 420, a time window sliding analysis is performed on the initial matching degree sequence to identify candidate data segments whose similarity scores of multiple consecutive time series nodes in the time series data stream exceed the dynamic threshold, and a set of candidate data segments is generated. The time window is a time period of fixed length. This window can be slid on the time series data stream to view the similarity scores of multiple consecutive time series nodes in the window. For example, the time window is set to 10 time series nodes. If the similarity scores of the 10 time series nodes in a certain window exceed the dynamic threshold, it means that the data in this time period has a high correlation with the detection instruction, and the data in this time period will be identified as a candidate data segment. The dynamic threshold will be adjusted according to the overall distribution of the data and the requirements of the detection instruction to ensure that suitable candidate data segments are screened out.

[0055] In step 430, based on the task type priority in the detection instruction feature, the length weight of each candidate data segment in the candidate data segment set is adjusted, wherein the length weight of the candidate data segment corresponding to the high-priority task is increased, and the length weight of the candidate data segment corresponding to the low-priority task is reduced, thereby generating an optimized candidate data segment set. In the example of detecting production equipment in workshop A, if temperature detection is a high-priority task in the detection instruction, the length weight of the candidate data segment related to temperature detection will increase to ensure that these key data are not missed. On the contrary, the length weight of the candidate data segment corresponding to the low-priority task will be reduced to avoid interference from too much irrelevant data. Through this length weight adjustment, data segments related to high-priority tasks can be more accurately screened out.

[0056] In step 440, the optimized candidate data segment set is input into the pre-trained time series association model, the context dependency of each candidate data segment in the time series data stream is analyzed, isolated candidate data segments whose correlation with the previous and next time series nodes is less than the correlation threshold are eliminated, and candidate data segments with time overlap are merged to generate a set of associated optimized data segments. The time series association model can learn the time order and mutual dependency between candidate data segments.

[0057] In step 450, density clustering is performed on the associated optimized data segment set according to the distribution density of each data segment in the time series data stream, and the data segments with a distribution density higher than the preset clustering threshold are screened out as the final key data segment set, wherein each key data segment in the key data segment set contains a preset minimum number of continuous time series nodes and covers the set of entity objects specified in the detection instruction feature. The density clustering algorithm can cluster adjacent and high-density data segments together according to the distribution of the data segments in the time series data stream. For example, if the data segments in a certain area are densely distributed and the density exceeds the preset clustering threshold, then these data segments will be clustered into one category. These high-density data segments are screened out as the final key data segment set. At the same time, each key data segment needs to contain a preset minimum number of continuous time series nodes to ensure the continuity and integrity of the data, and to cover the set of entity objects specified in the detection instruction feature, such as the production equipment in workshop A, so as to ensure the correlation between the key data segment and the detection instruction.

[0058] As an implementation method, step 410, based on the dynamic matching result between the detection instruction feature vector and the security feature vector of each timing node, calculates the similarity score between each timing node and the detection instruction feature vector, and generates an initial matching degree sequence, including: Step 411: normalizing the detection instruction feature vector and the security feature vector of each timing node respectively to generate a normalized detection instruction feature vector and a normalized timing node security feature vector; Step 412: Based on the semantic coverage of the entity distribution vector in the detection instruction feature vector and the dynamic distribution of the task weight matrix, dynamically adjust the weight of the normalized detection instruction feature vector to generate a dynamically weighted detection instruction feature vector; Step 413: performing a multi-dimensional similarity comparison between the dynamically weighted detection instruction feature vector and the normalized security feature vector of each timing node, wherein the multi-dimensional similarity comparison includes three independent calculation paths: semantic coverage overlap, entity distribution direction consistency, and task weight sensitivity; Step 414: for the semantic coverage overlap calculation path, extract the entity association network topology features in the dynamically weighted detection instruction feature vector, perform regional overlap analysis with the entity space distribution features in the normalized security feature vector, and generate a semantic coverage overlap score; Step 415: for the entity distribution direction consistency calculation path, perform directional projection calculation on the entity distribution vector in the dynamically weighted detection instruction feature vector and the entity distribution vector in the normalized security feature vector to generate an entity distribution direction consistency score; Step 416: for the task weight sensitivity calculation path, perform an element-by-element product operation on the task weight matrix in the dynamically weighted detection instruction feature vector and the risk sensitivity matrix in the normalized safety feature vector to generate a task weight sensitivity score; Step 417: input the semantic coverage overlap score, entity distribution direction consistency score and task weight sensitivity score into the pre-trained fusion decision model, dynamically weight the three scores according to the task type priority in the detection instruction feature, and generate a comprehensive similarity score for each time series node; Step 418: calling the timing dependency analysis model to perform context association correction on the comprehensive similarity score, wherein the timing dependency analysis model performs smooth correction on the comprehensive similarity score of the current timing node based on the dynamic matching relationship between the change trend of the security feature vector of the adjacent timing nodes and the detection instruction feature vector; Step 419: Arrange the corrected comprehensive similarity scores in chronological order to generate an initial matching sequence, wherein the value of each element in the initial matching sequence reflects the dynamic matching degree between the corresponding timing node and the detection instruction feature vector, and the value changes of adjacent elements follow the context association constraints output by the timing dependency analysis model.

[0059] In step 411, the detection instruction feature vector and the safety feature vector of each timing node are normalized respectively to generate a normalized detection instruction feature vector and a normalized timing node safety feature vector.

[0060] In step 412, based on the semantic coverage of the entity distribution vector in the detection instruction feature vector and the dynamic distribution of the task weight matrix, the weight of the normalized detection instruction feature vector is dynamically adjusted to generate a dynamically weighted detection instruction feature vector.

[0061] In step 413, a multi-dimensional similarity comparison is performed between the dynamically weighted detection instruction feature vector and the normalized security feature vector of each timing node, wherein the multi-dimensional similarity comparison includes three independent calculation paths: semantic coverage overlap, entity distribution direction consistency, and task weight sensitivity. These three calculation paths measure the similarity between the detection instruction feature vector and the security feature vector from different angles, and can more comprehensively reflect the degree of match between the two. Semantic coverage overlap focuses on the overlap between the entities and concepts involved in the detection instruction and the entities and concepts contained in the security feature vector; entity distribution direction consistency examines whether the direction of entity distribution in the detection instruction feature vector and the security feature vector is consistent; task weight sensitivity considers the degree of match between the priority of the task in the detection instruction and the risk sensitivity in the security feature vector. Through these three independent calculation paths, the similarity can be evaluated from multiple dimensions to avoid the limitations of single-dimensional evaluation.

[0062] In step 414, for the semantic coverage overlap calculation path, the entity association network topology features in the dynamically weighted detection instruction feature vector are extracted, and the entity space distribution features in the normalized security feature vector are analyzed for regional overlap to generate a semantic coverage overlap score. The entity association network topology features describe the relationship and structure between entities in the detection instruction, such as the connection relationship between devices, the distribution relationship of equipment in different areas of the workshop, etc. The entity space distribution features reflect the spatial distribution of entities in the security feature vector. These two features are subjected to regional overlap analysis to determine the degree of overlap between the entity area involved in the detection instruction and the area where the entity is located in the security feature vector.

[0063] In step 415, for the entity distribution direction consistency calculation path, the entity distribution vector in the dynamically weighted detection instruction feature vector and the entity distribution vector in the normalized security feature vector are subjected to directional projection calculation to generate an entity distribution direction consistency score. The entity distribution vector represents the distribution of the entity in space or feature dimension. By performing directional projection calculation on two entity distribution vectors, it can be determined whether their distribution directions are consistent. If the directions of the two vectors are relatively consistent, it means that the entity distribution is similar and the entity distribution direction consistency score is high.

[0064] In step 416, for the task weight sensitivity calculation path, the task weight matrix in the dynamically weighted detection instruction feature vector and the risk sensitivity matrix in the normalized safety feature vector are multiplied element by element to generate a task weight sensitivity score. In production safety inspections, the inspection instructions will clarify the importance of different tasks, which is reflected by the task weight matrix, and the risk sensitivity matrix in the safety feature vector reflects the risk sensitivity of different tasks at each timing node. If a task has a large weight in the task weight matrix and a high value in the risk sensitivity matrix, then the product result corresponding to the task is large, and the task weight sensitivity score is also high, which indicates that the task has a good match with the inspection instruction and is more sensitive to risk.

[0065] In step 417, the semantic coverage overlap score, entity distribution direction consistency score and task weight sensitivity score are input into the pre-trained fusion decision model, and the three scores are dynamically weighted and fused according to the task type priority in the detection instruction feature to generate a comprehensive similarity score for each time series node. These three scores reflect the degree of matching between the detection instruction feature vector and the security feature vector from different angles, but their importance may vary due to the different task type priorities of the detection instruction. For example, if the temperature detection of high-temperature equipment in the detection instruction is a high-priority task, then when fused, the weights of the semantic coverage overlap score, entity distribution direction consistency score and task weight sensitivity score related to the temperature of the high-temperature equipment will be relatively high. The pre-trained fusion decision model is trained based on a large amount of historical data, and can assign appropriate weights to the three scores according to the task type priority in the detection instruction feature, and then perform weighted summation.

[0066] In step 418, the timing dependency analysis model is called to perform context-related correction on the comprehensive similarity score, wherein the timing dependency analysis model performs smooth correction on the comprehensive similarity score of the current timing node based on the dynamic matching relationship between the change trend of the security feature vectors of the adjacent timing nodes and the detection instruction feature vector. In the actual production process, the data of each timing node does not exist in isolation, and there is temporal continuity and correlation between them. For example, when detecting the operating status of the production equipment in workshop A, the equipment status of the current timing node is usually related to the status of the adjacent timing nodes. If the security feature vectors of the adjacent timing nodes show that the equipment operating status is relatively stable, and the comprehensive similarity score of the current timing node suddenly fluctuates greatly, then there may be data errors or abnormalities. The timing dependency analysis model analyzes the change trend of the security feature vectors of the adjacent timing nodes and their matching relationship with the detection instruction feature vector, and adjusts the comprehensive similarity score of the current timing node. If the matching degree of adjacent nodes is high and the changing trend is stable, while the score of the current node is too low, the model can appropriately increase its score; conversely, if the matching degree of adjacent nodes is low and shows a deteriorating trend, and the score of the current node is too high, the model can reduce its score. This can make the comprehensive similarity score more consistent with the actual production situation and avoid misjudgment caused by individual abnormal data.

[0067] In step 419, the corrected comprehensive similarity scores are arranged in time sequence to generate an initial matching sequence, wherein the value of each element in the initial matching sequence reflects the dynamic matching degree between the corresponding time node and the detection instruction feature vector, and the value change of the adjacent elements follows the context association constraint condition output by the time dependency analysis model. The comprehensive similarity scores corrected by the context association are arranged in time sequence to form a sequence. This sequence intuitively shows the change of the matching degree between each time node and the detection instruction over time. For example, in the equipment inspection of workshop A, the initial matching sequence can clearly show the degree of compliance between the equipment operation status and the detection instruction requirements at different time points. Since the value change of the adjacent elements follows the context association constraint condition output by the time dependency analysis model, the sequence not only reflects the independent matching of each time node, but also reflects the continuity and correlation of the data in the time dimension. By analyzing this initial matching sequence, the key data segments related to the detection instruction can be screened out more accurately, providing strong data support for subsequent abnormal pattern recognition and risk assessment.

[0068] As an implementation method, step 500 calls a pre-trained security decision model to perform abnormal pattern recognition and risk assessment on key data segments, and generates a structured security report corresponding to the detection instruction, including: Step 510: Slice the key data segment at multiple granularities to generate hierarchical data blocks containing global state features and local abnormal features; Step 520: Input the hierarchical data block into the convolutional recursive network of the security decision model to extract spatiotemporal features; Step 530: performing a similarity comparison between the extracted spatiotemporal features and a preset abnormal pattern library through a pattern matching layer in the security decision model; Step 540: Generate risk assessment indicators including risk level, abnormality type and impact range based on the comparison results; Step 550: Generate a structured safety report including visual charts and handling suggestions based on the risk assessment indicators and the task types in the detection instruction characteristics.

[0069] In step 510, the key data segment is sliced ​​at multiple granularities to generate hierarchical data blocks containing global state features and local abnormal features. The key data segment contains various data of equipment operation over a period of time, and these data are analyzed from different granularities. For example, for the key data segment of a production equipment in workshop A, it may contain continuous data of multiple parameters such as temperature, pressure, and rotation speed of the equipment. These data are sliced ​​according to different time scales. For example, data over a longer period of time can be analyzed as global state features to understand the overall operation trend of the equipment; at the same time, data over a shorter period of time can be mined as local abnormal features to discover possible sudden abnormal situations. Through this multi-granularity slicing, data blocks at different levels can be generated, and each data block contains equipment operation information at a specific granularity. Multi-granularity slicing can be achieved using sliding window technology, and data blocks of different time scales can be obtained by adjusting the size and step size of the window.

[0070] In step 520, the hierarchical data block is input into the convolutional recurrent network of the safety decision model for spatiotemporal feature extraction. The convolutional recurrent network combines the advantages of convolutional neural networks (CNN) and recurrent neural networks (RNN), and can effectively process spatial features and time series information in the data. For the global state features in the hierarchical data block, the convolutional layer can extract the spatial features in the data, such as the correlation between different parameters of the equipment. The recursive layer can process the time series information of the data and capture the changing trend of the equipment operating status over time. For local abnormal features, the convolutional recursive network can more keenly detect sudden changes and abnormal patterns in the data. For example, when analyzing the temperature data of the equipment, the convolutional recursive network can identify abnormal situations such as a sudden increase or decrease in temperature. By inputting the hierarchical data block into the convolutional recursive network, key spatiotemporal features related to abnormal patterns and risks can be extracted.

[0071] In step 530, the extracted spatiotemporal features are compared with the preset abnormal pattern library for similarity through the pattern matching layer in the safety decision model. The preset abnormal pattern library stores various known safety abnormalities and their corresponding characteristic patterns. The extracted spatiotemporal features are compared with the patterns in the abnormal pattern library one by one, and the similarity between them is calculated. For example, the abnormal pattern library may contain characteristic patterns of abnormal conditions such as equipment overheating and excessive pressure. If the extracted spatiotemporal features have a high similarity with a certain abnormal pattern, then there may be a corresponding abnormal situation. Exemplarily, methods such as cosine similarity can be used to calculate the similarity between spatiotemporal features and abnormal patterns.

[0072] In step 540, a risk assessment index including risk level, abnormality type and impact range is generated based on the comparison results. According to the comparison results of the spatiotemporal characteristics and the abnormal pattern library, it can be determined whether there is an abnormal situation and further evaluate its risk level. If the similarity exceeds the preset threshold, it means that an abnormal situation exists. According to the severity of the abnormality and the possible impact, the risk level is determined, such as high, medium and low. At the same time, according to the matched abnormal pattern, the abnormality type is determined, such as equipment failure, operating error, etc. In addition, the scope of possible impact of the abnormal situation will be analyzed, such as whether it only affects a certain device, or whether it will affect surrounding equipment and personnel. By comprehensively considering these factors, a risk assessment index including risk level, abnormality type and impact range is generated to provide a basis for subsequent decision-making.

[0073] In step 550, a structured safety report including visual charts and disposal suggestions is generated according to the risk assessment indicators and the task type in the detection instruction characteristics. Visual charts can intuitively display the risk situation and help managers quickly understand the safety status of the equipment. For example, a line chart can be used to show the trend of equipment parameters over time, and a bar chart can be used to compare the distribution of different risk levels. Disposal suggestions will give specific countermeasures based on the risk assessment results and the requirements of the detection instructions. If the detection instruction requires a focus on the temperature and pressure of the equipment, and the risk assessment indicators show that the equipment temperature is too high, then the disposal suggestions may include reducing the equipment load, checking the cooling system, etc. The visual charts and disposal suggestions are integrated into the structured safety report, and the format and content of the report will be customized according to the task type of the detection instruction to meet the needs of different users.

[0074] As an implementation method, the collaborative optimization method of the aforementioned jointly trained multimodal feature extraction model and the security decision model includes the following steps 10 to 70: Step 10: Input the spatial feature sequence output by the visual encoder, the spectral feature sequence output by the acoustic encoder, and the working condition feature sequence output by the state encoder into the cross-modal attention network to generate a time-aligned multimodal joint feature sequence; Step 20: Timestamp matching is performed between the multimodal joint feature sequence and the equipment failure type label in the annotated safety event to generate a joint training sample set with time series annotations; Step 30: Extract features from the joint training sample set through the spatiotemporal convolution layer in the security decision model, and output the predicted risk level and probability distribution of abnormality type; Step 40: Compare the predicted risk level with the environmental risk level label in the annotated security event node by node to generate a first error gradient, and perform cross entropy calculation on the probability distribution of the abnormal type and the equipment failure type label to generate a second error gradient; Step 50: Input the multimodal joint feature sequence into the discriminator network, calculate the distribution difference between the multimodal joint feature sequence and the feature template extracted from the real security event, and generate a third error gradient; Step 60: generating a joint optimization direction of the multimodal feature extraction model and the safety decision model based on the weighted fusion result of the first error gradient, the second error gradient and the third error gradient; Step 70: Synchronously adjust the weight parameters of the cross-modal attention network, the kernel parameters of the spatiotemporal convolution layer, and the discrimination threshold of the discriminator network according to the joint optimization direction until the fusion value of the first error gradient and the second error gradient reaches a preset convergence condition.

[0075] In step 10, the spatial feature sequence output by the visual encoder, the spectral feature sequence output by the acoustic encoder, and the working condition feature sequence output by the state encoder are input into the cross-modal attention network to generate a time-aligned multimodal joint feature sequence. In this way, a time-aligned multimodal joint feature sequence is generated, providing a more accurate data basis for subsequent training and optimization.

[0076] In step 20, the multimodal joint feature sequence is timestamped with the equipment failure type label in the annotated safety event to generate a joint training sample set with time series annotations. The annotated safety event contains known equipment failure types and corresponding time information. Each time point in the multimodal joint feature sequence is compared with the timestamp in the annotated safety event, and the multimodal joint feature is associated with the corresponding equipment failure type label. For example, the annotated safety event records that a motor failure occurred in the equipment at a specific time point. The feature of the corresponding time point in the multimodal joint feature sequence can be found and associated with the label "motor failure". Through this timestamp matching, a joint training sample set with time series annotations is generated. These sample sets contain multimodal features and corresponding equipment failure type labels, providing supervised learning data for model training.

[0077] In step 30, the spatiotemporal convolution layer in the safety decision model extracts features from the joint training sample set, and outputs the predicted risk level and probability distribution of abnormal types. The spatiotemporal convolution layer can effectively process the spatiotemporal information in the joint training sample set and extract features related to risks and abnormalities. In the safety detection of workshop A, the spatiotemporal convolution layer can analyze the changes of multimodal joint feature sequences over time, as well as the spatial relationship between different modal features. For example, it can identify the changing trends of parameters such as equipment temperature and pressure over a period of time, as well as the association between these parameters and features such as equipment appearance and sound. By extracting and analyzing these features, the safety decision model outputs predicted risk levels, such as high, medium, and low, as well as the probability distribution of abnormal types, that is, the probability of each abnormal type occurring. In this way, the current safety status of the equipment and possible abnormal situations can be preliminarily judged.

[0078] In step 40, the predicted risk level is compared node by node with the environmental risk level label in the annotated safety event to generate a first error gradient, and the abnormal type probability distribution is cross-entropy calculated with the equipment failure type label to generate a second error gradient. The predicted risk level is the output result of the safety decision model, and the environmental risk level label in the annotated safety event is the known actual situation. The two are compared node by node to calculate the error at each node. For example, if the predicted risk level is low and the actual environmental risk level label is high, then there is an error. By statistically analyzing the errors of all nodes, a first error gradient is generated, which reflects the degree of difference between the predicted risk level and the actual risk level. At the same time, the abnormal type probability distribution is cross-entropy calculated with the equipment failure type label to generate a second error gradient, which reflects the difference between the predicted abnormal type probability distribution and the actual situation.

[0079] In step 50, the multimodal joint feature sequence is input into the discriminator network, and the distribution difference between it and the feature template extracted from the real security event is calculated to generate a third error gradient. The function of the discriminator network is to determine whether the input multimodal joint feature sequence comes from a real security event. The feature template extracted from the real security event represents the typical feature distribution under normal and abnormal conditions. The multimodal joint feature sequence is compared with these feature templates to calculate the distribution difference between them. For example, the maximum mean difference (MMD) and other methods can be used to measure the distribution difference. Through this comparison, a third error gradient is generated, which reflects the degree of difference between the multimodal joint feature sequence and the feature template of the real security event. This error gradient can help adjust the multimodal feature extraction model so that the features it generates are closer to the features of the real security event.

[0080] In step 60, based on the weighted fusion results of the first error gradient, the second error gradient and the third error gradient, the joint optimization direction of the multimodal feature extraction model and the security decision model is generated. These three error gradients reflect the performance problems of the model from different angles. The first error gradient focuses on the accuracy of the predicted risk level, the second error gradient focuses on the accuracy of the abnormal type prediction, and the third error gradient focuses on the degree of matching between the multimodal features and the real security event features. Different weights are assigned to these three error gradients, and then weighted fusion is performed. For example, if in the current security detection task, more attention is paid to the accuracy of the predicted risk level, then the weight of the first error gradient may be larger. Through weighted fusion, a comprehensive error gradient is obtained. According to the direction of this error gradient, the joint optimization direction of the multimodal feature extraction model and the security decision model is determined, that is, the direction in which the model parameters need to be adjusted.

[0081] In step 70, the weight parameters of the cross-modal attention network, the kernel parameters of the spatiotemporal convolution layer, and the discrimination threshold of the discriminator network are synchronously adjusted according to the joint optimization direction until the fusion value of the first error gradient and the second error gradient reaches the preset convergence condition. The cross-modal attention network is responsible for the fusion of multimodal features, the spatiotemporal convolution layer is used for feature extraction of the safety decision model, and the discriminator network is used to judge the authenticity of the features. According to the joint optimization direction, the parameters of these networks are synchronously adjusted. For example, when the joint optimization direction indicates that the fusion effect of multimodal features needs to be improved, the weight parameters of the cross-modal attention network can be adjusted so that the features of different modalities can be better integrated. During the adjustment process, the fusion value of the first error gradient and the second error gradient can be continuously calculated and compared with the preset convergence condition. The preset convergence condition can be an error threshold. When the fusion value is less than this threshold, it means that the performance of the model has reached a certain requirement and the training process can be stopped. Through this iterative adjustment, the performance of the multimodal feature extraction model and the safety decision model is continuously optimized so that they can work better together to improve the accuracy and reliability of production safety detection.

[0082] Throughout the collaborative optimization process, it is necessary to accurately process various data and error gradients, and reasonably adjust model parameters. For the calculation of error gradients, it is necessary to use appropriate mathematical methods and algorithms to ensure the accuracy of the calculation results. When adjusting model parameters, it is necessary to use optimization algorithms such as stochastic gradient descent (SGD) or its variants to improve the efficiency and stability of the adjustment. At the same time, it is also necessary to monitor and evaluate the training process to detect problems in a timely manner and make adjustments.

[0083] In order to improve the effect of collaborative optimization, some optimization strategies can be adopted. For example, batch normalization technology can be used to accelerate the convergence speed of the model and reduce the training time. In addition, the model fusion method can be used to combine multiple different models to improve the generalization ability and performance of the model. During the training process, data enhancement technology can be used to expand the joint training sample set, increase the diversity of samples, and thus improve the robustness of the model.

[0084] Steps 10-70 provide a complete set of methods and processes for collaborative optimization of the multimodal feature extraction model and the safety decision model. Through operations such as time alignment, sample annotation, feature extraction, error calculation, gradient fusion, and parameter adjustment, the performance and collaborative working ability of the two models can be continuously improved, providing more accurate and reliable support for safety production detection.

[0085] As an implementation manner, the method provided in the embodiment of the present invention may further include: Step 61: Divide the first error gradient into risk level intervals to generate a high-risk interval error component and a normal interval error component; Step 62: performing abnormal type cluster analysis on the second error gradient to generate a core fault type error component and a marginal abnormal type error component; Step 63: assigning enhanced weight coefficients to high-risk interval error components and core fault type error components according to the task type priority in the detection instruction feature; Step 64: input the weighted first error gradient and the second error gradient into the dynamic fusion controller to generate the main optimization direction of the model parameters; Step 65: Input the third error gradient into the distribution alignment controller to generate a sub-optimization direction for feature distribution optimization; Step 66: Coordinate the adjustment order of the main optimization direction and the secondary optimization direction through a priority scheduling strategy, wherein the secondary optimization direction is preferentially executed in the early stage of training to improve the feature alignment accuracy, and the main optimization direction is switched to in the later stage of training to enhance the risk prediction capability; Step 67: Dynamically control the parameter update frequency of the cross-modal attention network and the safety decision model based on the output results of the scheduling strategy to achieve the co-evolution of multimodal feature extraction and risk prediction.

[0086] In step 61, the first error gradient is divided into risk level intervals to generate a high-risk interval error component and a conventional interval error component. The first error gradient reflects the difference between the predicted risk level and the environmental risk level label in the annotated safety event. In the safety inspection of workshop A, the risk level can generally be divided into different levels such as high, medium and low. According to the preset risk level interval standard, the first error gradient is divided into a high-risk interval error component and a conventional interval error component. For example, when there is a difference between the predicted risk level and the actual environmental risk level label in the high-risk level interval, this part of the error will be classified as a high-risk interval error component; and the errors in the medium and low risk level intervals are classified as conventional interval error components. This division helps to deal with errors at different risk levels more specifically, because errors at high risk levels may have a more serious impact on safe production and require more attention.

[0087] In step 62, the second error gradient is subjected to abnormal type clustering analysis to generate a core fault type error component and a marginal abnormal type error component. The second error gradient is obtained by cross entropy calculation of the abnormal type probability distribution and the equipment fault type label, which reflects the difference between the predicted abnormal type and the actual abnormal type. In workshop A, there may be multiple abnormal types of equipment, such as motor failure, sensor failure, pipeline leakage, etc. These abnormal types are clustered and analyzed, and the common abnormal types with greater impact on production are classified as core fault types, and the relatively rare abnormal types with less impact are classified as marginal abnormal types. Then, according to the clustering results, the second error gradient is divided into a core fault type error component and a marginal abnormal type error component. For example, motor failure is a common and important abnormal type in workshop A, then the error related to motor failure will be classified as a core fault type error component; while some occasional small parts damage and other abnormal type errors are classified as marginal abnormal type error components. Such a division can make the identification and prediction of core fault types more emphasized in the optimization process.

[0088] In step 63, according to the task type priority in the detection instruction feature, the high-risk interval error component and the core fault type error component are given enhanced weight coefficients. The task type priority in the detection instruction feature reflects the importance of different detection tasks in safe production. In the safety inspection of workshop A, if the detection instruction focuses on the operating status of high-risk equipment and the identification of core fault types, then a higher weight coefficient will be given to the high-risk interval error component and the core fault type error component.

[0089] In step 64, the weighted first error gradient and the second error gradient are input into the dynamic fusion controller to generate the main optimization direction of the model parameters. The dynamic fusion controller is a module for comprehensively processing error gradients. It can fuse the weighted first error gradient and the second error gradient according to the characteristics and weights of different error gradients. In the scenario of workshop A, the high-risk interval error component, the conventional interval error component, the core fault type error component and the edge abnormal type error component after weight adjustment are comprehensively considered, and a comprehensive error gradient is generated by the dynamic fusion controller. This gradient represents the main optimization direction that the model parameters need to be adjusted. The dynamic fusion controller can adopt a multi-layer neural network structure. The input layer receives the weighted first error gradient and the second error gradient, and these two error gradients respectively reflect the error of the predicted risk level and the probability distribution of the abnormal type. Multiple hidden layers are set in the middle, and activation functions such as ReLU are used for nonlinear transformation to enhance the expression ability of the model. The output layer outputs the fused error gradient as the main optimization direction of the model parameters. When implementing, first initialize the network weights, then use the gradient descent method to iteratively update the weights based on the training data, and continuously adjust the fusion method so that the main optimization direction of the output can effectively guide the adjustment of model parameters and improve model performance.

[0090] In step 65, the third error gradient is input into the distribution alignment controller to generate a sub-optimal direction for feature distribution optimization. The third error gradient reflects the distribution difference between the multimodal joint feature sequence and the feature template extracted from the real security event. The function of the distribution alignment controller is to adjust the multimodal feature extraction model according to the third error gradient so that the distribution of the generated multimodal joint feature sequence is closer to the feature template of the real security event. In workshop A, different security events may correspond to different feature distributions. Through the distribution alignment controller, the direction of feature distribution optimization, i.e., the sub-optimal direction, is determined according to the third error gradient. For example, if the multimodal joint feature sequence has a large distribution difference from the real feature template in some feature dimensions, the distribution alignment controller will instruct the model to adjust the parameters so that the distribution in these feature dimensions is more consistent. The distribution alignment controller can be implemented using a neural network structure. Its input is the third error gradient, which reflects the distribution difference between the multimodal joint feature sequence and the feature template of the real security event. The network can be composed of an input layer, a hidden layer, and an output layer. The input layer receives the error gradient data, the hidden layer performs a nonlinear transformation, and increases the model expression capability through an activation function such as ReLU. The output layer outputs the direction and amplitude of the adjustment parameters. During implementation, the gradient descent method is used to update the network weights to minimize the distribution difference between the multimodal joint features and the true template. Continuous iterative training gradually aligns the multimodal feature distribution with the true template, improving the accuracy of model feature extraction.

[0091] In step 66, the adjustment order of the main optimization direction and the secondary optimization direction is coordinated through the priority scheduling strategy, wherein the secondary optimization direction is preferentially executed in the early stage of training to improve the feature alignment accuracy, and the main optimization direction is switched to in the later stage of training to enhance the risk prediction ability. The priority scheduling strategy is to reasonably arrange the adjustment order of the main optimization direction and the secondary optimization direction to achieve better optimization effect. In the early stage of training, the features generated by the multimodal feature extraction model may be greatly different from the feature template of the real security event. At this time, the secondary optimization direction is preferentially executed, that is, by adjusting the multimodal feature extraction model, the feature distribution is closer to the real template, and the feature alignment accuracy is improved. For example, at the beginning of the security detection model training, the focus is on the matching degree between the multimodal joint feature sequence and the real feature template, and the model parameters are adjusted by the distribution alignment controller. In the later stage of training, when the feature alignment accuracy reaches a certain level, it can be switched to the main optimization direction, that is, according to the weighted first error gradient and the second error gradient, the risk prediction ability of the model is enhanced, so that the model can more accurately predict the risk level and abnormal type.

[0092] In step 67, the parameter update frequency of the cross-modal attention network and the safety decision model is dynamically controlled based on the output result of the scheduling strategy to achieve the co-evolution of multimodal feature extraction and risk prediction. The scheduling strategy determines the adjustment order of the main optimization direction and the secondary optimization direction, and the parameter update frequency of the cross-modal attention network and the safety decision model can be dynamically controlled according to this order. In the early stage of training, since the secondary optimization direction is executed first, the parameters of the cross-modal attention network can be updated more frequently to improve the feature alignment accuracy. For example, by increasing the number of updates of the cross-modal attention network parameters, it can adjust the fusion mode of different modal features faster, so that the multimodal joint feature sequence is closer to the real feature template. In the later stage of training, when switching to the main optimization direction, more attention can be paid to the parameter update of the safety decision model to enhance the risk prediction ability. Through this dynamic control, the co-evolution of multimodal feature extraction and risk prediction can be achieved, so that both models can be effectively optimized at different training stages, and ultimately improve the performance of the entire safety production detection system.

[0093] As an implementation manner, the method provided in the embodiment of the present invention may further include: Step 71: after each parameter update, extract the updated multimodal joint feature sequence and risk prediction result; Step 72: Input the updated multimodal joint feature sequence into the discriminator network to generate a new distribution difference index; Step 73: Compare the updated risk prediction results with the annotated security events to generate new risk level errors and anomaly type errors; Step 74: Calculate the comprehensive optimization performance index of the current training stage based on the new distribution difference index, risk level error and anomaly type error; Step 75: When the comprehensive optimization efficiency index does not improve for N consecutive training cycles, trigger the weight coefficient reset mechanism of the dynamic fusion controller; Step 76: during the weight coefficient reset process, reallocate the fusion weights of the first error gradient and the second error gradient according to the error distribution pattern in the historical training data; Step 77: Synchronize the reset fusion weight to the priority scheduling strategy and redefine the switching conditions between the primary and secondary optimization directions; Step 78: The optimization status information is fed back to the cross-modal attention network through the feedback channel of the discriminator network, forming a closed-loop optimization loop from feature extraction to risk prediction to parameter adjustment.

[0094] In step 71, after each parameter update, the updated multimodal joint feature sequence and risk prediction results are extracted. The multimodal joint feature sequence is formed by the fusion of features of multiple modalities such as vision, acoustics and device status, and it contains rich information about the operating status of the device. The risk prediction result is the judgment of the risk level and abnormal type of the device output by the safety decision model based on these feature sequences. Each time the parameters of the cross-modal attention network, spatiotemporal convolution layer and discriminator network are updated, the updated information can be recalculated and extracted. For example, after updating the weight parameters of the cross-modal attention network, the fusion method of different modal features changes, and the multimodal joint feature sequence will also change accordingly, and the new sequence can be extracted in time. At the same time, the safety decision model performs risk prediction based on the new feature sequence to obtain the updated risk prediction results, which will be used for subsequent evaluation and optimization.

[0095] In step 72, the updated multimodal joint feature sequence is input into the discriminator network to generate a new distribution difference index. The function of the discriminator network is to determine whether the input multimodal joint feature sequence comes from a real security event. The feature templates extracted from real security events represent the typical feature distribution under normal and abnormal conditions. The updated multimodal joint feature sequence is compared with these feature templates to calculate the distribution difference between them. For example, the maximum mean difference (MMD) and other methods can be used to measure the distribution difference, and the formula is: , where P and Q are the distributions of multimodal joint feature sequences and true feature templates, respectively. is the mapping function that maps the data to the Reproducing Kernel Hilbert Space (RKHS), and m and n are the number of samples in the two distributions. Through this comparison, a new distribution difference index is generated, which reflects the degree of difference between the updated multimodal joint feature sequence and the real security event feature template.

[0096] In step 73, the updated risk prediction results are compared with the annotated safety events to generate new risk level errors and anomaly type errors. The annotated safety events contain known equipment failure types and environmental risk level labels, which are real situations. The updated risk prediction results are compared node by node with the annotated safety events. For the risk level, the difference between the predicted risk level and the actual environmental risk level label is calculated to obtain a new risk level error. For the anomaly type, the difference between the predicted anomaly type probability distribution and the equipment failure type label is calculated, and the new anomaly type error is obtained by cross entropy calculation. The cross entropy formula is: , where p(i) is the probability distribution of the real equipment failure type label, and q(i) is the predicted abnormal type probability distribution. Through this comparison, new risk level errors and abnormal type errors are obtained, which reflect the prediction accuracy of the updated safety decision model.

[0097] In step 74, the comprehensive optimization performance index of the current training stage is calculated based on the new distribution difference index, risk level error and abnormal type error. The comprehensive optimization performance index is an indicator that comprehensively measures the model optimization effect, which takes into account the accuracy of multimodal feature extraction and the prediction accuracy of the safety decision model. Different weights can be assigned to the new distribution difference index, risk level error and abnormal type error, and then the weighted sum is performed to obtain the comprehensive optimization performance index.

[0098] In step 75, when the comprehensive optimization efficiency index has not improved for N consecutive training cycles, the weight coefficient reset mechanism of the dynamic fusion controller is triggered. During the training process, the comprehensive optimization efficiency index should gradually improve as the training progresses. If the index has not improved for N consecutive training cycles, it means that the current optimization strategy may be trapped in a local optimum or other problems have occurred. At this time, the weight coefficient reset mechanism of the dynamic fusion controller can be triggered. The dynamic fusion controller is responsible for fusing the weighted first error gradient with the second error gradient to generate the main optimization direction of the model parameters. The weight coefficient reset mechanism will readjust the weights of these error gradients to change the optimization direction, try to jump out of the local optimum, and allow the model to continue to optimize in the direction of better performance.

[0099] In step 76, during the weight coefficient reset process, the fusion weights of the first error gradient and the second error gradient are reallocated according to the error distribution pattern in the historical training data. The historical training data records the error conditions of the model at different training stages, including risk level errors, abnormal type errors, etc. The distribution patterns of these errors are analyzed, such as which types of errors occur more frequently in certain stages, and which errors have a greater impact on model performance. Based on these analysis results, the fusion weights of the first error gradient and the second error gradient are reallocated. For example, if the historical data shows that the risk level error is large at a certain stage and has a significant impact on the model performance, then when the weight coefficient is reset, the weight of the first error gradient can be increased to make the model pay more attention to the prediction accuracy of the risk level.

[0100] In step 77, the reset fusion weights are synchronized to the priority scheduling strategy, and the switching conditions of the primary and secondary optimization directions are redefined. The priority scheduling strategy is responsible for coordinating the adjustment order of the primary optimization direction and the secondary optimization direction. In the early stage of training, the secondary optimization direction is prioritized to improve the feature alignment accuracy, and in the later stage of training, it is switched to the main optimization direction to enhance the risk prediction ability. The reset fusion weights will change the focus of the main optimization direction, so these weights need to be synchronized to the priority scheduling strategy. According to the new fusion weights, the switching conditions of the primary and secondary optimization directions are redefined. For example, if the weight of the first error gradient increases after the reset, it may switch to the main optimization direction in advance, so that the model can focus on improving the risk prediction ability earlier.

[0101] In step 78, the optimization state information is transmitted back to the cross-modal attention network through the feedback channel of the discriminator network, forming a closed-loop optimization loop from feature extraction to risk prediction and then to parameter adjustment. In the process of calculating the distribution difference index, the discriminator network accumulates information about the difference between the multimodal joint feature sequence and the real feature template. These optimization state information are transmitted back to the cross-modal attention network through the feedback channel. The cross-modal attention network is responsible for the fusion of multimodal features. According to these feedback information, it can adjust the weight distribution of different modal features so that the generated multimodal joint feature sequence is closer to the feature template of the real security incident. In this way, a closed-loop optimization loop is formed, from feature extraction to risk prediction, and then parameter adjustment is performed according to the prediction results. Finally, the adjustment information is fed back to the feature extraction link, and the optimization is continuously cycled to improve the overall performance of the model.

[0102] Steps 71-78 construct a closed-loop optimization loop in the collaborative optimization of the multimodal feature extraction model and the safety decision model. By extracting and updating information, calculating indicators, resetting weights, adjusting strategies, and feedback optimization, the performance of the model can be continuously optimized to improve the accuracy and reliability of production safety detection.

[0103] As an implementation mode, the method provided by the embodiment of the present invention may also include: pushing the generated structured security report to the target user terminal and recording the user operation log; extracting the user's modification content and confirmation response time of the security report as a feedback signal; converting the feedback signal into a model optimization reward value through a reinforcement learning framework; online fine-tuning the risk assessment strategy in the security decision-making model based on the reward value; and establishing a mapping relationship library between feedback data and original detection instructions for semantic parsing optimization of subsequent detection instructions.

[0104] This implementation constitutes a model optimization mechanism based on user feedback. By collecting and analyzing user feedback on structured security reports, the security decision model is fine-tuned online and the semantic parsing of subsequent detection instructions is optimized.

[0105] Specifically, the generated structured safety report is pushed to the target user terminal and the user operation log is recorded. In the safety detection scenario of workshop A, after completing the abnormal pattern recognition and risk assessment of the key data segments, a structured safety report containing visual charts and disposal suggestions is generated. The report presents in detail the risk level, abnormality type, impact range and other information of the equipment, as well as handling suggestions for different abnormal situations. This report is pushed to the target user terminals such as relevant managers and technicians. These terminals can be computers, mobile phones and other devices. At the same time, the user's operation log of the report can be recorded, including when the user opened the report, what content was viewed, whether printing or forwarding was performed, etc. These operation logs will provide data support for subsequent analysis of user behavior and feedback.

[0106] The generated structured safety report is pushed to the target user terminal and the user operation log is recorded. In the safety detection scenario of workshop A, after completing the abnormal pattern recognition and risk assessment of the key data segments, a structured safety report containing visual charts and disposal suggestions is generated. The report presents in detail the risk level, abnormality type, impact range and other information of the equipment, as well as handling suggestions for different abnormal situations. This report is pushed to the target user terminals such as relevant managers and technicians. These terminals can be computers, mobile phones and other devices. At the same time, the user's operation log of the report can be recorded, including when the user opened the report, what content was viewed, whether printing or forwarding was performed, etc. These operation logs will provide data support for subsequent analysis of user behavior and feedback.

[0107] Extract the user's modification content and confirmation response time of the security report as feedback signals. After viewing the structured security report, the user may modify the report content according to his or her own experience and actual situation, or confirm the risk assessment and disposal suggestions in the report. Extract these modifications, for example, the user may revise the judgment of certain abnormal types, or add some additional disposal suggestions. At the same time, the user's confirmation response time can be recorded, that is, the time interval from the report push to the user's confirmation. This time can reflect the user's recognition of the report and the efficiency of handling emergencies. For example, if the user quickly confirms the report in a short time, it means that the report content may be more in line with the actual situation; if the user does not confirm for a long time or makes a lot of modifications, it may indicate that the report is inaccurate or incomplete. These modifications and confirmation response times will serve as important feedback signals for subsequent model optimization.

[0108] The feedback signal is converted into a model optimization reward value through the reinforcement learning framework. Reinforcement learning is a machine learning method that learns the optimal strategy based on the reward signal fed back by the environment through the interaction between the intelligent agent and the environment. In this scenario, the user's feedback signal is used as environmental feedback and converted into a model optimization reward value through preset rules. For example, if the user makes fewer modifications to the report and the confirmation response time is short, it means that the model's risk assessment and disposal recommendations are relatively accurate, and a higher reward value can be given; conversely, if the user makes a lot of modifications and the confirmation response time is long, it means that the performance of the model needs to be improved, and a lower reward value can be given. In this way, user feedback is converted into a quantifiable reward signal to guide the optimization of the security decision-making model.

[0109] The risk assessment strategy in the security decision model is fine-tuned online based on the reward value. The risk assessment strategy of the security decision model determines how to judge the risk level and abnormality type of the device based on the key data segments input. These strategies are adjusted according to the reward value. If the reward value is high, it means that the current risk assessment strategy is relatively effective and these strategies can be maintained to a certain extent; if the reward value is low, the risk assessment strategy can be adjusted. For example, adjust the threshold in the abnormal pattern library, or change the weight of different features in risk assessment. Through this online fine-tuning, the security decision model can continuously adapt to the actual situation and improve the accuracy of risk assessment.

[0110] Establish a mapping relationship library between feedback data and original detection instructions for semantic parsing optimization of subsequent detection instructions. Feedback data includes information such as user modification content, confirmation response time, and model optimization reward value. The original detection instruction is the instruction initially issued to guide security detection. Associate the feedback data with the original detection instruction to establish a mapping relationship library. For example, if the security report corresponding to a detection instruction receives a high reward value from the user, it means that the semantic parsing of the instruction is relatively accurate, and this experience can be used as a reference when processing similar instructions in the future; if the report corresponding to an instruction receives a low reward value, it means that there may be problems with the semantic parsing of the instruction, and it needs to be adjusted in the subsequent parsing. Through this mapping relationship library, in the subsequent semantic parsing process of the detection instruction, historical feedback information can be combined to optimize the parsing results and improve the accuracy of instruction understanding.

[0111] Steps 600-1000 provide a model optimization and instruction parsing optimization mechanism based on user feedback. By pushing reports, collecting feedback, converting reward values, fine-tuning models, and establishing mapping relationship libraries, the performance of the safety decision model and the semantic parsing accuracy of the detection instructions can be continuously improved, thus better serving the safety production detection work.

[0112] As an implementation mode, the method provided by the embodiment of the present invention may further include: extracting keywords and labeling semantic roles on the modified content to identify risk dimensions that users are concerned about; analyzing the correlation between user confirmation response time and report complexity, and optimizing the presentation method of visual charts; correlating the risk dimensions that users are concerned about with the original detection instruction features to generate an instruction understanding preference vector; in the subsequent detection instruction parsing process, adjusting the expansion range of the entity association network in combination with the instruction understanding preference vector; and synchronizing the optimized parameters to the multimodal feature extraction model and the security decision-making model through an incremental learning algorithm.

[0113] Specifically, firstly, keyword extraction and semantic role labeling are performed on the user's modified content of the structured safety report to identify the risk dimensions that the user is concerned about. Keyword extraction is to find the words that best represent the core content of the text from the text. The TF-IDF (term frequency-inverse document frequency) algorithm can be used. The formula of the algorithm is TF-IDF(t, d, D) = TF(t, d) × IDF(t, D), where TF(t, d) represents the frequency of word t in document d, and IDF(t, D) represents the inverse document frequency, which is used to measure the general importance of word t in the entire document set D. For example, in a structured safety report on chemical production safety, the user modified the content to "the temperature of the reactor is too high recently, which requires special attention". The TF-IDF algorithm is used to extract "reactor" and "temperature is too high" as keywords. Semantic role labeling is to analyze the role played by each word in the sentence at the semantic level, such as agent, patient, time, place, etc., which can be achieved with the help of a pre-trained semantic role labeling model. For the above example, "reactor" is the target role, and "temperature is too high" describes the status. Through such analysis, we can identify that the user is concerned about the risk dimension of the reactor temperature.

[0114] Next, analyze the correlation between user confirmation response time and report complexity to optimize the presentation of visualization charts. Report complexity can be measured by factors such as the number of charts, data dimensions, and length of text descriptions. You can record the confirmation response time of multiple users for reports of different complexity, and then use statistical methods such as regression analysis to find the relationship between the two. For example, if it is found that users have a long confirmation response time for reports with a large number of data dimensions and complex charts, you can use data dimensionality reduction techniques, such as principal component analysis (PCA), to reduce data dimensions and simplify charts. Or use a layered display method to display key information and main conclusions first, and then expand detailed data when users need it. This can improve the efficiency of users' understanding of reports and reduce confirmation response time.

[0115] Then, the risk dimension that the user is concerned about is associated with the original detection instruction features to generate an instruction understanding preference vector. The original detection instruction features contain information such as entity distribution vectors and task weight matrices. The correlation between the risk dimension that the user is concerned about and the entity objects and task types in the detection instructions can be calculated, and the correlation values ​​are used as elements of the vector. For example, the detection instruction is to detect the equipment safety and environmental safety of the chemical production workshop. The risk dimension that the user is concerned about is that the temperature of the reactor is too high. The analysis found that the reactor belongs to the entity object in the equipment safety detection, and the high temperature is closely related to the operating status of the equipment. By calculating the correlation, it is found that the correlation between this risk dimension and the equipment safety detection task is 0.8, and the correlation with the environmental safety detection task is 0.2. In this way, the instruction understanding preference vector is generated to reflect the user's attention preference for different detection tasks and entity objects.

[0116] In the subsequent detection instruction parsing process, the expansion range of the entity association network is adjusted in combination with the instruction understanding preference vector. The entity association network is generated after the entity object set is input into the pre-trained domain knowledge graph for semantic expansion when the detection instruction is semantically parsed. According to the instruction understanding preference vector, for entity objects with high user attention, a wider semantic expansion is performed in the domain knowledge graph to increase the number of entities and relationship types associated with them. For example, since the user was concerned about the temperature of the reactor before, in the subsequent analysis of the detection instructions for chemical production safety, for the entity object of the reactor, more related entities such as the heating system, cooling system, temperature sensor, and the causal relationship and dependency relationship between them are expanded in the domain knowledge graph, so that the entity association network can more comprehensively cover the information that the user may be concerned about.

[0117] Finally, the optimized parameters are synchronized to the multimodal feature extraction model and the safety decision model through the incremental learning algorithm. Incremental learning means that when new data arrives, the model can learn and adapt to the new data without forgetting the previous learning results, based on the fact that the model has already learned a part of the data. The online gradient descent algorithm can be used to update the weight parameters of the multimodal feature extraction model and the safety decision model according to the optimized parameters. For example, in the multimodal feature extraction model, the feature distribution of the input data changes due to the adjustment of the expansion range of the entity association network. The online gradient descent algorithm is used to gradually adjust the weight parameters of components such as the visual encoder, acoustic encoder, and state encoder according to the new feature distribution and optimized parameters, so that the model can better process the new input data. For the safety decision model, the parameters of components such as the convolutional recurrent network and the pattern matching layer are also updated according to the optimized parameters to improve the model's abnormal pattern recognition and risk assessment capabilities. Through such an incremental learning process, the co-evolution of multimodal feature extraction and risk prediction is achieved, so that the model can better meet the needs of users and improve the accuracy and efficiency of safety production detection data processing.

[0118] By executing the above steps, we can make full use of users' feedback on structured safety reports, optimize the model and detection instruction parsing process, form a closed-loop optimization system, and continuously improve the performance of production safety detection data processing and user experience.

[0119] An embodiment of the present invention provides a computer system, such as Figure 2As shown, the computer system 100 includes: a processor 101 and a memory 103. The processor 101 and the memory 103 are connected, such as through a bus 102. Optionally, the computer system 100 may also include a transceiver 104. It should be noted that in actual applications, the transceiver 104 is not limited to one, and the structure of the computer system 100 does not constitute a limitation on the embodiments of the present invention.

[0120] An embodiment of the present invention provides a computer system. The computer system in the embodiment of the present invention includes: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors. When the one or more programs are executed by the processor, the above-mentioned method provided by the embodiment of the present invention is implemented.

Claims

1. A digital-based safety production detection data processing method, characterized in that: include: Acquire safety detection data in a target scenario, and receive detection instructions for the safety detection data, wherein the safety detection data includes a continuous time series data stream collected by a multi-dimensional sensor; Performing semantic analysis on the detection instruction to generate detection instruction features including task type and entity distribution relationship; Performing cross-modal fusion processing on the security detection data through a jointly trained multi-modal feature extraction model to generate a security feature vector for each time series node; Based on the dynamic matching result of the detection instruction feature and the security feature vector of each timing node, the key data segment is screened out from the timing data stream; The pre-trained security decision model is called to perform abnormal pattern recognition and risk assessment on the key data segment, and a structured security report corresponding to the detection instruction is generated.

2. The method according to claim 1, characterized in that The semantic parsing of the detection instruction to generate detection instruction features including the relationship between the task type and the entity distribution includes: Performing syntax tree parsing on the detection instruction to extract the entity object set and the operation predicate set in the detection instruction; Inputting the entity object set into the pre-trained domain knowledge graph for semantic expansion to generate an expanded entity association network; Performing intent classification on the operation predicate set to determine the task type priority of the detection instruction; The topological features of the entity association network and the task type priority are integrated to generate a multi-dimensional detection instruction feature, wherein the detection instruction feature includes an entity distribution vector and a task weight matrix; The multi-dimensional detection instruction features are dynamically weighted through an attention mechanism to generate an optimized detection instruction feature vector.

3. The method according to claim 2, characterized in that The performing intent classification on the operation predicate set to determine the task type priority of the detection instruction includes: Performing standardization processing on the operation predicate set to generate a standardized operation predicate set; Calling a pre-trained intent recognition model to perform multi-granular semantic analysis on the standardized set of operation predicates, extracting logical dependencies and action triggering conditions between the operation predicates, and generating an initial intent label set; Performing semantic verification on the initial intention label set based on a preset safety production domain knowledge base, verifying the matching degree of each label in the initial intention label set with the standard operation scenario in the domain knowledge base, and generating a confidence level of the verified intention label; Dynamically weighting the initial intention tag set according to the verified intention tag confidence, wherein the intention tag whose matching degree with the standard operation scenario is higher than a preset verification threshold is assigned a high confidence weight, and a weighted intention tag set is generated; Inputting the weighted intention label set into a predefined task priority rule base, and mapping each intention label to a corresponding priority weight according to a mapping relationship between a security response level and an operation urgency defined in the rule base; Combined with the contextual features of the detection instruction, the entity object set and operation timeliness constraint conditions involved in the detection instruction are extracted, and the priority weight is dynamically modified; The weighted intent label set is sorted based on the corrected priority weight to generate a task type priority queue, wherein the arrangement order of each element in the queue reflects the execution urgency of the corresponding operation predicate in production safety detection, and the head element of the queue is marked as the core task type of the detection instruction.

4. The method according to claim 2, characterized in that The multimodal feature extraction model trained jointly performs cross-modal fusion processing on the security detection data to generate a security feature vector for each time series node, including: Splitting the safety detection data into visual modal data, acoustic modal data and device status modal data; Calling a visual encoder to extract spatial features from the visual modality data to generate visual features including regional thermal distribution and motion trajectory; Calling an acoustic encoder to perform spectrum decomposition and abnormal waveform detection on the acoustic modal data to generate acoustic energy distribution features; Inputting the equipment state modal data into a state encoder for operating mode recognition to generate an equipment operation state vector; The visual features, acoustic energy distribution features and device operation state vectors are aligned and fused through a cross-modal attention network to generate a security feature vector for each time series node.

5. The method according to claim 4, characterized in that The step of screening out key data segments from the time series data stream based on the dynamic matching result between the detection instruction feature and the security feature vector of each time series node includes: Based on the dynamic matching result between the detection instruction feature vector and the security feature vector of each timing node, a similarity score between each timing node and the detection instruction feature vector is calculated to generate an initial matching degree sequence; Performing a time window sliding analysis on the initial matching degree sequence, identifying candidate data segments whose similarity scores of a plurality of consecutive time series nodes in the time series data stream all exceed a dynamic threshold, and generating a set of candidate data segments; Based on the task type priority in the detection instruction feature, adjusting the length weight of each candidate data segment in the candidate data segment set to generate an optimized candidate data segment set; Input the optimized candidate data segment set into the pre-trained time series association model, analyze the context dependency of each candidate data segment in the time series data stream, remove isolated candidate data segments whose correlation with the previous and next time series nodes is less than the correlation threshold, and merge the candidate data segments with time overlap to generate a set of associated optimized data segments; According to the distribution density of each data segment in the associated optimized data segment set in the time series data stream, the associated optimized data segment set is density clustered, and the data segments with distribution density higher than a preset clustering threshold are screened out as the final key data segment set, wherein each key data segment in the final key data segment set contains a preset minimum number of continuous time series nodes and covers the set of entity objects specified in the detection instruction feature.

6. The method according to claim 5, characterized in that The calling of the pre-trained security decision model to perform abnormal pattern recognition and risk assessment on the key data segment and generate a structured security report corresponding to the detection instruction includes: Slicing the key data segments at multiple granularities to generate hierarchical data blocks including global state features and local abnormal features; Inputting the hierarchical data block into the convolutional recursive network of the security decision model to extract spatiotemporal features; The extracted spatiotemporal features are compared with a preset abnormal pattern library for similarity through a pattern matching layer in the safety decision model; Generate risk assessment indicators including risk level, abnormality type and impact range based on the comparison results; A structured safety report including visual charts and disposal suggestions is generated according to the risk assessment indicators and the task types in the detection instruction characteristics.

7. The method according to claim 6, characterized in that The collaborative optimization method of the jointly trained multimodal feature extraction model and the security decision model includes: Inputting the spatial feature sequence output by the visual encoder, the spectral feature sequence output by the acoustic encoder, and the working condition feature sequence output by the state encoder into a cross-modal attention network to generate a time-aligned multimodal joint feature sequence; Timestamp matching is performed on the multimodal joint feature sequence and the equipment failure type label in the annotated safety event to generate a joint training sample set with time series annotations; Extract features from the joint training sample set through the spatiotemporal convolution layer in the security decision-making model, and output predicted risk level and probability distribution of abnormality type; The predicted risk level is compared node by node with the environmental risk level label in the annotated safety event to generate a first error gradient, and the abnormal type probability distribution is cross-entropy calculated with the equipment failure type label to generate a second error gradient; Inputting the multimodal joint feature sequence into the discriminator network, calculating the distribution difference between the multimodal joint feature sequence and the feature template extracted from the real security event, and generating a third error gradient; Based on the weighted fusion result of the first error gradient, the second error gradient and the third error gradient, generating a joint optimization direction of the multimodal feature extraction model and the safety decision model; The weight parameters of the cross-modal attention network, the kernel parameters of the spatiotemporal convolution layer, and the discrimination threshold of the discriminator network are synchronously adjusted according to the joint optimization direction until the fusion value of the first error gradient and the second error gradient reaches a preset convergence condition.

8. The method according to claim 5, characterized in that The method of calculating the similarity score between each timing node and the detection instruction feature vector based on the dynamic matching result between the detection instruction feature vector and the security feature vector of each timing node, and generating an initial matching degree sequence, comprises: Normalizing the detection instruction feature vector and the safety feature vector of each timing node respectively to generate a normalized detection instruction feature vector and a normalized safety feature vector of the timing node; Based on the semantic coverage of the entity distribution vector in the detection instruction feature vector and the dynamic distribution of the task weight matrix, dynamically adjusting the weight of the normalized detection instruction feature vector to generate a dynamically weighted detection instruction feature vector; Perform a multi-dimensional similarity comparison between the dynamically weighted detection instruction feature vector and the normalized security feature vector of each timing node, wherein the multi-dimensional similarity comparison includes three independent calculation paths: semantic coverage overlap, entity distribution direction consistency, and task weight sensitivity; For the semantic coverage overlap calculation path, extract the entity association network topology features in the dynamically weighted detection instruction feature vector, perform regional overlap analysis with the entity space distribution features in the normalized security feature vector, and generate a semantic coverage overlap score; For the entity distribution direction consistency calculation path, a directional projection calculation is performed on the entity distribution vector in the dynamically weighted detection instruction feature vector and the entity distribution vector in the normalized security feature vector to generate an entity distribution direction consistency score; For the task weight sensitivity calculation path, perform an element-by-element product operation on the task weight matrix in the dynamically weighted detection instruction feature vector and the risk sensitivity matrix in the normalized safety feature vector to generate a task weight sensitivity score; Input the semantic coverage overlap score, entity distribution direction consistency score and task weight sensitivity score into the pre-trained fusion decision model, dynamically weight the three scores according to the task type priority in the detection instruction feature, and generate a comprehensive similarity score for each time series node; Calling a timing dependency analysis model to perform context-related correction on the comprehensive similarity score, wherein the timing dependency analysis model performs smooth correction on the comprehensive similarity score of the current timing node based on a dynamic matching relationship between a change trend of a security feature vector of an adjacent timing node and a detection instruction feature vector; The corrected comprehensive similarity scores are arranged in chronological order to generate an initial matching sequence, wherein the value of each element in the initial matching sequence reflects the dynamic matching degree between the corresponding timing node and the detection instruction feature vector, and the value changes of adjacent elements follow the context association constraints output by the timing dependency analysis model.

9. The method according to claim 1, characterized in that The method further comprises: Push the generated structured security report to the target user terminal and record the user operation log; Extracting the user's modification content of the security report and the confirmation response time as a feedback signal; Converting the feedback signal into a model optimization reward value through a reinforcement learning framework; Online fine-tuning of the risk assessment strategy in the safety decision model based on the reward value; A mapping relationship library between feedback data and original detection instructions is established for semantic parsing optimization of subsequent detection instructions.

10. A computer system, characterized in that: include: one or more processors; Memory; one or more computer programs; The one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Vector planar element matching method based on OCSVM

    CN115238800A

  • Data processing method and device, electronic equipment and readable storage medium

    CN118035217A

  • Industrial intelligent detection method and system based on multi-modal large model

    CN118503832A

  • Visual detection method based on visual large model

    CN119580163A

  • Application service reliability anomaly detection method and related equipment

    CN119806937A

Cited By

  • Shared tray data tracking method based on unmanned loading scene

    CN120612033A

  • A shared pallet data tracking method based on unmanned loading scenarios

    CN120612033B

  • AI-based security risk analysis method

    CN120634819A

  • An AI-based security risk analysis method

    CN120634819B

  • Information security detection method

    CN120750614A