Method and System for Identifying and Detecting Illegal Behaviors Based on Multi-Source Data Fusion Learning

By applying multi-source data fusion learning technology on safety helmets at construction sites, the real-time nature of violation monitoring and single data source problems in the existing technology are solved, and more accurate and rapid violation identification and detection are achieved.

CN119723681BActive Publication Date: 2025-07-01PING YANG XIAN CHANG TAI DIAN LI SHI YE YOU XIAN GONG SI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510234875.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-01
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing technology has problems such as high false alarm rate, missed rate and slow response in monitoring of violations at construction sites, which are mainly due to insufficient real-time and single data source.

Method used

Violation behavior identification and detection technology based on multi-source data fusion learning is adopted, and the video content description generation module, multi-modal fusion module and Pareto optimal gradient update module are enhanced to enhance the feature richness and semantic integrity of the real-time input video of the hard helmet.

Benefits of technology

The model's ability to extract semantic information of character behavior in videos is improved, the performance of identifying and detecting violations is enhanced, the false alarm and missed response rate is reduced, and the response speed is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723681B_ABST
    Figure CN119723681B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for identifying and detecting illegal behaviors based on multi-source data fusion learning, belonging to the technical field of anomaly detection. The steps include: (1) for the input video information, using a video content description generation module to generate an overall text description of the video content; (2) performing multi-source and multi-modal information fusion on the input video, key video pictures, and video description text; (3) calculating the Pareto optimality of different gradient combinations and using this combined gradient to update the entire model. By extracting and fusing the multi-modal information of the input video, the present invention improves the model's ability to extract semantic information of human behaviors in the video, thereby greatly enhancing the model's performance in identifying and detecting illegal behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for identifying and detecting illegal behaviors based on multi-source data fusion learning, and belongs to the technical field of anomaly detection. Background Art

[0002] In the construction scenarios of the power industry, the working environment is complex and full of potential risks. Illegal behaviors of construction workers, such as not wearing safety helmets, entering high-risk areas, and working fatigued, may lead to serious safety accidents. To reduce potential safety hazards, construction departments usually adopt a combination of monitoring equipment and safety management systems to monitor illegal behaviors. With the increasing complexity of construction tasks and the diversity of construction scenarios, traditional monitoring methods have gradually revealed limitations and are difficult to meet the safety requirements of modern construction sites. In recent years, the rapid development of computer vision technology has provided new ideas for intelligent safety monitoring. As a portable monitoring device, the intelligent safety helmet based on computer vision can be directly applied to on-site construction workers.

[0003] Although much work has been dedicated to researching efficient intelligent safety helmets, they often have limitations in two aspects, which may lead to problems such as high false alarm rates, missed alarm rates, and slow response in monitoring illegal behaviors at construction sites:

[0004] (1) In terms of real-time performance: Construction sites often rely on security personnel for manual inspections or use static cameras for monitoring. However, this method has the disadvantages of limited coverage and poor real-time performance. For example, it is difficult to achieve round-the-clock and non-blind-spot monitoring through manual inspections, and due to the fixity of static cameras, it is difficult to capture potential hazard factors in dynamic construction scenarios. In addition, manual monitoring is easily affected by subjective human judgment, resulting in missed alarms or false alarms.

[0005] (2) In terms of single data source: Currently, most construction sites mainly adopt an intelligent monitoring technology based on a single data source. For example, video analysis is performed through cameras, and then illegal behaviors are monitored or risk judgments are made. However, this method shows obvious limitations in complex construction scenarios. When there is insufficient light, environmental occlusion, or multi-person interaction scenarios, the monitoring algorithms based on a single data source are prone to failure. In addition, a single data source often cannot comprehensively perceive environmental changes at construction sites or the subtle behavioral characteristics of construction workers, resulting in inaccurate monitoring results. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention proposes a technology and system for identifying and detecting illegal behaviors based on multi-source data fusion learning, which enhances the feature richness and semantic integrity of the real-time input video of safety helmets. First, the present invention proposes a video content description generation module to extract the description information in text modality from the input monitoring video, so as to avoid the lack of detailed semantics in visual information. Secondly, the present invention proposes an efficient multi-modal fusion module, which understands and encodes the target video from multiple dimensions by fusing the features of video, picture and text modalities. It can not only effectively capture the mutual dependence between modalities, but also optimize the interaction and fusion of information by dynamically adjusting the attention weights. Finally, based on the cross-entropy loss of the original action recognition, the present invention realizes the combined update of multi-source supervision information by calculating the Pareto direction of the multi-source feature optimization gradient, ensuring that there is no conflict between multiple supervision information.

[0007] Term Explanation:

[0008] LaViLa Video Description Generation Model: LaViLa (Language-Augmented Video-Language Pretraining) is a novel model that learns video-language representations by leveraging large language models (LLMs). This model redesigned the pre-trained LLMs to be able to process visual inputs and generate automatic video narratives through fine-tuning. The automatic narratives generated by LaViLa perform well in multiple aspects, including dense coverage of long videos, better temporal synchronization of visual information and text, and significantly improved text diversity. By contrastive learning with these additional narratives, LaViLa outperforms previous state-of-the-art methods in multiple first-person and third-person video tasks, achieving significant improvements especially on the EGTEA classification and Epic-Kitchens-100 multi-instance retrieval benchmarks.

[0009] Self-attention Mechanism: It is a powerful computational method that allows the model to focus on the relationships between different parts when processing input data. By calculating the weights for each element in the input sequence, self-attention can capture long-range dependencies, thereby improving the representation ability. In the self-attention mechanism, each element is transformed into query, key, and value vectors. The weights are generated by calculating the similarity between the query and the key, and then the weighted sum of the values is calculated. This mechanism is widely used in natural language processing, computer vision and other fields, especially in the Transformer architecture, which greatly improves the performance and flexibility of the model, making it more efficient to handle complex tasks. The self-attention mechanism can not only process variable-length sequences, but also effectively capture context information, providing strong support for various applications.

[0010] Cross-attention mechanism: An extended attention mechanism designed to enhance information interaction between different modalities or sequences. By associating the queries of one sequence with the keys and values of another sequence, cross-attention can effectively capture the relationship between the two. During the calculation, each element of the query sequence is compared with the elements of the key sequence to generate attention weights, and then the value sequence is weighted and summed to extract relevant information. This mechanism is widely used in multi-modal learning tasks, such as the combination of video and text, image caption generation, etc., and can improve the model's understanding and generation ability of complex contexts. Cross-attention helps to achieve more precise information integration, enabling the model to demonstrate stronger expressiveness and flexibility when processing diverse data.

[0011] Pareto optimality: An important concept in economics and decision theory used to describe the efficiency of resource allocation. In a resource allocation scenario, if it is impossible to reallocate resources to make at least one person better off without making others worse off, the state is called Pareto optimal. This means that in the Pareto optimal state, the allocation of resources reaches a balance point and cannot be improved without harming the interests of other individuals. Pareto optimality does not mean that everyone's welfare is optimal, but rather emphasizes achieving a balance between efficiency and fairness under limited resources.

[0012] UCF101 dataset: An extension of UCF50, containing 13,320 video clips divided into 101 categories. These 101 categories can be further divided into 5 types (body actions, interpersonal interactions, human-object interactions, playing musical instruments, and sports). The total duration of these video clips exceeds 27 hours. All videos are from YouTube, with a fixed frame rate of 25 FPS and a resolution of 320×240.

[0013] The technical solution of the present invention is as follows:

[0014] A method for identifying and detecting illegal behaviors based on multi-source data fusion learning, the steps are as follows:

[0015] (1) For the input video information, use the video content description generation module to generate an overall text description of the video content;

[0016] (2) Perform multi-source multi-modal information fusion on the input video, video key pictures, and video description text to capture high-quality video semantic information;

[0017] (3) For the hybrid representation obtained from multi-source multi-modal information fusion, refer to the idea of multi-objective optimization, calculate the Pareto optimality of different gradient combinations, and use this combined gradient to update the entire model.

[0018] Preferably according to the present invention, in step (1), it specifically includes:

[0019] (11) Based on the input video information, use the LaViLa video description generation model, and use the prompt text "Please describe the content of the video in detail" to generate a text description of the overall content of the video;

[0020] (12) Based on the input video information, use the LaViLa video description generation model, and use the prompt text "Please describe the actions of the people in the video" to generate a text description of the changes in the actions of the people in the video;

[0021] (13) Based on the two obtained text descriptions, use a large language model to combine and summarize the text. The combined text is more semantically coherent than direct splicing and can obtain context information and human behavior information from the video.

[0022] Preferably according to the present invention, in step (2), the key video image is the middle frame of the video segment.

[0023] As a further preferred solution, in step (2), the specific steps of multi-source multi-modal information fusion include:

[0024] (21) Based on the input video, key video image, and the video description text generated in step (1), use the pre-trained video encoder, image encoder, and text encoder to perform feature encoding respectively, and obtain video representations, picture representations, and text representations respectively;

[0025] (22) Based on the three modal representations obtained by encoding, input them in pairs into a multi-modal feature fusion network with shared weights;

[0026] (23) Based on the input dual-modal information, the multi-modal feature fusion network uses one modality as the query of the attention mechanism, and the other modality as the key and value. After encoding through the self-attention layer, cross-attention layer, and feed-forward network layer respectively, a dual-modal hybrid representation is obtained.

[0027] Preferably according to the present invention, in step (3), the specific steps include:

[0028] (31) Based on different dual-modal hybrid representations, calculate the cross-entropy loss function with the standard answers of human actions respectively to obtain three loss results;

[0029] (32) Based on the three losses, calculate the gradient updates caused by them on the multi-modal feature fusion network respectively, and define the loss function of the minimum convex hull problem of the three losses;

[0030] (33) Based on the defined minimum convex hull problem, the analytical solution is obtained using the Frank Wolfe Solver, and the overall gradient is updated using the resulting gradient.

[0031] Preferably according to the present invention, an optimization function is used to solve the parameters in the multi-modal feature fusion network, and the optimization function is the adam optimizer function in PyTorch.

[0032] A violation behavior recognition and detection system based on multi-source data fusion learning, comprising:

[0033] A video text description generation module, which, for the input video information, uses the video content description generation module to generate an overall text description of the video content;

[0034] A multi-source multi-modal feature fusion module, which performs multi-source multi-modal information fusion on the input video, video key pictures, and video description text;

[0035] A Pareto optimal gradient update module, which is used to calculate the Pareto optimality of different gradient combinations and update the entire model using the combined gradient.

[0036] Compared with the prior art, the beneficial effects of the present invention are:

[0037] 1. By extracting and fusing the multi-modal information of the input video, the present invention improves the ability of the model to extract semantic information of human behaviors in the video, thereby greatly enhancing the performance of the model in recognizing and detecting violation behaviors.

[0038] 2. The present invention designs a multi-source multi-modal feature hybrid network, and at the same time realizes efficient feature learning of the network by calculating the optimal Pareto combined gradient of the gradients caused by multi-modal features in different dimensions.

[0039] 3. The present invention conducts extensive experiments on the UCF101 Human Actions Dataset and verifies its effectiveness compared with the current state-of-the-art methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is the model diagram proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0041] The present invention will be further limited below in conjunction with the specification drawings and embodiments, but not limited thereto.

[0042] Embodiment 1:

[0043] This embodiment provides a method for recognizing and detecting violation behaviors based on multi-source data fusion learning, and the steps are as follows:

[0044] (1) For the input video information, use the video content description generation module to generate an overall text description of the video content. Specifically:

[0045] S11: For a video clip input v, first use the LaViLa video description generation model F LaViLa to generate the context description of the overall video. The prompt text p1 used is "Please describe the content of the video in detail", and the text description t1 is obtained:

[0046] t1 = F LaViLa (v, p1)

[0047] S12: Similar to step S11, use the prompt text p2 "Please describe the behavior of the people in the video in detail" to obtain the text description t2:

[0048] t2 = F LaViLa (v, p2)

[0049] S13: Use a large language model to summarize the two text descriptions. According to the prompt text p3 "Please summarize the content of this video based on the following two texts", obtain the overall video text description t:

[0050] t = LLM(t1, t2, p3)

[0051] Among them, LLM is a large language model, and GPT-4o is adopted.

[0052] (2) Perform multi-source and multi-modal information fusion on the input video, video key images, and video description text to capture high-quality video semantic information. Specifically:

[0053] S21: Based on the input video clip v, video key image i, and the text description obtained in step S13, use the corresponding encoders to generate single-modal representations respectively:

[0054] f v = VideoEncoder(v)

[0055] f i = ImageEncoder(i)

[0056] f t = TextEncoder(t)

[0057] Among them, f v , f i and f t are the video, image, and text features obtained through the video encoder, image encoder, and text encoder respectively;

[0058] S22: Combine the three obtained unimodal representations in pairs and input them into the multimodal feature fusion network, which consists of a self-attention layer, a cross-attention layer, and a feed-forward network layer. Specifically, for any two modal representations f1 and f2, first encode them through the self-attention layer respectively. Taking f1 as an example, the encoding process of the self-attention layer is as follows:

[0059] Q1 = W Q f1

[0060] K1 = W K f1

[0061] V1 = W V f1

[0062] Among them, W Q 、W K and W V are learnable weight matrices, and Q1, K1, and V1 are the query variable, key variable, and value variable in the self-attention layer respectively;

[0063] The output self-attention feature is:

[0064]

[0065] Similarly, obtain the self-attention feature O2 of f2. Use O1 as the query of the cross-attention layer, and O2 as the key and value of the cross-attention layer, and calculate the cross-attention value O c :

[0066] Q = W Q O1

[0067] K = W K O2

[0068] V = W V O2

[0069]

[0070] Among them, W Q 、W K and W V are learnable weight matrices, and Q, K, and V are the query variable, key variable, and value variable in the cross-attention layer respectively.

[0071] S23: The result of the cross-attention layer passes through the feed-forward network layer to obtain the final multimodal fusion feature output:

[0072] O = Linear(O c )

[0073] Among them, Linear is the feed-forward network layer.

[0074] (3) For the hybrid representation obtained by multi-source and multi-modal information fusion, referring to the idea of multi-objective optimization, calculate the Pareto optimality of different gradient combinations, and use this combined gradient to update the entire model;

[0075] It mainly includes:

[0076] S31: For any pairwise multi-modal fusion feature O, use a feed-forward network layer to map it to the predicted action category distribution

[0077]

[0078] where Linear is the feed-forward network layer.

[0079] S32: Use cross-entropy loss to calculate the gap between the human action recognition result and the standard answer. The cross-entropy loss is defined as:

[0080]

[0081] where C is the total number of defined actions, is the predicted action category after the feature passes through the mapping layer, y i is the one-hot form of the true action category. The cross-entropy loss function optimizes the network parameters by minimizing the gap between the predicted action category distribution and the true action category distribution;

[0082] S33: According to the loss calculation method in step S32, the losses obtained from pairwise combinations of different modal features can be calculated. Denote the losses obtained from different bimodal fusion features as L i , and the parameters of the multi-modal feature fusion network are θ. Define the loss function for the minimum convex hull problem of the three losses:

[0083]

[0084] Use the Frank Wolfe Solver solution method to solve the analytical solution of the above minimum convex hull problem:

[0085] α1, α2, α3 = FRANKWOLFESOLVER(θ)

[0086] where α1, α2, and α3 are the update weights of the gradients of the three losses;

[0087] S33: For the gradient changes caused by the three losses, multiply them by the solved weights respectively to update the multi-modal feature fusion network:

[0088]

[0089] where μ is the learning rate of the network.

[0090] In this embodiment, a classification experiment is carried out on the UCF101 human action recognition and detection data set. Table 1 shows the performance comparison between this embodiment and other models, and Table 2 shows the sources of each comparison model. It can be found that, overall, this embodiment can achieve the best performance.

[0091] Table 1: Performance comparison between this embodiment and other models

[0092] Algorithm Accuracy (%) R3D-18 73.16 HalluciNet 79.83 C3D 82.30 ActionFlowNet 83.90 MV-CNN 86.40 P3D 88.60 The present invention 89.10

[0093] Table 2: Sources of each comparison model

[0094]

[0095] Embodiment 2:

[0096] A violation behavior recognition and detection system based on multi-source data fusion learning, comprising:

[0097] A video text description generation module, which, for the input video information, uses the video content description generation module to generate an overall text description of the video content;

[0098] A multi-source multi-modal feature fusion module, which performs multi-source multi-modal information fusion on the input video, video key pictures, and video description text;

[0099] A Pareto optimal gradient update module, which is used to calculate the Pareto optimum of different gradient combinations and update the entire model using this combined gradient.

Claims

1. A method for identifying and detecting illegal behaviors based on multi-source data fusion learning, characterized in that: Here are the steps: (1) Based on the input video information, a video content description generation module is used to generate an overall text description of the video content; (2) Multi-source and multi-modal information fusion for input video, video key images, and video description text; S21: Based on the input video clip v, video key picture i and text description t, use the corresponding encoder to generate a unimodal representation: f v =VideoEncoder(v) f i =ImageEncoder(i) f t =TextEncoder(t) Among them, f v 、f i and f t They are the video, picture and text features obtained by the video encoder, picture encoder and text encoder respectively; S22: The three obtained single-modal representations are combined in pairs and input into the multimodal feature fusion network. The multimodal feature fusion network consists of a self-attention layer, a cross-attention layer and a feedforward network layer. Specifically, for any modal representations f1 and f2, they are first encoded by the self-attention layer respectively. The encoding process of the f1 self-attention layer is: Q1=W Q f1 K1=W K f1 V1=W V f1 Among them, W Q , W K With W V is a learnable weight matrix, Q1, K1, and V1 are the query variable, key variable, and value variable in the self-attention layer, respectively; The output self-attention features are: Similarly, we get the self-attention feature O2 of f2, use O1 as the query of the cross attention layer, and O2 as the key and value of the cross attention, and calculate the cross attention value O c : Q=W Q O1 K=W K O2 V=W V O2 Among them, W Q , W K With W V is a learnable weight matrix, Q, K, and V are the query variable, key variable, and value variable in the cross-attention layer, respectively; S23: The result of the cross attention layer passes through the feedforward network layer to obtain the final multimodal fusion feature output: O=Linear(O c ) Among them, Linear is the feedforward network layer; (3) Calculate the Pareto optimality of different gradient combinations and use this combined gradient to update the entire model; S31: For any pair of multimodal fusion features O, a feedforward network layer is used to map it to the predicted action category distribution Among them, Linear is the feed-forward network layer; S32: Use cross entropy loss to calculate the gap between the character action recognition result and the standard answer. The cross entropy loss is defined as: Where C is the total number of actions defined, is the predicted action category after the feature passes through the mapping layer, y i It is the one-hot form of the real action category. The cross entropy loss function optimizes the network parameters by minimizing the gap between the predicted action category distribution and the real action category distribution. S33: According to the loss calculation method in step S32, the loss obtained by combining two or more different modal features is calculated, and the loss obtained by combining different bimodal fusion features is recorded as L i , the parameter of the multimodal feature fusion network is θ, and the loss function of the minimum convex hull problem of the three losses is defined as: Use Frank Wolfe Solver to solve the above minimum convex hull problem: α1,α2,α3=FRANKWOLFESOLVER(θ) Among them, α1, α2, α3 are the update weights of the gradients of the three losses; S34: For the gradient changes caused by the three losses, multiply them by the solved weights respectively to update the multimodal feature fusion network: Among them, μ is the learning rate of the network.

2. The method for identifying and detecting illegal behaviors based on multi-source data fusion learning as claimed in claim 1, characterized in that: Step (1) specifically includes: (11) Based on the input video information, the LaViLa video description generation model is used, and the prompt text "Please describe the content of the video in detail" is used to generate a text description of the overall content of the video; (12) Based on the input video information, the LaViLa video description generation model is used, and the prompt text "Please describe the actions of the characters in the video" is used to generate a text description of the behavior changes of the characters in the video; (13) Based on the two text descriptions obtained, a large language model is used to perform a combined summary of the text.

3. The method for identifying and detecting illegal behaviors based on multi-source data fusion learning as claimed in claim 2, characterized in that: In step (2), the video key picture is the middle frame of the video clip.

4. The method for identifying and detecting illegal behaviors based on multi-source data fusion learning as claimed in claim 3 is characterized in that: The optimization function is used to solve the parameters in the multimodal feature fusion network.

5. A system for identifying and detecting illegal behaviors based on multi-source data fusion learning, applied to the method for identifying and detecting illegal behaviors based on multi-source data fusion learning according to claim 1, characterized in that: include: The video text description generation module generates an overall text description of the video content using the video content description generation module for the input video information; Multi-source multi-modal feature fusion module, which performs multi-source multi-modal information fusion on input video, video key images and video description text; The Pareto optimal gradient update module is used to calculate the Pareto optimality of different gradient combinations and use the combined gradient to update the entire model.

Citation Information

Patent Citations

  • Protective device for preventing falling during aerial work of distributed photovoltaic power station

    CN118047335A

  • Monitoring scene content code stream storage method and device, electronic equipment and medium

    CN118784791A