Highway-oriented multi-modal event understanding method and system

Through the multimodal event understanding method, combined with the image encoder and the LLM large language model, event detection and description generation in highway scenarios are carried out, which solves the problems of insufficient accuracy and insufficient dynamic semantic adaptation in the prior art, and achieves more accurate event understanding and description.

CN120198835APending Publication Date: 2025-06-24SHANDONG HI SPEED GRP CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510253859.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient accuracy, missed detection and lack of dynamic semantic adaptation in event detection and description generation in highway scenarios, especially in complex environments.

Method used

The multimodal event understanding method is adopted to acquire and segment feature images and real-time images, and use image encoder to enhance the prompt information of local image sub-blocks, and generate real-time text description semantic information in combination with the LLM large language model. The similarity matrix and weight matrix are constructed based on local image sub-blocks and text description semantic information, visual text cross-alignment, weighted score optimization, and finally generated event understanding description information through the PDVC model.

Benefits of technology

It improves event analysis accuracy and description accuracy, can accurately identify and describe key events in high-speed scenarios, and enhances the model's understanding and generation ability of complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198835A_ABST
    Figure CN120198835A_ABST
Patent Text Reader

Abstract

The invention belongs to the cross technical field of computer vision and natural language processing, and particularly relates to a multi-modal event understanding method and system oriented to an expressway, and the method comprises the steps: obtaining a data set and a real-time image formed in the driving process of a vehicle on the expressway, and the data set comprises feature images and feature text description semantic information; segmenting the image, and enhancing prompt information of segmented image sub-blocks; aligning the text description semantic information of the local image sub-blocks; pre-training a PDVC model based on the processed data set related data, and adjusting model parameters to reduce the learning rate; and inputting the processed real-time image data into the PDVC after the model parameters are adjusted, and outputting an event description text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of computer vision and natural language processing, and more particularly relates to a multi - modal event understanding method and system for highways. Background Art

[0002] With the rapid development of intelligent transportation systems, real - time event detection and understanding in highway scenarios have become a key technology for improving the efficiency of road safety management. Traditional methods mainly rely on a single modality (such as vision or text) for event analysis, but face the following limitations in complex traffic scenarios:

[0003] Existing vision - based event detection (such as YOLO, Faster R - CNN) usually relies on image feature extraction in fixed regions, making it difficult to dynamically capture local details of unexpected events on highways (such as local damage to vehicles, foreign objects on the road surface). In addition, in complex environments such as changes in lighting and occlusion, relying solely on visual information is prone to false detection or missed detection, resulting in insufficient description accuracy.

[0004] Traditional text generation models (such as LSTM, Transformer) usually generate descriptions based on predefined event templates, lacking dynamic semantic adaptation to real - time video content.

[0005] Current technologies mostly regard event detection and description generation as independent modules, lacking collaborative optimization. For example, the literature "Multi - Modal Accident Detection on Highways" (ACM MM, 2022) proposes a two - stage framework that first detects anomalies and then generates text. However, its text generation module does not utilize local features in the detection stage, resulting in descriptions lacking details (such as the lack of association with specific lane positions in the description of "a truck rollover"). In addition, existing video description models (such as PDVC) are not adapted to the long - term temporal dependence characteristics of traffic events, and the coherence and accuracy of the generated text are insufficient. Summary of the Invention

[0006] The present invention is precisely proposed based on the above - mentioned requirements of the existing technology. The technical problem to be solved by the present invention is to provide a multi - modal event understanding method and system for highways to improve event analysis accuracy and description accuracy.

[0007] To solve the above problems, the technical solutions provided by the present invention include:

[0008] A multi-modal event understanding method for highways is provided, including: obtaining a data set; the data set includes a set of feature images containing feature images and a set of feature text description semantic information corresponding to the feature images and containing feature text description semantic information; obtaining a video formed in real time during the driving of a vehicle on a highway, and using the images of the video as a set of real-time images containing real-time images; segmenting the feature images and real-time images respectively, enhancing the prompt information of the segmented image sub-blocks through an image encoder, and obtaining multiple corresponding local image sub-blocks respectively, forming a set of feature local image sub-blocks and a set of real-time local image sub-blocks; obtaining and matching real-time text description semantic information corresponding to the images from a text library according to the real-time images, and forming a set of real-time text description semantic information based on the large language model of LLM; constructing a similarity matrix and a weight matrix based on the set of local image sub-blocks and the corresponding set of text description semantic information, and obtaining a visual-text cross-alignment weighted score based on the similarity matrix and the weight matrix, so as to obtain the corresponding relationship between the most optimal local image sub-blocks and the text description semantic information; identifying the feature local image sub-blocks through supervised contrast learning to obtain first feature data; using the CLIP model to process the feature text description semantic information corresponding to the optimal local image sub-blocks to obtain corresponding second feature data; inputting the multi-modal information of the first data feature and the second data feature into the PDVC model to generate event understanding description information to realize pre-training of the multi-modal event understanding method on a public data set, and adjusting the model parameters to reduce the learning rate; identifying the image information in the set of feature local image sub-blocks through supervised contrast learning, and performing supervised learning to classify abnormal traffic conditions in the data set obtained in a high-speed scenario, comparing the abnormal traffic conditions with an accident type database to form first information, the first information including accident types; using the CLIP model and the Transformer architecture model to process the text description semantic information corresponding to the image information in the set of local image sub-blocks in sequence to obtain corresponding second information; inputting the first information and the second information into the PDVC with adjusted model parameters, and outputting an event description text.

[0009] Preferably, the enhancing the prompt information of the segmented image sub-blocks through an image encoder includes: the image editor is expressed as H(x) = {x i = χ(x, Φ i min(w,h))|i = 1,2,…n}, where Φ i is a random variable sampled from a uniform distribution U(α,β), α and β are predefined parameters of the lower and upper limits of the image sub-block cropping size, and the output result H(x) is a set of n cropped local images x i, each image highlights different aspects of the semantic information of the visual cues in the original image. χ(x, y) is a function for segmenting image x. A random position is selected to crop image x, where y controls the output of the local patch image to a specified size. The image x ∈ R h*w*n , where h and w represent the height and width of the image respectively.

[0010] Preferably, forming a real-time text description semantic information set based on the LLM large language model includes using the LLM large language model to generate rich descriptive text descriptions {y1, y2, y3... y m} for the given real-time text description semantic information y, forming a real-time text description semantic information set, which includes the features and details of category y understood from multiple perspectives. The text description semantic information set is expressed as: where m represents the total number of generated descriptions.

[0011] Preferably, a similarity matrix is constructed based on the local image patch set H(x) and the text description semantic information set J(y), expressed as: where, s(x i , y j ) = cos(H(x i ), J(y j )) In the similarity matrix, the elements in the same row represent the similarity scores between the same local image patch and all text description semantic information, and the elements in the same column represent the similarity scores between all local image patches and the same text description semantic information.

[0012] Preferably, a weight matrix is obtained based on the image patch weights and text weights. The image patch weights are expressed as U = {u1, u2, u3... u n}, and the text weights are expressed as V = {v1, v2, v3... v m}. The weight matrix is expressed as: where The value of u i indicates the weight of the local image patch relative to the relevance of the entire image. When the value of u i is higher, it means that the local image patch contains more key information about the core content of the image.

[0013] Preferably, the visual-text cross-alignment weighted score s wca (x, y) is obtained based on the correlation matrix and the weight matrix as: The description semantic information y that maximizes the cross-alignment score k indicates that it is the most suitable description semantic information for the local image patch x k .

[0014] Preferably, the data set includes a first data set and a second data set. The first data set includes a first image set containing first images and a first text description semantic information set containing semantic information of first text descriptions corresponding to the first images. The second data set includes a second image set containing second images and a second text description semantic information set containing semantic information of second text descriptions corresponding to the second images. The first data set includes the BDD-X data set, and the second data set includes the Highway Traffic Videos data set.

[0015] Preferably, a first local image sub-block is identified through supervised contrastive learning to obtain first feature data; the CLIP model is used to process the first text description semantic information optimally corresponding to the local image sub-block to obtain corresponding second feature data; the multi-modal information such as the first data feature and the second data feature is input into the PDVC to generate event understanding description information, realizing the first pre-training of the multi-modal event understanding method on the public data set, and adjusting the model parameters to reduce the learning rate.

[0016] Preferably, a second local image sub-block is identified through supervised contrastive learning to obtain third feature data; the CLIP model is used to process the second text description semantic information optimally corresponding to the local image sub-block to obtain corresponding fourth feature data; the multi-modal information such as the third data feature and the fourth data feature is input into the PDVC to generate event understanding description information, realizing the second pre-training of the multi-modal event understanding method on the public data set, and adjusting the model parameters to reduce the learning rate.

[0017] A multi-modal event understanding system for highways is also provided, including: a dataset acquisition module for acquiring a dataset, which includes a set of feature images containing feature images and a set of feature text description semantic information corresponding to the feature images and containing semantic information of the feature text description; a data acquisition module for acquiring a video formed in real time during the driving of a vehicle on a highway, and forming the images of the video as a set of real-time images containing real-time images; an image segmentation module for segmenting the feature images and real-time images respectively, enhancing the hint information of the segmented image sub-blocks through an image encoder, and obtaining respective corresponding multiple local image sub-blocks, forming a set of feature local image sub-blocks and a set of real-time local image sub-blocks; a text description semantic information generation module for acquiring and matching, according to the real-time images, real-time text description semantic information corresponding to the images from a text library, and forming a set of real-time text description semantic information based on the large language model of LLM; a corresponding association module for constructing a similarity matrix and a weight matrix based on the set of local image sub-blocks and the corresponding set of text description semantic information, and obtaining a visual-text cross-alignment weighted score based on the similarity matrix and the weight matrix, so as to obtain the corresponding relationship between the most optimal local image sub-blocks and the text description semantic information; a PDVC model training module for identifying the feature local image sub-blocks through supervised contrast learning to obtain first feature data; using the CLIP model to process the feature text description semantic information optimally corresponding to the local image sub-blocks to obtain corresponding second feature data; inputting the multi-modal information of the first data feature and the second data feature into the PDVC model to generate event understanding description information to realize the pre-training of the multi-modal event understanding method on a public dataset, and adjusting the model parameters to reduce the learning rate; an output module for identifying the image information in the set of feature local image sub-blocks through supervised contrast learning, and performing supervised learning to classify the abnormal traffic conditions in the dataset obtained in the high-speed scenario, comparing the abnormal traffic conditions with an accident type database to form first information, where the first information includes accident types; using the CLIP model and the Transformer architecture model to process the text description semantic information corresponding in sequence to the image information in the set of local image sub-blocks to obtain corresponding second information; and inputting the first information and the second information into the PDVC after adjusting the model parameters to output an event description text.

[0018] Compared with the prior art, the research in the cross field of computer vision and natural language processing of the present invention realizes the efficient cross-alignment of vision-text through multi-modal learning and cross-modal understanding, combines different forms of data such as image vision and text information, comprehensively improves the model's understanding and generation capabilities for complex scenarios, and can accurately identify, understand, and describe key events in high-speed scenarios. By considering local visual hints, it strengthens the interactive understanding of visual information and text information in high-speed scenarios, provides accurate and comprehensive event understanding, and provides a good theoretical basis method for event understanding in highway scenarios.

[0019] Specifically:

[0020] Process the video image and text, segment the video image into blocks for local recognition, and use the large language model (LLM) to enhance the text information to generate a more refined text description.

[0021] Process of multi-modal data: Cross-align the local block images with the processed text information to improve the accuracy of accident recognition, which is more conducive to the large model to sense and recognize multi-modal information.

[0022] Analyze the data: Adopt the PDVC model, integrate visual cues and text information to generate a comprehensive understanding of the event, and output text cues.

[0023] Enhance the model effect: Adopt the method of dataset knowledge transfer time description enhancement, use similar datasets to train the model, and then perform fine-tuning to improve the application effect of the model in high-speed scenarios. Description of the Drawings

[0024] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of this specification. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0025] Figure 1 It is the method flow chart of a multi-modal event understanding method for highways in the present invention;

[0026] Figure 2 It is the schematic diagram of multi-modal data cross-alignment backbone network data processing in the present invention;

[0027] Figure 3 It is the schematic diagram of the PDVC framework in the present invention. Detailed Embodiments

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0029] In the description of the embodiments of the present invention, it should be noted that unless otherwise clearly specified and limited, the term "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection. It can be a mechanical connection or an electrical connection. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0030] Throughout the description, the terms "top", "bottom", "above", "below", and "on" are relative positions with respect to the components of the device, such as the relative positions of the top and bottom substrates inside the device. It is understood that the device is multifunctional and independent of its orientation in space.

[0031] For the convenience of understanding the embodiments of the present invention, the following will further explain with specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation to the embodiments of the present invention.

[0032] This embodiment provides a multi-modal event understanding method for highways, as Figures 1-3 shown.

[0033] The multi-modal event understanding method for highways includes:

[0034] Obtain a data set; the data set includes a feature image set containing feature images and a feature text description semantic information set containing semantic information corresponding to the feature images in the feature text description.

[0035] The data set includes a first data set and a second data set. The first data set includes a first image set containing first images and a first text description semantic information set containing semantic information corresponding to the first images in the first text description. The second data set includes a second image set containing second images and a second text description semantic information set containing semantic information corresponding to the second images in the second text description.

[0036] The first data set includes the BDD-X data set, which is constructed based on the Berkeley DeepDrive data set and focuses on generating text descriptions and explanations. It not only contains more than 77 hours of driving videos but also contains 3-4 driving behaviors, such as accelerating, decelerating, turning, etc. All these behaviors are accompanied by descriptions and explanations, and there is relatively rich vehicle-end information.

[0037] The second dataset includes the Highway Traffic Videos dataset, provided by the City University of Hong Kong, which contains highway traffic video data captured from roadside surveillance cameras, covering various weather conditions, time periods, and traffic congestion situations. At the same time, the videos in the dataset are accompanied by annotations, including vehicles, road signs, and lane markings, etc.

[0038] Both of the above two datasets have a high correlation with the local visual dataset of the studied highway scenario. Therefore, the local visual knowledge enhancement method for the dataset is feasible.

[0039] Obtain the video formed in real time during the driving process of the vehicle on the highway, and form the images of the video as a real-time image set containing real-time images.

[0040] Segment the feature image and the real-time image respectively, and enhance the prompt information of the segmented image sub-blocks through an image encoder to obtain their respective corresponding multiple local image sub-blocks, forming a feature local image sub-block set and a real-time local image sub-block set. Among them, the feature local image sub-block set includes a first local image sub-block set and a second local image sub-block set.

[0041] Specifically, the image is represented as x ∈ R h*w*n , where h and w respectively represent the height and width of the image, and n represents the number of local image sub-blocks into which the image x is divided. Correspondingly, the set of segmented local image sub-blocks is represented as {x1, x2, x3... x n}.

[0042] Segment the image x through the function χ(x, y), and randomly select a position to crop the image x, where y controls the output of the local block image to a specified size.

[0043] Enhance the prompt information extracted from the local image sub-block set through an image encoder.

[0044] The image editor is represented as:

[0045] H(x) = {x i = χ(x, Φ i min(w, h))|i = 1, 2,... n}

[0046] In the formula, Φ i is a random variable sampled from the uniform distribution U(α, β). α and β are predefined parameters for the lower and upper limits of the cropping size of the image sub-blocks. By randomly selecting the parameter Φ i, to obtain image sub - blocks that can cover different parts of the image, so as to retrieve different semantic information in diverse local image sub - blocks, which helps to enhance the model's comprehensive understanding and generalization ability of image content. The final output result H(x) is a set of n cropped local images x i , and each image highlights different aspects of the semantic information of the visual cues in the original image.

[0047] Identifying local image sub - blocks in the query image through local visual cue embedding can better integrate with the fine - grained text descriptions of each category, improving the model's understanding and processing ability of specific regions in the image.

[0048] According to the real - time image, obtain the real - time text description semantic information corresponding to the image from the text library, and form a real - time text description semantic information set based on the LLM large - language model.

[0049] The text library includes descriptions of traffic elements and judgment of event types such as traffic accidents.

[0050] Using the LLM large - language model, that is, the text encoder J(*), generate rich descriptive text descriptions {y1, y2, y3…y m} for the given real - time text description semantic information y, forming a real - time text description semantic information set. This set includes the features and details of category y understood from multiple perspectives. The real - time text description semantic information set is represented as: where m represents the total number of generated descriptions.

[0051] Utilize the powerful language understanding and generation ability of the LLM to generate better responses or refined text descriptions by inputting prompt examples or relevant knowledge, enhancing the semantic information of the input labels and realizing the enrichment of the text description semantic information.

[0052] Construct a similarity matrix and a weight matrix based on the local image sub - block set and the corresponding text description semantic information set, and obtain the visual - text cross - alignment weighted score based on the similarity matrix and the weight matrix, so as to obtain the corresponding relationship between the most optimal local image sub - blocks and the text description semantic information.

[0053] The local image sub - block set includes the first local image sub - block set, the second local image sub - block set, and the real - time local image sub - block set; the text description semantic information set includes the first text description semantic information set, the second text description semantic information set, and the real - time text description semantic information set.

[0054] Construct a similarity matrix based on the local image sub - block set H(x) and the text description semantic information set J(y), which is represented as:

[0055]

[0056] Among them, s(x i , y j ) = cos(H(x i ), J(y j ))

[0057] In the similarity matrix, the elements in the same row represent the similarity scores between the same local image sub - block and all text description semantic information, and the elements in the same column represent the similarity scores between all local image sub - blocks and the same text description semantic information.

[0058] Based on the image patch weights and text weights, a weight matrix is obtained. Among them, the image patch weights are represented as U = {u1, u2, u3... u n}}, and the text weights are represented as V = {v1, v2, v3... v m}}, and the weight matrix is represented as:

[0059]

[0060] Among them The value of u i indicates the weight of the local image sub - block with respect to the relevance of the entire image. When the value of u i is higher, it means that the local image sub - block contains more key information of the core content of the image.

[0061] v j represents the correlation between the j - th text description y j and the text description y. Similarly, when the value of v j is higher, it means that the text description has a stronger association with the label of the image and can better explain the content of the image. Essentially, the magnitude of the weight represents the importance of a specific image sub - block or text description in their respective wholes.

[0062] Based on the above - mentioned correlation matrix and weight matrix, the visual - text cross - alignment weighted score s wca (x, y) is: The text description y that maximizes the cross - alignment score k indicates that it is the most suitable text description for the local image sub - block x k .

[0063] The first local image sub - block is identified through supervised contrastive learning to obtain the first feature data; the CLIP model is used to process the semantic information of the first text description that is optimally corresponding to the local image sub - block to obtain the corresponding second feature data. Among them, the CLIP model is Contrastive Language - Image Pre - training.

[0064] Input multimodal information such as the first data feature and the second data feature into the PDVC to generate event understanding description information, and implement the first pre-training of the multimodal event understanding method on the public dataset, and adjust the model parameters to reduce the learning rate. Among them, PDVC is Progressive Dense Video Captioning, which is a deep learning model for automatically generating natural language descriptions of videos.

[0065] Identify the second local image sub-block through supervised contrast learning to obtain the third feature data; use the CLIP model to process the semantic information of the second text description that best corresponds to the local image sub-block to obtain the corresponding fourth feature data.

[0066] Input multimodal information such as the third data feature and the fourth data feature into the PDVC to generate event understanding description information, and implement the second pre-training of the multimodal event understanding method on the public dataset, and adjust the model parameters to reduce the learning rate.

[0067] Identify the image information in the real-time local image sub-block set through supervised contrast learning to obtain the corresponding feature data, and conduct supervised learning in sequence to classify the abnormal traffic conditions in the dataset obtained in the high-speed scenario, compare the abnormal conditions with the databases of different accident types to clarify the accident type, and encode the input and store it to form the first information.

[0068] Use the CLIP model and the Transformer architecture to process the semantic information of the text descriptions corresponding to the image information in the local image sub-block set in sequence to obtain the corresponding second information.

[0069] Input the first information and the second information into the PDVC with adjusted parameters to correct the PDVC framework for processing the first information and the second information, and output the event description text.

[0070] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multimodal event understanding method for highways, characterized in that: include: Acquire a data set; the data set includes a feature image set containing a feature image and a feature text description semantic information set containing feature text description semantic information corresponding to the feature image; Acquire a video formed in real time during the driving of the vehicle on the highway, and form an image of the video as a real-time image set including the real-time image; The feature image and the real-time image are segmented respectively, and the prompt information of the segmented image sub-blocks is enhanced by an image encoder to obtain a plurality of local image sub-blocks corresponding to each other, thereby forming a feature local image sub-block set and a real-time local image sub-block set; According to the real-time image, real-time text description semantic information corresponding to the image is obtained and matched from the text library, and a real-time text description semantic information set is formed based on the LLM large language model; A similarity matrix and a weight matrix are constructed based on a set of local image sub-blocks and a set of text description semantic information corresponding thereto, and a visual text cross-alignment weighted score is obtained based on the similarity matrix and the weight matrix, thereby obtaining the most preferred correspondence between the local image sub-blocks and the text description semantic information; Identify the characteristic local image sub-block through supervised contrast learning to obtain the first preset characteristic data; use the CLIP model to process the characteristic text description semantic information that best corresponds to the local image sub-block to obtain the corresponding second preset characteristic data; input the first preset data feature and the second preset data feature multimodal information into the PDVC model to generate event understanding description information to realize the pre-training of the multimodal event understanding method on the public data set, and adjust the model parameters to reduce the learning rate; The image information in the feature local image sub-block set is identified through supervised contrast learning, and supervised learning is performed to classify the abnormal traffic conditions in the data set obtained in the high-speed situation, and the abnormal traffic conditions and the accident type database are compared to form the first information, which includes the accident type; the text description semantic information corresponding to the image information in the local image sub-block set is processed in sequence using the CLIP model and the Transformer architecture model to obtain the corresponding second information; the first information and the second information are input into the PDVC after adjusting the model parameters, and the event description text is output.

2. The multimodal event understanding method for highways according to claim 1 is characterized in that: The prompt information of the image sub-block after segmentation by the image encoder includes: the image encoder is represented as H(x)={x i =χ(x,Φ i min(w,h))|i=1,2,…n} Where Φ i is a random variable sampled from a uniform distribution U(α,β), α,β are predefined parameters for the lower and upper limits of the image sub-block cropping size, and the output H(x) is a set of n cropped local images x i , each image highlights different aspects of the original image visual cue semantic information, χ(x, y) is the function of segmenting image x, and randomly selects positions to crop image x, where y controls the local block image output to a specified size, and the image x∈R h*w*n , where h and w represent the height and width of the image respectively.

3. The multimodal event understanding method for highways according to claim 1 is characterized in that: The formation of a real-time text description semantic information set based on the LLM large language model includes using the LLM large language model to generate a rich descriptive text description {y1, y2, y3…y m }, forming a real-time text description semantic information set, which includes the features and details of category y understood from multiple perspectives. The text description semantic information set is expressed as: Where m represents the total number of generated descriptions.

4. The multimodal event understanding method for highways according to claim 1, characterized in that: The similarity matrix is ​​constructed based on the local image sub-block set H(x) and the text description semantic information set J(y), which is expressed as: Among them, s(x i ,y j )=cos(H(x i ),J(y j )), in the similarity matrix, the elements in the same row represent the similarity scores between the same local image sub-block and all text description semantic information, and the elements in the same column represent the similarity scores between all local image sub-blocks and the same text description semantic information.

5. The multimodal event understanding method for highways according to claim 4 is characterized in that: The weight matrix is ​​obtained based on the image patch weight and text weight, where the image patch weight is represented as U = {u1,u2,u3…u n }, the text weight is expressed as V = {v1,v2,v3…v m }, the weight matrix is ​​expressed as: in u i The value of indicates the weight of the local image sub-block relative to the entire image. i The higher the value, the more key information the local image sub-block contains about the core content of the image.

6. The multimodal event understanding method for highways according to claim 5 is characterized in that: Based on the correlation matrix and weight matrix, the visual text cross alignment weighted score s is obtained. wca (x,y) is: The description semantic information y that maximizes the cross alignment score k Indicates that it is a local image sub-block x k The most appropriate description of semantic information.

7. The multimodal event understanding method for highways according to claim 1 is characterized in that: The data set includes a first data set and a second data set, the first data set includes a first image set containing a first image and a first text description semantic information set containing first text description semantic information corresponding to the first image, the second data set includes a second image set containing a second image and a second text description semantic information set containing second text description semantic information corresponding to the second image, the first data set includes the BDD-X data set, and the second data set includes the Highway Traffic Videos data set.

8. The multimodal event understanding method for highways according to claim 7 is characterized in that: Through supervised contrast learning, the first local image sub-block is identified to obtain the first feature data; the first text description semantic information that best corresponds to the local image sub-block is processed using the CLIP model to obtain the corresponding second feature data; multimodal information such as the first data feature and the second data feature is input into PDVC to generate event understanding description information, realize the first pre-training of the multimodal event understanding method on a public data set, and adjust the model parameters to reduce the learning rate.

9. The multimodal event understanding method for highways according to claim 7, characterized in that: The second local image sub-block is identified through supervised contrast learning to obtain the third feature data; the second text description semantic information that best corresponds to the local image sub-block is processed using the CLIP model to obtain the corresponding fourth feature data; multimodal information such as the third data feature and the fourth data feature is input into PDVC to generate event understanding description information, realize the second pre-training of the multimodal event understanding method on the public data set, and adjust the model parameters to reduce the learning rate.

10. A multimodal event understanding system for highways, characterized in that: include: A data set acquisition module, used for acquiring a data set, wherein the data set includes a feature image set containing a feature image and a feature text description semantic information set containing feature text description semantic information corresponding to the feature image; A data acquisition module is used to acquire a real-time video of a vehicle while it is traveling on a highway, and to form an image of the video as a real-time image set including the real-time image; The image segmentation module segments the feature image and the real-time image respectively, and enhances the prompt information of the segmented image sub-blocks through the image encoder to obtain a plurality of local image sub-blocks corresponding to each other, thereby forming a feature local image sub-block set and a real-time local image sub-block set; A text description semantic information generation module, which obtains and matches the real-time text description semantic information corresponding to the image from the text library according to the real-time image, and forms a real-time text description semantic information set based on the LLM large language model; The corresponding association module constructs a similarity matrix and a weight matrix based on the local image sub-block set and the corresponding text description semantic information set, and obtains the visual text cross alignment weighted score based on the similarity matrix and the weight matrix, so as to obtain the most preferred correspondence between the local image sub-block and the text description semantic information; The PDVC model training module identifies the characteristic local image sub-block through supervised contrast learning to obtain the first characteristic data; uses the CLIP model to process the characteristic text description semantic information that best corresponds to the local image sub-block to obtain the corresponding second characteristic data; inputs the first data feature and the second data feature multimodal information into the PDVC model to generate event understanding description information to realize the pre-training of the multimodal event understanding method on the public data set, and adjusts the model parameters to reduce the learning rate; The output module recognizes the image information in the feature local image sub-block set through supervised contrast learning, and performs supervised learning to classify the abnormal traffic conditions in the data set obtained under the high-speed situation, and compares the abnormal traffic conditions with the accident type database to form the first information, wherein the first information includes the accident type; uses the CLIP model and the Transformer architecture model to process the text description semantic information corresponding to the image information in the local image sub-block set in sequence to obtain the corresponding second information; inputs the first information and the second information into the PDVC after adjusting the model parameters, and outputs the event description text.

Citation Information

Cited By

  • Target detection method based on frame image and event stream feature fusion

    CN120953942A