Special vehicle task state identification method and device based on cross-modal alignment

Through the special vehicle task state recognition method with image-audio cross-modal alignment, the hierarchical two-way alignment feature representation module and multimodal feature fusion network are used to solve the problem of insufficient accuracy and reliability of special vehicle recognition in complex environments, and the effective avoidance of autonomous driving in intelligent transportation systems is achieved.

CN120472681APending Publication Date: 2025-08-12CHANGAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510610009.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art lacks accuracy and reliability of special vehicle task status recognition in complex environments, the cross-modal method has rough particle size and poor dynamic adaptability, making it difficult to meet the real-time needs of intelligent transportation systems.

Method used

The special vehicle task state recognition method based on image-audio cross-modal alignment is adopted. Through a hierarchical bidirectional alignment feature representation module and multimodal feature fusion network, deep interaction and semantic consistency model between image and audio features are realized. Modal features are processed using forward and backward sequence to sequence models, and meaningful features are selected for recognition.

Benefits of technology

It significantly improves the accuracy and reliability of special vehicle identification in complex scenarios, provides technical support for autonomous driving to avoid special vehicles in intelligent traffic, and solves the problems of insufficient reliability of traditional single-modal recognition and the rough particle size and poor dynamic adaptability of existing cross-modal methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472681A_ABST
    Figure CN120472681A_ABST
Patent Text Reader

Abstract

The invention discloses a special vehicle task state recognition method based on cross-modal alignment, and the method comprises the steps: constructing an image-audio cross-modal retrieval model, and achieving the deep interaction of image features and audio features through the hierarchical extraction of features; performing multi-granularity space-time correlation analysis on vehicle appearance features in the image and siren spectrum features in the audio by using a forward and backward sequence-to-sequence model two-way architecture, and dynamically screening out cross-modal redundant information; based on attention weight adaptive fusion effective features, outputting a special vehicle task state discrimination result and cross-modal confidence assessment; according to the invention, accurate identification of the task state of the special vehicle in a complex environment is realized, and good technical support is provided for automatic driving to avoid the special vehicle which is executing the task in intelligent traffic; the problems that traditional single-mode recognition is insufficient in reliability in a complex environment, and an existing cross-mode method is rough in feature alignment granularity and poor in dynamic adaptability are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent identification and coordinated dispatch of special vehicles in intelligent transportation, and specifically relates to a method and device for special vehicle mission status identification based on image-audio cross-modal alignment. Background Art

[0002] In intelligent transportation systems, rapid and accurate identification of special vehicles (such as ambulances, fire trucks, and police cars) performing emergency missions is key to allocating road priority and dynamically regulating traffic flow. Traditional methods for identifying the mission status of special vehicles rely primarily on single-modal perception: the visual modality determines the mission status by identifying vehicle appearance features (such as warning lights, paint, and body logos), but its reliability significantly decreases at night, in rainy and snowy conditions, or in complex scenarios with occlusion. The audio modality identifies the mission status based on the detection of sirens in specific frequency bands. However, this method also faces the following challenges: environmental noise interference (such as sirens, construction noise, and traffic noise) can lead to false triggering, and a single audio modality makes it difficult to accurately extract semantic information related to the mission status. In recent years, scholars have begun to attempt to fuse image and audio data, using feature alignment of multimodal information to improve the accuracy of mission status recognition. However, existing cross-modal methods still suffer from the following major bottlenecks: Feature alignment is coarse-grained, processing image and audio signals through simple concatenation or shallow correlation, failing to deeply explore the spatiotemporal correlation between the two at a fine-grained level; They also have poor adaptability to dynamic scenarios, lacking a lightweight real-time inference framework designed for instantaneous changes in the mission status of special vehicles (such as sirens on / off, lights on / off), and struggle to meet low-latency requirements at the edge. Furthermore, radar, as an important means of traffic perception, can provide vehicle motion status (such as speed and acceleration) and environmental information (such as distance and obstacle detection), but research on its relevance to mission status recognition is still in its early stages. The semantic alignment between the spatiotemporal features of radar data and mission status has not yet been fully explored, and radar signals are easily affected by environmental noise (such as reflection interference in rainy and foggy weather), affecting recognition reliability.

[0003] Therefore, there is an urgent need for an image-audio cross-modal fusion technology for special vehicle mission status recognition. Through end-to-end deep feature interaction, multimodal semantic consistency modeling can be achieved, and ultimately a breakthrough in the accuracy and robustness of "visual-auditory" joint discrimination in complex environments can be achieved. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a method and device for special vehicle mission status recognition based on image-audio cross-modal alignment.

[0005] The present invention adopts the following technical solutions:

[0006] A special vehicle mission status recognition method based on cross-modal alignment specifically includes the following steps:

[0007] S1, using electronic equipment to receive special vehicle images and scene audio;

[0008] S2: Input the special vehicle image and audio into the trained special vehicle task state recognition model based on image-audio cross-modal alignment to output whether the image and audio match.

[0009] The training process of the special vehicle mission status recognition model based on image-audio cross-modal alignment includes:

[0010] S201, obtaining a sample set including multiple images of fire trucks, ambulances, police cars, and engineering vehicles and multiple audio clips of fire trucks, ambulances, police cars, and engineering vehicles, extracting image features and audio features to construct a multimodal feature dataset;

[0011] S202, using a multimodal feature dataset to train a pre-built special vehicle mission state recognition model to obtain a trained special vehicle mission state recognition model based on image-audio cross-modal alignment;

[0012] S3, using an electronic device to display the matching results of the image and audio obtained in step 2, so that the autonomous driving vehicle can complete the avoidance by referring to the results.

[0013] Furthermore, the specific practices of S201 include:

[0014] a1, for each special vehicle image in the sample set, dividing the special vehicle image into blocks of the same size;

[0015] b1, using the Patch Embedding in the pre-trained ViT network to obtain the patch features of each tile and the pre-trained Faster-rcnn model to obtain the visual area features of each special vehicle image;

[0016] c1, for each audio segment in the sample set, use the ResNet18 network to obtain the audio features of each audio segment;

[0017] f1, using the image block features and visual area features to form image features, and constructing a multimodal feature dataset of the sample set with the image features and audio features.

[0018] Furthermore, the specific steps of S202 include:

[0019] a2. Define the hierarchical bidirectional alignment special vehicle feature representation module, multimodal feature fusion network, modal retrieval module, and optimization loss function based on image-audio cross-modal alignment;

[0020] b2, in the current iteration cycle, selecting a set of image features of special vehicles from the multimodal feature dataset as input, inputting the features into the hierarchical bidirectionally aligned special vehicle feature representation module, and extracting hierarchical image features and hierarchical audio features respectively;

[0021] c2, input the hierarchical image features and hierarchical audio features generated in b2 into the multimodal feature fusion network, fuse the image features and audio features through the cross-modal interaction mechanism, and generate a joint multimodal representation;

[0022] d2, inputting the joint multimodal representation into the modal retrieval module, performing cross-modal similarity measurement on image features and audio features, and generating cross-modal aligned retrieval results;

[0023] e2, based on the cross-modal retrieval results of d2, combines hierarchical graph features, hierarchical audio features, and joint multimodal representation, and uses the optimized loss function based on image-audio cross-modal alignment to calculate the cross-modal alignment loss value of the current iteration;

[0024] f2, adjusting the parameters of the special vehicle mission state recognition model according to the cross-modal alignment loss value through the back-propagation algorithm, and reselecting new special vehicle image features and corresponding audio feature pairs from the multimodal feature dataset as input;

[0025] g2, repeat b2 to f2 until the special vehicle mission status recognition model converges or reaches the preset number of iterations, and obtain the trained special vehicle mission status recognition model based on image-audio cross-modal alignment.

[0026] Furthermore, b2 specifically includes:

[0027] b21, in the current iteration, the special vehicle mission status recognition model receives the preprocessed image features F i and audio feature F a As input, in the hierarchical bidirectional alignment representation module of the special vehicle mission state recognition model, the image feature F i and audio feature F a First, through the first Mamba encoding, deep image features i are extracted deep and audio feature a deep , then, the deep image feature i deep and deep audio features a deep The two sequences are then processed by forward and reverse bidirectional SSM to obtain the deep semantic features of the two modalities i c1 and a c1After obtaining the deep semantic features, the special vehicle mission status recognition model performs a second Mamba encoding on the deep semantic features to extract the mid-level image features and mid-level audio features. Similarly, through the forward and reverse bidirectional SSM processing, the hierarchical bidirectional alignment representation module obtains the mid-level semantic features of the two modalities. c2 and a c2 Finally, the special vehicle mission status recognition model performs a third Mamba processing on the features after the second Mamba encoding to extract shallow image features and shallow audio features, and obtains the shallow semantic features of the two modalities through forward and reverse bidirectional SSM processing. c3 and a c3 ;

[0028] b22, the image feature F preprocessed by b21 i and audio feature F a The special vehicle mission state recognition model is input respectively, and the single-modal image feature i is obtained through a mamba encoding. intra and single-mode audio features a intra .

[0029] Furthermore, c2 specifically includes:

[0030] c21, the deep semantic features i described in b21 c1 and a c1 , middle-level semantic features i c2 and a c2 , shallow semantic features i c3 and a c3 The cross-modal image features i of each modality are obtained by weighted summation cross and cross-modal audio features a cross , set weights α1, α2, α3, expressed as:

[0031] i cross =α1i c1 +α2i c2 +α3i c3

[0032] a cross =α1a c1 +α2a c2 +α3a c3

[0033] c22, using a multimodal fusion strategy to combine cross-modal image features i cross With the single-modal image feature i intra Fusion is performed to obtain the final image feature I. Similarly, the cross-modal audio feature a cross With unimodal audio features a intra are also fused to obtain the final audio feature A.

[0034] A special vehicle mission status identification device based on cross-modal alignment includes: a receiving module, a cross-modal retrieval module and a display module arranged on an electronic device;

[0035] The receiving module is configured to receive a special vehicle image and audio in a current driving scenario;

[0036] The cross-modal retrieval module is configured to input the received special vehicle image and audio into a trained special vehicle mission state recognition model based on image-audio cross-modal alignment to output a matching result;

[0037] The display module is configured to use an electronic device to display the image and audio matching results so that the autonomous driving vehicle can complete the avoidance by referring to the results.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] (1) This paper designs a cross-modal image-audio retrieval technology for special vehicle mission status recognition. By deeply integrating visual and auditory information, it significantly improves the accuracy and reliability of special vehicle recognition in complex scenarios.

[0040] (2) The present invention designs a hierarchical bidirectional alignment representation module, which enables the model to capture and utilize this contextual information at different levels, thereby significantly improving the accuracy and reliability of special vehicle mission status recognition. By using forward and backward sequence-to-sequence models (SSM) to process modal features, it is possible to more effectively filter out audio-irrelevant or redundant features in the image, retaining more meaningful features for subsequent semantic alignment, thereby improving the utilization of cross-modal features.

[0041] (3) The present invention solves the problems of insufficient reliability of traditional single-modal recognition in complex environments, coarse granularity of feature alignment of existing cross-modal methods, and poor dynamic adaptability, and achieves the accuracy of special vehicle task status recognition in complex environments, providing good technical support for automatic driving to avoid special vehicles that are performing tasks in intelligent transportation. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 Schematic diagram of the process of the special vehicle mission status recognition method based on image-audio cross-modal alignment provided by the present invention;

[0043] Figure 2 This is a special vehicle classification diagram and corresponding audio example diagram provided by the present invention;

[0044] Figure 31 is a schematic diagram of the training of a special vehicle mission state recognition model based on image-audio cross-modal alignment provided by the present invention;

[0045] Figure 4 This is a schematic diagram of a special vehicle mission status recognition device based on image-audio cross-modal alignment provided by the present invention. DETAILED DESCRIPTION

[0046] The present invention will be described in further detail below with reference to the accompanying drawings and specific examples, but the embodiments of the present invention are not limited thereto.

[0047] like Figure 1 As shown, a special vehicle mission status recognition method based on cross-modal alignment specifically includes the following steps:

[0048] S1, using electronic equipment to receive special vehicle images and scene audio;

[0049] S2: Input the special vehicle image and audio into the trained special vehicle task state recognition model based on image-audio cross-modal alignment to output whether the image and audio match.

[0050] The training process of the special vehicle mission status recognition model based on image-audio cross-modal alignment includes:

[0051] S201, obtaining a sample set including multiple images of fire trucks, ambulances, police cars, and engineering vehicles and multiple audio clips of fire trucks, ambulances, police cars, and engineering vehicles, extracting image features and audio features to construct a multimodal feature dataset;

[0052] S202, using a multimodal feature dataset to train a pre-built special vehicle mission state recognition model to obtain a trained special vehicle mission state recognition model based on image-audio cross-modal alignment;

[0053] For example, Figure 2 As shown, first, any special vehicle image is selected, and the corresponding special vehicle alarm audio is annotated for each special vehicle image to create a sample set of special vehicle images. The sample set contains 1078 characteristic vehicle images, and each special vehicle image corresponds to 1 special vehicle audio;

[0054] This step feeds the special vehicle audio into a trained special vehicle task state recognition model using image-audio cross-modal alignment. The model outputs a probability distribution of the previous k audio segments corresponding to the current special vehicle image. The specific process is consistent with the probability distribution of the output audio during training. Finally, this step selects the audio with the highest probability as the corresponding match.

[0055] S3, using an electronic device to display the matching results of the image and audio obtained in step 2, so that the autonomous driving vehicle can complete the avoidance by referring to the results.

[0056] In an optional embodiment of the present invention, extracting image features and audio features to construct a multimodal feature dataset includes:

[0057] a1, for each special vehicle image in the sample set, dividing the special vehicle image into blocks of the same size;

[0058] b1, using the Patch Embedding in the pre-trained ViT network to obtain the patch features of each tile and the pre-trained Faster-rcnn model to obtain the visual area features of each special vehicle image;

[0059] c1, for each audio segment in the sample set, use the ResNet18 network to obtain the audio features of each audio segment;

[0060] f1, using the image block features and visual area features to form image features, and constructing a multimodal feature dataset of the sample set with the image features and audio features.

[0061] In an optional embodiment of the present invention, Figure 3 As shown, the special vehicle mission state recognition model trained by hierarchical bidirectional alignment and multimodal fusion includes:

[0062] a2. Define the hierarchical bidirectional alignment special vehicle feature representation module, multimodal feature fusion network, modal retrieval module, and optimization loss function based on image-audio cross-modal alignment;

[0063] b2, in the current iteration cycle, selecting a set of special vehicle image features and corresponding audio features from the multimodal feature dataset as input, inputting the inputs into the hierarchical bidirectionally aligned special vehicle feature representation module, and extracting hierarchical image features and hierarchical audio features respectively;

[0064] c2, input the hierarchical image features and hierarchical audio features generated in b2 into the multimodal feature fusion network, fuse the image features and audio features through the cross-modal interaction mechanism, and generate a joint multimodal representation;

[0065] d2, inputting the joint multimodal representation into the modal retrieval module, performing cross-modal similarity measurement on image features and audio features, and generating cross-modal aligned retrieval results;

[0066] e2, based on the cross-modal retrieval results of d2, combines hierarchical graph features, hierarchical audio features, and joint multimodal representation, and uses the optimized loss function based on image-audio cross-modal alignment to calculate the cross-modal alignment loss value of the current iteration;

[0067] f2, adjusting the parameters of the special vehicle mission state recognition model according to the cross-modal alignment loss value through the back-propagation algorithm, and reselecting new special vehicle image features and corresponding audio feature pairs from the multimodal feature dataset as input;

[0068] g2, repeat b2 to f2 until the special vehicle mission status recognition model converges or reaches the preset number of iterations, and obtain the trained special vehicle mission status recognition model based on image-audio cross-modal alignment.

[0069] In an optional embodiment of the present invention, b2 includes:

[0070] b21, in the current iteration, the special vehicle mission status recognition model receives the preprocessed image features F i and audio feature F a As input, in the hierarchical bidirectional alignment representation module of the special vehicle mission state recognition model, the image feature F i and audio feature F a First, through the first Mamba encoding, deep image features i are extracted deep and audio feature a deep , then, the deep image feature i deep and deep audio features a deep The two sequences are then processed by forward and reverse bidirectional SSM to obtain the deep semantic features of the two modalities i c1 and a c1 After obtaining the deep semantic features, the special vehicle mission status recognition model performs a second Mamba encoding on the deep semantic features to extract the mid-level image features and mid-level audio features. Similarly, through the forward and reverse bidirectional SSM processing, the hierarchical bidirectional alignment representation module obtains the mid-level semantic features of the two modalities. c2 and a c2 Finally, the special vehicle mission status recognition model performs a third Mamba processing on the features after the second Mamba encoding to extract shallow image features and shallow audio features, and obtains the shallow semantic features of the two modalities through forward and reverse bidirectional SSM processing. c3 and a c3 ;

[0071] b22, the preprocessed image feature F mentioned in b21 i and audio feature F a The special vehicle mission state recognition model is input respectively, and the single-modal image feature i is obtained through a mamba encoding. intra and single-mode audio features a intra .

[0072] In an optional embodiment of the present invention, c2 includes:

[0073] c21, the deep semantic features i described in b21 c1 and a c1 , middle-level semantic features i c2 and a c2 , shallow semantic features i c3 and a c3 The cross-modal image features i of each modality are obtained by weighted summation cross and cross-modal audio features a cross , set weights α1, α2, α3, expressed as:

[0074] i cross =α1i c1 +α2i c2 +α3i c3

[0075] a cross =α1a c1 +α2a c2 +α3a c3

[0076] c22, using a multimodal fusion strategy to combine cross-modal image features i cross With the single-modal image feature i intra Fusion is performed to obtain the final image feature I. Similarly, the cross-modal audio feature a cross With unimodal audio features a intra are also fused to obtain the final audio feature A.

[0077] like Figure 4 As shown, the present invention also provides a special vehicle mission status identification device based on cross-modal alignment, comprising: a receiving module, a cross-modal retrieval module and a display module arranged on an electronic device;

[0078] A receiving module configured to receive a special vehicle image and audio in a current driving scenario;

[0079] a cross-modal retrieval module configured to input the received special vehicle image and audio into a trained special vehicle mission state recognition model based on image-audio cross-modal alignment to output a matching result;

[0080] The display module is configured to display the image and audio matching results using an electronic device so that the autonomous driving vehicle can refer to the results to complete the avoidance.

[0081] It is worth noting that the terms "first" and "second" in this disclosure are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of this disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0082] Although the present application is described herein with reference to various embodiments, those skilled in the art will be able to understand and implement other variations of the disclosed embodiments in practicing the claimed application by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality.

[0083] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A special vehicle mission status recognition method based on cross-modal alignment, characterized in that: The specific steps include: S1, using electronic equipment to receive special vehicle images and scene audio; S2: Input the special vehicle image and audio into the trained special vehicle task state recognition model based on image-audio cross-modal alignment to output whether the image and audio match. The training process of the special vehicle mission status recognition model based on image-audio cross-modal alignment includes: S201, obtaining a sample set including multiple images of fire trucks, ambulances, police cars, and engineering vehicles and multiple audio clips of fire trucks, ambulances, police cars, and engineering vehicles, extracting image features and audio features to construct a multimodal feature dataset; S202, using a multimodal feature dataset to train a pre-built special vehicle mission state recognition model to obtain a trained special vehicle mission state recognition model based on image-audio cross-modal alignment; S3, using an electronic device to display the matching results of the image and audio obtained in step 2, so that the autonomous driving vehicle can complete the avoidance by referring to the results.

2. A special vehicle mission status recognition method based on cross-modal alignment according to claim 1, characterized in that: Specific practices of S201 include: a1, for each special vehicle image in the sample set, dividing the special vehicle image into blocks of the same size; b1, using the Patch Embedding in the pre-trained ViT network to obtain the patch features of each tile and the pre-trained Faster-rcnn model to obtain the visual area features of each special vehicle image; c1, for each audio segment in the sample set, use the ResNet18 network to obtain the audio features of each audio segment; f1, using the image block features and visual area features to form image features, and constructing a multimodal feature dataset of the sample set with the image features and audio features.

3. The method for identifying special vehicle mission status based on cross-modal alignment according to claim 2, characterized in that: Specific practices of S202 include: a2. Define the hierarchical bidirectional alignment special vehicle feature representation module, multimodal feature fusion network, modal retrieval module, and optimization loss function based on image-audio cross-modal alignment; b2, in the current iteration cycle, selecting a set of image features of special vehicles from the multimodal feature dataset as input, inputting the features into the hierarchical bidirectionally aligned special vehicle feature representation module, and extracting hierarchical image features and hierarchical audio features respectively; c2, input the hierarchical image features and hierarchical audio features generated in b2 into the multimodal feature fusion network, fuse the image features and audio features through the cross-modal interaction mechanism, and generate a joint multimodal representation; d2, inputting the joint multimodal representation into the modal retrieval module, performing cross-modal similarity measurement on image features and audio features, and generating cross-modal aligned retrieval results; e2, based on the cross-modal retrieval results of d2, combines hierarchical graph features, hierarchical audio features, and joint multimodal representation, and uses the optimized loss function based on image-audio cross-modal alignment to calculate the cross-modal alignment loss value of the current iteration; f2, adjusting the parameters of the special vehicle mission state recognition model according to the cross-modal alignment loss value through the back-propagation algorithm, and reselecting new special vehicle image features and corresponding audio feature pairs from the multimodal feature dataset as input; g2, repeat b2 to f2 until the special vehicle mission status recognition model converges or reaches the preset number of iterations, and obtain the trained special vehicle mission status recognition model based on image-audio cross-modal alignment.

4. The method for identifying special vehicle mission status based on cross-modal alignment according to claim 3 is characterized in that: b2 specifically includes: b21, in the current iteration, the special vehicle mission status recognition model receives the preprocessed image features F i and audio feature F a As input, in the hierarchical bidirectional alignment representation module of the special vehicle mission state recognition model, the image feature F i and audio feature F a First, through the first Mamba encoding, deep image features i are extracted deep and audio feature a deep , then, the deep image feature i deep and deep audio features a deep The two sequences are then processed by forward and reverse bidirectional SSM to obtain the deep semantic features of the two modalities i c1 and a c1 After obtaining the deep semantic features, the special vehicle mission status recognition model performs a second Mamba encoding on the deep semantic features to extract the mid-level image features and mid-level audio features. Similarly, through the forward and reverse bidirectional SSM processing, the hierarchical bidirectional alignment representation module obtains the mid-level semantic features of the two modalities. c2 and a c2 Finally, the special vehicle mission status recognition model performs a third Mamba processing on the features after the second Mamba encoding to extract shallow image features and shallow audio features, and obtains the shallow semantic features of the two modalities through forward and reverse bidirectional SSM processing. c3 and a c3 ; b22, the image feature F preprocessed by b21 i and audio feature F a The special vehicle mission state recognition model is input respectively, and the single-modal image feature i is obtained through a mamba encoding. intra and single-mode audio features a intra .

5. The method for identifying special vehicle mission status based on cross-modal alignment according to claim 4, characterized in that: C2 specifically includes: c21, the deep semantic features i described in b21 c1 and a c1 , middle-level semantic features i c2 and a c2 , shallow semantic features i c3 and a c3 The cross-modal image features i of each modality are obtained by weighted summation cross and cross-modal audio features a cross , set weights α1, α2, α3, expressed as: I cross =α1i c1 +α2i c2 +α3i c3 a cross =α1a c1 +α2a c2 +α3a c3 c22, using a multimodal fusion strategy to combine cross-modal image features i cross With the single-modal image feature i intra Fusion is performed to obtain the final image feature I. Similarly, the cross-modal audio feature a cross With unimodal audio features a intra are also fused to obtain the final audio feature A.

6. A special vehicle mission status recognition device based on cross-modal alignment, characterized in that: include: A receiving module, a cross-modal retrieval module, and a display module are provided on the electronic device; The receiving module is configured to receive a special vehicle image and audio in a current driving scenario; The cross-modal retrieval module is configured to input the received special vehicle image and audio into a trained special vehicle mission state recognition model based on image-audio cross-modal alignment to output a matching result; The display module is configured to use an electronic device to display the image and audio matching results so that the autonomous driving vehicle can complete the avoidance by referring to the results.