Traffic image processing method and device, equipment and storage medium

By performing feature extraction and multi-sensor data filtering on traffic images, the target traffic image is obtained and concise description text is generated, which solves the problem of low efficiency and accuracy of traffic image description model in the prior art, and achieves more efficient and accurate traffic image description.

CN119992472APending Publication Date: 2025-05-13文远京行(北京)科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411983985.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing traffic image description models are difficult to effectively identify key information when processing traffic images, resulting in a decrease in the efficiency and accuracy of the generated description text.

Method used

The preset image processing sub-model is used to extract features of traffic images, filter non-critical information with multi-sensor data to obtain target traffic images, and generate concise and focused description text through text generation sub-model.

Benefits of technology

It significantly reduces the interference of irrelevant information on the description text, improves the accuracy and relevance of the description text, making the generated description text more concise and focus on key information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992472A_ABST
    Figure CN119992472A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of traffic image processing, and discloses a traffic image processing method and device, equipment and a storage medium, which are used for improving the description capability of a traffic image description text on key information and reducing the interference of non-invalid information. The traffic image processing method comprises the following steps: performing feature extraction on a traffic image to be processed by adopting a preset image processing sub-model to obtain an object feature map; filtering non-key information based on the object feature map and multi-sensor data to obtain a target traffic image; and inputting the target traffic image and the multi-sensor data into a preset text generation sub-model to obtain a traffic image description text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of traffic image processing, and in particular to a traffic image processing method, device, equipment and storage medium. Background Art

[0002] As people pay more attention to traffic safety, road monitoring systems and driving recorders have become important components of road traffic safety monitoring. Automatically, quickly and accurately identifying and describing traffic scenes from multiple angles not only helps emergency response work, but also promotes the development of safer and smarter traffic solutions.

[0003] In related technologies, the description model of traffic images usually translates the image content into a natural language to describe the interaction between pedestrians, vehicles and other elements in the image, so as to facilitate users to quickly understand the image content, quickly search and locate images, and other intelligent interactive functions. However, in practical applications, traffic images often contain a large amount of information that is irrelevant to the target. This information not only increases the amount of data, but also may cause interference to the generated description text, reducing the efficiency and accuracy of description text generation. Summary of the invention

[0004] The present application provides a traffic image processing method, device, equipment and storage medium, which are used when the traditional traffic image description model cannot effectively identify the valid information in the traffic image, and the description of all information reduces the efficiency and accuracy of description text generation.

[0005] In a first aspect, the present application provides a traffic image processing method, comprising: using a preset image processing sub-model to extract features of a traffic image to be processed, to obtain an object feature map;

[0006] Filtering non-critical information based on the object feature map and multi-sensor data to obtain a target traffic image;

[0007] The target traffic image and the multi-sensor data are input into a preset text generation sub-model to obtain a traffic image description text.

[0008] A second aspect of the present application provides a traffic image processing device, comprising: an extraction module, configured to extract features of a traffic image to be processed using a preset image processing sub-model, and obtain an object feature map;

[0009] A filtering module, used for filtering non-critical information based on the object feature map and multi-sensor data to obtain a target traffic image;

[0010] The generation module is used to input the target traffic image and the multi-sensor data into a preset text generation sub-model to obtain a traffic image description text.

[0011] The third aspect of the present application provides a traffic image processing device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the traffic image processing device executes the above-mentioned traffic image processing method.

[0012] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned traffic image processing method.

[0013] In the technical solution provided by the present application, the target traffic image is obtained by extracting each road object in the traffic image to be processed and filtering out irrelevant information in the traffic image to be processed in combination with the multi-sensor data of the main vehicle. The descriptive text is generated based on the target traffic image. This not only reduces the processing amount of the text generation sub-model, but also significantly reduces the interference of irrelevant information on the descriptive text, making the generated descriptive text more concise and focused on key information, thereby improving the accuracy and relevance of the descriptive text. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 A schematic diagram of an embodiment of the traffic image processing method in this application;

[0015] Figure 2 is a schematic diagram of another embodiment of the traffic image processing method in the present application;

[0016] Figure 3 A schematic diagram of an embodiment of a traffic image processing device in the present application;

[0017] Figure 4 is a schematic diagram of another embodiment of the traffic image processing device in the present application;

[0018] Figure 5 It is a schematic diagram of an embodiment of a traffic image processing device in the present application. DETAILED DESCRIPTION

[0019] The present application provides a traffic image processing method, device, equipment and storage medium for improving the description capability of traffic image description text for key information and reducing the interference of non-invalid information.

[0020] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0021] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 , an embodiment of the traffic image processing method in the present application includes:

[0022] 101. Use the preset image processing sub-model to extract features of the traffic image to be processed to obtain an object feature map.

[0023] It is understandable that the execution subject of the present application may be a traffic image processing device, or a vehicle terminal, an automatic driving system or a server, which is not limited here. The present application embodiment is described by taking a vehicle terminal as the execution subject as an example.

[0024] The traffic image to be processed in this embodiment can be collected from a fixed high-altitude camera or a vehicle-mounted camera, or can be processed by intercepting key frames in a traffic video.

[0025] In this embodiment, the image processing sub-model is used to distinguish between key information and non-key information in the traffic image to be processed, and filter the non-key information in the traffic image to be processed to obtain a target traffic image including the key information in the traffic image to be processed. The image processing sub-model can be based on the network architecture of a convolutional neural network (CNN) or a recurrent neural network (RNN), such as ResNet50, to improve the accuracy of object extraction in the traffic image to be processed.

[0026] In this embodiment, the object feature map includes feature information of each object in the traffic image to be processed, such as the position, shape, size, label, etc. of key objects and non-key objects, which includes that the key object may be the main vehicle and its associated object features, wherein the associated objects are used to indicate objects that interact with the main vehicle or affect the behavior of the main vehicle, such as traffic signs in the lane where the main vehicle is located, other vehicles, pedestrians and other traffic participants, obstacles, and weather that affects the driving of the main vehicle, etc. Non-key objects are used to indicate objects that do not interact with the operation of the main vehicle or have no effect on the behavior of the main vehicle, such as buildings in the background, distant vehicles, and climate that has no obvious effect on driving.

[0027] Exemplarily, the image processing submodel includes extracting features of the traffic image to be processed through the first convolution layer to obtain a shallow feature map; performing a pooling operation on the shallow feature map through the first pooling layer to obtain a first feature map; inputting the first feature map and the shallow feature map into the second convolution layer for feature extraction to obtain a middle feature map; performing a pooling operation on the middle feature map through the second pooling layer to obtain a second feature map; inputting the second feature map and the middle feature map into the third convolution layer for feature extraction to obtain a deep feature map; processing the shallow feature map, the middle feature map and the deep feature map through the fully connected layer to obtain an object feature map. This embodiment can efficiently identify the feature information corresponding to each object in the traffic image to be processed to adapt to the influence of complex traffic scenes and various camera shooting environments. Features at different levels contain information with different focuses. The underlying features mainly reflect high-fidelity shallow feature information such as light and dark, edges, and brightness, while the deep features contain richer semantic information, providing an accurate feature description basis for subsequent text descriptions.

[0028] 102. Non-critical information is filtered based on the object feature map and multi-sensor data to obtain the target traffic image.

[0029] In this embodiment, multi-sensor data can be collected by various vehicle-mounted sensors or estimated by corresponding algorithms, such as collecting vehicle speed through a vehicle-mounted speed sensor, collecting point cloud data through a radar, collecting steering angle through an angle sensor, and collecting vehicle acceleration through an acceleration sensor.

[0030] Specifically, non-critical information in the object feature map is filtered out based on the object feature map, multi-sensor data filtering and preset screening rules to obtain the target traffic image.

[0031] Among them, the preset screening rules are used to indicate the recognition rules of objects associated with the driving behavior of the main vehicle, so as to retain important objects related to the traffic scene and remove background and irrelevant objects. For example, objects that are too far away and other non-critical information such as the background of the driving behavior of the main vehicle, the road surface, the sky, etc. are filtered out.

[0032] Exemplarily, the main vehicle and its associated objects are determined based on the object feature map and the multi-sensor data, and a cropping operation is performed to obtain a target traffic image, which includes a target area where each key object is located.

[0033] The above-mentioned determining the main vehicle and its associated objects based on the object feature map and multi-sensor data, and performing a cropping operation to obtain a target traffic image includes: determining the center point coordinates of the main vehicle and its associated objects, cropping according to the center point coordinates and the scale features corresponding to each object to obtain a candidate traffic image; and performing a standardization operation on the candidate traffic image to obtain a target traffic image. The standardization operation is to enlarge, reduce, rotate, etc. the candidate traffic image to obtain a target traffic image of a target scale, for example, the image is enlarged to 240x240.

[0034] 103. Input the target traffic image and multi-sensor data into a preset text generation sub-model to obtain a traffic image description text.

[0035] In this embodiment, the text generation sub-model can fuse image features with multi-sensor data to capture more details and contextual information. After extracting and fusing multimodal features, natural language processing technology is used to convert these features into natural language descriptions. A decoder structure with an attention mechanism can be adopted, such as a long short-term memory network (LSTM) and a Transformer, to improve the accuracy and relevance of the description.

[0036] Specifically, the key object features in the target traffic image are fused with the multi-sensor data to obtain fused feature data; description word vectors are generated one by one based on the key object features; a nonlinear activation function is used to determine the attention score corresponding to each description word vector according to each key object feature, wherein the attention score is used to characterize the importance of the word vector; and a traffic image description text is generated based on context information, each description word vector and the attention score corresponding to each description word vector.

[0037] Exemplarily, the training process of the model is: calling the image processing sub-initial model to identify the key information in each image sample in the training sample, and filtering out non-key information from each image sample to obtain a target image sample including the key information in the image sample; performing back-propagation optimization on the image processing sub-initial model through a preset first loss function until the image processing sub-initial model meets the corresponding preset stop training condition, and obtaining a set of images to be described; using the set of images to be processed to train the text generation sub-initial model, and performing back-propagation optimization on the text generation sub-initial model through a preset second loss function until the trained text generation sub-initial model meets the corresponding preset stop training condition.

[0038] In this embodiment, by extracting each road object in the traffic image to be processed and filtering irrelevant information in the traffic image to be processed in combination with the multi-sensor data of the main vehicle, a target traffic image is obtained, and a description text is generated based on the target traffic image. This not only reduces the processing amount of the text generation sub-model, but also significantly reduces the interference of irrelevant information on the description text, making the generated description text more concise and focused on key information, thereby improving the accuracy and relevance of the description text.

[0039] In practical applications, the traffic image processing method of the present application can also be further used to generate traffic video description text, see Figure 2 Another embodiment of the traffic image processing method in the present application includes:

[0040] 201. Comprehensively evaluate each frame in the traffic video to be processed based on multi-sensor data and preset weights of each sensor to obtain an external indicator score corresponding to each frame.

[0041] In this embodiment, the external indicator score is used to indicate the importance and criticality of the frame image. By screening image frames with higher external indicator scores, it can be ensured that important time periods and events in the traffic video are covered, thereby completely and accurately describing the scene of the traffic video.

[0042] Specifically, the sensor data of the two frames before and after in the traffic video to be processed are compared to obtain the difference of each sensor corresponding to each frame; weighted summation is performed based on the difference of each sensor corresponding to each frame and the preset weight of each sensor to obtain the external indicator score corresponding to each frame.

[0043] It is understandable that in video data scenarios where vehicles need to turn, meet, or brake suddenly, there are usually jumps in acceleration and steering angle. In some accident scenarios, obstacles are too close. Therefore, corresponding external indicators can be set to assist in determining key frames in a video. Among them, sensor weights can be set according to actual conditions, such as the highest weight for acceleration jumps and lower weight for steering angle jumps.

[0044] For example, assuming that in a traffic video, the acceleration change of the current frame is larger than that of the previous frame, and its acceleration score is higher than that of other image frames whose acceleration has not changed significantly, the scores of various dimensions such as vehicle speed, low steering, and obstacle distance are further evaluated, and the scores of each dimension are weighted and summed based on the weights of each sensor to screen out key frames.

[0045] 202. Select a preset number of image frames as key frames according to the external indicator score corresponding to each frame to obtain a traffic image to be processed.

[0046] Specifically, the image frames are sorted based on the size of the external indicator score corresponding to each frame, and a preset number of image frames before scoring are selected as key frames to obtain the traffic image to be processed.

[0047] The preset number can be set according to the length of the video and different model architectures, for example, the top 12 to 36 frames are taken for description text reasoning.

[0048] 203. Use a preset image processing sub-model to extract features from the traffic image to be processed to obtain an object feature map.

[0049] Step 203 can be performed with reference to step 101 and will not be described in detail here.

[0050] 204. Non-critical information is filtered based on the object feature map and multi-sensor data to obtain a target traffic image.

[0051] Specifically, a main vehicle and its associated objects are determined based on an object feature map and multi-sensor data; a connected area of ​​the main vehicle and its associated objects is determined as a key image area, and the remaining image area is downsampled to obtain a target traffic image.

[0052] Optionally, the above-mentioned determining the main vehicle and its associated objects based on the object feature map and multi-sensor data includes: determining the target position of the main vehicle based on the object feature map, and determining each road object within a first preset distance range with the target position as the center; determining the associated objects in each road object according to the steering angle, speed and acceleration of the main vehicle.

[0053] This embodiment searches for nearby road objects with the target position of the main vehicle as the center, and further analyzes whether these road objects are associated with the movement of the main vehicle, such as whether they are located on the driving trajectory of the main vehicle or have a risk of collision with the main vehicle. For example, if the main vehicle is turning left, then objects in the turning direction of the main vehicle (such as another vehicle on the left) can be used as associated objects.

[0054] Optionally, the above-mentioned downsampling of the remaining image area to obtain the target traffic image includes: setting the pixel value of the remaining image area to the target value to obtain the target traffic image. The target value can be set to 0 or other values, which can effectively reduce the amount of data while retaining the pixel data of the key area to ensure the accuracy of the description.

[0055] Optionally, the downsampling of the remaining image area to obtain the target traffic image includes: reducing the resolution of the remaining image area to obtain the target traffic image.

[0056] 205. Input the target traffic image and multi-sensor data into a preset text generation sub-model to obtain a traffic image description text.

[0057] Specifically, the target traffic image and multi-sensor data are fused to obtain fused feature data; corresponding descriptive words for the main vehicle and each associated object are generated based on the fused feature data, and a traffic image description text is generated through an attention mechanism and context information.

[0058] Optionally, key object features in the target traffic image and multi-sensor data are spliced ​​to obtain fused feature data.

[0059] It is understandable that the text generation sub-model can further integrate the audio features in the traffic video to be processed, splice the key object features, multi-sensor data and audio features in the target traffic image, and obtain fused feature data to capture more details and contextual information in the video.

[0060] 206. Perform text adjustment based on the time sequence corresponding to each key frame and the similarity of the traffic image description text corresponding to each key frame to obtain a video description text.

[0061] Specifically, the traffic image description text corresponding to each key frame is converted into a text vector, and the similarity corresponding to each text vector is calculated; the text is adjusted based on the time sequence corresponding to each key frame and the similarity corresponding to each text vector to obtain the video description text.

[0062] Use natural language processing technologies such as Word2Vec or BERT to convert the description text of each frame into a vector representation, calculate the similarity between the description text vectors of each key frame, deduplicate and clean the text with a similarity greater than a preset value, and adjust the structure and order of the text based on the timing information of each key frame in the video so that the final description text is fluent and consistent with the time process of the video.

[0063] In this embodiment, each frame image and multi-sensor data in the traffic video is comprehensively evaluated to screen out key frames, and the object feature map corresponding to each key frame is extracted; and the irrelevant information in the traffic image to be processed is filtered out in combination with the multi-sensor data of the main vehicle to obtain the target traffic image, and the description text is generated based on the target traffic image, which not only reduces the processing amount of the text generation sub-model, but also significantly reduces the interference of irrelevant information on the description text, making the generated description text more concise and focused on key information, thereby improving the accuracy and relevance of the description text; text adjustment is performed through timing and text similarity, ensuring that the finally generated video description text is logical and coherent, thereby improving the convenience for users to understand the video content.

[0064] The above describes the traffic image processing method in this application. The following describes the traffic image processing device in this application. Figure 3 , an embodiment of the traffic image processing device in the present application includes:

[0065] An extraction module 301 is used to extract features from the traffic image to be processed using a preset image processing sub-model to obtain an object feature map;

[0066] A filtering module 302, for filtering non-critical information based on the object feature map and the multi-sensor data to obtain a target traffic image;

[0067] The generation module 303 is used to input the target traffic image and multi-sensor data into a preset text generation sub-model to obtain a traffic image description text.

[0068] In this embodiment, by extracting each road object in the traffic image to be processed and filtering irrelevant information in the traffic image to be processed in combination with the multi-sensor data of the main vehicle, a target traffic image is obtained, and a description text is generated based on the target traffic image. This not only reduces the processing amount of the text generation sub-model, but also significantly reduces the interference of irrelevant information on the description text, making the generated description text more concise and focused on key information, thereby improving the accuracy and relevance of the description text.

[0069] See also Figure 4 Another embodiment of the traffic image processing device in the present application includes:

[0070] An extraction module 301 is used to extract features from the traffic image to be processed using a preset image processing sub-model to obtain an object feature map;

[0071] A filtering module 302, for filtering non-critical information based on the object feature map and the multi-sensor data to obtain a target traffic image;

[0072] The generation module 303 is used to input the target traffic image and multi-sensor data into a preset text generation sub-model to obtain a traffic image description text.

[0073] Optionally, the traffic image processing device further includes an evaluation module 304, which is used to perform a comprehensive evaluation on each frame in the traffic video to be processed according to the multi-sensor data and the preset weights of each sensor to obtain an external indicator score corresponding to each frame;

[0074] The selection module 305 is used to select a preset number of image frames as key frames according to the external indicator score corresponding to each frame to obtain the traffic image to be processed.

[0075] Optionally, the evaluation module 304 is specifically used to compare the sensor data of the two frames before and after in the traffic video to be processed to obtain the difference amount of each sensor corresponding to each frame;

[0076] The external indicator score corresponding to each frame is obtained by performing a weighted summation based on the difference of each sensor corresponding to each frame and the preset weight of each sensor.

[0077] Optionally, the traffic image processing device further includes an adjustment module 306 for performing text adjustment based on the time sequence corresponding to each key frame and the similarity of the traffic image description text corresponding to each key frame to obtain a video description text.

[0078] Optionally, the filtering module 302 is specifically used for: a determination unit 3021, used for determining a host vehicle and its associated objects based on the object feature map and the multi-sensor data;

[0079] The processing unit 3022 is used to determine the connected area of ​​the main vehicle and its associated objects as the key image area, and perform downsampling processing on the remaining image area to obtain a target traffic image.

[0080] Optionally, the processing unit 3022 is specifically configured to: determine the target position of the host vehicle based on the object feature map, and determine each road object within a first preset distance range with the target position as the center;

[0081] The associated object is determined in each road object according to the steering angle, speed and acceleration of the host vehicle.

[0082] Optionally, the generating module 303 is specifically used to: fuse the target traffic image and the multi-sensor data to obtain fused feature data;

[0083] Based on the fused feature data, the corresponding descriptive words for the main vehicle and each associated object are generated, and the traffic image description text is generated through the attention mechanism and context information.

[0084] In this embodiment, each frame image and multi-sensor data in the traffic video is comprehensively evaluated to screen out key frames, and the object feature map corresponding to each key frame is extracted; and the irrelevant information in the traffic image to be processed is filtered out in combination with the multi-sensor data of the main vehicle to obtain the target traffic image, and the description text is generated based on the target traffic image, which not only reduces the processing amount of the text generation sub-model, but also significantly reduces the interference of irrelevant information on the description text, making the generated description text more concise and focused on key information, thereby improving the accuracy and relevance of the description text; text adjustment is performed through timing and text similarity, ensuring that the finally generated video description text is logical and coherent, thereby improving the convenience for users to understand the video content.

[0085] above Figure 3 and Figure 4 The traffic image processing device in the present application is described in detail from the perspective of modular functional entities, and the traffic image processing device in the present application is described in detail from the perspective of hardware processing.

[0086] See also Figure 5As shown, the traffic image processing device includes a processor 500 and a memory 501 . The memory 501 stores machine executable instructions that can be executed by the processor 500 . The processor 500 executes the machine executable instructions to implement the above-mentioned traffic image processing method.

[0087] Further, Figure 5 The traffic image processing device shown further includes a bus 502 and a communication interface 503 , and the processor 500 , the communication interface 503 and the memory 501 are connected via the bus 502 .

[0088] The memory 501 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), for example, at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 503 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 502 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0089] The processor 500 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 500. The above processor 500 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as a hardware decoding processor to be executed, or a combination of hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 501 , and the processor 500 reads the information in the memory 501 and completes the method steps of the above-mentioned embodiment in combination with its hardware.

[0090] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the traffic image processing method.

[0091] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., and other media that can store program codes.

[0093] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A traffic image processing method, characterized in that: The traffic image processing method comprises: Use the preset image processing sub-model to extract features from the traffic image to be processed to obtain an object feature map; Filtering non-critical information based on the object feature map and multi-sensor data to obtain a target traffic image; The target traffic image and the multi-sensor data are input into a preset text generation sub-model to obtain a traffic image description text.

2. The traffic image processing method according to claim 1, characterized in that: The filtering of non-critical information based on the object feature map and multi-sensor data to obtain a target traffic image includes: Determine a host vehicle and its associated objects based on the object feature map and the multi-sensor data; The connected area of ​​the main vehicle and its associated objects is determined as a key image area, and the remaining image area is downsampled to obtain a target traffic image.

3. The traffic image processing method according to claim 2, characterized in that: The multi-sensor data includes steering angle, speed and acceleration; The determining the host vehicle and its associated objects based on the object feature map and the multi-sensor data includes: Determine the target position of the host vehicle based on the object feature map, and determine each road object within a first preset distance range with the target position as the center of the circle; An associated object is determined in each of the road objects according to the steering angle, speed and acceleration of the host vehicle.

4. The traffic image processing method according to claim 2, characterized in that: The step of inputting the target traffic image and multi-sensor data into a preset text generation sub-model to obtain a traffic image description text includes: Fusing the target traffic image and the multi-sensor data to obtain fused feature data; Based on the fused feature data, corresponding descriptive words for the main vehicle and each of the associated objects are generated, and a traffic image description text is generated through an attention mechanism and context information.

5. The traffic image processing method according to any one of claims 1 to 4, characterized in that: Before extracting features from the traffic image to be processed by using the preset image processing sub-model to obtain the object feature map, the method further includes: Comprehensively evaluate each frame in the traffic video to be processed based on multi-sensor data and preset weights of each sensor to obtain the external indicator score corresponding to each frame; A preset number of image frames are selected as key frames according to the external indicator scores corresponding to each frame to obtain the traffic image to be processed.

6. The traffic image processing method according to any one of claims 1 to 5, characterized in that: The method of comprehensively evaluating each frame in the traffic video to be processed based on the multi-sensor data and the preset weights of each sensor to obtain the external indicator score corresponding to each frame includes: The external index score corresponding to each frame of the traffic video to be processed is obtained by comprehensively evaluating each frame according to the multi-sensor data and the preset weights of each sensor, including: Compare the sensor data of the two frames before and after in the traffic video to be processed to obtain the difference of each sensor corresponding to each frame; The external indicator score corresponding to each frame is obtained by performing a weighted summation based on the difference of each sensor corresponding to each frame and the preset weight of each sensor.

7. The traffic image processing method according to claim 5, characterized in that: After inputting the target traffic image and the multi-sensor data into a preset text generation sub-model to obtain a traffic image description text, the method further includes: The text is adjusted based on the time sequence corresponding to each key frame and the similarity of the traffic image description text corresponding to each key frame to obtain the video description text.

8. A traffic image processing device, characterized in that: The traffic image processing device comprises: An extraction module, used for extracting features from the traffic image to be processed by using a preset image processing sub-model to obtain an object feature map; A filtering module, used for filtering non-critical information based on the object feature map and multi-sensor data to obtain a target traffic image; The generation module is used to input the target traffic image and the multi-sensor data into a preset text generation sub-model to obtain a traffic image description text.

9. A traffic image processing device, characterized in that: The traffic image processing device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory to enable the traffic image processing device to execute the traffic image processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instruction is read and executed, the traffic image processing method according to any one of claims 1 to 7 is executed.