Road side-based traffic information processing method, unit and system

Through fine-tuning of visual language model of roadside equipment and lightweight model training, the problem of insufficient understanding of complex traffic scenarios by roadside perception equipment is solved, and an in-depth understanding and expression of traffic information is achieved, supporting vehicle-road collaboration and traffic management.

CN120299238APending Publication Date: 2025-07-11CONTINENTAL HOLDING CHINA CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510408694.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Roadside perception equipment lacks the ability to understand and express complex traffic scenarios, making it difficult to achieve efficient traffic information processing.

Method used

The visual language model is used to fine-tune the roadside equipment, combined with lightweight model training, to form a new model that can understand various traffic scenarios, and to process traffic information through the roadside edge computing unit.

Benefits of technology

It improves the understanding and expression ability of complex traffic scenarios, can output traffic information related to prompts, and supports vehicle-road collaboration and traffic management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299238A_ABST
    Figure CN120299238A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a road side-based traffic information processing method, unit and system, which can improve the ability to understand and express complex traffic scenes. The method comprises the following steps: acquiring image data of a traffic scene acquired from a roadside; obtaining a first prompt word; and according to the image data and a first prompt word, utilizing a first model to understand a traffic scene represented by the image data, and outputting traffic information related to the first prompt word, wherein the first model is a model capable of understanding various traffic scenes after a visual language model is adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to intelligent transportation technologies, and particularly to a method, unit, and system for processing traffic information based on the roadside. Background Art

[0002] Roadside perception devices (such as roadside radars, cameras, etc.) can break through the vision limitations of in-vehicle sensors and achieve wider and unobstructed road monitoring due to their high-position deployment advantages. For example, roadside devices can cover multiple lanes and intersection areas and detect traffic participants in hidden areas in real time. However, the perception at the roadside end is usually simple object perception and prediction, lacking the ability to understand and express complex traffic scenarios. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a method, unit, and system for processing traffic information based on the roadside, which can improve the ability to understand and express complex traffic scenarios.

[0004] A method for processing traffic information based on the roadside according to an embodiment of the present invention includes: acquiring image data of a traffic scenario collected from the roadside; acquiring a first prompt; and using a first model to understand the traffic scenario represented by the image data according to the image data and the first prompt, and outputting traffic information related to the first prompt; wherein the first model is a model obtained by adjusting a vision-language model and capable of understanding various traffic scenarios.

[0005] In some embodiments, the first model is a model obtained by fine-tuning the vision-language model using a lightweight model and capable of understanding various traffic scenarios.

[0006] In some embodiments, during the fine-tuning, a second prompt, a third prompt, and labeled image data are input into the vision-language model and the lightweight model to train the lightweight model.

[0007] In some embodiments, the second prompt is a prompt related to traffic rules, and through the training, the second prompt becomes the knowledge base of the vision-language model, and the third prompt is similar to the first prompt.

[0008] In some embodiments, the labeled image data is an image reflecting a traffic scenario collected from the roadside.

[0009] In some embodiments, the traffic information includes at least one of the following: road hazard situations, violation situations, and traffic flow situations.

[0010] In some embodiments, the method further includes: sending the traffic information to the corresponding vehicle by wireless communication.

[0011] A roadside edge computing unit according to an embodiment of the present invention includes: a first receiving module for receiving image data of a traffic scene collected by sensors installed on the roadside; a second receiving module for receiving an input first prompt; and a pre-deployed first model, which is a model obtained by adjusting a vision-language model and capable of understanding various traffic scenes, for understanding the traffic scene represented by the image data according to the first prompt and outputting traffic information related to the first prompt based on the understanding.

[0012] In some embodiments, the first model is specifically a model obtained by fine-tuning the vision-language model using a lightweight model and capable of understanding various traffic scenes.

[0013] In some embodiments, during the fine-tuning, a second prompt, a third prompt, and labeled image data are input into the vision-language model and the lightweight model to train the lightweight model.

[0014] In some embodiments, the second prompt is a prompt related to traffic rules. Through the training, the second prompt becomes the knowledge base of the vision-language model, and the third prompt is similar to the first prompt.

[0015] In some embodiments, the labeled image data is an image reflecting a traffic scene collected from the roadside.

[0016] A vehicle-road collaborative system according to an embodiment of the present invention includes: a vehicle and the above-mentioned roadside edge computing unit; the roadside edge computing unit further includes: a wireless sending unit for sending the traffic information to the vehicle; the vehicle includes: a wireless receiving unit for receiving the traffic information; and a processing unit for controlling the vehicle according to the received traffic information.

[0017] A computer device / system according to an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method according to the embodiment of the present invention.

[0018] A computer program product according to an embodiment of the present invention includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method according to the embodiment of the present invention are implemented.

[0019] Advantageous effects of the embodiments of the present invention:

[0020] In an embodiment of the present invention, a first model (a model obtained by adjusting a vision-language model and capable of understanding various traffic scenarios) is used to understand the traffic scenario that the image data can reflect according to a prompt and the image data of the traffic scenario collected from the roadside, and based on the understanding of the traffic scenario, traffic information related to the prompt is output, such as violation information therein. Thus, compared with traditional object perception and prediction algorithms, the ability to understand and express complex traffic scenarios is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Other details and advantages of the present invention will become apparent from the following detailed description. It should be understood that the following drawings are merely illustrative and thus cannot be considered as limiting the present invention. The following will be described in detail with reference to the drawings, where:

[0022] Figure 1 is a schematic structural diagram of an embodiment of the vehicle-road collaborative system of the present invention;

[0023] Figure 2 is a schematic flowchart of an embodiment of fine-tuning the basic vision-language model of the present invention;

[0024] Figure 3 is a schematic flowchart of an embodiment of the method for processing roadside-based traffic information of the present invention;

[0025] Figure 4 is a schematic structural diagram of an embodiment of the computer device / equipment / system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not used to limit the present invention.

[0027] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. Moreover, the terms "first", "second", etc. are applicable to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein.

[0028] In the perception solution at the roadside end, it is usually simple target perception and prediction. For example, an image recognition algorithm is used to recognize targets (such as vehicles) in the images collected by roadside sensors, and their future trajectories are predicted based on the recognition results to achieve functions such as risk warning. However, such a solution lacks the ability to understand and express complex traffic scenarios. For example, it lacks the understanding of special traffic control, traffic police gestures, traffic rules and laws, etc.

[0029] Vision-Language Models (VLMs) are a branch of large language models (such as ChatGPT, Wenxin Yiyan), belonging to multi-modal large language models, which combine two modalities: vision and language. Common vision-language models can include, for example, CLIP, BLIP, and LLaVA, etc. Vision-language models can not only input text information but also input image information, and then understand user needs and scenarios based on the dual characteristics of text and image, and output results that meet user needs. Vision-language models are based on a very large knowledge reserve and will have a deeper understanding of complex scenarios shown in images than general image recognition algorithms.

[0030] In the embodiments of the present invention, a vision-language model is added to the perception at the roadside end to utilize the vision-language model to enhance the roadside end's understanding and expression of traffic scenarios, especially the understanding and expression of complex or special traffic scenarios. Since general vision-language models cannot directly achieve the understanding of traffic scenarios or the understanding is inaccurate, before deploying the vision-language model in this embodiment, the vision-language model can be fine-tuned to enable it to have the ability to accurately understand various traffic scenarios.

[0031] The traffic information obtained by the embodiments of the present invention through the vision-language model can be used in scenarios such as vehicle-road collaboration or traffic management. For example, in vehicle-road collaboration, it can be used for vehicle autonomous driving, violation or danger reminder. The solution of the embodiments of the present invention will be illustrated by taking vehicle-road collaboration as an example below.

[0032] As Figure 1 shown, it is a schematic structural diagram of an embodiment of the vehicle-road collaboration system of the present invention. It includes: a roadside edge computing unit 10 located at the roadside, and a vehicle 20.

[0033] Among them, the roadside edge computing unit 10 includes:

[0034] The first receiving module 101 is configured to receive image data and provide it to the vision-language model 103. Among them, the image data can be, for example, an image of a road traffic scene collected by a vision sensor located on the roadside, such as a camera. For example, cameras can be respectively set in different directions at an intersection, and then images of the intersection are collected from different directions, and the images collected by cameras in different directions are fused to obtain a panoramic image of the intersection, which can reflect the current traffic conditions at the intersection (such as traffic flow, etc.). Finally, the panoramic image is provided to the first receiving module 101.

[0035] The second receiving module 102 is configured to receive a prompt and provide the prompt to the vision-language model 103. Among them, in large language models such as vision-language models, a prompt is generally a piece of text input by a user, which is used to guide the model to generate an output that meets the requirements. For example, the prompt can be "Please analyze the parking behavior in the current camera image and mark all vehicles violating the parking regulations in Shanghai". In addition, the prompt provided to the vision-language model 103 is generally in text form. If the prompt received by the second receiving module 102 is in voice form, it is converted into text and then provided to the vision-language model 103.

[0036] The vision-language model 103 is configured to understand the traffic scene that the image data can reflect according to the image data provided by the first receiving module 101 and the prompt provided by the second receiving module 102, and output traffic information related to the first prompt, such as illegal parking information.

[0037] Among them, the vision-language model 103 is an adjusted vision-language model that can understand various traffic scenes. That is to say, this vision-language model is a model retrained on the basis of a basic vision-language model for the understanding and expression of traffic scenes. Among them, directly training a vision-language model usually requires high costs and time, so it is not cost-effective. Therefore, in this embodiment, a lightweight model can be used to fine-tune the basic vision-language model, so as to obtain a new vision-language model composed of the lightweight model and the basic vision-language model. The process of fine-tuning can be referred to Figure 2 as shown and will be described later.

[0038] The wireless sending unit 104 is configured to send the traffic information output by the vision-language model 103 to the vehicle 20.

[0039] Among them, traffic information can, for example, include: road hazard situations and / or violations, etc., which can be used to alert or prompt the vehicle 20. For example, it can prompt that the vehicle currently has an illegal parking behavior. Additionally, in scenarios such as traffic management, traffic information can further include: traffic flow conditions, so as to facilitate traffic management departments to grasp road congestion and other situations. Of course, traffic information can include different contents based on different application scenarios or requirements. The above are only examples and not limitations on the embodiments of the present invention. Additionally, in some applications, the wireless transmission unit 104 may not be an essential module.

[0040] Among them, the vehicle 20 includes:

[0041] A wireless receiving unit 201, configured to receive traffic information from the roadside edge computing unit 10. Among them, the wireless receiving unit 201 and the wireless transmission unit 104 can implement wireless communication through technologies such as V2X (Vehicle to Everything) to interact data. And

[0042] A processing unit 202, configured to control the vehicle according to the traffic information received by the wireless receiving unit 201. Among them, controlling the vehicle can, for example, include: controlling the driving strategy of an autonomous vehicle, or controlling to provide a danger warning or a traffic violation reminder to the driver, and so on.

[0043] In this embodiment, a vision-language model is used to process images of road traffic scenarios to improve the understanding and expression ability of road traffic scenarios.

[0044] As Figure 2 shown, it is a schematic flowchart of an embodiment for fine-tuning a basic vision-language model of the present invention. Specifically, a lightweight model can be used to fine-tune the basic vision-language model to obtain a new vision-language model that can understand various traffic scenarios. Among them, the lightweight model can, for example, be a lightweight Transformer model, a GhostNet model, a ShuffleNet model, and so on.

[0045] During fine-tuning, first, prompts and annotated images are prepared. Among them, the prompts include two categories. One category is related to regulations such as traffic rules, for example, it can be specific traffic rules. The other category is prompts related to user needs, such as "Please find the vehicle occupying the road illegally in the image?". The images are road images collected by roadside sensors in various traffic situations, and these images are manually annotated, where the annotated content includes but is not limited to: behaviors that do not conform to specific traffic rules, traffic scenarios, traffic signs, traffic police gestures, temporary road signs, and so on.

[0046] Next, the prompt and the image related to the prompt are input into the basic vision-language model and the lightweight model respectively. The basic vision-language model and the lightweight model process the prompt and the image respectively to obtain output results. The encoding result is obtained after integrating the output results of both.

[0047] Next, the encoding result is compared with the expected result (obtained based on annotation), the loss value is calculated through the loss function, and then the model parameters of the vehicle model are adjusted by backpropagation until the loss value of the loss function is less than a specific threshold.

[0048] In the above, using prompts reflecting different user needs and a large number of images corresponding to the prompts, the lightweight model is trained, and finally a new vision-language model composed of the lightweight model and the basic vision-language model is formed. When this new vision-language model is deployed on the roadside, it can deeply understand and express the traffic scenes reflected in the road traffic images.

[0049] As Figure 3 shown, it is a schematic flowchart of an embodiment of the method for processing roadside traffic information of the present invention, which includes:

[0050] Step S30: Obtain the first prompt and the image data of the traffic scene collected from the roadside.

[0051] Step S32: According to the image data and the first prompt, use the new vision-language model to understand the traffic scene represented by the image data and output the traffic information related to the first prompt.

[0052] Among them, the new vision-language model is a model obtained by adjusting the basic vision-language model and capable of understanding various traffic scenes. Specifically, it can be a new vision-language model obtained by fine-tuning (training) the basic vision-language model with a lightweight model and capable of understanding various traffic scenes.

[0053] Among them, during fine-tuning, the second prompt, the third prompt, and the annotated image data are input into the vision-language model and the lightweight model to train the lightweight model. Among them, the second prompt is a prompt related to traffic rules. The third prompt is similar to the first prompt. In other words, the first prompt is a subset of the third prompt. Specifically, during training, various prompts will be used to train the model, and when using, the prompt input to the model should be the prompt or a combination of prompts in the prompts used during training. Of course, in large language, it is not required that the prompts be exactly the same, as long as they express similar meanings.

[0054] The method of this embodiment uses the vision-language model to understand and express the traffic information in the image, so it can improve the ability to understand and express complex traffic scenes.

[0055] As shown Figure 4 in the figure, it is a schematic structural diagram of an embodiment of the computer device / equipment / system 4 of the present invention, which includes a memory 40, a processor 42, and a computer program stored on the memory 40. The processor 42 executes the computer program to implement the steps in the method of the embodiment of the present invention.

[0056] In addition, an embodiment of the present invention further provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method of the embodiment of the present invention are implemented.

[0057] The descriptions of the above program product and the device / equipment / system embodiment are similar to the descriptions of the above method embodiment and / or the roadside edge computing unit, and have similar beneficial effects. For the technical details not disclosed in the embodiments of the program product and the device / equipment / system of the present application, please refer to the descriptions of the method and / or the roadside edge computing unit embodiment of the present application for understanding.

[0058] The above processor may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, a microprocessor, etc. It can be understood that other electronic devices capable of implementing the functions of the above processor are also possible, and the embodiments of the present application do not make specific limitations.

[0059] The above computer storage medium / memory can be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; it can also be various terminals including one or any combination of the above memories.

[0060] It should be noted that the above description is only an example, rather than a limitation of the present invention. In other embodiments of the present invention, the method may have more, fewer, or different steps, and the order, inclusion, and functional relationships between the steps may be different from those described and illustrated. For example, usually multiple steps can be combined into a single step, and a single step can also be split into multiple steps. For those of ordinary skill in the art, without creative work, the sequence changes of the steps are also within the protection scope of the present invention.

[0061] The technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which can be a personal computer, server, or network device, etc.), or a processor or microcontroller, to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0062] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments.

[0063] Although the present invention has been disclosed above with preferred embodiments, the present invention is not limited thereto. Any changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention shall be included within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the scope defined by the claims.

Claims

1. A method for processing roadside-based traffic information, characterized in that, including: obtaining image data of a traffic scene collected from the roadside; obtaining a first prompt; and according to the image data and the first prompt, using a first model to understand the traffic scene represented by the image data, and outputting traffic information related to the first prompt; wherein the first model is a model obtained by adjusting a vision-language model and capable of understanding various traffic scenes.

2. The method for processing roadside-based traffic information according to claim 1, characterized in that Specifically, the first model is a model obtained by fine-tuning the vision-language model using a lightweight model and capable of understanding various traffic scenes.

3. The method for processing roadside-based traffic information according to claim 2, wherein During the fine-tuning, a second prompt, a third prompt, and labeled image data are input into the vision-language model and the lightweight model to train the lightweight model; wherein the second prompt is a prompt related to traffic rules, and through the training, the second prompt becomes the knowledge base of the vision-language model, and the third prompt is similar to the first prompt; wherein the labeled image data is an image collected from the roadside and reflecting the traffic scene.

4. The method for processing roadside-based traffic information according to claim 1, wherein The traffic information includes at least one of the following: road hazard situations, violation situations, and traffic flow situations.

5. The method for processing roadside-based traffic information according to claim 1, wherein The method further includes: wirelessly transmitting the traffic information to a corresponding vehicle.

6. A roadside edge computing unit, characterized in that, including: a first receiving module, configured to receive image data of a traffic scene collected by a sensor installed on the roadside; a second receiving module, configured to receive an input first prompt; and a pre-deployed first model, which is a model obtained by adjusting a vision-language model and capable of understanding various traffic scenes, and is configured to understand the traffic scene represented by the image data according to the first prompt, and output traffic information related to the first prompt based on the understanding.

7. The roadside edge computing unit according to claim 6, characterized in that, Specifically, the first model is a model obtained by fine-tuning the vision-language model using a lightweight model and capable of understanding various traffic scenes; During the fine-tuning, a second prompt, a third prompt, and labeled image data are input into the vision-language model and the lightweight model to train the lightweight model; wherein the second prompt is a prompt related to traffic rules, and through the training, the second prompt becomes the knowledge base of the vision-language model, and the third prompt is similar to the first prompt; wherein the labeled image data is an image collected from the roadside and reflecting the traffic scene.

8. A vehicle-road collaborative system, characterized in that, including: a vehicle and a roadside edge computing unit as claimed in claim 6 or 7; the roadside edge computing unit further includes: a wireless transmitting unit, configured to transmit the traffic information; the vehicle includes: a wireless receiving unit, configured to receive the traffic information; and a processing unit, configured to control the vehicle according to the received traffic information.

9. A computer device / apparatus / system, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method as claimed in any one of claims 1 to 5.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method as claimed in any one of claims 1 to 5 are implemented.

Citation Information

Cited By

  • Method, apparatus, device, storage medium and program product for vehicle control

    CN121305906A