Traffic scene understanding method, traffic scene management method and related products

By comprehension analysis of the global and local images of traffic scenes, high-quality understanding results are generated, and scheduling instruction sets are generated through machine learning models and small models are called, the problem of excessive resource burden in the existing intelligent traffic management system is solved, and the effective allocation of resources and system efficiency is improved.

CN120048105APending Publication Date: 2025-05-27CHINA MOBILE SHANGHAI ICT CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510129356.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

There is a lack of understanding and analysis of traffic scenarios in the existing intelligent traffic management system, which leads to all small models running simultaneously, increasing the system resource burden.

Method used

By obtaining the global image of the traffic scene, segmenting it into local images, and using a multimodal big model to understand the global and local images, generating high-quality understanding results. Then, the machine learning model is used to analyze and understand the results, generate a scheduling instruction set, and call a small model to reduce the resource burden.

Benefits of technology

It improves the accuracy of understanding of traffic scenarios, avoids all small models running simultaneously, reduces the system's resource burden, realizes effective allocation of resources, and improves the system's response speed and task processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048105A_ABST
    Figure CN120048105A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic scene understanding method, a traffic scene management method and related products, which understand a current traffic scene by using one or more multi-modal large models, improve the understanding accuracy and provide a scientific and reasonable decision basis for subsequent traffic scene management, thereby avoiding simultaneous operation of all small models and improving the traffic scene management efficiency. And the system resource burden is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of transportation technologies, and in particular, to a method for traffic scene understanding, a method for traffic scene management, and related products. Background Art

[0002] An intelligent (road) traffic management system is a system that uses intelligent traffic technologies and methods to establish for traffic control, traffic management, and traffic decision-making, and realizes comprehensive, efficient, and scientific management of traffic scenes. Currently, the management of traffic scenes lacks understanding and analysis of traffic scenes, resulting in all small models related to traffic scene management in the system running simultaneously, thus increasing the burden on system resources. Summary of the Invention

[0003] This application provides a method for traffic scene understanding, a method for traffic scene management, and related products to solve the problem in the prior art that the lack of understanding and analysis of traffic scenes leads to all small models related to traffic scene management in the system running simultaneously, thus increasing the burden on system resources.

[0004] To achieve the above object, an embodiment of this application provides a method for traffic scene understanding, including:

[0005] Obtain a global image of the current traffic scene;

[0006] Segment the global image to obtain a plurality of local images;

[0007] Use one or more multimodal large models to understand the global image and the plurality of local images to obtain a current understanding result.

[0008] The method for traffic scene understanding provided by the embodiment of this application improves the understanding accuracy by combining the global image and local images of the traffic scene for understanding and analysis of the traffic scene, can generate high-quality and multi-level understanding results, provides a scientific and reasonable decision-making basis for subsequent traffic scene management, thereby avoiding all small models from running simultaneously, and further reducing the burden on system resources.

[0009] To achieve the above object, an embodiment of this application provides a method for traffic scene management, including:

[0010] Use one or more multimodal large models to understand the current traffic scene to obtain a current understanding result;

[0011] Use a machine learning model to analyze the current understanding result to obtain a scheduling instruction set for the current traffic scene; wherein, the scheduling instruction set includes a plurality of scheduling instructions, and the machine learning model is obtained by using a machine learning algorithm to learn the historical understanding results of historical traffic scenes;

[0012] Invoke the small model through the said scheduling instruction set.

[0013] The traffic scenario management method provided by the embodiments of the present application makes full use of the understanding ability of the large model and the real-time operation ability of the small model. Through the multi-modal large model, the traffic scenario is understood and analyzed, and a high-quality understanding result is output. Furthermore, the small model can be called more precisely, not only avoiding the simultaneous invocation of all small models, reducing the burden on system resources and achieving effective allocation of resources, but also ensuring that the small model can quickly respond and execute tasks, improving the reaction speed and task processing efficiency of the entire system. The embodiments of the present application do not require manual interaction, greatly reducing the dependence on manual intervention and enhancing the automated response ability.

[0014] To achieve the above object, the embodiments of the present application further provide a traffic scenario understanding device, including:

[0015] An acquisition module, configured to acquire a global image of the current traffic scenario;

[0016] A segmentation module, configured to segment the global image to obtain a plurality of local images;

[0017] A first understanding module, configured to use one or more multi-modal large models to understand the global image and a plurality of the local images to obtain a current understanding result.

[0018] To achieve the above object, the embodiments of the present application further provide a traffic scenario management device, including:

[0019] A second understanding module, configured to use one or more multi-modal large models to understand the current traffic scenario to obtain a current understanding result;

[0020] A decision module, configured to use a machine learning model to analyze the current understanding result to obtain a scheduling instruction set for the current traffic scenario; wherein, the scheduling instruction set includes a plurality of scheduling instructions, and the machine learning model is obtained by using a machine learning algorithm to learn the historical understanding results of historical traffic scenarios;

[0021] An invocation module, configured to invoke the small model through the scheduling instruction set.

[0022] To achieve the above object, the embodiments of the present application further provide a traffic scenario management system, including: a scenario understanding module and a scenario management module:

[0023] The scenario understanding module is configured to execute the traffic scenario understanding method as described above;

[0024] The scenario management module is configured to execute the traffic scenario management method as described above.

[0025] To achieve the above object, an embodiment of the present application further provides a traffic scenario management device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the traffic scenario management method as described above is implemented.

[0026] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the traffic scenario management method as described above.

[0027] To achieve the above object, an embodiment of the present application further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the traffic scenario management method as described above is implemented. Description of the Drawings

[0028] Figure 1 is a flowchart of a traffic scenario understanding method provided by an embodiment of the present application;

[0029] Figure 2 is a schematic diagram of a multi-modal large model training method provided by an embodiment of the present application;

[0030] Figure 3 is a schematic flowchart of obtaining a current understanding result provided by an embodiment of the present application;

[0031] Figure 4 is a flowchart of a traffic scenario management method provided by an embodiment of the present application;

[0032] Figure 5 is a structural block diagram of a traffic scenario understanding device provided by an embodiment of the present application;

[0033] Figure 6 is a structural block diagram of a traffic scenario management device provided by an embodiment of the present application;

[0034] Figure 7 is a structural block diagram of a traffic scenario management system provided by an embodiment of the present application.

[0035] Figure 8 is a structural block diagram of a traffic scenario management device provided by an embodiment of the present application. Detailed Embodiments

[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0037] See Figure 1 , Figure 1 which is a flowchart of a traffic scene understanding method provided by an embodiment of the present application. The traffic scene understanding method includes:

[0038] S11. Obtain a global image of the current traffic scene;

[0039] S12. Segment the global image to obtain a number of local images;

[0040] S13. Use one or more multimodal large models to understand the global image and the number of local images to obtain the current understanding result.

[0041] The traffic scene understanding method provided by the embodiment of the present application analyzes and understands the traffic scene by combining the global image and local images of the traffic scene, improves the understanding accuracy, can generate high-quality and multi-level understanding results, provides a scientific and reasonable decision-making basis for subsequent traffic scene management, thereby avoiding the simultaneous operation of all small models, and further reducing the system resource burden.

[0042] Optionally, the multimodal large model is an open-source multimodal large model (for example, Qwen, InterVL) or a closed-source multimodal large model (for example, ChatGPT4).

[0043] To enhance the understanding ability of the large model, the multimodal large model is trained, and the trained multimodal large model is used for scene understanding. The multimodal large model tasks involved in the embodiments of the present application are mainly image understanding and do not rely on multi-round dialogue capabilities and question answering. Combining Figure 2 , for understanding and generation tasks, the multimodal large model is trained as follows:

[0044] (1) Pre-training (optional): The llama3.2 method is used for step-by-step pre-training: First, collect text-image pairs (traffic knowledge and images of historical traffic scenes), freeze the text model (no directed text input), and train the MLP (multi-layer perceptron model) and the image model. Then, freeze the image model and perform pre-training on a large scale of data. Among them, the text model processes text information and can adopt the Chinese-LLaMA network; the image model processes image information and can adopt the EVA-ViT-G network.

[0045] (2) Fine-tuning of directional text (i.e., preset prompt words) input: Input the concatenation result of the image + directional text into a pre-trained multi-modal large model for fine-tuning. This can save memory usage, does not require training with multiple rounds of dialogue text, focuses more on image-to-text generation, and the solution is more concise. Among them, the directional text is tokenized through a Tokenizer.

[0046] Preferably, during the fine-tuning process, a total of 2 to 3 epochs (epochs in model training refer to the number of times the entire training dataset is used by the model) of training can strengthen the training of the multi-modal large model's ability to understand traffic scenarios.

[0047] In an alternative embodiment, the splitting of the global image to obtain a plurality of local images includes:

[0048] Identify the global image to obtain the number of traffic participants in the global image;

[0049] If the number of traffic participants is greater than 0, then split the global image according to the size of the global image to obtain a plurality of local images.

[0050] Exemplarily, the global image comes from an image collected on the roadside / vehicle side. The global image is input into an open-set detection large model for identification to obtain the number of traffic participants in the global image. For example, traffic participants include elements such as vehicles, pedestrians, cones, traffic signs, etc. If the number of traffic participants is 0, then this global image is excluded and there is no need to manage this global image; if the number of traffic participants is greater than 0, then split the global image according to the size of the global image to obtain a plurality of local images. For example, if the size of the global image is 1080*1080, then the global image is divided into nine equal parts.

[0051] In an alternative embodiment, the use of one or more multi-modal large models to understand the global image and a plurality of the local images to obtain a current understanding result includes:

[0052] Input the global image into one or more of the multi-modal large models to obtain a global description text of the global image;

[0053] For each local image, input the local image into one or more of the multi-modal large models to obtain a local description text of the local image;

[0054] Use each local description text to correct the global description text to obtain a current understanding result.

[0055] The embodiments of this application introduce the concept of synthesizing the overall semantics (i.e., the current understanding result) from local semantics (i.e., local descriptive text), which greatly improves the quality and details of the current understanding result for the traffic scene presented in the image and also ensures the accuracy and consistency of the description.

[0056] In the embodiments of this application, the global description text is obtained by inputting the global image into one or more of the multimodal large models; for example, by inputting the global image into one or more of the multimodal large models, the descriptive texts output by each multimodal large model are obtained, and these descriptive texts are merged and de-duplicated to obtain the global description text.

[0057] Similarly, for each local image, the local description text is obtained by inputting the local image into one or more of the multimodal large models; for example, by inputting the local image into one or more of the multimodal large models, the descriptive texts output by each multimodal large model are obtained, and these descriptive texts are merged and de-duplicated to obtain the local description text; this operation is repeated for each local image to obtain each local description text.

[0058] In an alternative embodiment, the step of inputting the global image into one or more of the multimodal large models to obtain the global description text of the global image includes:

[0059] Obtaining several first preset prompt words;

[0060] For each of the first preset prompt words, the first preset prompt word and the global image are input together into one or more of the multimodal large models to obtain candidate global description texts output by each multimodal large model;

[0061] The several candidate global description texts are merged and de-duplicated to obtain the global description text.

[0062] In the embodiments of this application, by setting several first preset prompt words, candidate global description texts corresponding to each first preset prompt word are output, and then by combining several candidate global description texts, the global description text is obtained, which greatly improves the quality and details of the global description text for the traffic scene presented in the global image and also ensures the accuracy and consistency of the description.

[0063] In the embodiments of this application, the merging and de-duplication process includes:

[0064] (1) Merge several of the candidate global description texts to obtain a merged description text; delete the semantically similar content in the merged description text through semantic similarity detection to generate the global description text. For example, calculate the text similarity through BERT embedding vectors, and judge whether there is semantically similar content based on the text similarity. If so, delete it;

[0065] (2) Further, invalid information in the global description text obtained in the previous step can also be deleted to obtain the final global description text; for example, using the capabilities of a language large model, with the prompt "Remove unnecessary background and surrounding environment descriptions, delete the image summary part and unimportant detail descriptions (such as weather, roadside buildings, etc.), and ensure that the final text only retains the key information related to the traffic scene" to process the global description text obtained in the previous step to obtain the final global description text.

[0066] Preferably, in order to reduce the computational amount while improving the description accuracy and processing efficiency, two first preset prompt words are set, namely the first first preset prompt word and the second first preset prompt word; correspondingly, two candidate global description texts are generated, namely the first candidate global description text and the second candidate global description text; the step of inputting the global image into one or more of the multimodal large models to obtain the global description text of the global image includes:

[0067] Input the global image and the first first preset prompt word into one or more of the multimodal large models to obtain the first candidate global description text output by each multimodal large model;

[0068] Input the global image and the second first preset prompt word into one or more of the multimodal large models to obtain the second candidate global description text output by each multimodal large model;

[0069] Perform a merge and deduplication process on several of the first candidate global description texts and several of the second candidate global description texts to obtain the global description text.

[0070] Exemplarily, the first first preset prompt word is "Please describe the traffic scene in this image", and input this first first preset prompt word and the image of the current traffic scene (Image original ) into the multimodal large model A to obtain the first candidate global description text output by the multimodal large model A; this first candidate global description text is used as the basic description text of the global image;

[0071] The second first preset prompt word is "As a vehicle about to enter, is there anything special to note in the current traffic scene?", and input this second first preset prompt word and the image of the current traffic scene (Image original)Input into the multimodal large model B to obtain the second candidate global description text output by the multimodal large model B, and use this second candidate global description text as the enhanced description text of the local image;

[0072] Combining the first candidate global description text and the second candidate global description text results in a global description text with higher accuracy and better quality. The merging and deduplication process in the embodiments of this application is similar to the above and will not be elaborated here.

[0073] In an optional embodiment, the process of performing merging and deduplication on a number of the first candidate global description texts and a number of the second candidate global description texts to obtain the global description text includes:

[0074] Perform merging and deduplication on a number of the first candidate global description texts to obtain the first global description text;

[0075] Perform merging and deduplication on a number of the second candidate global description texts to obtain the second global description text;

[0076] Perform merging and deduplication on the first global description text and the second global description text to obtain the global description text.

[0077] In the embodiments of this application, the merging and deduplication are first performed separately on the first candidate global description texts and the second candidate global description texts of the same type, and then the results of this merging and deduplication process are further merged and deduplicated, which can improve the processing efficiency. The merging and deduplication process in the embodiments of this application is similar to the above and will not be elaborated here.

[0078] In an optional embodiment, for each of the local images, inputting the local image into one or more of the multimodal large models to obtain the local description text of the local image includes:

[0079] Obtain a number of second preset prompt words;

[0080] For each of the local images and each of the second preset prompt words, input the local image and the second preset prompt word together into one or more of the multimodal large models to obtain the candidate local description text output by each of the multimodal large models, and perform merging and deduplication on a number of the candidate local description texts to obtain the local description text.

[0081] In the embodiments of this application, by setting a number of second preset prompt words, candidate local description texts corresponding to each of the second preset prompt words are output, and then the local description text is obtained by combining a number of the candidate local description texts, which greatly improves the quality and details of the local description text for the traffic scene presented in the local image, and also ensures the accuracy and consistency of the description.

[0082] Preferably, in order to reduce the computational complexity while improving the description accuracy and processing efficiency, two second preset prompt words are set, namely the first second preset prompt word and the second second preset prompt word; correspondingly, two candidate local description texts are generated, namely the first candidate local description text and the second candidate local description text; for each of the local images, inputting the local image into one or more of the multimodal large models to obtain the local description text of the local image includes:

[0083] For each of the local images, perform the following steps:

[0084] Input the local image and the first second preset prompt word into one or more of the multimodal large models to obtain the first candidate local description text output by each of the multimodal large models;

[0085] Input the local image and the second second preset prompt word into one or more of the multimodal large models to obtain the second candidate local description text output by each of the multimodal large models;

[0086] Perform a merging and deduplication process on a number of the first candidate local description texts and a number of the second candidate local description texts to obtain the local description text.

[0087] Exemplarily, the first second preset prompt word is "Please describe the traffic scene in this image". Input this first second preset prompt word and the local image (Image pached ) of the current traffic scene into the multimodal large model A to obtain the first candidate local description text output by the multimodal large model A; this first candidate global description text serves as the basic description text of the local image;

[0088] The second second preset prompt word is "As a vehicle about to enter, is there anything special to note in the current traffic scene?". Input this second second preset prompt word and the local image (Image pached ) of the current traffic scene into the multimodal large model B to obtain the second candidate local description text output by the multimodal large model B, and this second candidate local description text serves as the enhanced description text of the local image;

[0089] Combining the first candidate local description text and the second candidate local description text results in a local description text with higher accuracy and better quality. The merging and deduplication process in the embodiments of the present application is similar to the above and will not be elaborated here.

[0090] In an alternative embodiment, the performing a merging and deduplication process on a number of the first candidate local description texts and a number of the second candidate local description texts to obtain the local description text includes:

[0091] Merge and deduplicate several of the first candidate local description texts to obtain a first local description text;

[0092] Merge and deduplicate several of the second candidate local description texts to obtain a second local description text;

[0093] Merge and deduplicate the first local description text and the second local description text to obtain the local description text.

[0094] In the embodiments of the present application, the first candidate local description texts and the second candidate local description texts of the same type are first merged and deduplicated respectively, and then the results of the merge and deduplication process are merged and deduplicated again, which can improve the processing efficiency. The merge and deduplication process in the embodiments of the present application is similar to the above, and will not be elaborated here.

[0095] In an alternative embodiment, the using each of the local description texts to correct the global description text to obtain the current understanding result includes:

[0096] Compare each of the local description texts with the global description text in turn. If the local description text is not included in the global description text, correct the global description text according to the local description text until the comparison of the last local description text is completed to obtain the current understanding result.

[0097] In the embodiments of the present application, by comparing each of the local description texts with the global description text in turn to correct the global description text. For example, if the local description text is included in the global description text, the global description text is not corrected; if the local description text is not included in the global description text, the local description text is added to the global description text until the comparison of the last local description text is completed, and all corrections are completed. The finally corrected global description text is the current understanding result.

[0098] In a specific example, such as Figure 3 , use two first preset prompt words and one second preset prompt word for scene understanding:

[0099] Input the global image and the first first preset prompt word into two multimodal large models respectively to obtain the first candidate global description texts output by the two multimodal large models respectively; deduplicate and merge these two first candidate global description texts to obtain a first global description text;

[0100] Input the global image and the second first preset prompt into two multi-modal large models respectively to obtain the second candidate global description texts output by these two multi-modal large models; deduplicate and merge these two second candidate global description texts to obtain the second global description text;

[0101] Deduplicate and merge the first global description text and the second global description text to obtain the global description text;

[0102] Divide the global image into nine equal parts to obtain nine local images. Each local image and the second preset prompt are respectively input into a multi-modal large model to obtain the local description text of each local image;

[0103] Use each local description text to correct the global description text to obtain the current understanding result.

[0104] In an optional embodiment, after using one or more multi-modal large models to understand the global image and several local images to obtain the current understanding result, the traffic scene understanding method further includes:

[0105] Use the traffic light data to improve the current understanding result to obtain the improved current understanding result;

[0106] Among them, the traffic light data is obtained through the following steps:

[0107] Obtain the V2X data of each device in the current traffic scene;

[0108] Parse the V2X data to obtain the traffic light data.

[0109] In the embodiment of the present application, the V2X (vehicle to everything) data of each device in the current traffic scene comes from the V2X data received by the roadside RSU (Road side Unit) and / or the on-board OBU (On board Unit) at the intersection. By parsing the SPAT (Signal Phase and Timing) information in the V2X data, the traffic light data is obtained, and then the current understanding result is improved in combination with the traffic light data to obtain the improved current understanding result, so as to analyze the improved current understanding result to obtain a scheduling instruction set for the current traffic scene. Among them, the traffic light data includes the status of the current traffic lights (such as red light, green light, yellow light) and the duration of the traffic light status.

[0110] In the embodiments of the present application, during the process of scene understanding, traffic light data is dynamically embedded to provide accurate traffic light states and their impacts on traffic scenes, further improving the ability to understand traffic scenes.

[0111] See Figure 4 , Figure 4 which is a flowchart of a traffic scene management method provided by the embodiments of the present application. The traffic scene management method includes:

[0112] S21. Use one or more multimodal large models to understand the current traffic scene and obtain the current understanding result;

[0113] S22. Use a machine learning model to analyze the current understanding result and obtain a scheduling instruction set for the current traffic scene; wherein, the scheduling instruction set includes several scheduling instructions, and the machine learning model is obtained by using a machine learning algorithm to learn the historical understanding results of historical traffic scenes;

[0114] S23. Invoke a small model through the scheduling instruction set.

[0115] In the embodiments of the present application, first, use one or more multimodal large models to understand the current traffic scene and obtain the current understanding result; for example, by directly inputting the image of the current traffic scene and preset prompt words into one or more of the multimodal large models, obtaining the description text output by each multimodal large model, and performing merging and duplicate removal processing on these description texts to obtain the current understanding result. Of course, the traffic scene understanding method of the above embodiments can also be applied, and no specific limitation is made here. In addition, the multimodal large model can refer to the multimodal large model of the above embodiments and will not be elaborated here.

[0116] Next, use a machine learning model to analyze the current understanding result and obtain a scheduling instruction set for the current traffic scene; wherein, the scheduling instruction set includes several scheduling instructions, and the machine learning model is obtained by using a machine learning algorithm to learn the historical understanding results of historical traffic scenes. Specifically, input the current understanding result into the machine learning model to obtain the scheduling instruction set, and generate scheduling instructions through machine learning, improving the generation efficiency and accuracy of the scheduling instructions.

[0117] Since the large model pays more attention to the understanding ability, after the scene is understood and output, the machine learning model disassembles the understanding of the large model and coordinates with the small model to perform various tasks. Generally, general machine learning models are based on language large models, generating detailed instructions through language large models, or generating the required instructions through reinforcement learning models, or generating the required instructions through decision tree models, and no specific limitation is made here.

[0118] Finally, the small model is called through the scheduling instruction set. For example, various small models related to traffic scenario management are stored in the system, and a corresponding relationship between various small models and each scheduling instruction is preset. Based on the corresponding relationship, the small model corresponding to each scheduling instruction is called. This not only avoids the simultaneous invocation of all small models, reduces the burden on system resources, but also ensures that the small model can quickly respond and execute tasks, such as quickly locating and tracking vehicles, signal light coordination, etc., ultimately improving the reaction speed and task processing efficiency of the entire system.

[0119] In addition, the scheduling instruction set further includes: the invocation order of each scheduling instruction. In this way, the corresponding small model is called based on the invocation order, further enhancing the scientific nature of management;

[0120] The embodiment of the present application makes full use of the understanding ability of the large model and the real-time operation ability of the small model. The proposed cooperation mechanism between the large and small models realizes the effective allocation of resources. The large model focuses on scenario parsing, while the small model is responsible for quickly and accurately executing tasks, forming complementary advantages and improving the reaction speed and task processing efficiency of the entire system. In addition, the entire process does not require manual interaction, greatly reducing the dependence on manual intervention and enhancing the automated response ability.

[0121] In an optional embodiment, the machine learning model is a decision tree model; the decision tree model includes several decision trees, and the root node of each decision tree is a key element in the historical understanding result.

[0122] The embodiment of the present application adopts a decision tree model. Its main advantage lies in traffic scenarios, especially traffic intersection scenarios, which require quick decision-making and low-cost calculation. Compared with large models or reinforcement learning models, the execution efficiency and response time of the decision tree model are significantly improved, and decisions can be made within milliseconds without complex computing resources or long-term training. Especially in terms of interpretability, compared with black-box deep learning models, the decision tree algorithm naturally has higher interpretability and greater flexibility. The embodiment of the present application performs instruction conversion based on the decision tree model, which not only further improves the generation efficiency and accuracy of scheduling instructions, but also increases the interpretability and flexibility of the system.

[0123] The embodiment of the present application is especially aimed at traffic scenarios where the intersection scenario is complex and changes frequently, such as traffic light changes, entry / exit of traffic participants such as vehicles, etc. To cope with these changes, the embodiment of the present application decomposes the current understanding result into a series of conditional judgments through the decision tree model, so as to determine how to generate specific scheduling instructions.

[0124] Exemplarily, a decision tree model is constructed and trained through the following steps: 1. Obtain historical traffic scenarios, and use one or more multimodal large models to understand the historical traffic scenarios to obtain historical understanding results; 2. Identify key elements in the historical understanding results; among them, the key elements include at least one of traffic participants and elements affecting traffic, such as: traffic light status, vehicle and pedestrian positions, road congestion conditions, weather, etc.; 3. Use each key element as the root node of each decision tree to construct and train the decision tree model; that is, establish a decision tree for each key element to form a decision tree model; each node of the decision tree represents a scenario judgment, and based on each root node, a scheduling instruction set is generated.

[0125] For example, a decision tree is constructed for the traffic light status: Root node: traffic light status; Second layer node: road traffic signs; Third layer node: obstacle situation ahead; Fourth layer node: scheduling instruction. The specific process is to judge the traffic light status through the root node, judge the road traffic signs through the second layer node, and judge the obstacle situation ahead through the third layer node, such as whether there are obstacles or pedestrians ahead. Finally, through the last layer node, that is, the fourth layer node, a scheduling instruction is generated through decision-making.

[0126] The embodiment of the present application uses the decision tree model to make layer-by-layer judgments on each element in the historical traffic scenario to generate a clear scheduling instruction set. These scheduling instructions can not only quickly adapt to traffic changes, but also each decision-making process is interpretable, which is beneficial to the improvement in the development and debugging stages. In addition, the decision tree can be combined with other algorithms (such as NLP models) to further improve the accuracy and adaptability of instruction generation.

[0127] The decision tree model is a dynamic adaptive tree structure, ensuring that according to the real-time changes in the traffic scenario (such as traffic light changes, vehicles entering / leaving), the branch paths of the nodes can be automatically adjusted. This not only reduces unnecessary calculations, but also makes the generated scheduling instructions more real-time and flexible, and is especially suitable for traffic scenarios with frequent dynamic changes such as traffic lights, pedestrians, and vehicles at intersections.

[0128] The scheduling instructions generated by the decision tree model can precisely match the computing power of the small model. For example, for a small model with limited computing resources, the scheduling instructions will be optimized to only focus on specific areas or specific categories of traffic participants (such as pedestrians, vehicles, etc.) to reduce the load on the small model.

[0129] Optionally, the nodes of the decision tree in the decision tree model have priorities. In this way, during the generation process of the scheduling instructions, the decision tree can handle different nodes with different priorities, especially in the case of emergencies (such as emergency vehicles), to respond quickly. Through the priority weight mechanism, it is ensured that emergency instructions can be processed within milliseconds, thereby improving the real-time response ability of the system and achieving efficient resource allocation.

[0130] Specific traffic scenario examples are as follows, which are used to illustrate how to generate an actionable scheduling instruction set from complex traffic scenarios:

[0131] 1. Vehicle detection and tracking scenario

[0132] Output understood by the large model: Through analyzing images and V2X data, the multi-modal large model concludes that there are 3 cars, 2 pedestrians, and a traffic light in the traffic scenario. One of the red cars is about to turn left, the current traffic light is red, and the pedestrians are standing at the intersection waiting.

[0133] Scheduling instruction set after passing through the decision tree model:

[0134] Scheduling instruction 1: Identify and track the red car.

[0135] Scheduling instruction 2: Ignore stationary pedestrians and other non-moving objects.

[0136] Scheduling instruction 3: Under the red light state, there is no need to predict the actions of pedestrians.

[0137] 2. Traffic light interaction scenario

[0138] Output understood by the large model: Through analyzing images and V2X data, the multi-modal large model obtains that the traffic light ahead is red, there are pedestrians crossing the road at the intersection, the traffic flow is large, and there is no idle lane ahead on the driving route.

[0139] Scheduling instruction set after passing through the decision tree model:

[0140] Scheduling instruction 1: Monitor the traffic light state through V2X data.

[0141] Scheduling instruction 2: Predict the trajectory of nearby pedestrians.

[0142] Scheduling instruction 3: Ignore the actions of vehicles or pedestrians in the distance.

[0143] Through the decomposition of scheduling instructions, the multi-modal large model provides an understanding of all elements in the traffic scenario, while the small model only needs to focus on the positioning and tracking tasks of vehicles, or perform tasks related to traffic lights and intersection situations, avoiding unnecessary scenario analysis and reducing the amount of computation.

[0144] See Figure 5 , Figure 5 which is the structural block diagram of a traffic scenario understanding device provided by an embodiment of this application. The traffic scenario understanding device includes:

[0145] An acquisition module 11, configured to acquire a global image of the current traffic scenario;

[0146] The splitting module 12 is used to split the global image to obtain a number of local images;

[0147] The first understanding module 13 is used to understand the global image and a number of the local images by using one or more multimodal large models to obtain the current understanding result.

[0148] Optionally, the first understanding module 13 is further used for:

[0149] Input the global image into one or more of the multimodal large models to obtain the global description text of the global image;

[0150] For each of the local images, input the local image into one or more of the multimodal large models to obtain the local description text of the local image;

[0151] Use each of the local description texts to correct the global description text to obtain the current understanding result.

[0152] Optionally, the first understanding module 13 is further used for:

[0153] Obtain a number of first preset prompt words;

[0154] For each of the first preset prompt words, input the first preset prompt word and the global image into one or more of the multimodal large models to obtain the candidate global description text output by each of the multimodal large models;

[0155] Perform merging and deduplication processing on a number of the candidate global description texts to obtain the global description text.

[0156] Optionally, the first understanding module 13 is further used for:

[0157] Obtain a number of second preset prompt words;

[0158] For each of the local images and each of the second preset prompt words, input the local image and the second preset prompt word into one or more of the multimodal large models to obtain the candidate local description text output by each of the multimodal large models, and perform merging and deduplication processing on a number of the candidate local description texts to obtain the local description text.

[0159] Optionally, the first understanding module 13 is further used for:

[0160] Compare each of the local description texts with the global description text in turn. If the local description text is not included in the global description text, correct the global description text according to the local description text until the comparison of the last local description text is completed to obtain the current understanding result.

[0161] Optionally, the traffic scene understanding device further includes:

[0162] A parsing module, configured to obtain V2X data of each device in the current traffic scene; parse the V2X data to obtain the traffic light data;

[0163] A refinement module, configured to refine the current understanding result by using the traffic light data to obtain a refined current understanding result.

[0164] It should be noted that the working processes of the various modules in the traffic scene understanding device described in the embodiments of the present application may refer to the working processes of the traffic scene understanding method described in the above embodiments, and achieve the same beneficial effects, which will not be elaborated here.

[0165] See Figure 6 , Figure 6 is a structural block diagram of a traffic scene management device provided by an embodiment of the present application. The traffic scene management device includes:

[0166] A second understanding module 21, configured to understand the current traffic scene by using one or more multimodal large models to obtain a current understanding result;

[0167] A decision-making module 22, configured to analyze the current understanding result by using a machine learning model to obtain a scheduling instruction set for the current traffic scene; wherein, the scheduling instruction set includes several scheduling instructions, and the machine learning model is obtained by learning the historical understanding results of historical traffic scenes by using machine learning algorithms;

[0168] An invocation module 23, configured to invoke a small model through the scheduling instruction set.

[0169] Optionally, the machine learning model is a decision tree model; the decision tree model includes several decision trees, and the root node of each decision tree is a key element in the historical understanding result.

[0170] Optionally, the nodes of the decision trees in the decision tree model have priorities.

[0171] It should be noted that the working processes of the various modules in the traffic scene management device described in the embodiments of the present application may refer to the working processes of the traffic scene management method described in the above embodiments, and achieve the same beneficial effects, which will not be elaborated here.

[0172] See Figure 7 , Figure 7It is a structural block diagram of a traffic scenario management system provided by an embodiment of the present application; the traffic scenario management system includes: a scenario understanding module 31 and a scenario management module 32: the scenario understanding module 31 is used to execute the traffic scenario understanding method described in any of the above embodiments; the scenario management module 32 is used to execute the traffic scenario management method described in any of the above embodiments.

[0173] An embodiment of the present application also provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the traffic scenario understanding method or the traffic scenario management method described in any of the above embodiments.

[0174] An embodiment of the present application also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, they implement the traffic scenario understanding method or the traffic scenario management method described in any of the above embodiments.

[0175] See Figure 8 , Figure 8 It is a structural block diagram of a traffic scenario management device provided by an embodiment of the present application. The traffic scenario management device includes: a processor 41, a memory 42, and a computer program stored in the memory 42 and executable on the processor 41. When the processor 41 executes the computer program, it implements the steps in the above embodiments of the traffic scenario understanding method or the traffic scenario management method. Or, when the processor 41 executes the computer program, it implements the functions of each module / unit in the above embodiments of each device.

[0176] Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 42 and executed by the processor 41 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the traffic scenario management device.

[0177] The traffic scenario management device may include, but is not limited to, a processor 41 and a memory 42. Those skilled in the art can understand that the schematic diagram is only an example of the traffic scenario management device, and does not constitute a limitation on the traffic scenario management device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the traffic scenario management device may also include input / output devices, network access devices, buses, etc.

[0178] The processor 41 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 41 is the control center of the traffic scene management device, and connects all parts of the entire traffic scene management device through various interfaces and lines.

[0179] The memory 42 can be used to store the computer programs and / or modules. The processor 41 realizes various functions of the traffic scene management device by running or executing the computer programs and / or modules stored in the memory 42, and by calling the data stored in the memory 42. The memory 42 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 42 may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0180] Among them, if the modules / units integrated in the traffic scene management device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 41, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0181] It should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the accompanying drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0182] The above is the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.

Claims

1. A traffic scene understanding method, characterized in that: include: Get a global image of the current traffic scene; Segmenting the global image to obtain a number of local images; The global image and the plurality of local images are understood using one or more multimodal large models to obtain a current understanding result.

2. The traffic scene understanding method according to claim 1, characterized in that: The using one or more multimodal large models to understand the global image and the plurality of local images to obtain a current understanding result includes: Inputting the global image into one or more of the multimodal large models to obtain a global description text of the global image; For each of the partial images, input the partial image into one or more of the multimodal large models to obtain a partial description text of the partial image; The global description text is modified using each of the local description texts to obtain a current understanding result.

3. The traffic scene understanding method according to claim 2, characterized in that: The step of inputting the global image into one or more of the multimodal large models to obtain a global description text of the global image comprises: Obtaining a plurality of first preset prompt words; For each of the first preset prompt words, input the first preset prompt word and the global image together into one or more of the multimodal large models to obtain a candidate global description text output by each of the multimodal large models; A plurality of candidate global description texts are merged and deduplicated to obtain the global description text.

4. The traffic scene understanding method according to claim 2, characterized in that: For each of the partial images, inputting the partial image into one or more of the multimodal large models to obtain a partial description text of the partial image includes: Obtaining several second preset prompt words; For each of the local images and each of the second preset prompt words, the local image and the second preset prompt words are input together into one or more of the multimodal large models to obtain candidate local description texts output by each of the multimodal large models, and several of the candidate local description texts are merged and deduplicated to obtain the local description text.

5. The traffic scene understanding method according to claim 2, characterized in that: The using each of the local description texts to modify the global description text to obtain a current understanding result includes: Each of the local description texts is compared with the global description text in turn. If the local description text is not included in the global description text, the global description text is modified according to the local description text until the last local description text is compared to obtain the current understanding result.

6. The traffic scene understanding method according to claim 1, characterized in that: After the global image and the local images are understood by using one or more multimodal large models to obtain a current understanding result, the traffic scene understanding method further includes: The current understanding result is improved by using the traffic light data to obtain the improved current understanding result; The traffic light data is obtained by the following steps: Obtain V2X data of each device in the current traffic scene; The V2X data is parsed to obtain the traffic light data.

7. A traffic scene management method, characterized in that: include: Use one or more multimodal large models to understand the current traffic scene and obtain the current understanding results; The current understanding result is analyzed by using a machine learning model to obtain a dispatch instruction set for the current traffic scene; wherein the dispatch instruction set includes a plurality of dispatch instructions, and the machine learning model is obtained by learning the historical understanding results of historical traffic scenes by using a machine learning algorithm; The small model is called through the dispatch instruction set.

8. The traffic scene management method according to claim 7, characterized in that: The machine learning model is a decision tree model; the decision tree model includes a plurality of decision trees, and the root node of each decision tree is a key element in the historical understanding result.

9. The traffic scene management method according to claim 8, characterized in that: The nodes of the decision tree in the decision tree model have priorities.

10. A traffic scene understanding device, characterized in that: include: An acquisition module is used to obtain a global image of the current traffic scene; A segmentation module, used for segmenting the global image to obtain a plurality of local images; The first understanding module is used to understand the global image and the plurality of local images using one or more multimodal large models to obtain a current understanding result.

11. A traffic scene management device, characterized in that: include: The second understanding module is used to understand the current traffic scene using one or more multimodal large models to obtain the current understanding result; A decision module, used to analyze the current understanding results using a machine learning model to obtain a dispatch instruction set for the current traffic scene; wherein the dispatch instruction set includes a number of dispatch instructions, and the machine learning model is obtained by learning the historical understanding results of historical traffic scenes using a machine learning algorithm; The calling module is used to call the small model through the scheduling instruction set.

12. A traffic scene management system, characterized in that: include: Scene understanding module and scene management module: The scene understanding module is used to execute the traffic scene understanding method according to any one of claims 1 to 6; The scene management module is used to execute the traffic scene management method as described in any one of claims 7 to 9.

13. A traffic scene management device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When executing the computer program, the processor implements the traffic scene understanding method according to any one of claims 1 to 6, or the traffic scene management method according to any one of claims 7 to 9.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located controls the device to execute the traffic scene understanding method as described in any one of claims 1 to 6, or the traffic scene management method as described in any one of claims 7 to 9.

15. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements the traffic scene understanding method as described in any one of claims 1 to 6, or the traffic scene management method as described in any one of claims 7 to 9.