Traffic control and control model training methods, devices, equipment, media and products

By introducing traffic information templates and a critic model scoring mechanism into the large language model and optimizing its parameters, the problem of insufficient generalization ability in complex traffic signal control scenarios is solved, and efficient decision-making is achieved in different traffic scenarios.

CN119005253BActive Publication Date: 2025-09-19HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411090601.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2025-09-19
Estimated Expiration
2044-08-08

AI Technical Summary

Technical Problem

Large language models have weak generalization capabilities in complex traffic signal control scenarios and cannot make efficient decisions.

Method used

By obtaining text information template samples of traffic information, inputting them into GPT-4 to obtain reference results, and using the target critic model to score the pending results, the parameters of the large language model are adjusted to increase the probability of the output results as the training goal, and the critic model is introduced to guide strategy optimization to form a target large language model.

Benefits of technology

It achieves efficient decision-making in various traffic scenarios, improves the generalization ability of large language models, and produces more cost-effective and effective control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119005253B_ABST
    Figure CN119005253B_ABST
Patent Text Reader

Abstract

The present application discloses a traffic control and control model training method, apparatus, equipment, medium and product, which are applied to the field of machine learning. This method obtains a reference result through GPT‑4, and then uses improving the probability of outputting the reference result as a training goal to obtain a second large language model. Next, the target critic model is used as a training goal, with the output probability corresponding to the pending result having a higher score, to obtain a target large language model. This solution applies the large language model to the traffic control scenario, so that the large language model imitates the high-quality decision-making and reasoning trajectory generated by learning GPT‑4, and at the same time introduces the critic model to guide the strategy optimization of the large language model, so that it evaluates and improves the control decision of the large language model. The target large language model finally obtained can produce a more cost-effective and effective control strategy than GPT‑4.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of machine learning, and in particular relates to a traffic control and control model training method, device, equipment, medium and product. Background Art

[0002] Traffic signal control manages road traffic flow through traffic lights, which are set up at intersections and other key areas to indicate when vehicles should stop, go, turn, or change lanes. The Big Language Model is a neural network-based natural language processing technology. Trained on massive amounts of text data, it can learn and predict patterns and regularities in natural language text. It can not only generate natural language text but also deeply understand its meaning and handle various natural language tasks such as text summarization, question-answering, and translation. The basic idea of ​​the Big Language Model is to treat natural language text as a sequence of data. By inputting this sequence data and then performing computation and transformation through multiple layers of neurons, the corresponding output sequence is generated.

[0003] However, the current large language models have weak generalization capabilities and are unable to make efficient decisions in various complex traffic signal control scenarios. Summary of the Invention

[0004] The embodiments of the present application provide a traffic control and control model training method, device, equipment, medium and product, which can make efficient decisions in various complex traffic signal control scenarios.

[0005] In one aspect, an embodiment of the present application provides a traffic control model training method, comprising:

[0006] Obtain a sample text information template about traffic information;

[0007] Inputting the text information template sample into the fourth-generation pre-trained language model (Fourth-Generative Pre-trained Transformer, GPT-4) to obtain a reference result for controlling traffic lights;

[0008] Taking the text information template sample as input and increasing the probability of outputting the reference result as a training goal, adjusting the parameters of the pre-established first language model to obtain a second language model;

[0009] Inputting the text information template sample into the second language model multiple times to obtain all pending results output by the second language model; the reference results and the pending results both include decision and reasoning trajectories;

[0010] Scoring all pending results based on a pre-trained target critic model;

[0011] The text information template sample is used as input, and the higher the score of the pending result, the higher the output probability corresponding to the pending result is used as a training goal, and the parameters of the second large language model are adjusted to obtain a target large language model for traffic control.

[0012] On the other hand, before scoring all the pending results based on the pre-trained target critic model, the method further includes:

[0013] Determining a corresponding target score based on the number of vehicles queuing in a future target time period corresponding to each decision for controlling the traffic light; under the corresponding decision, the fewer the number of vehicles queuing in the target time period, the higher the corresponding target score;

[0014] Input each decision into the pre-established initial critic model to obtain the corresponding prediction score;

[0015] Calculating a first loss function value between the target score and the predicted score;

[0016] If the first loss function value does not satisfy the first iteration stopping condition, adjusting the parameters of the initial critic model and returning to the step of inputting each decision into the pre-established initial critic model to obtain the corresponding prediction score;

[0017] When the first loss function value satisfies the first iteration stopping condition, the initial critic model under the current parameters is determined as the target critic model.

[0018] On the other hand, the text information template sample is input into GPT-4 to obtain a reference result for controlling a traffic light, including:

[0019] Input the text information template sample into GPT-4 to obtain an initial reference result;

[0020] Based on the target critic model, scoring each possible output result;

[0021] If the initial reference result has the highest score among the output results, determining the initial reference result as the reference result;

[0022] If the initial reference result does not have the highest score among the output results, the initial reference result is discarded.

[0023] On the other hand, the text information template sample is used as input, and the parameters of the pre-established first language model are adjusted with increasing the probability of outputting the reference result as the training goal to obtain a second language model, including:

[0024] Inputting the text information template sample into the first large language model to obtain a target probability of the first large language model outputting the reference result;

[0025] Determining a second loss function value of the first logarithmic function according to the target probability; the greater the target probability, the smaller the second loss function value;

[0026] If the second loss function value does not satisfy the second iteration stopping condition, adjusting the parameters of the first large language model, and returning to the step of inputting the text information template sample into the first large language model to obtain a target probability of the first large language model outputting the reference result;

[0027] When the second loss function value satisfies the second iteration stopping condition, the first large language model under the current parameters is determined as the second large language model.

[0028] On the other hand, the text information template sample is used as input, and the output probability corresponding to the pending result with a higher score is set as a training goal, and the parameters of the second large language model are adjusted to obtain a target large language model for traffic control, including:

[0029] Inputting the text information template sample into the second language model to obtain the probability of each of the pending results;

[0030] Obtaining a probability difference between the probability of the pending result with a lower score and the probability of the pending result with a higher score;

[0031] Determining a third loss function value based on the probability difference and the target difference;

[0032] If the third loss function value does not satisfy the third iteration stopping condition, adjusting the parameters of the second language model, and returning to the step of inputting the text information template sample into the second language model to obtain the probabilities of each of the pending results;

[0033] When the third loss function value satisfies the third iteration stopping condition, the second large language model under the current parameters is determined as the target large language model.

[0034] On the other hand, the text information template includes any one or any combination of traffic status text information, traffic scene description information, traffic light control task description information, traffic priority information, and traffic light control action space.

[0035] On the other hand, the embodiment of the present application further provides a traffic control method, comprising:

[0036] Get the text information template about traffic information;

[0037] Inputting the text information template into a target large language model to obtain a target output result; the target large language model is trained by the above-mentioned traffic control model training method;

[0038] The traffic light is controlled according to the target result.

[0039] In another aspect, an embodiment of the present application further provides a traffic control model training device, the device comprising:

[0040] An acquisition module, used for acquiring a text information template sample about traffic information;

[0041] A first input module is used to input the text information template sample into GPT-4 to obtain a reference result for controlling traffic lights;

[0042] a first adjustment module, configured to take the text information template sample as input, and adjust the parameters of the pre-established first language model with increasing the probability of outputting the reference result as a training goal, so as to obtain a second language model;

[0043] a second input module, configured to input the text information template sample multiple times into the second language model to obtain all pending results output by the second language model; wherein the reference results and the pending results both include decision and reasoning trajectories;

[0044] A scoring module, configured to score all pending results based on a pre-trained target critic model;

[0045] The second adjustment module is used to take the text information template sample as input and adjust the parameters of the second large language model with the output probability corresponding to the pending result with a higher score as a training goal to obtain a target large language model for traffic control.

[0046] In another aspect, an embodiment of the present application further provides an electronic device, comprising: a processor and a memory storing computer program instructions;

[0047] When the processor executes the computer program instructions, the traffic control model training method or the traffic control method as described above is implemented.

[0048] On the other hand, an embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the traffic control model training method or traffic control method as described above is implemented.

[0049] On the other hand, an embodiment of the present application further provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the traffic control model training method or traffic control method as described above.

[0050] A traffic control model training method according to an embodiment of the present application has the following process: after obtaining a text information template sample about traffic information, input it into GPT-4 to obtain a reference result. Then, the text information template sample is used as input, and with the improvement of the probability of outputting the reference result as the training goal, the parameters of the pre-established first language model are adjusted to obtain the second language model. Next, the text information template sample is input into the second language model multiple times to obtain all the pending results output by the second language model; and they are scored based on the target critic model respectively. The text information template sample is used as input, and with the higher output probability corresponding to the pending result with a higher score as the training goal, the parameters of the second language model are adjusted to obtain the target large language model for traffic control. The solution proposed in the embodiment of the present application applies the large language model to the traffic control scenario, so that the large language model imitates and learns the high-quality decision and reasoning trajectories produced by GPT-4, and at the same time introduces the critic model to guide the strategy optimization of the large language model, so that it evaluates and improves the control decision of the large language model. The resulting target large language model can produce more cost-effective and effective control strategies than GPT-4, and enables it to demonstrate excellent generalization capabilities in different traffic flow scenarios, enabling it to make efficient decisions in various traffic scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0052] Figure 1 A flow chart of a traffic control model training method provided by one embodiment of the present application is shown;

[0053] Figure 2 An architectural diagram of a traffic control model training method provided in an embodiment of the present application;

[0054] Figure 3 is a structural diagram of a traffic control model training device provided by another embodiment of the present application;

[0055] Figure 4 This is a structural diagram of an electronic device provided in yet another embodiment of the present application. DETAILED DESCRIPTION

[0056] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.

[0057] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.

[0058] The primary purpose of traffic signal control is to ensure the safe, efficient, and orderly passage of vehicles and pedestrians on the road. Traffic signal control systems typically consist of traffic lights, traffic signal controllers, traffic detectors, and communication networks. Traffic lights typically display red, yellow, and green, representing stop, ready to go, and go, respectively. Traffic signal controllers are devices that control traffic lights, adjusting them based on traffic flow and time. Traffic detectors monitor the presence and movement of vehicles and pedestrians to help the controller make appropriate adjustments. Communication networks connect traffic signal controllers, allowing them to receive and send information as needed. Traffic signal control systems adjust traffic lights based on road traffic conditions to maximize efficiency, minimize congestion, and ensure traffic safety.

[0059] Large language models are trained using deep learning models on massive amounts of text data. They not only generate natural language text but also deeply understand its meaning and handle a variety of natural language tasks, including text summarization, question answering, and translation. This application applies large language models to traffic control scenarios. However, the original large language model has weak generalization capabilities and cannot make efficient decisions in various complex traffic signal control scenarios.

[0060] In order to solve the problems of traditional solutions, the embodiments of the present application provide a traffic control and control model training method, device, equipment, medium and product. The traffic control model training method provided by the embodiments of the present application is first introduced below. Figure 1 FIG. 1 shows a flow chart of a traffic control model training method provided by an embodiment of the present application; FIG. Figure 1 As shown, the method includes the following steps:

[0061] S101: Obtain a text information template sample related to traffic information.

[0062] S102: Input the text information template sample into GPT-4 to obtain a reference result for controlling traffic lights.

[0063] S103: Taking the text information template sample as input and taking improving the probability of outputting the reference result as the training goal, adjusting the parameters of the pre-established first language model to obtain a second language model.

[0064] S104: Input the text information template sample into the second largest language model multiple times to obtain all pending results output by the second largest language model.

[0065] Both the reference results and the pending results include decision and reasoning traces.

[0066] S105: Score all pending results based on the pre-trained target critic model.

[0067] S106: Using the text information template sample as input and taking the higher output probability corresponding to the pending result with a higher score as the training goal, adjusting the parameters of the second largest language model to obtain the target large language model for traffic control.

[0068] For S101, it is necessary to obtain a text information template sample about traffic information, which can be based on the real-time traffic status text information o collected at the intersection, traffic scene description information d scene , Traffic light control task description information d task , traffic priority information d know And the traffic light control action space A, with the strategy π of the large language model agent θ Control the signal configuration of the target intersection. The output of the large language model is the analytical reasoning trajectory Y and the control action a (i.e., decision) for adjusting the intersection signal. Its goal is to optimize the traffic efficiency of the intersection in the long term. It can be formally expressed as:

[0069] (Y,a)=π θ (Prompt(o,d scene ,d task ,dknow ,A)) (1)

[0070] Where X=Prompt(o,d scene ,d task ,d know ,A) is the text information template (or template sample) input into the large language model.

[0071] For S102, after obtaining the above-mentioned text information template sample, it is input into GPT-4, and the initial reference result output by GPT-4 can be obtained. The initial reference result can be directly used as the reference result; it can also be filtered through the target critic model to retain only the initial reference result with the highest score among all possible results. The reference result includes a reference decision and a corresponding reference reasoning trajectory. The decision and the reasoning trajectory are one-to-one corresponding. Whether it is GPT-4 or a large language model, the corresponding decision is obtained through an reasoning trajectory. The decision is the control action of adjusting the intersection signal, and the reasoning trajectory is the process of obtaining the corresponding decision based on the input text information template.

[0072] For S103, after GPT-4 outputs the reference decision and reference reasoning trajectory, this application uses the text information template sample as input, and uses the reference decision and reference reasoning trajectory as the expected output to adjust the parameters of the pre-established first language model to obtain the second language model. It should be noted that training usually requires a large number of text information template samples. In this case, each text information template sample performs the above steps, and may obtain a corresponding reference decision and reference reasoning trajectory, which are used to train the first language model to obtain the second language model.

[0073] This application does not limit the specific adjustment of the parameters of the first language model. The first language model can be gradually made to imitate and learn the decision and reasoning trajectory of GPT-4. First, the text information template sample is used as input, and the probability of outputting reference decision and reference reasoning trajectory is increased as the training goal. The parameters of the first language model are adjusted to obtain the second language model. At this time, the second language model imitates and learns the reference decision and reference reasoning trajectory of GPT-4, and can produce high-quality decision and reasoning trajectories.

[0074] In S104, after obtaining the second largest language model, the text information template sample is fed into the second largest language model multiple times, resulting in multiple sets of pending results, including pending decisions and pending reasoning trajectories. The large language model generates responses by sampling the next token and obtaining the corresponding output. However, when the temperature parameter is introduced into the model, the model does not necessarily sample the token with the highest probability; instead, it has a certain probability of sampling tokens with slightly lower probabilities. Therefore, for the same text information template sample, the model may output different first pending decisions and first pending reasoning trajectories each time. Multiple inputs of the text information template sample can yield multiple sets of pending decisions and pending reasoning trajectories.

[0075] In S105, after obtaining multiple sets of pending decisions and pending inference trajectories output by the second largest language model, all pending results are scored based on a pre-trained target critic model. The critic model is a scoring mechanism that assigns a higher target score to vehicles queueing in a future target time period under the corresponding decision. The number of vehicles queueing in the future target time period under each pending decision is then determined, and the corresponding score is assigned. The critic model further guides and optimizes the control strategy of the target language model, enabling it to make efficient decisions in various traffic scenarios.

[0076] For S106, the text information template sample is used as input, and the higher the score, the higher the output probability corresponding to the pending result is used as the training goal, and the parameters of the second largest language model are adjusted to obtain the target large language model for traffic control.

[0077] As mentioned above, there may be a large number of text information template samples. Each text information template sample will have corresponding multiple sets of pending decisions and pending inference trajectories. In the final training process of the second largest language model to obtain the target large language model, training is performed based on each text information template sample.

[0078] A traffic control model training method according to an embodiment of the present application has the following process: after obtaining a text information template sample about traffic information, input it into GPT-4 to obtain a reference result. Then, the text information template sample is used as input, and with the improvement of the probability of outputting the reference result as the training goal, the parameters of the pre-established first language model are adjusted to obtain the second language model. Next, the text information template sample is input into the second language model multiple times to obtain all the pending results output by the second language model; and they are scored based on the target critic model respectively. The text information template sample is used as input, and with the higher output probability corresponding to the pending result with a higher score as the training goal, the parameters of the second language model are adjusted to obtain the target large language model for traffic control. The solution proposed in the embodiment of the present application applies the large language model to the traffic control scenario, so that the large language model imitates and learns the high-quality decision and reasoning trajectories produced by GPT-4, and at the same time introduces the critic model to guide the strategy optimization of the large language model, so that it evaluates and improves the control decision of the large language model. The resulting target large language model can produce more cost-effective and effective control strategies than GPT-4, and enables it to demonstrate excellent generalization capabilities in different traffic flow scenarios, enabling it to make efficient decisions in various traffic scenarios.

[0079] In practical applications, scoring all pending results based on a pre-trained target critic model may require filtering the initial reference decision and initial reference reasoning trajectory output by GPT-4 to obtain the reference decision and reference reasoning trajectory. Therefore, it is necessary to train the pre-established initial critic model to obtain a target critic model that meets the requirements.

[0080] This application does not limit how the target critic model selects target decisions and target reasoning trajectories. A specific implementation scheme is provided herein. Before scoring all pending results based on the pre-trained target critic model, the corresponding target score is determined based on the number of vehicles queuing in the target time period corresponding to each decision to control the traffic light. The scoring is determined based on the actual situation. For example, the fewer the number of vehicles queuing in the future target time period under the corresponding decision, the higher the corresponding target score. The future target time period can also be determined based on the actual situation.

[0081] Then, each decision is input into the pre-established initial critic model to obtain the corresponding predicted score, and the first loss function value between the target score and the predicted score is calculated. If the first loss function value does not meet the first iteration stopping condition, the parameters of the initial critic model are adjusted, and the step of inputting each decision into the pre-established initial critic model to obtain the corresponding predicted score is continued. If the first loss function value meets the first iteration stopping condition, the initial critic model under the current parameters is determined as the target critic model. It should be noted that the first iteration stopping condition here can be that the first loss function value tends to remain unchanged (that is, the change value of the first loss function value is less than a set value), or that the first loss function value no longer decreases after several consecutive training cycles; the specific content of the first iteration stopping condition can be set according to actual conditions. The specific calculation formula of the first loss function value can refer to formula (2) below.

[0082] The critic model training method provided in the embodiments of this application determines the scores corresponding to different decisions by predicting the number of vehicles queuing in a future target time period. The fewer vehicles in the queue, the higher the score. This then yields the desired target critic model, enabling accurate scoring.

[0083] The above embodiment mentioned that the target critic model is used to score the pending results of the second language model. In a specific implementation, the target critic model can also be used to verify the initial reference decision and initial reference reasoning trajectory output by GPT-4, ensuring that the selected reference decision and reference reasoning trajectory minimize the number of vehicles in the queue. These are then used to train the first language model, allowing it to imitate and learn from the reference decision and reference reasoning trajectory.

[0084] Specifically, the process of inputting a text information template sample into GPT-4 to obtain a reference decision for controlling a traffic light and a corresponding reference reasoning trajectory includes: inputting a text information template sample into GPT-4 to obtain an initial reference decision and an initial reference reasoning trajectory; based on a pre-trained target critic model, scoring each possible decision and reasoning trajectory, and the initial reference decision and initial reference reasoning trajectory output by GPT-4 are one of the possible decisions and reasoning trajectories.

[0085] After each decision and reasoning trajectory receives a corresponding score, the score corresponding to the initial reference decision and initial reference reasoning trajectory output by GPT-4 is found and compared with the scores of the remaining decisions and reasoning trajectories. If the initial reference decision and initial reference reasoning trajectory have the highest score among all decisions and reasoning trajectories, they are determined as the reference decision and reference reasoning trajectory; if the initial reference decision and initial reference reasoning trajectory do not have the highest score among all decisions and reasoning trajectories, they are discarded.

[0086] This application models the traffic signal control task as a partially observable Markov process, enabling a large language model to make decisions based on the current traffic status of the intersection. Figure 2 The architecture diagram of the traffic control model training method provided in the embodiment of the present application; Figure 2 As shown in the figure, it includes three steps: collection and screening of reasoning trajectories, imitation learning fine-tuning, and policy optimization guided by the critic model.

[0087] like Figure 2 In step 1, the reasoning trajectory is collected and screened. First, the constructed text information template sample is used to let GPT-4 interact with the simulated traffic environment, and the initial reference decision and initial reference reasoning trajectory outputted are collected. In order to ensure the quality of the collected data, this solution trains the action-value network to find the reference decision and reference reasoning trajectory that best match the long-term goal of traffic signal control (such as minimizing the future queue length). The above-mentioned action-value network is the critic model. First, the initial critic model is constructed, and then the target critic model is trained. This training process can be completed by optimizing the Bellman Equation in a simulated environment. The calculation formula of the first loss function value is as follows:

[0088]

[0089] Among them t and a t is the observation and control action at the signal switching time step t, and γ∈[0,1] is the reward discount factor. R(o,a) is the reward function, which provides feedback (e.g., the negative value of the number of vehicles in the queue) for executing action a under observation o. Q(o,a) is the action-value function (i.e., the critic model) that estimates the future cumulative reward (score) obtained after executing a.

[0090] Subsequently, the trained action-value function (i.e., target critic model) is used to evaluate the initial results of GPT-4. Only the initial result with the highest score among all the results is retained and used as the reference result. Formally:

[0091]

[0092] Where T is the simulation duration, X t is the agent prompt, Y t is the reasoning trajectory of GPT-4.

[0093] The screening scheme provided by the embodiment of the present application can screen the initial reference decision and initial reference reasoning trajectory output by GPT-4. Only the initial reference decision output by GPT-4 with the highest score will be retained and determined as the reference decision. Then, the large language model is imitated and learned based on the reference decision to ensure better imitation learning effect.

[0094] After the reference decision and reference reasoning trajectories are screened in the above embodiment, the probability of outputting the reference result can be increased as a training goal, and the first language model can be trained to obtain the second language model. Specifically, the text information template sample can be used as input, and the parameters of the pre-established first language model can be adjusted with the probability of outputting the reference result as the training goal to obtain the second language model. The specific process is as follows:

[0095] Input the text information template sample into the first large language model to obtain a target probability of the first large language model outputting a reference result; and determine the second loss function value of the first logarithmic function based on the target probability. The larger the target probability, the smaller the second loss function value.

[0096] If the second loss function value does not satisfy the second iterative stopping condition, the parameters of the first largest language model are adjusted, and the process returns to the step of inputting the text information template sample into the first largest language model to obtain a target probability of the first largest language model outputting a reference result. If the second loss function value also satisfies the second iterative stopping condition, the first largest language model with the current parameters is determined as the second largest language model.

[0097] It should be noted that the second iteration stopping condition here can be that the second loss function value tends to remain unchanged (that is, the change in the second loss function value is less than a set value), or that the second loss function value no longer decreases after several consecutive training cycles; the specific content of the second iteration stopping condition can be set according to actual conditions.

[0098] like Figure 2Step 2 in the above is the imitation learning fine-tuning process. With the goal of imitating the reference results of GPT-4, the first language model is trained. This model is then made to imitate the reference decision and reference inference trajectory of GPT-4 to obtain the second language model. The text information template sample X is regarded as a fine-tuning instruction, including the control action a (decision) selected by GPT-4 and the inference trajectory Y = [Y; a] as the expected answer. The negative log-likelihood (NLL) is used as the loss function. The second loss function value is calculated as follows:

[0099]

[0100] in To generate the answer character y when the prompt is X w probability.

[0101] Through the training method provided in the embodiment of the present application, the first largest language model can imitate and learn the reference decision and reference reasoning trajectory of GPT-4, thereby obtaining a second largest language model that can produce high-quality decisions and reasoning trajectories.

[0102] As mentioned above, given the same text information template sample, the model may output different first pending decisions and first pending reasoning trajectories each time. Therefore, multiple inputs of the text information template sample can generate multiple sets of pending decisions and pending reasoning trajectories. This embodiment uses a pre-trained target critic model to score multiple sets of first pending decisions and first pending reasoning trajectories.

[0103] Next, the text information template sample is used as input, and the higher the score, the higher the output probability corresponding to the pending result is used as the training goal, and the parameters of the second largest language model are adjusted to obtain the target large language model for traffic control.

[0104] First, a text information template sample is input into the second largest language model to obtain the probability of each pending result. Then, the probability difference between the probability of the pending result with a lower score and the probability of the pending result with a higher score is obtained, and a third loss function value is determined based on the probability difference and the target difference (which can be set to 0). If the third loss function value does not meet the third iteration stopping condition, the parameters of the second largest language model are adjusted, and the process returns to the step of inputting the text information template sample into the second largest language model to obtain the probability of each pending result. If the third loss function value meets the third iteration stopping condition, the second largest language model under the current parameters is determined as the target large language model.

[0105] Because there may be more than two pending results, the above steps need to be performed between each two pending results. It should be noted that the third iteration stopping condition here can be that the third loss function value tends to remain unchanged (i.e., the change in the third loss function value is less than a set value), or it can be that the third loss function value no longer decreases after several consecutive training cycles; the specific content of the third iteration stopping condition can be set according to the actual situation.

[0106] like Figure 2 Step three in the ,is policy optimization guided by the critic model. To further improve the effectiveness of the control policy output by the large language model, this embodiment proposes a policy optimization method that adjusts the inference trajectory of the large language to derive more reasonable control decisions. Similarly, a pre-trained action-value function can be used as the target critic model to evaluate the future rewards obtained by taking the control action selected by the target large language. Subsequently, an alignment fine-tuning algorithm is used to adjust the inference trajectory, ultimately guiding the large language model to adopt the target decision that produces higher future rewards.

[0107] Specifically, this scheme samples k strategies π under the prompt X. θ The first undetermined inference trajectory generated The target critic model gives each trajectory Y i The fraction q of the derived control action i =Q(o,a i ). Then, follow the trajectory Y i The average log-likelihood value of the characters represents their generation probability:

[0108]

[0109] Afterwards, the second language model is trained with the goal of increasing the output probability for the higher-scoring pending results. A ranking feedback loss with a bounded boundary constraint (RBC) is used for optimization to guide the large language model in deriving the inference trajectory that produces the highest-scoring control action. The calculation process for the third loss function is as follows:

[0110]

[0111] in Is the lowest ratio Y j The probability of the higher scoring inference trajectory, β is a hyperparameter, is the alignment term used to improve trajectories that produce higher-scoring control actions, It is a constraint used to prevent performance degradation.

[0112] After completing the above steps, the target large language model applied to traffic signal control can be obtained.

[0113] Through the solution provided in the embodiments of the present application, a critic model is introduced to guide the strategy optimization of the large language model, so that it can evaluate and improve the control decisions of the large language model. The higher the score, the higher the output probability corresponding to the pending result, thereby further improving the performance of the target large language model and enabling it to make efficient control decisions in various traffic environments.

[0114] In actual applications, the specific content of the text information template can be set according to actual needs. This embodiment provides a specific solution. The text information template may include: traffic status text information, traffic scene description information, traffic light control task description information, traffic priority information and any one or any combination of the traffic light control action space.

[0115] The traffic scene description information is a description of the actual traffic scene. For example, it can be expressed in the following text: "The traffic light controls an intersection with four directions: east, south, west and north."

[0116] Traffic light control task descriptions describe the tasks the model needs to complete. These can be expressed in text such as "Which is the next optimal traffic light configuration?", "Which is the next optimal lane to allow for traffic?", and "Which direction is allowed next?" The purpose of these control task descriptions is to enable the large language model to understand the task to be completed, namely, to control the traffic light.

[0117] Traffic priority information mainly consists of common sense knowledge in the field of traffic, such as "lanes with more queued vehicles should be given priority" and "vehicles that are far away from the intersection should not be paid too much attention."

[0118] The traffic light control action space includes specific control actions for traffic lights, such as "North-South Straight (NTST)", "North-South Left Turn (NLSL)", "North-South Right Turn (NRSR)," and "East-West Straight (ETWT). Other expressions can also be used in practice. The information represented by these texts controls the color displayed by each traffic light. Ultimately, the large language model outputs these actions, and traffic lights are controlled accordingly.

[0119] Traffic status text information is collected by sensors at intersections. This text information can include information such as the number of vehicles queued in each lane, the number of vehicles traveling in each lane, and the number of vehicles in each lane at different distances from the intersection. Different traffic statuses can be collected based on actual conditions and converted into corresponding text information.

[0120] Using the text information template provided in the embodiment of the present application as a sample can enable GPT-4 or a large language model to better understand the task to be completed, thereby outputting the required information and realizing the control of traffic lights.

[0121] To solve the above technical problems, the present application also provides a traffic control method, which specifically includes the following steps:

[0122] First, a text template about traffic information is obtained and fed into a target large language model to obtain the target output. Finally, traffic lights are controlled based on the target output. The target large language model is trained using the aforementioned traffic control model training method.

[0123] Because the traffic control method provided in the embodiment of the present application corresponds to the above-mentioned traffic control model training method, it has the same embodiments and beneficial effects and will not be repeated here.

[0124] In order to solve the above technical problems, an embodiment of the present application provides a traffic control model training device. Figure 3 FIG. 1 is a structural diagram of a traffic control model training device provided by another embodiment of the present application; Figure 3 As shown, the device includes the following modules:

[0125] An acquisition module 301 is used to acquire a text information template sample about traffic information;

[0126] A first input module 302 is used to input a text information template sample into GPT-4 to obtain a reference result for controlling traffic lights;

[0127] A first adjustment module 303 is configured to take a text information template sample as input, and adjust the parameters of the pre-established first language model with the goal of improving the probability of outputting a reference result, so as to obtain a second language model;

[0128] A second input module 304 is configured to input the text information template sample into the second largest language model multiple times to obtain all pending results output by the second largest language model; the reference results and the pending results both include decision and reasoning trajectories;

[0129] A scoring module 305 is used to score all pending results based on a pre-trained target critic model;

[0130] The second adjustment module 306 is configured to take the text information template sample as input and, with the higher output probability corresponding to the pending result having a higher score as a training objective, adjust the parameters of the second largest language model to obtain a target large language model for traffic control.

[0131] In one embodiment, the traffic control model training device further includes: a determination module configured to determine a corresponding target score based on the number of queued vehicles in a future target time period corresponding to each decision for controlling a traffic light before scoring all pending results based on a pre-trained target critic model; the smaller the number of queued vehicles in the target time period under the corresponding decision, the higher the corresponding target score;

[0132] The third input module is used to input each decision into the pre-established initial critic model to obtain the corresponding prediction score;

[0133] A calculation module, configured to calculate a first loss function value between a target score and a predicted score;

[0134] a third adjustment module, configured to adjust the parameters of the initial critic model if the first loss function value does not satisfy the first iteration stopping condition, and return to the step of inputting each decision into the pre-established initial critic model to obtain the corresponding prediction score;

[0135] The determination module is further configured to determine the initial critic model under current parameters as the target critic model when the first loss function value satisfies a first iteration stopping condition.

[0136] In one embodiment, the first input module 302 is specifically configured to:

[0137] Input the text information template sample into GPT-4 to obtain the initial reference result;

[0138] Based on the target critic model, score each possible output result;

[0139] If the initial reference result has the highest score among the output results, the initial reference result is determined as the reference result;

[0140] If the initial reference result does not have the highest score among the output results, the initial reference result will be discarded.

[0141] In one embodiment, the first adjustment module 303 is specifically configured to:

[0142] Inputting the text information template sample into the first language model to obtain a target probability of the first language model outputting a reference result;

[0143] Determine the second loss function value of the first logarithmic function according to the target probability; the greater the target probability, the smaller the second loss function value;

[0144] If the second loss function value does not satisfy the second iteration stopping condition, adjusting the parameters of the first language model, and returning to the step of inputting the text information template sample into the first language model to obtain the target probability of the first language model outputting the reference result;

[0145] When the second loss function value satisfies the second iteration stopping condition, the first largest language model under the current parameters is determined as the second largest language model.

[0146] In some embodiments, the second adjustment module 306 is specifically configured to:

[0147] Input the text information template sample into the second largest language model to obtain the probability of each pending result;

[0148] Obtain the probability difference between the probability of the pending result with a lower score and the probability of the pending result with a higher score;

[0149] Determining a third loss function value based on the probability difference and the target difference;

[0150] If the third loss function value does not satisfy the third iteration stopping condition, adjusting the parameters of the second largest language model, and returning to the step of inputting the text information template sample into the second largest language model to obtain the probabilities of each pending result;

[0151] When the third loss function value satisfies the third iteration stopping condition, the second largest language model under the current parameters is determined as the target large language model.

[0152] Since the device embodiment corresponds to the embodiment of the above-mentioned traffic control model training method and thus has the same beneficial effects, it will not be described in detail here.

[0153] Figure 4 : is a structural diagram of an electronic device provided in another embodiment of the present application; Figure 4 As shown, the electronic device may include a processor 401 and a memory 402 storing computer program instructions.

[0154] Specifically, the processor 401 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0155] Memory 402 may include a large capacity memory for data or instructions. By way of example and not limitation, memory 402 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, memory 402 is a non-volatile solid-state memory.

[0156] The memory 402 may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0157] The processor 401 reads and executes computer program instructions stored in the memory 402 to implement any one of the traffic control model training methods or traffic control methods in the above embodiments.

[0158] In one example, the electronic device may further include a communication interface 403 and a bus 404. The processor 401, the memory 402, and the communication interface 403 are connected via the bus 404 and communicate with each other.

[0159] The communication interface 403 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0160] Bus 404 comprises hardware, software or both, couples the parts of electronic equipment to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus 404 can comprise one or more buses.Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.

[0161] In addition, in conjunction with the traffic control model training method or traffic control method in the above embodiments, embodiments of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the traffic control model training method or traffic control method in the above embodiments is implemented.

[0162] An embodiment of the present application also provides a computer program product, including a computer program, which, when processed and executed, implements any one of the traffic control model training methods or traffic control methods in the above embodiments.

[0163] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0164] The functional blocks shown in the above structured block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card or the like. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0165] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0166] The above describes various aspects of the present disclosure with reference to the flowcharts and / or block diagrams of the traffic control and control model training methods, devices, equipment, media and products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more boxes in the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0167] The above content is only a specific implementation method of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of the present application.

Claims

1. A traffic control model training method, characterized in that: include: Obtain a sample text information template about traffic information; Inputting the text information template sample into the fourth-generation pre-trained language model GPT-4 to obtain a reference result for controlling traffic lights; Taking the text information template sample as input and increasing the probability of outputting the reference result as a training goal, adjusting the parameters of the pre-established first language model to obtain a second language model; Inputting the text information template sample into the second language model multiple times to obtain all pending results output by the second language model; the reference results and the pending results both include decision and reasoning trajectories; Determine a target score based on the number of vehicles queuing in the future target time period corresponding to each decision for controlling the traffic light; Under the corresponding decision, the fewer the number of vehicles queuing in the target time period, the higher the corresponding target score; Input each decision into the pre-established initial critic model to obtain the corresponding prediction score; Calculating a first loss function value between the target score and the predicted score; If the first loss function value does not satisfy the first iteration stopping condition, adjusting the parameters of the initial critic model and returning to the step of inputting each decision into the pre-established initial critic model to obtain the corresponding prediction score; When the first loss function value satisfies the first iteration stopping condition, determining the initial critic model under the current parameters as the target critic model; Scoring all pending results based on the target critic model; Inputting the text information template sample into the second language model to obtain the probability of each of the pending results; Obtaining a probability difference between the probability of the pending result with a lower score and the probability of the pending result with a higher score; Determining a third loss function value based on the probability difference and the target difference; If the third loss function value does not satisfy the third iteration stopping condition, adjusting the parameters of the second language model, and returning to the step of inputting the text information template sample into the second language model to obtain the probabilities of each of the pending results; When the third loss function value satisfies the third iteration stopping condition, the second large language model under the current parameters is determined as the target large language model for traffic control.

2. The traffic control model training method according to claim 1, characterized in that: The text information template sample is input into GPT-4 to obtain a reference result for controlling traffic lights, including: Input the text information template sample into GPT-4 to obtain an initial reference result; Based on the target critic model, scoring each possible output result; If the initial reference result has the highest score among the output results, determining the initial reference result as the reference result; If the initial reference result does not have the highest score among the output results, the initial reference result is discarded.

3. The traffic control model training method according to claim 2, characterized in that: Taking the text information template sample as input and taking improving the probability of outputting the reference result as a training goal, adjusting the parameters of the pre-established first language model to obtain a second language model, including: Inputting the text information template sample into the first large language model to obtain a target probability of the first large language model outputting the reference result; Determining a second loss function value of the first logarithmic function according to the target probability; the greater the target probability, the smaller the second loss function value; If the second loss function value does not satisfy the second iteration stopping condition, adjusting the parameters of the first large language model, and returning to the step of inputting the text information template sample into the first large language model to obtain a target probability of the first large language model outputting the reference result; When the second loss function value satisfies the second iteration stopping condition, the first large language model under the current parameters is determined as the second large language model.

4. The traffic control model training method according to claim 1, characterized in that: The text information template includes: any one or any combination of traffic status text information, traffic scene description information, traffic light control task description information, traffic priority information and traffic light control action space.

5. A traffic control method, characterized in that: include: Get the text information template about traffic information; Inputting the text information template into the target large language model to obtain the target output result; The target large language model is obtained by training the traffic control model training method according to any one of claims 1 to 4; The traffic light is controlled according to the target result.

6. A traffic control model training device, characterized in that: The device comprises: An acquisition module, used for acquiring a text information template sample about traffic information; A first input module is used to input the text information template sample into GPT-4 to obtain a reference result for controlling traffic lights; a first adjustment module, configured to take the text information template sample as input, and adjust the parameters of the pre-established first language model with increasing the probability of outputting the reference result as a training goal, so as to obtain a second language model; a second input module, configured to input the text information template sample multiple times into the second language model to obtain all pending results output by the second language model; wherein the reference results and the pending results both include decision and reasoning trajectories; a determination module for determining a corresponding target score based on the number of vehicles queuing in a future target time period corresponding to each decision for controlling the traffic light; wherein under the corresponding decision, the fewer the number of vehicles queuing in the target time period, the higher the corresponding target score; The third input module is used to input each decision into the pre-established initial critic model to obtain the corresponding prediction score; a calculation module, configured to calculate a first loss function value between the target score and the predicted score; a third adjustment module, configured to adjust the parameters of the initial critic model if the first loss function value does not satisfy the first iteration stopping condition, and return to the step of inputting each decision into the pre-established initial critic model to obtain the corresponding prediction score; The determining module is further configured to determine the initial critic model under current parameters as a target critic model when the first loss function value satisfies the first iteration stopping condition; A scoring module, configured to score all the pending results based on the target critic model; The second adjustment module is configured to input the text information template sample into the second large language model to obtain the probability of each of the pending results; obtain the probability difference between the probability of the pending result with a lower score and the probability of the pending result with a higher score; determine a third loss function value based on the probability difference and the target difference; if the third loss function value does not satisfy a third iteration stopping condition, adjust the parameters of the second large language model and return to the step of inputting the text information template sample into the second large language model to obtain the probability of each of the pending results; if the third loss function value satisfies the third iteration stopping condition, determine the second large language model under the current parameters as the target large language model for traffic control.

7. An electronic device, characterized in that: The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the traffic control model training method according to any one of claims 1 to 4 or the traffic control method according to claim 5 is implemented.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the traffic control model training method according to any one of claims 1 to 4 or the traffic control method according to claim 5.

9. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the traffic control model training method according to any one of claims 1 to 4 or the traffic control method according to claim 5.

Citation Information

Patent Citations

  • Generative large language model training method and model-based man-machine voice interaction method

    CN116127046A

  • Generative large language model training method and model-based man-machine voice interaction method

    CN116244416A