Decision-making device and decision-making method for new type multivariate perception
By adopting a decision-making device based on a generative content model in driving assistance technology, using driving images and status information to generate driving prediction information, the shortcomings in the judgment of input data types and marginal situations in the prior art are solved, and flexible prediction and accurate driving decisions are achieved.
Patent Information
- Application Number
- CN202311581663.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-16
AI Technical Summary
The existing driving assistance technology is based on general machine learning models and cannot expand beyond the input data types and categories defined in the training stage. It lacks flexibility in marginal situation judgment, making it difficult to predict risks in driving situations.
Using a decision-making device based on a generative content model, by receiving the driving images and states of the vehicle, the image identification model and prediction model are used to generate driving prediction information, including object motion prediction and vehicle motion prediction, and driving decisions are generated based on this. The generative content model can process multiple input data types by training in historical driving data and sensing data, thereby improving judgment flexibility.
It realizes flexibility and accuracy in risk prediction and driving decision-making in driving situations, can provide effective driving suggestions in a variety of input data and marginal situations, and improves the flexibility of driving assistance technology.
Smart Images

Figure CN120014570A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a decision-making device and method, and more particularly to a decision-making device and method based on a generative content model. Background Art
[0002] The main function of the Advanced Driver Assistance Systems (ADAS) installed on vehicles today is to warn and / or control the vehicle in emergency situations, but it fails to predict possible risks in advance based on the vehicle's driving situation and driving operations and provide driving decision suggestions.
[0003] In addition, most driving assistance technologies are based on general machine learning models, and the input data type and category of the model are defined during the training of the machine learning model, making it impossible to expand the input data when the model is used later. On the other hand, for scenarios not included in the training phase, the machine learning model can only make judgments based on the closest summary category in the training data. It can be seen that when using machine learning models as the basis of driving assistance technology, there is a lack of flexibility in the type of input data and the judgment of marginal scenarios.
[0004] In view of this, how to predict risks in driving scenarios and increase the flexibility of driving assistance technology is a goal that the industry urgently needs to work on. Summary of the invention
[0005] In order to solve the above problems, the present disclosure proposes a decision-making device, comprising a transceiver interface and a processor. The processor is electrically connected to the transceiver interface and is used to perform the following operations: receiving a driving image of a vehicle and a driving state corresponding to the driving image from the transceiver interface; performing an image recognition on the driving image using an image recognition model to generate multiple object information of multiple objects in the driving image; generating a driving prediction information based on the multiple object information and the driving state using a prediction model, the driving prediction information includes multiple object motion predictions of the multiple objects and a motion prediction of the vehicle, and the prediction model is generated after training based on a generative content model; and generating a driving decision based on the driving prediction information.
[0006] In one embodiment of the present invention, the decision-making device further includes a storage device electrically connected to the processor, and the storage device is used to store multiple historical driving images and multiple historical driving states corresponding to the multiple historical driving images, and the prediction model is generated by the following operations: using the image recognition model to perform image recognition on the multiple historical driving images to generate multiple historical object information of multiple historical objects in the multiple historical driving images; and training the generative content model based on the multiple historical object information and the multiple historical driving states to generate the prediction model.
[0007] In one embodiment of the present invention, the operation of generating the prediction model further includes: performing a text processing on the multiple historical object information and the multiple historical driving states to convert them into a training text; and training the generative content model based on the training text to generate the prediction model.
[0008] In one embodiment of the present invention, the storage further stores a plurality of historical sensing data corresponding to the plurality of historical driving images, and the prediction model is generated by the following operation: training the generative content model based on the plurality of historical object information, the plurality of historical driving states and the plurality of historical sensing data to generate the prediction model.
[0009] In one embodiment of the present invention, the operation of generating the driving prediction information further includes: performing a text processing on the multiple object information and the driving status to convert them into an input text; and inputting the input text into the prediction model to generate the driving prediction information in a text format.
[0010] In one embodiment of the present invention, the operation of generating the driving prediction information further includes: generating the driving prediction information using the prediction model based on the multiple object information, the driving status, a positioning information of the vehicle and a driving assistance information, wherein the multiple object information, the driving status, the positioning information and the driving assistance information are subjected to a text processing to be converted into an input text.
[0011] In one embodiment of the present invention, the processor is further used to perform the following operations: comparing the driving decision and an actual driving operation of the vehicle to generate an operation difference; and in response to the operation difference being greater than a threshold, fine-tuning the prediction model based on the driving image, the driving status and the actual driving operation corresponding to the driving decision.
[0012] In one embodiment of the present invention, the processor further receives a situation text from the transceiver interface, and is further used to perform the following operation: based on the situation text, an image generation model is used to generate a time series situation image described by the situation text.
[0013] In one embodiment of the present invention, the processor further generates a control signal to control a power system of the vehicle, wherein the control signal is generated based on the driving decision.
[0014] In one embodiment of the present invention, the motion predictions of the multiple objects include future trajectory data of the multiple objects at a future time, and the motion prediction of the vehicle includes the future trajectory data of the vehicle at the future time.
[0015] The present disclosure also provides a decision-making method, which is applicable to a processor, and its steps include: receiving a driving image of a vehicle and a driving state corresponding to the driving image; using an image recognition model to perform an image recognition on the driving image to generate multiple object information of multiple objects in the driving image; based on the multiple object information and the driving state, using a prediction model to generate a driving prediction information, the driving prediction information includes multiple object motion predictions of the multiple objects and a motion prediction of the vehicle, and the prediction model is generated after training based on a generative content model; and generating a driving decision based on the driving prediction information.
[0016] In one embodiment of the present invention, the processor is further electrically connected to a storage device for storing a plurality of historical driving images and a plurality of historical driving states corresponding to the plurality of historical driving images, and the prediction model is generated by the following steps: performing image recognition on the plurality of historical driving images using the image recognition model to generate a plurality of historical object information of a plurality of historical objects in the plurality of historical driving images; and training the generative content model based on the plurality of historical object information and the plurality of historical driving states to generate the prediction model.
[0017] In one embodiment of the present invention, the step of generating the prediction model further includes: performing text processing on the multiple historical object information and the multiple historical driving states to convert them into a training text; and training the generative content model based on the training text to generate the prediction model.
[0018] In one embodiment of the present invention, the storage further stores multiple historical sensing data corresponding to the multiple historical driving images, and the prediction model is generated by the following steps: training the generative content model based on the multiple historical object information, the multiple historical driving states and the multiple historical sensing data to generate the prediction model.
[0019] In one embodiment of the present invention, the step of generating the driving prediction information further includes: performing a text processing on the multiple object information and the driving status to convert them into an input text; and inputting the input text into the prediction model to generate the driving prediction information in a text format.
[0020] In one embodiment of the present invention, the step of generating the driving prediction information further includes: generating the driving prediction information with the prediction model based on the multiple object information, the driving status, a positioning information of the vehicle and a driving assistance information, wherein the multiple object information, the driving status, the positioning information and the driving assistance information are subjected to a text processing to be converted into an input text.
[0021] In one embodiment of the present invention, the decision-making method further includes: comparing the driving decision with an actual driving operation of the vehicle to generate an operation difference; and in response to the operation difference being greater than a threshold, fine-tuning the prediction model based on the driving image, the driving status and the actual driving operation corresponding to the driving decision.
[0022] In one embodiment of the present invention, the decision method further comprises: generating a time series situation image described by the situation text using an image generation model based on the situation text.
[0023] In one embodiment of the present invention, the decision-making method further comprises: generating a control signal to control a power system of the vehicle, wherein the control signal is generated based on the driving decision.
[0024] In one embodiment of the present invention, the motion predictions of the multiple objects include future trajectory data of the multiple objects at a future time, and the motion prediction of the vehicle includes the future trajectory data of the vehicle at the future time.
[0025] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are intended to provide further explanation of the disclosure as claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to make the above and other objects, features, advantages and embodiments of the present disclosure more clearly understood, the accompanying drawings are described as follows:
[0027] Figure 1 This is a schematic diagram of a decision-making device in a first embodiment of the present disclosure;
[0028] Figure 2 This is a flowchart of the decision-making device training the prediction model in the first embodiment of the present disclosure;
[0029] Figure 3A schematic diagram of a decision-making device generating driving prediction information based on a prediction model in a first embodiment of the present disclosure; and
[0030] Figure 4 Flow chart of the decision making method in the second embodiment of the present disclosure.
[0031]
Explanation of symbols
[0032] 1: Decision-making device
[0033] 12: Processor
[0034] 14: Transceiver interface
[0035] S201~S203, S301~S302, S401~S404: Steps
[0036] HDI: Historical Driving Images
[0037] IRM: Image Recognition Model
[0038] HOI: Historical Object Information
[0039] HDS: Historical driving status
[0040] TT: Training text
[0041] PM: Predictive Model
[0042] DI: Driving Image
[0043] OI: Object Information
[0044] DS: Driving status
[0045] PI: Positioning Information
[0046] DAI: Driving Assistance Information
[0047] IT: Input Text
[0048] PT: Driving prediction information
[0049] DD: Driving Decision
[0050] 400: Decision-making Methods DETAILED DESCRIPTION
[0051] In order to make the description of the present disclosure more detailed and complete, reference may be made to the attached drawings and various embodiments described below, in which the same numbers in the drawings represent the same or similar elements.
[0052] Please refer to Figure 1 , which is a schematic diagram of a decision device 1 in the first embodiment of the present disclosure. The decision device 1 comprises a processor 12 and a transceiver interface 14, wherein the processor 12 is coupled to the transceiver interface 14.
[0053] In some embodiments, the processor 12 may include a central processing unit (CPU), a graphics processing unit (GPU), a multi-processor, a distributed processing system, an application specific integrated circuit (ASIC) and / or a suitable computing unit.
[0054] The transceiver interface 14 is used to receive vehicle data, such as: external images, vehicle speed, steering wheel rotation angle, turn signal on / off status, throttle and brake operation status, etc. In some embodiments, the transceiver interface 14 is used to communicate with a camera, a speedometer, a microphone, a positioning element, a vehicle computer and / or other sensors on the vehicle. In some embodiments, the transceiver interface 14 may also include a network interface such as Ethernet and Wi-Fi.
[0055] The decision device 1 is used to generate driving decision suggestions based on the driving image and relevant information of the vehicle. For example, when the vehicle deviates from the lane, drives too fast, or is too close to other vehicles, the decision device 1 can judge the objects and the states of the objects in the environment based on the driving image, and generate driving decision suggestions (for example, changing the driving direction, slowing down). In some embodiments, the decision device 1 can be a vehicle-mounted device installed on the vehicle.
[0056] Please refer to further Figure 2 , which is a flow chart of the decision-making device 1 training the prediction model PM in the first embodiment disclosed herein.
[0057] In some embodiments, the decision device 1 further includes a storage device, which is electrically connected to the processor 12 and is used to store a plurality of historical driving images HDI and a plurality of historical driving states HDS corresponding to the plurality of historical driving images.
[0058] Specifically, the historical driving image HDI is a record of images captured by a camera installed on the vehicle when the vehicle was driving on the road in the past, and the historical driving status HDS is a record of the vehicle status such as the vehicle speed, throttle and brake operation when the historical driving image HDI was captured. The historical driving status HDS further includes a plurality of sensing data corresponding to the recording of a plurality of sensors installed on the vehicle when the vehicle was driving on the road in the past. The plurality of sensors include, but are not limited to, Lidar, radar, thermal camera, and sonar sensor. The sensing data includes driving orientation data, driving steering information, vehicle status data (for example, corresponding to the object manipulation of the historical driving image HDI, such as changing lanes, turning, etc.) and / or driving speed, etc.
[0059] In some embodiments, the storage may include semiconductor or solid-state memory, magnetic tape, removable computer disk, random access memory (RAM), read-only memory (ROM), hard disk, and / or optical disk.
[0060] like Figure 2 As shown, the decision device 1 first executes step S201 to perform image recognition on the historical driving image HDI using the image recognition model IRM to generate a plurality of historical object information HOIs of a plurality of historical objects in the historical driving image HDI.
[0061] Specifically, the processor 12 of the decision device 1 uses the image recognition model IRM to identify the movement speed, coordinates in the historical driving image HDI, object attributes (such as vehicle type, marking type) or vehicle body surrounding environment (such as weather, light and shadow, number of lanes) of objects such as vehicles, pedestrians, road markings, obstacles, etc. in the historical driving image HDI (i.e., historical object information HOI).
[0062] In some embodiments, the processor 12 may use the image recognition model IRM to mark the coordinates of the object in the historical driving image HDI and the vector used to represent the moving direction and speed of the object (ie, historical object information HOI).
[0063] Next, the decision device 1 trains a generative content model (Generative-Content Model) based on the historical object information HOI and the historical driving state HDS to generate a prediction model PM. For example, the generative content model can be a generative pre-trained transformer (GPT) or a generative adversarial network (GAN). Specifically, the decision device 1 can train the generative content model based on the vehicle driving environment recorded in the historical object information HOI and the vehicle driving state recorded in the historical driving state HDS as training data, and use the trained generative content model as the prediction model PM.
[0064] It should be noted that the generative content model has higher flexibility in input data compared to general machine learning models. More specifically, the effective input data that a general machine learning model can recognize when applied needs to be the same or similar to the type of training data set used in the training phase. In other words, if the training data is set to include object information in the image and a specified feature vector format for the corresponding vehicle operation data, the trained model can only output vehicle operation recommendations based on the feature vector format of the object information collected in the image, and cannot receive other information beyond the feature vector format of the training data set. In contrast, even if the trained generative content model receives data types that have not been learned in the training phase (for example: driving advice instructions, navigation routes), the model can also make a certain degree of judgment accordingly.
[0065] In some embodiments, the operation of generating the prediction model PM further includes the decision device 1 executing step S202, the processor 12 performing a text processing on the historical object information HOI and the historical driving status HDS to convert them into a training text TT; and the decision device 1 executing step S203, the processor 12 training the generative content model based on the training text TT to generate the prediction model PM.
[0066] In step S202, the processor 12 may tokenize the historical object information HOI and the historical driving status HDS to segment the description of the object and the vehicle status in the historical object information HOI and the historical driving status HDS into smaller units (i.e., training text TT) that can be processed by the generative content model. The tokens generated after segmentation may be single words, characters, subwords, or symbols, and the form of the tokens may be determined based on the type and size of the model. The tokenization operation may enable the model to more efficiently process input data of different languages, vocabularies, and formats to reduce computational and memory costs.
[0067] In some embodiments, the processor 12 may utilize a tokenization algorithm such as Rule-based Tokenization, Byte Pair Encoding (BPE), Unigram, WordPiece, etc. to complete the tokenization operation.
[0068] Next, in step S203 , the processor 12 may use the training text TT as training data to train the generative content model, and use the trained generative content model as the prediction model PM.
[0069] After completing the training of the prediction model PM, the decision-making device 1 can use the trained prediction model PM to predict possible driving risks and provide driving decision suggestions based on the current environment and vehicle information of the vehicle. Therefore, the processor 12 can receive a driving image DI of a vehicle and a driving status DS corresponding to the driving image DI from the transceiver interface 14. Specifically, the processor 12 can obtain the image (i.e., driving image DI) captured by the camera installed on the vehicle when the vehicle is driving on the road through the transceiver interface 14. At the same time, the processor 12 can also obtain information such as the vehicle speed, satellite positioning, throttle and brake operation conditions of the vehicle when the image is taken through the transceiver interface 14 (i.e., driving status DS).
[0070] After obtaining the driving image DI and the driving state DS, the operation of the decision device 1 generating the driving prediction information PT based on the prediction model PM can be referred to in Figure 3 .
[0071] like Figure 3 As shown, first, in step S301, the processor 12 uses the image recognition model IRM to perform an image recognition on the driving image DI to generate a plurality of object information OI of a plurality of objects in the driving image. Similar to the aforementioned step S201, the processor 12 can recognize the object information OI in the driving image DI through the same operation.
[0072] Next, the processor 12 generates driving prediction information PT using the prediction model PM based on the object information OI and the driving state DS. The driving prediction information PT includes a plurality of object motion predictions of the plurality of objects and a motion prediction of the vehicle.
[0073] Specifically, the processor 12 can input the object information OI and the driving state DS into the prediction model PM, and the prediction model PM can predict the future trajectory of the objects in the environment and the vehicle itself based on the current driving environment and vehicle state (i.e., object motion prediction and motion prediction). Among them, the motion prediction further includes the future trajectory data of the vehicle itself at a future time; correspondingly, the multiple object motion predictions further include the future trajectory data of the multiple objects at a future time. Specifically, the trajectory data further includes driving orientation data (orientation), driving steering information (steering information), vehicle status data (for example, corresponding to the object manipulation of the historical driving image HDI, such as changing lanes, turning, etc.) and / or driving speed, etc. Preferably, the processor 12 can also input a situational information such as: the current driving weather, light, driving environment (highway, flat road, or narrow alley) into the prediction model PM.
[0074] In some embodiments, the processor 12 may execute step S302 to perform text processing on the object information OI and the driving status DS to convert them into input text IT. Similar to the aforementioned step S202, the processor 12 may tokenize the object information OI and the driving status DS through the same operation to generate the input text IT. Furthermore, after generating the input text IT, the processor 12 may input the input text IT into the prediction model PM to generate driving prediction information PT. In this embodiment, the driving prediction information PT is in a text format, but the types of prediction models PM are different (e.g., text generation, image generation, or voice generation), and the driving prediction information PT may also be in an image format or a voice format.
[0075] In some embodiments, in addition to the object information OI and the driving status DS, the processor 12 can also generate driving prediction information PT based on the vehicle's positioning information PI and driving assistance information DAI using a prediction model PM, wherein the multiple object information, the driving status, the positioning information and the driving assistance information are subjected to text processing to be converted into an input text.
[0076] For example, based on the characteristics of the generative content model that can expand the input data, the processor 12 can also input the satellite positioning information of the vehicle at the moment of driving (i.e., positioning information PI) into the prediction model PM. In addition, the processor 12 can also input the current navigation route and / or the user's driving instructions (e.g., avoiding highways) (i.e., driving assistance information DAI) into the prediction model PM. In this way, the prediction model PM can refer to various different aspects of information when the vehicle is driving to generate driving prediction information PT, which can then be used as a reference for driving decision recommendations. It should be noted that if the positioning information PI and the driving assistance information DAI are not text data, the positioning information PI and the driving assistance information DAI must first be processed into text to convert them into input text, and then the input text is input into the prediction model PM.
[0077] Finally, the processor 12 generates a driving decision DD based on the driving prediction information PT.
[0078] Specifically, after the processor 12 generates the driving prediction information PT, it can determine whether the vehicle has an accident risk (e.g., collision with other vehicles or obstacles, speeding out of control, rollover). Therefore, further, the processor 12 can provide driving decision suggestions (i.e., driving decision DD) based on the prediction results (i.e., driving prediction information PT). For example, if the driving prediction information PT indicates that there is a vehicle approaching quickly from the left and may enter the future travel path of the vehicle, the processor 12 can generate a driving decision suggestion to step on the brakes accordingly.
[0079] In some embodiments, the processor 12 can also use the prediction model PM to generate a driving decision DD. Since the prediction model PM has obtained relevant information about the driving situation and generated the driving prediction information PT, the prediction model PM can also generate a driving decision DD based on the prediction result of the driving prediction information PT. In other embodiments, the driving decision DD can be a display text, an explanation voice, or the processor 12 controls the vehicle's driving computer.
[0080] In some embodiments, after generating the driving decision DD, the decision device 1 may further compare the driving decision DD with the actual driving operation of the driver. If there is a certain difference between the driving decision suggestion generated by the prediction model PM and the driver's driving operation (for example, the driving decision DD is to dodge an obstacle to the left, but the driver actually brakes to stop to avoid colliding with the obstacle), the decision device 1 may use the current driving image, driving status and the driver's driving operation as a reference to fine-tune the prediction model PM.
[0081] Specifically, the processor 12 of the decision device 1 compares the driving decision and an actual driving operation of the vehicle to generate an operation difference; and in response to the operation difference being greater than a threshold, fine-tunes the prediction model based on the driving image, the driving state and the actual driving operation corresponding to the driving decision.
[0082] In some embodiments, the decision device 1 can also generate an image of the driving situation based on the text describing the driving situation. Specifically, the processor 12 generates a time series situation image described by the situation text based on a situation text using an image generation model, wherein the image generation model can be a generation model that generates images based on text, such as a stable diffusion model.
[0083] Generally speaking, the aforementioned driving images refer to images taken from the perspective of the vehicle (eg, a driving recorder). Therefore, even though surveillance camera images from a third-person perspective can be collected for many traffic accidents, driving images from a first-person perspective are more difficult to obtain.
[0084] Therefore, for certain driving scenarios where driving images are relatively rare or difficult to collect (for example, accident driving images), the decision-making device 1 can generate a simulated continuous time series of driving images (i.e., situational images) based on text records such as statements of the parties involved and police accident records, where the situational text can include information such as objects in the driving scenario, road types, the movement status of objects, and the environment surrounding the vehicle body.
[0085] Specifically, the decision-making device 1 can be as follows: Figure 2 The illustrated step S201 and Figure 3 The same operation as step S301 is performed to generate the plurality of context object information corresponding to the context image. Figure 2 The illustrated step S202 and Figure 3 The same operation of step S302 is shown to perform text processing on the plurality of context object information to tokenize the plurality of context object information. Finally, the decision device 1 can use the plurality of context object information and the context text as training data to fine-tune the prediction model PM.
[0086] In some embodiments, the decision device 1 can be applied to a self-driving car. Specifically, the processor 12 can generate a control signal based on the driving decision DD, and the control signal is used to control the power system of the vehicle. For example, when the driving decision DD suggests turning right, the processor 12 can generate a control signal to turn the steering wheel of the vehicle to the right; or when the driving decision DD suggests deceleration, the processor 12 can generate a control signal to control the brake and throttle of the vehicle to decelerate the vehicle.
[0087] Please refer to Figure 4 , which is a flow chart of the decision method 400 in the second embodiment of the present disclosure. The decision method 400 includes steps S401 to S404. The decision method 400 is used to generate driving decision suggestions based on the driving image and the relevant information of the vehicle. The decision method 400 can be executed by a processor (for example: Figure 1 The processor 12 is shown.
[0088] First, in step S401, the processor receives a vehicle driving image and a vehicle driving status corresponding to the vehicle driving image.
[0089] Next, in step S402 , the processor uses an image recognition model to perform an image recognition on the driving image to generate a plurality of object information of a plurality of objects in the driving image.
[0090] Next, in step S403, the processor generates a driving prediction information based on the multiple object information and the driving status using a prediction model, the driving prediction information includes multiple object motion predictions of the multiple objects and a motion prediction of the vehicle, and the prediction model is generated after training based on a generative content model.
[0091] Finally, in step S404, the processor generates a driving decision based on the driving prediction information.
[0092] In some embodiments, the processor is electrically connected to a storage device (for example, the storage device in the first embodiment), which is used to store multiple historical driving images and multiple historical driving states corresponding to the multiple historical driving images, and the processor generates the prediction model through the following steps: the processor uses the image recognition model to perform image recognition on the multiple historical driving images to generate multiple historical object information of multiple historical objects in the multiple historical driving images; and the processor trains the generative content model based on the multiple historical object information and the multiple historical driving states to generate the prediction model.
[0093] In some embodiments, the step of generating the prediction model further includes: the processor performs a text processing on the multiple historical object information and the multiple historical driving states to convert them into a training text; and the processor trains the generative content model based on the training text to generate the prediction model.
[0094] In some embodiments, step S403 further includes the processor performing a text processing on the plurality of object information and the driving status to convert them into an input text; and the processor inputting the input text into the prediction model to generate the driving prediction information.
[0095] In some embodiments, step S403 also includes the processor generating the driving prediction information based on the multiple object information, the driving status, a positioning information of the vehicle and a driving assistance information using the prediction model, wherein the multiple object information, the driving status, the positioning information and the driving assistance information are subjected to text processing to be converted into an input text.
[0096] In some embodiments, the decision-making method 400 further includes the processor comparing the driving decision and an actual driving operation of the vehicle to generate an operation difference; and in response to the operation difference being greater than a threshold, the processor fine-tuning the prediction model based on the driving image corresponding to the driving decision, the driving status and the actual driving operation.
[0097] In some embodiments, the method of fine-tuning the prediction model further includes calculating the driving decision and an actual driving operation of the vehicle based on a loss function to generate an operation difference.
[0098] In some embodiments, the decision method 400 further includes the processor generating a situational image described by the situational text using an image generation model based on a situational text; and the processor fine-tuning the prediction model based on the situational image and the situational text.
[0099] In some embodiments, the step of fine-tuning the prediction model further includes generating a time series situation image described by the situation text using an image generation model based on the situation text.
[0100] In some embodiments, the decision method 400 further includes the processor generating a control signal to control a power system of the vehicle, wherein the control signal is generated based on the driving decision.
[0101] In some embodiments, the motion predictions of the multiple objects include future trajectory data of the multiple objects at a future time, and the motion prediction of the vehicle includes the future trajectory data of the vehicle at the future time.
[0102] In summary, the decision-making device and method proposed in the present disclosure can predict whether the vehicle will have the risk of an accident next based on the vehicle's driving images, and then provide suggestions for driving decisions. In addition, since the decision suggestions are generated by the generative content model, the decision-making device 1 can refer to driving assistance information such as navigation routes and user preferences in addition to the driving images and the driving status of the vehicle to generate corresponding driving decisions. Furthermore, the decision-making device 1 can also convert the situational text into images, and fine-tune the prediction model accordingly, and supplement certain driving situations where it is difficult to obtain driving images, so that the training data can be more comprehensive, and the prediction model can make driving decisions for driving emergencies more accurate.
[0103] Although several embodiments are described above as examples, the decision-making device and method proposed in the present disclosure may also be implemented by other systems, hardware, software, storage media or a combination thereof. Therefore, the protection scope of the present disclosure should not be limited to the specific implementation methods described in the embodiments of the present disclosure, but should be determined by the definition of the attached claims.
[0104] It is obvious to those with ordinary knowledge in the technical field to which the present disclosure belongs that various modifications and changes can be made to the structure of the present disclosure without departing from the scope or spirit of the present disclosure. In view of the foregoing, the protection scope of the present disclosure also covers the modifications and changes made within the attached claims.
Claims
1. A decision-making device, characterized in that: Include: a transceiver interface; and A processor is electrically connected to the transceiver interface and is used to perform the following operations: Receive a vehicle driving image and a vehicle driving status corresponding to the vehicle driving image from the transceiver interface; Performing image recognition on the driving image using an image recognition model to generate a plurality of object information of a plurality of objects in the driving image; Generate a vehicle prediction information based on the plurality of object information and the vehicle driving state using a prediction model, the vehicle driving prediction information including a plurality of object motion predictions of the plurality of objects and a motion prediction of the vehicle, and the prediction model is generated after training based on a generative content model; and Based on the driving prediction information, a driving decision is generated.
2. The decision-making device according to claim 1, characterized in that: The system further comprises a storage device, the storage device is electrically connected to the processor, the storage device is used to store a plurality of historical driving images and a plurality of historical driving states corresponding to the plurality of historical driving images, and the prediction model is generated by the following operations: Using the image recognition model to perform the image recognition on the plurality of historical driving images to generate a plurality of historical object information of a plurality of historical objects in the plurality of historical driving images; and The generative content model is trained based on the multiple historical object information and the multiple historical driving states to generate the prediction model.
3. The decision-making device according to claim 2, characterized in that: The operation of generating the prediction model further comprises: Performing text processing on the plurality of historical object information and the plurality of historical driving states to convert them into a training text; and The generative content model is trained based on the training text to generate the prediction model.
4. The decision-making device according to claim 2, characterized in that: The storage device further stores a plurality of historical sensing data corresponding to the plurality of historical driving images, and the prediction model is generated by the following operations: The generative content model is trained based on the plurality of historical object information, the plurality of historical driving states, and the plurality of historical sensing data to generate the prediction model.
5. The decision-making device according to claim 1, characterized in that: The operation of generating the driving prediction information further includes: Performing a text processing on the plurality of object information and the driving status to convert them into an input text; and The input text is input into the prediction model to generate the driving prediction information in a text format.
6. The decision-making device according to claim 1, characterized in that: The operation of generating the driving prediction information further includes: The driving prediction information is generated by the prediction model based on the multiple object information, the driving status, a positioning information of the vehicle and a driving assistance information, wherein the multiple object information, the driving status, the positioning information and the driving assistance information are subjected to a text processing to be converted into an input text.
7. The decision-making device according to claim 1, characterized in that: The processor is further configured to perform the following operations: comparing the driving decision with an actual driving operation of the vehicle to generate an operation difference; and In response to the operation difference being greater than a threshold, the prediction model is fine-tuned based on the driving image, the driving state, and the actual driving operation corresponding to the driving decision.
8. The decision-making device according to claim 1, characterized in that: The processor further receives a context text from the transceiver interface, and is further configured to perform the following operations: Based on the situation text, an image generation model is used to generate a time series situation image described by the situation text.
9. The decision-making device according to claim 1, characterized in that: The processor further generates a control signal to control a power system of the vehicle, wherein the control signal is generated based on the driving decision.
10. The decision-making device according to claim 1, characterized in that: The motion predictions of the multiple objects include future trajectory data of the multiple objects at a future time, and the motion prediction of the vehicle includes the future trajectory data of the vehicle at the future time.
11. A decision-making method, characterized in that: Applicable to a processor, the steps include: Receiving a vehicle line image of a vehicle and a vehicle line status corresponding to the vehicle line image; Performing image recognition on the driving image using an image recognition model to generate a plurality of object information of a plurality of objects in the driving image; Generate a vehicle prediction information based on the plurality of object information and the vehicle driving state using a prediction model, the vehicle driving prediction information including a plurality of object motion predictions of the plurality of objects and a motion prediction of the vehicle, and the prediction model is generated after training based on a generative content model; and A driving decision is generated based on the driving prediction information.
12. The decision-making method according to claim 11, characterized in that: The processor is further electrically connected to a storage device, the storage device is used to store a plurality of historical driving images and a plurality of historical driving states corresponding to the plurality of historical driving images, and the prediction model is generated by the following steps: Using the image recognition model to perform the image recognition on the plurality of historical driving images to generate a plurality of historical object information of a plurality of historical objects in the plurality of historical driving images; and The generative content model is trained based on the multiple historical object information and the multiple historical driving states to generate the prediction model.
13. The decision-making method according to claim 12, characterized in that: The step of generating the prediction model further comprises: Performing text processing on the plurality of historical object information and the plurality of historical driving states to convert them into a training text; and The generative content model is trained based on the training text to generate the prediction model.
14. The decision-making method according to claim 12, characterized in that: The storage device further stores a plurality of historical sensing data corresponding to the plurality of historical driving images, and the prediction model is generated by the following steps: The generative content model is trained based on the plurality of historical object information, the plurality of historical driving states, and the plurality of historical sensing data to generate the prediction model.
15. The decision-making method according to claim 11, characterized in that: The step of generating the driving prediction information further comprises: Performing a text processing on the plurality of object information and the driving status to convert them into an input text; and The input text is input into the prediction model to generate the driving prediction information in a text format.
16. The decision-making method according to claim 11, characterized in that: The step of generating the driving prediction information further comprises: The driving prediction information is generated by the prediction model based on the multiple object information, the driving status, a positioning information of the vehicle and a driving assistance information, wherein the multiple object information, the driving status, the positioning information and the driving assistance information are subjected to a text processing to be converted into an input text.
17. The decision-making method according to claim 11, characterized in that: Further including: comparing the driving decision with an actual driving operation of the vehicle to generate an operation difference; and In response to the operation difference being greater than a threshold, the prediction model is fine-tuned based on the driving image, the driving state, and the actual driving operation corresponding to the driving decision.
18. The decision-making method according to claim 11, characterized in that: Further including: Based on a situation text, an image generation model is used to generate a time series situation image described by the situation text.
19. The decision-making method according to claim 11, characterized in that: Further including: A control signal is generated to control a power system of the vehicle, wherein the control signal is generated based on the driving decision.
20. The decision-making method according to claim 11, characterized in that: The motion predictions of the multiple objects include future trajectory data of the multiple objects at a future time, and the motion prediction of the vehicle includes the future trajectory data of the vehicle at the future time.