Intelligent driving method and system
By judging triggering conditions in the intelligent driving system, reducing the number of interactions with remote devices, and combining the results of remote devices and local vehicle processing for path planning and control, the safety and reliability issues of intelligent driving caused by network congestion are solved, and more efficient intelligent driving is achieved.
Patent Information
- Application Number
- PCT/CN2025/082382
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2025-03-13
- Publication Date
- 2025-10-30
AI Technical Summary
Existing intelligent driving systems are prone to congestion during network interactions between vehicles and remote devices in congested periods, leading to delays in control signals and affecting the safety and reliability of intelligent driving.
By determining whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, multi-terminal collaborative intelligent driving is only carried out when the conditions are met, reducing the number of interactions with remote devices, and combining the processing results of remote devices and the vehicle's local processing for path planning and control.
It alleviated network congestion, improved the safety and reliability of intelligent driving, reduced network latency, and enhanced the system's responsiveness.
Smart Images

Figure CN2025082382_30102025_PF_FP_ABST
Abstract
Description
Intelligent driving methods and systems Technical Field
[0001] This application relates to the field of vehicle technology, and in particular to an intelligent driving method and system. Background Technology
[0002] Currently, intelligent driving technology has been widely applied in the automotive field, which can greatly assist drivers (especially novice drivers) in driving and parking, improve driving / parking safety, and alleviate traffic congestion on roads / parking lots.
[0003] With the development of large-scale model technology, its application in intelligent driving systems is gradually becoming an important trend. Although large-scale models can improve intelligent driving capabilities, they require chips with high bandwidth and computing power, which still poses a significant challenge for current in-vehicle devices. Therefore, large-scale models can be deployed on remote devices, allowing the vehicle to achieve intelligent driving through interaction with these remote devices.
[0004] In existing technologies, each time an intelligent driving function is activated, the vehicle sends data to a remote device, which processes this data and then returns a control signal to the vehicle. The vehicle can then control itself based on the control signal returned by the remote device, thus achieving intelligent driving. However, each intelligent driving process involves a relatively long interaction between the vehicle and the remote device, which can easily lead to network congestion during peak traffic hours. This can prevent the vehicle from receiving the control signal from the remote device in a timely manner, resulting in the intelligent driving system being unable to control the vehicle promptly, leading to poor safety and reliability of intelligent driving. Summary of the Invention
[0005] In view of this, this application provides an intelligent driving method and system. This intelligent driving method can, when appropriate, utilize remote devices to improve the safety and reliability of intelligent vehicle driving, while also alleviating network congestion between vehicles and remote devices during peak traffic periods.
[0006] It should be noted that the intelligent driving involved in this application includes intelligent navigation (or intelligent cruise control) and intelligent parking; wherein, intelligent navigation (or intelligent cruise control) may include automatic navigation (or automatic cruise control) and assisted navigation (or assisted cruise control), and intelligent parking may include automatic parking and assisted parking.
[0007] In a first aspect, this application provides an intelligent driving method, which includes: first, determining whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein the current situation includes at least one of the following: the scene in which the current vehicle is located, the network parameters between the current vehicle and the remote device, the reliability of the first processing result obtained by the current vehicle in processing the first sensor data, or the system state of the current vehicle; when it is determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, sending multi-terminal collaborative intelligent driving data to the remote device; wherein the multi-terminal collaborative intelligent driving data includes second sensor data, which is collected by the current vehicle after collecting the first sensor data or the second sensor data is the first sensor data; next, receiving the second processing result sent by the remote device, the second processing result being obtained by the remote device in processing the multi-terminal collaborative intelligent driving data; and then, performing path planning and control on the current vehicle based on the second processing result.
[0008] In other words, this application first determines whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving during each intelligent driving process. If the triggering conditions are met, multi-terminal collaborative intelligent driving is then implemented. If the triggering conditions are not met, the current vehicle independently performs intelligent driving. Compared to existing technologies where the current vehicle performs multi-terminal collaborative intelligent driving in every intelligent driving process, this application reduces the number of interactions between the current vehicle and remote devices, as well as the number of vehicles interacting with remote devices, thereby alleviating the computational burden on remote devices. Furthermore, it can alleviate network congestion during peak traffic periods, reducing the latency for the current vehicle to receive data from remote devices. This allows the current vehicle's intelligent driving system to promptly plan and control the vehicle's path based on the data sent by the remote devices, thereby improving the safety and reliability of intelligent driving.
[0009] For example, remote devices may include, but are not limited to, servers, other vehicles, mobile terminals, flight equipment, etc., and this application does not limit them.
[0010] For example, the specific implementation of the server in this application can be a cloud server, a physical (standalone) server, a server cluster, etc., and this application does not limit this. For example, when the remote device is a cloud server, multi-terminal collaborative intelligent driving can be called vehicle-cloud collaborative intelligent driving, that is, cloud server and vehicle collaborative intelligent driving.
[0011] For example, other vehicles may have higher permissions than the current vehicle.
[0012] For example, mobile terminals may include, but are not limited to, mobile phones, tablets, laptops, wearable devices (such as smartwatches), etc., and this application does not limit them.
[0013] For example, flying equipment can refer to devices such as drones that have low-altitude flight capabilities and can be used as relays to connect to the network. When road congestion leads to a decrease in communication capabilities, or when a network in a certain area is unable to allow vehicles to access the network normally due to an accident, flying equipment can be controlled to fly to that area / road and temporarily establish network capabilities, providing the necessary communication bandwidth and stability guarantees.
[0014] For example, the first sensing data may include N (N is a positive integer, which can be set as needed, for example, N=3 or 5) seconds of sensing data collected by the current vehicle's sensors.
[0015] For example, the first sensing data may include, but is not limited to: image data, lidar data, millimeter wave data, ultrasonic data, vehicle motion status (such as vehicle speed, vehicle direction, vehicle acceleration, etc.), etc., and this application does not limit it.
[0016] For example, the second sensing data may include N (N is a positive integer, which can be set as needed, such as N=3 or 5) seconds of sensing data collected by the current vehicle's sensors.
[0017] For example, both the first and second sensor data are sensor data from 20:00:00 to 20:00:03.
[0018] For example, the first sensing data is the sensing data from 20:00:00 to 20:00:03, and the second sensing data is the sensing data from 20:00:04 to 20:00:06.
[0019] For example, the second sensing data may include, but is not limited to: image data, lidar data, millimeter wave data, ultrasonic data, vehicle motion status (such as vehicle speed, vehicle direction, vehicle acceleration, etc.), etc., and this application does not limit it.
[0020] Exemplary examples show that the current vehicle may include, but is not limited to, sensors, display modules, and intelligent driving systems, etc., and this application does not impose any limitations on this. The sensors may collect the aforementioned first and second sensing data, and the intelligent driving system executes the aforementioned intelligent driving method. Exemplary examples show that the intelligent driving system may include, but is not limited to, a perception module, a prediction module, a fusion module, and a control module.
[0021] It should be understood that the triggering conditions for multi-terminal collaborative intelligent driving may also include other conditions, and this application does not impose any restrictions on them.
[0022] According to the first aspect, the triggering conditions for multi-terminal collaborative intelligent driving include at least one of the following: the current vehicle is in a preset scenario; the network parameters between the current vehicle and the remote device meet preset network conditions; the reliability of the first processing result meets preset reliability conditions; and the current vehicle's system state meets preset state.
[0023] For example, a preset scenario can refer to a scenario where the current vehicle cannot identify anomalies. Examples include scenarios with standing water on the road (where road surface reflection makes it difficult to identify water depth, obstacles in the water, etc.), outdoor grassy scenarios (where it's difficult to identify abnormal obstacles in the grass, such as tree stumps, branches, bricks, etc.), intersections with limited visibility (where it's difficult to determine if pedestrians or vehicles will suddenly appear), scenarios where the vehicle is close to obstacles or curbs, scenarios with long-tailed obstacles (long-tailed obstacles can refer to low-lying objects that are difficult for the current vehicle to identify (such as tree stumps in haystacks, broken glass, metal debris, etc.), small perforated objects (such as fences, wire mesh, etc.), suspended obstacles (height restriction poles, surveillance cameras, tilted signs, etc.)), and blurry scenarios. It should be understood that preset scenarios can also include other scenarios, and this application does not limit this.
[0024] For example, the network parameters between the current vehicle and the remote device may include at least one of the following: current data transmission latency, current packet loss rate, and current data transmission rate. It should be understood that the network parameters may also include other parameters, and this application does not limit this. The network parameters between the current vehicle and the remote device can be used to characterize (or describe) the network state between the current vehicle and the remote device.
[0025] For example, preset network conditions can be set as needed; for instance, data transmission latency is less than a latency threshold (e.g., 5 seconds); data packet loss rate is less than a packet loss rate threshold (e.g., 3%); and data transmission rate is greater than a transmission rate threshold (e.g., 1 Mbit / s), etc. It should be understood that preset network conditions may also include other conditions, and this application does not limit them. For example, preset network conditions can be used to determine whether the network status between the current vehicle and the remote device is good.
[0026] For example, the reliability of the first processing result can be represented by the probability, confidence level, or confidence level variance of the first processing result. Preset reliability conditions may include at least one of the following: the probability of the first processing result is greater than a probability threshold; the confidence level of the first processing result is less than a confidence level threshold; the confidence level variance of the first processing result is greater than a variance threshold. It should be understood that the reliability of the first processing result can also be represented by other information, and preset reliability conditions may also include other conditions, which can be set according to requirements, and this application does not impose any limitations on this.
[0027] For example, the preset state may include at least one of the following: the system is in a frozen state, or the intelligent driving system is in a function takeover state. It should be understood that the preset state may also include other states, and this application does not limit this.
[0028] According to the first aspect, or any implementation of the first aspect above, the method further includes: identifying the scene in which the vehicle is currently located based on the second sensing data.
[0029] For example, the second sensor data can be input into a grass scene classifier. The grass scene classifier can output the probability of "1" and the probability of "0". "1" indicates that the current vehicle is in a grass scene, and "0" indicates that the current vehicle is in a non-grass scene. When the probability of "1" output by the grass scene classifier is greater than (or equal to) the probability of "0", it can be determined that the current vehicle is in a grass scene, that is, the current vehicle is in a preset scene. When the probability of "1" output by the grass scene classifier is less than (or equal to) the probability of "0", it can be determined that the current vehicle is in a non-grass scene, that is, the current vehicle is in a non-preset scene.
[0030] For example, the second sensor data can be input into a flooded road scene classifier. The classifier can output a probability of "1" and a probability of "0". "1" indicates that the current vehicle is in a flooded road scene, and "0" indicates that the current vehicle is not in a flooded road scene. When the probability of "1" output by the flooded road scene classifier is greater than (or equal to) the probability of "0", it can be determined that the current vehicle is in a flooded road scene, meaning it is in a preset scene. When the probability of "1" output by the flooded road scene classifier is less than (or equal to) the probability of "0", it can be determined that the current vehicle is not in a flooded road scene, meaning it is not in a preset scene.
[0031] For example, the second sensor data can be input into a fuzzy scene classifier. The fuzzy scene classifier can output the probability of "1" and the probability of "0". "1" indicates that the current vehicle is in a fuzzy scene (such as a foggy scene, or a blurry image), and "0" indicates that the current vehicle is in a non-fuzzy scene. When the probability of "1" output by the fuzzy scene classifier is greater than (or equal to) the probability of "0", it can be determined that the current vehicle is in a fuzzy scene, that is, the scene the current vehicle is in is a preset scene. When the probability of "1" output by the fuzzy scene classifier is less than (or equal to) the probability of "0", it can be determined that the current vehicle is in a non-fuzzy scene, that is, the scene the current vehicle is in is not a preset scene.
[0032] For example, the second sensor data can be input into an intersection scene classifier. The intersection scene classifier can output the probability of "1" and the probability of "0". "1" indicates that the current vehicle is in an intersection scene with limited visibility, and "0" indicates that the current vehicle is in an intersection scene without limited visibility. When the probability of "1" output by the intersection scene classifier is greater than (or equal to) the probability of "0", it can be determined that the current vehicle is in an intersection scene with limited visibility, that is, the current vehicle is in a preset scene. When the probability of "1" output by the intersection scene classifier is less than (or equal to) the probability of "0", it can be determined that the current vehicle is in an intersection scene without limited visibility, that is, the current vehicle is in a preset scene.
[0033] For example, the second sensor data can be input into a long-tail obstacle classifier. The long-tail obstacle classifier can output a probability of "1" and a probability of "0". "1" indicates that the current vehicle is in a scene with long-tail obstacles, and "0" indicates that the current vehicle is in a scene without long-tail obstacles. When the probability of the long-tail obstacle classifier outputting "1" is greater than (or equal to) the probability of "0", it can be determined that the current vehicle is in a scene with long-tail obstacles, meaning the current vehicle is in a preset scene. When the probability of the long-tail obstacle classifier outputting "1" is less than (or equal to) the probability of "0", it can be determined that the current vehicle is in a scene without long-tail obstacles, meaning the current vehicle is in a scene that is not a preset scene.
[0034] According to the first aspect, or any implementation of the first aspect above, the preset reliability conditions include at least one of the following: the probability of the first processing result is greater than the probability threshold; the confidence level of the first processing result is less than the confidence level threshold; the variance of the confidence level of the first processing result is greater than the variance threshold.
[0035] According to the first aspect, or any implementation of the first aspect above, the method further includes: before determining whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, in response to the first user operation, enabling the multi-terminal collaborative processing function; in response to the second user operation, enabling the intelligent driving function.
[0036] In other words, in some examples, users need to enable multi-device collaborative processing and intelligent driving functions in advance before the vehicle can collaborate with remote devices for intelligent driving.
[0037] It should be understood that this application does not restrict whether the vehicle should first enable intelligent driving functions or multi-terminal collaborative processing functions.
[0038] According to the first aspect, or any implementation of the first aspect above, the method further includes: before determining whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, when the current vehicle parking failure is detected, enabling the intelligent driving function and displaying the multi-terminal collaborative processing function as a switch control; in response to a third user operation on the switch control, enabling the multi-terminal collaborative processing function.
[0039] In other words, in some examples, when a user's parking failure is detected, the current vehicle activates the intelligent driving function and prompts the user to perform an operation to enable the multi-terminal collaborative processing function.
[0040] According to the first aspect, or any implementation of the first aspect above, the method further includes: before determining whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, when it is detected that the current vehicle is in a parking state, enabling the intelligent driving function and displaying the multi-terminal collaborative processing function as a switch control; in response to a fourth user operation on the switch control, enabling the multi-terminal collaborative processing function.
[0041] In other words, in some examples, when the vehicle is detected to be in a parked state, the vehicle then activates the intelligent driving function and prompts the user to perform operations to enable the multi-terminal collaborative processing function.
[0042] It should be understood that users can also interact with the in-vehicle terminal via voice and gestures to instruct the in-vehicle terminal to enable multi-terminal collaborative processing and intelligent driving functions, and this application does not impose any restrictions on this.
[0043] According to the first aspect, or any implementation of the first aspect above, based on the second processing result, path planning and control of the current vehicle are performed, including: processing the third sensor data to obtain the third processing result; wherein the third sensor data is collected after the current vehicle collects the second sensor data; fusing the second processing result and the third processing result to obtain the target processing result; and performing path planning and control of the current vehicle based on the target processing result.
[0044] Since remote devices can identify difficult scenarios, hidden risks, long-tail obstacles, etc., this application combines the processing results generated by the remote devices with the processing results generated locally by the current vehicle to perform path planning and control of the current vehicle, which can improve the safety and reliability of intelligent driving and the user experience.
[0045] For example, the third sensing data may include N (N is a positive integer, which can be set as needed, such as N=3 or 5) seconds of sensing data collected by the current vehicle's sensors.
[0046] For example, the first sensor data is the sensor data from 20:00:00 to 20:00:03, the second sensor data is the sensor data from 20:00:04 to 20:00:06, and the third sensor data is the sensor data from 20:00:07 to 20:00:09.
[0047] For example, the third sensing data may include, but is not limited to: image data, lidar data, millimeter wave data, ultrasonic data, vehicle motion status (such as vehicle speed, vehicle direction, vehicle acceleration, etc.), etc., and this application does not limit it.
[0048] According to the first aspect, or any implementation of the first aspect above, the second processing result includes the detection information of the first object, and the third processing result includes the detection information of the second object; fusing the second processing result and the third processing result to obtain the target processing result includes: spatiotemporally aligning the detection information of the first object and the detection information of the second object to obtain the target processing result.
[0049] Because the current vehicle transmits multi-terminal collaborative intelligent driving data to a remote device, and the remote device processes the multi-terminal collaborative intelligent driving data to obtain the second processing result, as well as transmitting the second processing result back to the current vehicle, both require a certain amount of time. Therefore, after the current vehicle uploads the second sensor data collected in time period 1 (times A1 to A2, with A1 preceding A2) to the remote device at time A2, the current vehicle can only receive the second processing result at time B2; where time B2 is after time A2. Since the current vehicle's sensors continuously collect sensor data, and the current vehicle's intelligent driving system also continuously processes the collected sensor data, at time B2, the current vehicle's perception module processes the sensor data collected in time period 2 (times B1 to B2, with B1 preceding B2) (i.e., the third sensor data). In other words, the second processing result is obtained by processing the sensor data collected in time period 1, and the third processing result is obtained by processing the sensor data collected in time period 2. Therefore, the prediction module can align the detection information of the first object and the detection information of the second object in time and space to obtain the target processing result.
[0050] According to the first aspect, or any implementation of the first aspect above, intelligent driving is intelligent parking. The second processing result also includes parking space indication information. Based on the target processing result, path planning is performed on the current vehicle, including: determining a virtual parking space based on the target processing result and parking space indication information; determining the virtual parking space as the target parking space; and performing path planning on the current vehicle based on the target processing result and the target parking space. In this way, path planning can still be achieved even when a real target parking space cannot be identified based on sensor data; thus, the user does not need to drive the current vehicle to the vicinity of an empty parking space, and the intelligent driving system can still achieve automatic / assisted parking. This is especially suitable for scenarios where finding a target parking space is difficult in parking lots.
[0051] According to the first aspect, or any implementation of the first aspect above, the second processing result includes text description information of the current vehicle scene. The method further includes: generating and displaying prompt information based on the text description information of the current vehicle scene.
[0052] For example, a warning message can be generated based on the text description of the obstacle, such as "There is a tree stump protruding from the grass in the image, which may be rotten and its height may cause it to scrape against the vehicle chassis!".
[0053] For example, a warning message can be generated based on the text description of the obstacle, such as "There is a rock ahead of the road, and its height may cause it to scrape against the vehicle chassis." The message could include the text "Please be aware of the rock ahead of the road!" and an obstacle icon.
[0054] For example, a prompt message can be generated based on the text description "A vehicle may be entering the intersection from the right ahead" at the intersection, such as the text "Please note that a vehicle may be entering the intersection from the right ahead!" and a vehicle icon.
[0055] For example, a prompt message can be generated based on the text description of the parking space. For instance, it could indicate the number of available parking spaces on the left and right sides.
[0056] It should be understood that other prompts can also be generated based on the text description information of the current vehicle scene in the second processing result; this application does not limit this.
[0057] According to the first aspect, or any implementation of the first aspect above, the method further includes: displaying the current network connection status between the current vehicle and the remote device.
[0058] For example, the vehicle terminal can also display parking progress prompts.
[0059] It should be understood that the in-vehicle terminal in the current vehicle can also display other information, and this application does not impose any restrictions on this.
[0060] According to the first aspect, or any implementation of the first aspect above, the multi-terminal collaborative intelligent driving data further includes: at least one of the trigger condition identifier of multi-terminal collaborative intelligent driving satisfied by the current situation or the fourth processing result, wherein the fourth processing result is obtained by the current vehicle processing the second sensor data.
[0061] For example, the fourth processing result may include at least one of the following: the processing result obtained by the sensing module processing the second sensing data, the processing result obtained by the prediction module processing the processing result of the sensing module in the fourth processing result, the processing result obtained by the fusion module processing the processing results of the sensing module and the prediction module in the fourth processing result, and the processing result obtained by the planning and control module processing the processing result of the fusion module in the fourth processing result.
[0062] When multi-device collaborative intelligent driving data also includes a fourth processing result, the remote device can process the fourth processing result and the second sensor data to generate a second processing result. Compared to the prior art where the remote device generates control signals based solely on sensor data, the data relied upon in this application to generate the second processing result is richer, resulting in a more accurate second processing result. This can improve the accuracy of the control signals generated by the vehicle, thereby enhancing the safety and reliability of intelligent driving and improving the user experience.
[0063] For example, the trigger condition identifier for multi-terminal collaborative intelligent driving that is satisfied in the current situation is used to uniquely identify the trigger condition for multi-terminal collaborative intelligent driving.
[0064] For example, the triggering conditions are: the network parameters between the current vehicle and the remote device meet the preset network conditions, and the corresponding triggering condition is identified as D1. The triggering condition is: the current vehicle is in a grassy scene, and the corresponding triggering condition is identified as D2. The triggering condition is: the current vehicle is in a scene with standing water on the road, and the corresponding triggering condition is identified as D3. The triggering condition is: the current vehicle is in a blurred scene, and the corresponding triggering condition is identified as D4. ... and the triggering condition is: the confidence level of the first processing result is less than the confidence level threshold, and the corresponding triggering condition is identified as D10. The triggering condition is: the confidence level variance of the first processing result is greater than the variance threshold, and the corresponding triggering condition is identified as D11. The triggering condition is: the probability of the first processing result is greater than the probability threshold, and the corresponding triggering condition is identified as D12; and so on.
[0065] For example, the triggering condition is: the network parameters between the current vehicle and the remote device meet the preset network conditions, and the corresponding triggering condition identifier is the identifier of the first trigger. The triggering condition is: the current vehicle is in a grassy scene, and the corresponding triggering condition identifier is the identifier of the corresponding second trigger, and so on. The triggering condition is: the confidence level of the first processing result is less than the confidence level threshold, and the corresponding triggering condition identifier is the identifier of the corresponding fourth trigger; and so on.
[0066] Furthermore, multi-terminal collaborative intelligent driving data may also include multi-terminal collaborative intelligent driving requests. These requests instruct remote devices to process other data (excluding the requests themselves) contained within the multi-terminal collaborative intelligent driving data to obtain a second processing result. In other words, after receiving the multi-terminal collaborative intelligent driving data, the remote device can respond to the requests using a large model service, invoking the large model to process the other data contained within the multi-terminal collaborative intelligent driving data and obtain a second processing result. The specific processing procedure of the remote device will be explained later. Afterward, the remote device can send the second processing result to the vehicle.
[0067] According to the first aspect, or any implementation of the first aspect above, intelligent driving is intelligent parking, and the second processing result includes the attribute information of the first object; according to the second processing result, path planning and control of the current vehicle are performed, including: selecting a target parking space according to the attribute information of the first object; generating a target path for the current vehicle according to the attribute information of the first object and the target parking space; and controlling the current vehicle according to the target path of the current vehicle.
[0068] For example, the attribute information of the first object may include people, animals, hard obstacles (such as stones, steel pipes, etc.), and soft obstacles (such as plastic bags, leaves, etc.).
[0069] Since the first object may be in an empty parking space or on the road leading to the parking space, the system plans and controls the path of the current vehicle based on the attribute information of the first object. This can effectively and promptly avoid collisions between the current vehicle and people, animals, or hard obstacles during the parking process (collisions with soft obstacles are allowed because they have little or no impact on the vehicle). This further improves the safety and reliability of intelligent parking and increases the probability of successful parking.
[0070] Secondly, this application provides an intelligent driving method, which includes: first, receiving multi-terminal collaborative intelligent driving data sent by the current vehicle; wherein the multi-terminal collaborative intelligent driving data includes a trigger condition identifier for multi-terminal collaborative intelligent driving satisfied by the current situation and second sensor data; next, determining a target large model from multiple large models based on the trigger condition identifier for multi-terminal collaborative intelligent driving satisfied by the current situation; then, determining a second processing result based on the target large model and the second sensor data; subsequently, sending the second processing result to the current vehicle, the second processing result being used for path planning and control of the current vehicle.
[0071] Since the large model that matches the triggering conditions of multi-terminal collaborative intelligent driving satisfied by the current situation identifies the result (i.e. the second processing result) more closely resembles the current situation, the second processing result obtained by the remote device of this application through processing the second sensing data using the large model that matches the triggering conditions of multi-terminal collaborative intelligent driving satisfied by the current situation is more accurate. This can improve the accuracy and reliability of intelligent driving, as well as enhance the user experience.
[0072] Secondly, by selectively calling large models to process the second sensor data, the accuracy of the second processing results can be ensured while saving resources of remote devices.
[0073] Furthermore, in the process of multi-terminal collaborative intelligent driving, remote devices can dynamically acquire necessary information, i.e., the fourth processing result, based on the complexity of the road scene. In other words, the types of data contained in the multi-terminal collaborative intelligent driving data sent by the vehicle to the remote device each time can be different. This can reduce the transmission of invalid data and alleviate network pressure.
[0074] For example, a large model can refer to a machine learning model with a large number of parameters and complex computational structures. These models are typically built from deep neural networks and have billions or even hundreds of billions of parameters. Large models are designed to improve the expressive power and predictive performance of the model, enabling them to handle more complex tasks and data. For example, multiple large models can be deployed on remote devices.
[0075] For example, large models deployed in remote devices may include Visual Language Models (VLMs). A Visual Language Model (VLM) is a neural network model that, based on a Large Language Model (LLM), adds image and video input encoding to the input side, and outputs results capable of understanding and regenerating images and videos. Specifically, the Large Language Model (LLM) is used for natural language generation and understanding, stacking numerous network layers and employing self-supervised learning to acquire dialogue and reasoning capabilities.
[0076] For example, large models deployed in remote devices can include multimodal large language models (MLLMs). A multimodal large language model refers to a neural network model that, based on a large language model, adds encodings of various sensor data such as visual images, videos, speech, LiDAR point cloud data, and millimeter-wave point cloud data to the input side, and outputs results that can understand and regenerate multiple sensor data.
[0077] For example, large models deployed in remote devices can include Visual Foundation Models (VFMs). A Visual Foundation Model is a neural network model that is pre-trained on a large amount of data for visual perception tasks (such as image detection, segmentation, depth estimation, etc.) and has the ability to generalize to multiple categories or scenes.
[0078] For example, large models deployed in remote devices can include task-specific large models. Task-specific large models can refer to neural network models used to perform any single task (such as depth estimation, image segmentation, image detection, etc.).
[0079] It should be understood that other large models can also be deployed on remote devices, and this application does not impose any restrictions on this.
[0080] For example, one or more models applied to the data processing can be pre-set for each of the various triggering conditions for multi-terminal collaborative intelligent driving.
[0081] For example, for the trigger condition of multi-terminal collaborative intelligent driving: "The current scene of the vehicle is a grass scene", the large model used to process the data is set as a visual language large model.
[0082] For example, regarding the triggering condition for multi-terminal collaborative intelligent driving: "the confidence level of the processing result output by the perception module is lower than the confidence level threshold", the large model applied to the data processing is set as the visual basic large model.
[0083] For example, regarding the triggering condition for multi-terminal collaborative intelligent driving: "The current vehicle is in a scenario where it is close to obstacles or curbs", the large models used to process the data are set as a visual language large model and a visual basic large model.
[0084] According to the second aspect, based on the triggering conditions of multi-terminal collaborative intelligent driving satisfied in the current situation, a target large model is determined from multiple large models, including: finding a mapping relationship based on the triggering condition identifier of multi-terminal collaborative intelligent driving satisfied in the current situation, so as to determine one or more target large models from multiple large models; wherein, the mapping relationship includes the relationship between each of the multiple triggering conditions of multi-terminal collaborative intelligent driving and the corresponding one or more large models.
[0085] It should be noted that multiple triggering conditions can correspond to a large model. That is, the mapping relationship can include the relationship between a triggering condition and a large model (i.e., one-to-one), the relationship between a triggering condition and multiple large models (i.e., one-to-many), and the relationship between multiple triggering conditions and a large model (i.e., many-to-one).
[0086] According to the second aspect, or any implementation of the second aspect above, the multi-terminal collaborative intelligent driving data also includes a fourth processing result, which is obtained by the current vehicle processing the second sensing data; based on the target large model and the second sensing data, the second processing result is determined, including: fusing the second sensing data and the fourth processing result to obtain an intermediate processing result; and using the target large model to process the intermediate processing result to obtain the second processing result.
[0087] Compared to simply inputting the second sensor data into the target large model, this approach allows for the use of more comprehensive and richer information to generate more accurate second processing results, thereby improving the reliability and safety of intelligent driving.
[0088] According to the second aspect, or any implementation of the second aspect above, the multi-terminal collaborative intelligent driving data also includes a fourth processing result, which is obtained by the current vehicle processing the second sensor data; based on the target large model and the second sensor data, the second processing result is determined, including: obtaining the level of the second sensor data and the level of the fourth processing result; according to the level of the second sensor data and the level of the fourth processing result, the target large model is used to process the second sensor data and the fourth processing result to obtain the second processing result.
[0089] According to the second aspect, or any implementation of the second aspect above, the fourth processing result includes N sets of data, each of which corresponds to one of the N levels. The second sensing data corresponds to one level, and N is a positive integer. Based on the level of the second sensing data and the level of the fourth processing result, the target large model is used to process the second sensing data and the fourth processing result to obtain the second processing result, including: cyclically executing the following steps until i equals N, where the initial value of i is 1; inputting the i-th set of data, the (i-1)-th intermediate result, and the second sensing data from the fourth processing result into the target large model to obtain the i-th intermediate result; where the i-th set of data corresponds to level i, and the 0th intermediate result is obtained by inputting the second sensing data into the target large model; incrementing i by 1; when i equals N, determining the second processing result based on the N intermediate results from the 1st intermediate result to the Nth intermediate result.
[0090] Specifically, if the (i-1)th intermediate result does not contain key information (such as obstacles, parking spaces, risk information, etc.), the data from the i-th group of the fourth processing result, the (i-1)th intermediate result, and the second sensor data can be input into the target large model to obtain the i-th intermediate result. This allows for the rapid identification of risks in the current vehicle's scenario, reducing processing steps, and ensures that the large model does not overlook risks in key areas of concern to the current vehicle, thus improving the reliability and safety of intelligent driving.
[0091] For example, the i-th group of data, the (i-1)-th group of intermediate results and the second sensing data in the fourth processing result are input into the target large model to obtain the i-th group of intermediate results. This can be understood as using the i-th group of data and the (i-1)-th group of intermediate results as auxiliary information, and the target large model combines this auxiliary information to process the second sensing data to obtain the second processing result.
[0092] According to the second aspect, or any implementation thereof, the method further includes: receiving sensing data sent by the field-end device; determining a second processing result based on the target large model and the second sensing data, including: processing the second sensing data and the sensing data sent by the field-end device using the target large model to obtain the second processing result. This allows for the use of more comprehensive and richer information to generate a more accurate second processing result, thereby improving the reliability and safety of intelligent driving.
[0093] For example, the second processing result is obtained by using the target large model to process the second sensing data sent by the current vehicle and the sensing data sent by the field device. This can be understood as: the sensing data sent by the field device is regarded as auxiliary information, and the target large model combines the auxiliary information to process the second sensing data to obtain the second processing result.
[0094] According to the second aspect, or any implementation thereof, a second processing result is determined based on the target large model and the second sensor data. This includes: processing the second sensor data and historical data using the target large model to obtain the second processing result; wherein the historical data includes historical sensor data and / or historical processing results. This provides more comprehensive and richer information for generating the second processing result, improving its accuracy and thus enhancing the reliability and safety of intelligent driving.
[0095] In addition, the method of "using the target large model to process the second sensor data and historical data to obtain the second processing result" can introduce historical data into the processing of the target large model. In this way, there is no need to use these historical data to retrain the large model, which can save training costs.
[0096] For example, the second processing result is obtained by using a target large model to process the second sensing data and historical data. This can be understood as treating the historical data as auxiliary information, and the target large model combines this auxiliary information to process the second sensing data to obtain the second processing result.
[0097] According to the second aspect, or any implementation of the second aspect above, the second processing result includes: the detection information of the first object, the text description information of the scene where the current vehicle is located, and the attribute information of the first object.
[0098] Thirdly, this application provides a multi-terminal cooperative intelligent driving system, which includes a vehicle and a remote device, wherein:
[0099] The vehicle is used to determine whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein the current situation includes at least one of the following: the scene in which the vehicle is located, the network parameters between the vehicle and the remote device, the reliability of the first processing result obtained by the vehicle from processing the first sensor data, or the system state of the vehicle.
[0100] The vehicle is also used to send multi-terminal collaborative intelligent driving data to remote devices when it is determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein, the multi-terminal collaborative intelligent driving data includes second sensor data, which is collected by the vehicle after collecting the first sensor data or the second sensor data is the first sensor data;
[0101] The remote device is used to process the second sensor data, obtain the second processing result, and send the second processing result to the vehicle.
[0102] The vehicle is also used for path planning and control based on the second processing result.
[0103] According to the third aspect, the multi-terminal collaborative intelligent driving data also includes the trigger condition identifier of multi-terminal collaborative intelligent driving that is satisfied in the current situation; the remote device is also used to determine the target large model from multiple large models based on the trigger condition identifier of multi-terminal collaborative intelligent driving that is satisfied in the current situation; and to determine the second processing result based on the target large model and the second sensing data.
[0104] It should be understood that vehicles in a multi-terminal collaborative intelligent driving system can also execute the steps in any of the implementation methods in the first aspect, which will not be elaborated here.
[0105] It should be understood that remote devices in a multi-terminal collaborative intelligent driving system can also execute the steps in any two of the second aspects, which will not be elaborated here.
[0106] The third aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the third aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.
[0107] The third aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects of the third aspect and any implementation thereof can be found in the technical effects of the second aspect and any implementation thereof, as described above, and will not be repeated here.
[0108] Fourthly, this application also provides a vehicle for:
[0109] Determine whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein, the current situation includes at least one of the following: the scene in which the vehicle is located, the network parameters between the vehicle and the remote device, the reliability of the first processing result obtained by the vehicle in processing the first sensor data, or the system state of the vehicle.
[0110] When it is determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, multi-terminal collaborative intelligent driving data is sent to the remote device; wherein, the multi-terminal collaborative intelligent driving data includes second sensor data, which is collected by the vehicle after collecting the first sensor data or the second sensor data is the first sensor data.
[0111] Receive the second processing result sent by the remote device. The second processing result is obtained by the remote device processing multi-terminal collaborative intelligent driving data.
[0112] Based on the second processing result, the vehicle's path is planned and controlled.
[0113] It should be understood that vehicles using multi-terminal collaborative intelligent driving systems can also execute the steps in any of the implementation methods in the first aspect, which will not be elaborated here.
[0114] The fourth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the fourth aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.
[0115] Fifthly, this application also provides a remote device for:
[0116] Receive multi-terminal collaborative intelligent driving data sent by the current vehicle; wherein, the multi-terminal collaborative intelligent driving data includes the trigger condition identifier of multi-terminal collaborative intelligent driving that is met in the current situation and the second sensor data;
[0117] Based on the triggering conditions for multi-terminal collaborative intelligent driving that are met in the current situation, the target large model is determined from multiple large models;
[0118] The second processing result is determined based on the target large model and the second sensor data;
[0119] Send the second processing result to the current vehicle. The second processing result is used to perform path planning and control for the current vehicle.
[0120] The fifth aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects of the fifth aspect and any implementation thereof are similar to those of the second aspect and any implementation thereof, and will not be repeated here.
[0121] In a sixth aspect, this application provides an in-vehicle terminal, including: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, which, when executed by the processor, cause the in-vehicle terminal to perform the method in the first aspect or any possible implementation thereof.
[0122] The sixth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the sixth aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.
[0123] In a seventh aspect, this application provides a remote device, including: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, which, when executed by the processor, cause the remote device to perform the method in the second aspect or any possible implementation thereof.
[0124] The seventh aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects corresponding to the seventh aspect and any implementation thereof are similar to those corresponding to the second aspect and any implementation thereof, and will not be repeated here.
[0125] Eighthly, this application provides a chip including one or more interface circuits and one or more processors; the one or more processors receive or transmit data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the first aspect or any possible implementation of the first aspect are executed.
[0126] The eighth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects corresponding to the eighth aspect and any implementation thereof are similar to those corresponding to the first aspect and any implementation thereof, and will not be repeated here.
[0127] Ninthly, this application provides a chip including one or more interface circuits and one or more processors; the one or more processors receive or transmit data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the second aspect or any possible implementation of the second aspect are executed.
[0128] The ninth aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects corresponding to the ninth aspect and any implementation thereof are similar to those corresponding to the second aspect and any implementation thereof, and will not be repeated here.
[0129] In a tenth aspect, this application provides a computer-readable storage medium storing a computer program that, when run on a computer or processor, causes the computer or processor to perform the method of the first aspect or any possible implementation thereof.
[0130] The tenth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects corresponding to the tenth aspect and any implementation thereof are similar to those corresponding to the first aspect and any implementation thereof, and will not be repeated here.
[0131] In one aspect, this application provides a computer-readable storage medium storing a computer program that, when run on a computer or processor, causes the computer or processor to perform the method in the second aspect or any possible implementation thereof.
[0132] The eleventh aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects corresponding to the eleventh aspect and any implementation thereof can be found in the technical effects corresponding to the second aspect and any implementation thereof, as described above, and will not be repeated here.
[0133] In a twelfth aspect, this application provides a computer program product including computer instructions that, when executed by a computer or processor, cause the computer or processor to perform the method in the first aspect or any possible implementation thereof.
[0134] The twelfth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the twelfth aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.
[0135] In a thirteenth aspect, this application provides a computer program product including computer instructions that, when executed by a computer or processor, cause the computer or processor to perform the method in the second aspect or any possible implementation thereof.
[0136] The thirteenth aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects corresponding to the thirteenth aspect and any implementation thereof are similar to those corresponding to the second aspect and any implementation thereof, and will not be repeated here. Attached Figure Description
[0137] Figure 1A is a schematic diagram illustrating an exemplary application scenario;
[0138] Figure 1B is a schematic diagram illustrating an exemplary application scenario;
[0139] Figure 1C is a schematic diagram of the structure of an exemplary multi-terminal cooperative intelligent driving system;
[0140] Figure 2 is a schematic diagram of an exemplary intelligent driving process 200;
[0141] Figure 3 is a schematic diagram illustrating an exemplary intelligent driving process 300;
[0142] Figure 4A is a schematic diagram of the interface of an exemplary vehicle terminal;
[0143] Figure 4B is a schematic diagram of the interface of an exemplary vehicle terminal;
[0144] Figure 4C is a schematic diagram of the interface of an in-vehicle terminal as an example;
[0145] Figure 5 is a schematic diagram of an exemplary intelligent driving process 500;
[0146] Figure 6A is a schematic diagram of the interface of an exemplary vehicle terminal;
[0147] Figure 6B is a schematic diagram of the interface of an exemplary vehicle terminal;
[0148] Figure 6C is an exemplary schematic diagram of the interface of an in-vehicle terminal;
[0149] Figure 6D is a schematic diagram of the interface of an exemplary vehicle terminal;
[0150] Figure 7 is a schematic diagram of an exemplary intelligent driving process 700;
[0151] Figure 8A is a schematic diagram of the interface of an exemplary vehicle terminal;
[0152] Figure 8B is a schematic diagram of the interface of an exemplary vehicle terminal;
[0153] Figure 8C is a schematic diagram of the interface of an exemplary vehicle terminal;
[0154] Figure 8D is a schematic diagram of the interface of an exemplary vehicle terminal;
[0155] Figure 8E is an exemplary schematic diagram of the interface of an in-vehicle terminal;
[0156] Figure 8F is an exemplary schematic diagram of the interface of the vehicle terminal;
[0157] Figure 8G is an exemplary schematic diagram of the interface of an in-vehicle terminal;
[0158] Figure 9 is a schematic diagram illustrating an exemplary intelligent driving process 900;
[0159] Figure 10 is a schematic diagram of the structure of an exemplary intelligent driving system;
[0160] Figure 11 is a schematic diagram of the structure of an exemplary device;
[0161] Figure 12 is a schematic diagram of the structure of an exemplary electronic device. Detailed Implementation
[0162] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0163] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0164] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.
[0165] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0166] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.
[0167] For example, the intelligent driving involved in this application includes intelligent navigation (or intelligent cruise control) and intelligent parking; wherein, intelligent navigation (or intelligent cruise control) may include automatic navigation (or automatic cruise control) and assisted navigation (or assisted cruise control), and intelligent parking may include automatic parking and assisted parking.
[0168] Figure 1A is a schematic diagram illustrating an exemplary application scenario. The application scenario shown in Figure 1A is an application scenario of automatic (or assisted) navigation.
[0169] Referring to Figure 1A, exemplarily, while the vehicle is driving on urban roads, the driver can activate the automatic (or assisted) navigation function of the vehicle's onboard terminal. In this way, the onboard terminal can execute the intelligent driving method involved in this application to plan and control the vehicle's path, thereby achieving automatic (or assisted) navigation. This can assist the driver, improve driving safety, and alleviate road congestion.
[0170] Figure 1B is a schematic diagram illustrating an exemplary application scenario. The application scenario shown in Figure 1B is an automatic (or assisted) parking application scenario.
[0171] Referring to Figure 1B, exemplarily, during the parking process of a vehicle in a parking lot, the driver can activate the automatic (or assisted) parking function of the vehicle's onboard terminal. In this way, the onboard terminal can execute the intelligent driving method involved in this application to plan and control the vehicle's path, thereby achieving automatic (or assisted) parking. This can assist the driver in parking, improve parking efficiency and safety, and alleviate congestion in parking lots.
[0172] Figure 1C is a schematic diagram of the structure of an exemplary multi-terminal collaborative intelligent driving system.
[0173] Referring to Figure 1C, exemplarily, in one possible scenario, the multi-terminal cooperative intelligent driving system may include a remote device and a vehicle. In this case, multi-terminal cooperative intelligent driving can be understood as two-terminal cooperative intelligent driving, that is, cooperative intelligent driving between a remote device and a vehicle.
[0174] For example, remote devices may include, but are not limited to, servers, edge devices, etc., and this application does not limit them.
[0175] For example, the specific implementation of the server in this application can be a cloud server, a physical (standalone) server, a server cluster, etc., and this application does not limit this. For example, when the remote device is a cloud server, multi-terminal collaborative intelligent driving can be called vehicle-cloud collaborative intelligent driving, that is, cloud server and vehicle collaborative intelligent driving.
[0176] For example, when road congestion leads to a decrease in communication capacity, or when a network in a certain area is unable to allow the current vehicle to access the network normally due to an accident, other vehicles, mobile terminals, or flying devices (which can be controlled to fly to that area / road) can be used as relay devices to provide relay capabilities, thereby building a network between the vehicle, the relay device, and the remote device, enabling the vehicle to connect with the remote device.
[0177] For example, other vehicles may have higher permissions than the current vehicle.
[0178] For example, mobile terminals may include, but are not limited to, mobile phones, tablets, laptops, wearable devices (such as smartwatches), etc., and this application does not limit them.
[0179] For example, flying equipment can refer to devices such as drones that have low-altitude flight capabilities and can provide relay capabilities for networking.
[0180] For example, in one possible scenario, the multi-terminal collaborative intelligent driving system may include remote devices, site-side devices, and vehicles. In this case, multi-terminal collaborative intelligent driving can be understood as three-terminal collaborative intelligent driving, namely, collaborative intelligent driving of remote devices, site-side devices, and vehicles.
[0181] For example, the field-end equipment can be deployed in public parking lots; it can also be deployed in garages of residential communities, office buildings, etc., near street parking spaces; and near intersections, congested roads, and roads prone to traffic accidents. The sensors in the field-end equipment can include, but are not limited to: network-enabled cameras, LiDAR, millimeter-wave radar, or flying devices with image and LiDAR data acquisition capabilities, etc., and this application does not impose any limitations on these. It should be understood that Figure 1C is only an example of a field-end equipment, and the field-end equipment can include more components than shown in Figure 1C, and this application does not impose any limitations on these.
[0182] It should be understood that field-end equipment is an optional component in multi-terminal collaborative intelligent driving systems.
[0183] Referring again to Figure 1C, exemplarily, a large model service and a large model are deployed in the remote device. The large model service can call the large model to process the data sent by the vehicle (as well as the data sent by the field device (in the case of a multi-terminal collaborative intelligent driving system, the field device is also included)), obtain the processing results, and return them to the vehicle.
[0184] For example, a large model can refer to a machine learning model with a large number of parameters and complex computational structures. These models are typically built from deep neural networks and have billions or even hundreds of billions of parameters. Large models are designed to improve the expressive power and predictive performance of the model, enabling them to handle more complex tasks and data. For example, multiple large models can be deployed on remote devices.
[0185] For example, large models deployed in remote devices may include Visual Language Models (VLMs). A Visual Language Model (VLM) is a neural network model that, based on a Large Language Model (LLM), adds image and video input encoding to the input side, and outputs results capable of understanding and regenerating images and videos. Specifically, the Large Language Model (LLM) is used for natural language generation and understanding, stacking numerous network layers and employing self-supervised learning to acquire dialogue and reasoning capabilities.
[0186] For example, large models deployed in remote devices can include multimodal large language models (MLLMs). A multimodal large language model refers to a neural network model that, based on a large language model, adds encodings of various sensor data such as visual images, videos, speech, LiDAR point cloud data, and millimeter-wave point cloud data to the input side, and outputs results that can understand and regenerate multiple sensor data.
[0187] For example, large models deployed in remote devices can include Visual Foundation Models (VFMs). A Visual Foundation Model is a neural network model that is pre-trained on a large amount of data for visual perception tasks (such as image detection, segmentation, depth estimation, etc.) and has the ability to generalize to multiple categories or scenes.
[0188] For example, large models deployed in remote devices can include task-specific large models. Task-specific large models can refer to neural network models used to perform any single task (such as depth estimation, image segmentation, image detection, etc.).
[0189] It should be understood that other large models can also be deployed in the remote device, and this application does not limit this. Furthermore, Figure 1C is only an example of a remote device, and the remote device may include more or fewer components than shown in Figure 1C, and this application does not limit this.
[0190] Referring again to Figure 1C, exemplarily, a vehicle may include, but is not limited to, sensors, a display module, and an intelligent driving system, etc., which this application does not limit. In one possible embodiment, the display module may be part of an in-vehicle terminal (the intelligent driving system is also part of the in-vehicle terminal). In another possible embodiment, the display module may also be a mobile terminal connected to the in-vehicle terminal, such as a mobile phone, tablet, smartwatch, etc. It should be understood that Figure 1C is only an example of a vehicle, and a vehicle may include more or fewer components than shown in Figure 1C, which this application does not limit.
[0191] For example, the sensors may include, but are not limited to, cameras, speed sensors, lidar, ultrasonic modules, accelerometers, steering wheel angle sensors, etc., and this application does not limit them.
[0192] For example, the display module may include a display screen, which may be used to display a human machine interface (HMI).
[0193] For example, an intelligent driving system may include an intelligent triggering module, a perception module, a prediction module, a fusion module, and a control module. It should be understood that the intelligent driving system of Figure 1C is merely an example of this application, and the intelligent driving system of this application may include more or fewer modules than shown in Figure 1C, may combine two or more components (or modules), or may have different component configurations. The various components (or modules) shown in Figure 1C can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits. For example, the intelligent triggering module, perception module, prediction module, fusion module, and control module can be implemented by a CPU or by a dedicated chip; this application does not limit this.
[0194] For example, the process of intelligent driving in a vehicle coordinating with remote devices can be as follows:
[0195] On the one hand, the perception module can acquire the sensing data collected by the sensor, process the sensing data, and output the processing results to the prediction module and the fusion module.
[0196] On the other hand, the intelligent triggering module acquires sensor data and other data collected by the sensors, and then determines whether to trigger collaborative intelligent driving between the vehicle and the remote device based on the sensor data and / or other data. When the intelligent triggering module determines to trigger collaborative intelligent driving between the vehicle and the remote device, it can send the sensor data to the remote device. The remote device can then process the sensor data sent by the vehicle and obtain the processing result; subsequently, the remote device can send the processing result to the prediction module and / or the fusion module.
[0197] Optionally, when the intelligent triggering module determines that the vehicle and the remote device are cooperating in intelligent driving, the intelligent triggering module can instruct any one or more of the perception module, prediction module, fusion module, and planning and control module to send the corresponding processing results to the remote device. In this case, the remote device can process the sensor data and processing results sent by the vehicle to obtain the processing result; then, the remote device can send the processing result to the prediction module and / or the fusion module.
[0198] Optionally, when the remote device, during the processing of sensor data sent by the vehicle, determines that it needs the processing results of any one or more modules among the perception module, prediction module, fusion module, and planning and control module, it can send a request to the vehicle. Upon receiving the request, any one or more modules among the perception module, prediction module, fusion module, and planning and control module can send the corresponding processing results to the remote device. In this scenario, the remote device can process the sensor data and processing results sent by the vehicle to obtain the processing results; subsequently, the remote device can send the processing results to the prediction module and / or fusion module.
[0199] Optionally, after receiving the sensor data sent by the vehicle (or when the remote device determines that it needs the sensor data from the field device while processing the sensor data sent by the vehicle), it can send a request to the field device. Upon receiving the request, the field device can send the corresponding sensor data to the remote device. In this case, the remote device can process the sensor data sent by the vehicle and the sensor data sent by the field device to obtain a processing result; subsequently, the remote device can send the processing result to the prediction module and / or the fusion module.
[0200] Optionally, after receiving the sensor data and processing results sent by the vehicle, as well as the sensor data sent by the field device, the remote device can process the sensor data and processing results sent by the vehicle and the sensor data sent by the field device to obtain the processing result. Then, the remote device can send the processing result to the prediction module and / or the fusion module.
[0201] In one possible scenario, the prediction module receives the processing result sent by the remote device. In this case, the prediction module can fuse the processing result output by the sensing module and the processing result sent by the remote device to obtain the target processing result; then, the prediction module can process the target processing result and output the processing result to the fusion module.
[0202] In one possible scenario, the prediction module may not receive the processing result from the remote device. In this case, the prediction module can process the processing result output by the perception module and output the result to the fusion module.
[0203] In one possible scenario, the fusion module receives the processing results sent by the remote device. In this case, the fusion module can process the processing results output by the prediction module, the perception module, and the remote device, and output the processing results to the planning and control module.
[0204] In one possible scenario, the fusion module may not receive the processing results from the remote device. In this case, the fusion module can process the processing results output by the prediction module and the perception module, and then output the processing results to the planning and control module.
[0205] For example, the planning and control module can perform path planning and control for the current vehicle based on the processing results output by the fusion module.
[0206] For example, when the intelligent triggering module determines not to trigger collaborative intelligent driving between the vehicle and remote devices, the intelligent driving system in the vehicle can perform intelligent driving independently. The specific process is as follows: The perception module acquires sensor data collected by sensors, processes the data, and outputs the processing results to the prediction module and the fusion module. The prediction module processes the processing results output by the perception module and outputs the processing results to the fusion module. The fusion module processes the processing results output by the prediction module and the perception module, and outputs the processing results to the planning and control module. The planning and control module can perform path planning and control for the current vehicle based on the processing results output by the fusion module.
[0207] Based on Figure 1C, the following describes the process of intelligent driving in collaboration between the vehicle and remote equipment.
[0208] Figure 2 is a schematic diagram illustrating an exemplary intelligent driving process 200. Process 200 is the vehicle processing procedure during multi-terminal collaborative intelligent driving.
[0209] S201, determine whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein, the current situation includes at least one of the following: the scene in which the current vehicle is located, the network parameters between the current vehicle and the remote device, the reliability of the first processing result obtained by the current vehicle in processing the first sensor data, or the system state of the current vehicle.
[0210] For example, during each intelligent driving process, the intelligent triggering module of the intelligent driving system can execute S201, which is to determine whether to trigger multi-terminal collaborative intelligent driving.
[0211] For example, trigger conditions for multi-terminal collaborative intelligent driving can be preset. The intelligent triggering module can then determine whether to trigger multi-terminal collaborative intelligent driving by judging whether the current situation meets these trigger conditions. When it is determined that the current situation meets the trigger conditions, multi-terminal collaborative intelligent driving can be triggered.
[0212] For example, the triggering conditions for multi-terminal collaborative intelligent driving may include at least one of the following: the current vehicle is in a preset scenario; the network parameters between the current vehicle and the remote device meet preset network conditions; the reliability of the first processing result meets preset reliability conditions; and the system state of the current vehicle meets preset states. It should be understood that the triggering conditions for multi-terminal collaborative intelligent driving may also include other conditions, which are not limited in this application.
[0213] For example, the first processing result is obtained by the current vehicle processing the first sensing data. It should be noted that the first processing result may be obtained by the perception module in the intelligent driving system processing the first sensing data, or it may be obtained by other modules in the vehicle (such as the risk identification module) processing the first sensing data; that is, this application does not limit which module in the vehicle processes the first sensing data to obtain the first processing result.
[0214] For example, the first sensing data may include N (N is a positive integer, which can be set as needed, for example, N=3 or 5) seconds of sensing data collected by the vehicle's sensors.
[0215] For example, the first sensing data may include, but is not limited to: image data, lidar data, millimeter wave data, ultrasonic data, vehicle motion status (such as vehicle speed, vehicle direction, vehicle acceleration, etc.), etc., and this application does not limit it.
[0216] For example, a preset scenario can refer to a scenario where the current vehicle cannot identify anomalies. Examples include scenarios with standing water on the road (where road surface reflection makes it difficult to identify water depth, obstacles in the water, etc.), outdoor grassy scenarios (where it's difficult to identify abnormal obstacles in the grass, such as tree stumps, branches, bricks, etc.), intersections with limited visibility (where it's difficult to determine if pedestrians or vehicles will suddenly appear), scenarios where the vehicle is close to obstacles or curbs, scenarios with long-tailed obstacles (long-tailed obstacles can refer to low-lying objects that are difficult for the current vehicle to identify (such as tree stumps in haystacks, broken glass, metal debris, etc.), small perforated objects (such as fences, wire mesh, etc.), suspended obstacles (height restriction poles, surveillance cameras, tilted signs, etc.)), and blurry scenarios. It should be understood that preset scenarios can also include other scenarios, and this application does not limit this.
[0217] For example, the network parameters between the current vehicle and the remote device may include at least one of the following: current data transmission latency, current packet loss rate, and current data transmission rate. It should be understood that the network parameters may also include other parameters, and this application does not limit this. The network parameters between the current vehicle and the remote device can be used to characterize (or describe) the network state between the current vehicle and the remote device.
[0218] For example, preset network conditions can be set as needed; for instance, data transmission latency is less than a latency threshold (e.g., 5 seconds); data packet loss rate is less than a packet loss rate threshold (e.g., 3%); and data transmission rate is greater than a transmission rate threshold (e.g., 1 Mbit / s), etc. It should be understood that preset network conditions may also include other conditions, and this application does not limit them. For example, preset network conditions can be used to determine whether the network status between the current vehicle and the remote device is good.
[0219] For example, the reliability of the first processing result can be represented by the probability, confidence level, or confidence level variance of the first processing result. Preset reliability conditions may include at least one of the following: the probability of the first processing result is greater than a probability threshold; the confidence level of the first processing result is less than a confidence level threshold; the confidence level variance of the first processing result is greater than a variance threshold. It should be understood that the reliability of the first processing result can also be represented by other information, and the preset reliability conditions may also include other conditions, which can be set according to requirements, and this application does not impose any limitations on this.
[0220] For example, the preset state may include at least one of the following: the system is in a frozen state, or the intelligent driving system is in a function takeover state. It should be understood that the preset state may also include other states, and this application does not limit this.
[0221] For example, the intelligent triggering module may include a first trigger. The first trigger can be implemented using a logical judgment algorithm, and can also be called a first judge. For example, the network parameters between the current vehicle and the remote device can be input to the first trigger; then, the first trigger can determine whether the network parameters between the current vehicle and the remote device meet preset network conditions. For example, the first trigger can determine whether the current data transmission rate is greater than (or equal to) a transmission rate threshold; for another example, the first trigger can determine whether the current data packet loss rate is less than (or equal to) a packet loss rate threshold; and for yet another example, the first trigger can determine whether the current data transmission delay is less than (or equal to) a delay threshold, etc. If the current data transmission delay is less than (or equal to) a delay threshold, and / or the current data packet loss rate is less than (or equal to) a packet loss rate threshold, and / or the current data transmission rate is greater than (or equal to) a transmission rate threshold, it can be determined that the network parameters between the current vehicle and the remote device meet the preset network conditions; that is, the network status between the current vehicle and the remote device is good.
[0222] For example, the intelligent processing module can acquire second sensing data collected by the sensors; then, based on the second sensing data, it can identify the current scene of the vehicle. Afterwards, it can determine whether the current scene of the vehicle is a preset scene. Here, the second sensing data is either collected by the vehicle after acquiring the first sensing data, or the second sensing data is the same as the first sensing data.
[0223] For example, the second sensing data may include N (N is a positive integer, which can be set as needed, for example, N=3 or 5) seconds of sensing data collected by the vehicle's sensors.
[0224] For example, both the first and second sensor data are sensor data from 20:00:00 to 20:00:03.
[0225] For example, the first sensing data is the sensing data from 20:00:00 to 20:00:03, and the second sensing data is the sensing data from 20:00:04 to 20:00:06.
[0226] For example, the second sensing data may include, but is not limited to: image data, lidar data, millimeter wave data, ultrasonic data, vehicle motion status (such as vehicle speed, vehicle direction, vehicle acceleration, etc.), etc., and this application does not limit it.
[0227] For example, the intelligent processing module may include multiple second triggers; each second trigger may be implemented using a scene classification algorithm, and the second trigger may also be called a scene classifier. The scene classifier can be used to determine whether the scene in which the current vehicle is located is a preset scene. For example, multiple classifiers may include a grass scene classifier, a waterlogged road scene separator, a fuzzy scene classifier, an intersection scene classifier, a long-tail obstacle classifier, a fuzzy scene classifier, etc., and this application does not limit this.
[0228] For example, the second sensor data can be input into a grass scene classifier; the grass scene classifier can output the probability of "1" and the probability of "0"; where "1" indicates that the current vehicle is in a grass scene, and "0" indicates that the current vehicle is in a non-grass scene. When the probability of "1" output by the grass scene classifier is greater than (or equal to) the probability of "0", it can be determined that the current vehicle is in a grass scene, that is, the current vehicle is in a preset scene. When the probability of "1" output by the grass scene classifier is less than (or equal to) the probability of "0", it can be determined that the current vehicle is in a non-grass scene, that is, the current vehicle is in a non-preset scene.
[0229] For example, the second sensor data can be input into a flooded road scene classifier. The flooded road scene classifier can output a probability of "1" and a probability of "0". "1" indicates that the current vehicle is in a flooded road scene, and "0" indicates that the current vehicle is not in a flooded road scene. When the probability of "1" output by the flooded road scene classifier is greater than (or equal to) the probability of "0", it can be determined that the current vehicle is in a flooded road scene, that is, the current vehicle is in a preset scene. When the probability of "1" output by the flooded road scene classifier is less than (or equal to) the probability of "0", it can be determined that the current vehicle is not in a flooded road scene, that is, the current vehicle is in a preset scene.
[0230] For example, the second sensor data can be input into a fuzzy scene classifier. The fuzzy scene classifier can output the probability of "1" and the probability of "0". "1" indicates that the current vehicle is in a fuzzy scene (such as a foggy scene, or a blurry image), and "0" indicates that the current vehicle is in a non-fuzzy scene. When the probability of "1" output by the fuzzy scene classifier is greater than (or equal to) the probability of "0", it can be determined that the current vehicle is in a fuzzy scene, that is, the current vehicle is in a preset scene. When the probability of "1" output by the fuzzy scene classifier is less than (or equal to) the probability of "0", it can be determined that the current vehicle is in a non-fuzzy scene, that is, the current vehicle is in a non-preset scene.
[0231] For example, the second sensor data can be input into an intersection scene classifier. The intersection scene classifier can output the probability of "1" and the probability of "0". "1" indicates that the current vehicle is in an intersection scene with limited visibility, and "0" indicates that the current vehicle is in an intersection scene without limited visibility. When the probability of "1" output by the intersection scene classifier is greater than (or equal to) the probability of "0", it can be determined that the current vehicle is in an intersection scene with limited visibility, that is, the current vehicle is in a preset scene. When the probability of "1" output by the intersection scene classifier is less than (or equal to) the probability of "0", it can be determined that the current vehicle is in an intersection scene without limited visibility, that is, the current vehicle is in a preset scene.
[0232] For example, the second sensor data can be input into a long-tail obstacle classifier. The long-tail obstacle classifier can output a probability of "1" and a probability of "0". "1" indicates that the current vehicle is in a scene with long-tail obstacles, and "0" indicates that the current vehicle is in a scene without long-tail obstacles. When the probability of "1" output by the long-tail obstacle classifier is greater than (or equal to) the probability of "0", it can be determined that the current vehicle is in a scene with long-tail obstacles, that is, the current vehicle is in a preset scene. When the probability of "1" output by the long-tail obstacle classifier is less than (or equal to) the probability of "0", it can be determined that the current vehicle is in a scene without long-tail obstacles, that is, the current vehicle is in a preset scene.
[0233] It should be understood that the intelligent triggering module can also use other methods to determine whether the current vehicle is in a preset scenario based on the second sensor data, and this application does not limit this.
[0234] For example, the intelligent processing module may further include a third trigger; wherein the third trigger can be implemented using a logical judgment algorithm, and the third trigger may also be called a third judge. Specifically, the second sensor data can be input to the third trigger; the third trigger determines, based on the second sensor data, whether the distance between the current vehicle and the obstacle / curb is less than a preset distance. If the distance between the current vehicle and the obstacle / curb is less than (or equal to) the preset distance, it can be determined that the current vehicle is in a scenario where it is close to the obstacle or curb, that is, the current vehicle is in a preset scenario. If the distance between the current vehicle and the obstacle / curb is greater than (or equal to) the preset distance, it can be determined that the current vehicle is not in a scenario where it is close to the obstacle or curb, that is, the current vehicle is in a preset scenario.
[0235] For example, the intelligent triggering module may also include multiple fourth triggers; wherein each fourth trigger may be implemented using a logical judgment algorithm, and the fourth trigger may also be called a fourth judge.
[0236] For example, a fourth trigger can determine the probability of the first processing result and determine whether the probability of the first processing result is greater than a probability threshold. For example, the first processing result is the risk identification result output by the vehicle risk identification module. If the risk identification result output by the vehicle risk identification module is greater than the probability threshold, it means that the reliability of the first processing result meets the preset reliability condition; otherwise, it means that the first processing result does not meet the preset reliability condition.
[0237] For example, a fourth trigger can determine the confidence level of the first processing result (confidence level can be used to describe reliability) and determine whether the confidence level of the first processing result is less than a confidence level threshold. For example, if the first processing result is the processing result output by the sensing module, and the confidence level of the processing result output by the sensing module is less than the confidence level threshold, it means that the reliability of the first processing result meets the preset reliability condition; otherwise, it means that the first processing result does not meet the preset reliability condition.
[0238] Suppose that the perception module outputs recognition results for multiple categories, and the probability of recognizing the correct category for each category is Pi (0≤i≤M, where M is the number of categories), then the confidence level is C=-∑P i log2P i .
[0239] For example, a fourth trigger can determine the confidence variance of the first processing result and determine whether the confidence variance of the first processing result is greater than a variance threshold. For example, the first processing result is the processing result output by the sensing module. If the confidence variance of the processing result output by the sensing module is greater than the variance threshold, it means that the reliability of the first processing result meets the preset reliability condition; otherwise, it means that the first processing result does not meet the preset reliability condition.
[0240] For example, the intelligent triggering module may include multiple fifth triggers, which may be existing triggers or monitors in the vehicle. For example, a fifth trigger may monitor system processes, functional processing units, or algorithm processing units; when a system process, functional processing unit, or algorithm processing unit is detected to be occupying CPU time for a long time with low utilization, it can be determined that the system is in a frozen state. For example, a fifth trigger may also determine whether the system is in a frozen state based on the perception results output by the perception module (e.g., ultrasonic detection results). For example, a fifth trigger may detect whether the intelligent driving system is in a functional takeover state (i.e., a state requiring user control of the vehicle).
[0241] For example, when the triggering condition for multi-terminal collaborative intelligent driving only includes the network parameters between the current vehicle and the remote device meeting preset network conditions, if it is determined that the network parameters between the current vehicle and the remote device meet the preset network conditions, then it can be determined that the current situation meets the triggering condition for multi-terminal collaborative intelligent driving; at this time, multi-terminal collaborative intelligent driving can be triggered by the first trigger. Otherwise, the first trigger does not trigger multi-terminal collaborative intelligent driving.
[0242] Similarly, when the triggering conditions for multi-terminal collaborative intelligent driving include only one type of triggering condition, it can be determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving when it is determined that the current situation meets that triggering condition; at this time, the corresponding trigger can trigger multi-terminal collaborative intelligent driving. Otherwise, the corresponding trigger will not trigger multi-terminal collaborative intelligent driving; this will not be elaborated further here.
[0243] For example, when the triggering conditions for multi-terminal collaborative intelligent driving include the current vehicle being in a preset scenario and the network parameters between the current vehicle and the remote device meeting preset network conditions, the intelligent triggering module can determine whether the current vehicle is in a preset scenario and whether the network parameters between the current vehicle and the remote device meet preset network conditions. When it is determined that the current vehicle is in a preset scenario and the network parameters between the current vehicle and the remote device meet preset network conditions, it can be determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; at this time, the first trigger and the second trigger can trigger multi-terminal collaborative intelligent driving. Otherwise, it can be determined that the current situation does not meet the triggering conditions for multi-terminal collaborative intelligent driving; at this time, the first trigger and the second trigger do not trigger multi-terminal collaborative intelligent driving.
[0244] Similarly, when the triggering conditions for multi-terminal collaborative intelligent driving include only two triggering conditions, it can be determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving if it is determined that the current situation meets these two triggering conditions; at this time, the corresponding trigger can trigger multi-terminal collaborative intelligent driving. Otherwise, the corresponding trigger will not trigger multi-terminal collaborative intelligent driving; this will not be elaborated further here.
[0245] For example, when the triggering conditions for multi-terminal collaborative intelligent driving include that the network parameters between the current vehicle and the remote device meet preset network conditions, the current vehicle is in a preset scenario, and the reliability of the first processing result meets preset reliability conditions, the intelligent triggering module can determine whether the network parameters between the current vehicle and the remote device meet preset network conditions, whether the current vehicle is in a preset scenario, and whether the reliability of the first processing result meets preset reliability conditions. When it is determined that the network parameters between the current vehicle and the remote device meet preset network conditions, the current vehicle is in a preset scenario, and the reliability of the first processing result meets preset reliability conditions, it can be determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; at this time, the first to fourth triggers can trigger multi-terminal collaborative intelligent driving. Otherwise, it can be determined that the current situation does not meet the triggering conditions for multi-terminal collaborative intelligent driving; at this time, the first to fourth triggers do not trigger multi-terminal collaborative intelligent driving.
[0246] Similarly, when the triggering conditions for multi-terminal collaborative intelligent driving include any three of the above triggering conditions, if the current situation meets any three of the above triggering conditions, it can be determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, and the corresponding trigger can trigger multi-terminal collaborative intelligent driving; otherwise, it is determined that the current situation does not meet the triggering conditions for multi-terminal collaborative intelligent driving, and the corresponding trigger may not trigger multi-terminal collaborative intelligent driving; this will not be elaborated further here.
[0247] For example, when the triggering conditions for multi-terminal collaborative intelligent driving include the network parameters between the current vehicle and the remote device meeting preset network conditions, the current vehicle being in a preset scenario, the reliability of the first processing result meeting preset reliability conditions, and the current vehicle's system state meeting preset states, the intelligent triggering module can determine whether the network parameters between the current vehicle and the remote device meet preset network conditions, whether the current vehicle being in a preset scenario, whether the reliability of the first processing result meeting preset reliability conditions, and whether the current vehicle's system state meeting preset states. When it is determined that the network parameters between the current vehicle and the remote device meet preset network conditions, the current vehicle being in a preset scenario, the reliability of the first processing result meeting preset reliability conditions, and the current vehicle's system state meeting preset states, it can be determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; at this time, the first to fifth triggers can trigger multi-terminal collaborative intelligent driving. Otherwise, it can be determined that the current situation does not meet the triggering conditions for multi-terminal collaborative intelligent driving; at this time, the first to fifth triggers do not trigger multi-terminal collaborative intelligent driving.
[0248] In one possible approach, the intelligent triggering module can determine whether the current vehicle is in a preset scenario, whether the network parameters between the current vehicle and the remote device meet preset network conditions, whether the reliability of the first processing result meets preset reliability conditions, and whether the current vehicle's system state meets preset states. If the current vehicle is in a preset scenario, or the network parameters between the current vehicle and the remote device meet preset network conditions, or the reliability of the first processing result meets preset reliability conditions, or the current vehicle's system state meets preset states, then the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; in this case, the first to fifth triggers can trigger multi-terminal collaborative intelligent driving. Otherwise, it can be determined that the current situation does not meet the triggering conditions for multi-terminal collaborative intelligent driving; in this case, the first to fifth triggers can trigger multi-terminal collaborative intelligent driving.
[0249] It should be noted that the second or third trigger for multi-terminal collaborative intelligent driving can refer to a second or third trigger that determines the current scenario of the vehicle as a preset scenario. The fourth trigger for multi-terminal collaborative intelligent driving can refer to a fourth trigger that determines the reliability of the first processing result meets a preset reliability condition. The fifth trigger for multi-terminal collaborative intelligent driving can refer to a fifth trigger that determines the current system state of the vehicle meets a preset state.
[0250] For example, when it is determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, S202 can be executed. When it is determined that the current situation does not meet the triggering conditions for multi-terminal collaborative intelligent driving, the vehicle's intelligent driving system can independently perform intelligent driving, as described above, and will not be repeated here.
[0251] S202, when it is determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, multi-terminal collaborative intelligent driving data is sent to the remote device; wherein, the multi-terminal collaborative intelligent driving data includes second sensor data.
[0252] For example, when the intelligent triggering module determines that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, it can instruct the vehicle to send multi-terminal collaborative intelligent driving data to the remote device.
[0253] For example, multi-terminal collaborative intelligent driving data may include second sensor data.
[0254] For example, the multi-terminal collaborative intelligent driving data may also include at least one of the following: the trigger condition identifier of multi-terminal collaborative intelligent driving that is satisfied in the current situation or the fourth processing result.
[0255] For example, the fourth processing result is obtained by the current vehicle processing the second sensing data; specifically, the fourth processing result may include at least one of the following: the processing result obtained by the perception module processing the second sensing data, the processing result obtained by the prediction module processing the processing result of the perception module in the fourth processing result, the processing result obtained by the fusion module processing the processing result of the perception module and the processing result of the prediction module in the fourth processing result, and the processing result obtained by the regulation control module processing the processing result of the fusion module in the fourth processing result.
[0256] For example, the trigger condition identifier for multi-terminal collaborative intelligent driving that is satisfied in the current situation is used to uniquely identify the trigger condition for multi-terminal collaborative intelligent driving.
[0257] For example, the triggering conditions are: the network parameters between the current vehicle and the remote device meet the preset network conditions, and the corresponding triggering condition is identified as D1. The triggering condition is: the current vehicle is in a grassy scene, and the corresponding triggering condition is identified as D2. The triggering condition is: the current vehicle is in a scene with standing water on the road, and the corresponding triggering condition is identified as D3. The triggering condition is: the current vehicle is in a blurred scene, and the corresponding triggering condition is identified as D4. ... and the triggering condition is: the confidence level of the first processing result is less than the confidence level threshold, and the corresponding triggering condition is identified as D10. The triggering condition is: the confidence level variance of the first processing result is greater than the variance threshold, and the corresponding triggering condition is identified as D11. The triggering condition is: the probability of the first processing result is greater than the probability threshold, and the corresponding triggering condition is identified as D12; and so on.
[0258] For example, the triggering condition is: the network parameters between the current vehicle and the remote device meet the preset network conditions, and the corresponding triggering condition identifier is the identifier of the first trigger. The triggering condition is: the current vehicle is in a grassy scene, and the corresponding triggering condition identifier is the identifier of the corresponding second trigger, and so on. The triggering condition is: the confidence level of the first processing result is less than the confidence level threshold, and the corresponding triggering condition identifier is the identifier of the corresponding fourth trigger; and so on.
[0259] Furthermore, multi-terminal collaborative intelligent driving data may also include multi-terminal collaborative intelligent driving requests. These requests instruct remote devices to process other data (excluding the requests themselves) contained within the multi-terminal collaborative intelligent driving data to obtain a second processing result. In other words, after receiving the multi-terminal collaborative intelligent driving data, the remote device can respond to the requests using a large model service, invoking the large model to process the other data contained within the multi-terminal collaborative intelligent driving data and obtain a second processing result. The specific processing procedure of the remote device will be explained later. Afterward, the remote device can send the second processing result to the vehicle.
[0260] S203, Receive the second processing result sent by the remote device. The second processing result is obtained by the remote device processing multi-terminal collaborative intelligent driving data.
[0261] For example, the second processing result may include textual description information of the current vehicle scene and detection information of the first object. The detection information of the first object may be represented using key-value pairs; this ensures that the second processing result returned by the remote device is consistent with the representation of the processing result output by the perception module or prediction module in the vehicle, facilitating processing by the intelligent driving system in the vehicle.
[0262] For example, the text description information of the current vehicle scene may include, but is not limited to: global (overall) text description information of the scene, text description information of the parking space, text description information of obstacles, text description information of the road, text description information of risks, and parking suggestions, etc., and this application does not limit it.
[0263] For example, the text description of the current scene in which the vehicle is located can be as follows:
[0264] [Scene Global Description]: This is an underground parking lot scene with bright lighting.
[0265] [Parking Space Description]: There are 3 empty parking spaces on the left side of the image, and a vehicle is parked in the parking space on the right side.
[0266] [Obstacle Description]: There is a pillar near the parking space, but no obvious obstacles at the entrance to the parking space.
[0267] [Road Description]: The road surface is made of cement, and it is shown as a two-lane road with no abnormalities.
[0268] [Risk Description]: There is currently no risk.
[0269] [Parking suggestion]: Please drive to the empty parking space near the left front of the image and park there.
[0270] For example, the text description of the current vehicle's location could be as follows:
[0271] [Scene Description]: This is a grassy parking lot scene, and the sky brightness suggests it is cloudy.
[0272] [Parking Space Description]: There are no obvious parking space signs around, but based on the surrounding vehicles, it can be determined that temporary parking is possible here.
[0273] [Obstacle Description]: A tree stump is visible on the grass in the image. It may be rotten and its height may cause it to scrape against the vehicle's chassis.
[0274] [Road Description]: This is an unpaved road surface. Please be aware of road safety.
[0275] [Parking Advice]: Based on the surrounding vehicles, it is assumed that this area is suitable for parking. However, there are tree stumps in the area, which pose a risk of scraping the undercarriage. Please be aware of parking safety and it is recommended to find a more suitable parking spot.
[0276] It should be noted that this application refers to identifying the location of available parking spaces through signs and indicator lights, and providing information on available parking spaces in general directions.
[0277] For example, the detection information of the first object may include the detection bounding box of the first object, the edge contour of the first object, the depth information of the image data contained in the second sensing data, the segmentation result of the first object, etc., and this application does not limit it. The first object may include a moving object and / or a stationary object.
[0278] For example, the detection information for the first object can be as follows:
[0279] Object1: Bbox: [100, 100, 200, 100] (These four values are the coordinates of the top left corner of the detection box, as well as its height and width), Class:
[0001] (Indicates the object category, such as "01" for plastic bag).
[0280] Object2: Bbox: [200, 150, 400, 200], Class:
[0002] ("02" represents a car).
[0281] Object3: Polygon: [(150,150),(100,200),(200,180),(250,100),(150,150)] (represents the coordinates of a series of vertices that constitute a closed polygon region), Class
[0001] .
[0282] Object4: Edge: [(180,180),(200,180),(210,190)] (represents the coordinates of multiple vertices that make up the edge contour), Class
[0003] ("03" represents the curb).
[0283] ...
[0284] Depth1 (Depth Information 1): Depth Map 1.
[0285] Depth2 (Depth Information 2): Depth Map 2.
[0286] ...
[0287] For example, the second processing result may also include attribute information of the first object. The attribute information of the first object may include people, animals, hard obstacles (such as stones, steel pipes, etc.), and soft obstacles (such as plastic bags, leaves, etc.).
[0288] S204. Based on the second processing result, perform path planning and control for the current vehicle.
[0289] For example, path planning can be performed on the current vehicle based on the second processing result to obtain the target path, speed, acceleration, etc. of the current vehicle; then, vehicle control signals (such as throttle control signals, steering wheel control signals, gear control signals, etc.) can be generated based on the target path, speed, acceleration, etc. of the current vehicle; then, the current vehicle is controlled according to the vehicle control signals so that the current vehicle travels to the target position (such as the target parking space) according to the target path.
[0290] For example, S204 may include: S2041 to S2043:
[0291] S2041, Process the third sensor data to obtain a third processing result; wherein the third sensor data is collected by the current vehicle after the second sensor data is collected.
[0292] S2042, merge the second processing result and the third processing result to obtain the target processing result.
[0293] S2043, based on the target processing results, performs path planning and control for the current vehicle.
[0294] In one possible approach, S2041 is executed by the perception module, S2042 by the prediction module, and S2043 by the prediction module, fusion module, and planning and control module. In this case, the second processing result of S2042 for fusion is the detection information of the first object.
[0295] For example, the third sensing data may include N (N is a positive integer, which can be set as needed, such as N=3 or 5) seconds of sensing data collected by the vehicle's sensors.
[0296] For example, the first sensor data is the sensor data from 20:00:00 to 20:00:03, the second sensor data is the sensor data from 20:00:04 to 20:00:06, and the third sensor data is the sensor data from 20:00:07 to 20:00:09.
[0297] For example, the third sensing data may include, but is not limited to: image data, lidar data, millimeter wave data, ultrasonic data, vehicle motion status (such as vehicle speed, vehicle direction, vehicle acceleration, etc.), etc., and this application does not limit it.
[0298] In this scenario, the third processing result is the output of the perception module. The processing result output by the perception module (i.e., the third processing result) includes, but is not limited to: detection information of the second object (which may include the detection bounding box of the second object, the edge contour of the second object, the depth information of the image data contained in the third sensing data, the segmentation result of the second object, etc. The second object may include moving objects and / or stationary objects), road conditions, traffic signs, etc.
[0299] For example, the perception module can output the third processing result to the prediction module and the fusion module, and the prediction module executes S2042. Specifically, since the current vehicle transmits multi-terminal collaborative intelligent driving data to the remote device, the remote device needs time to process the multi-terminal collaborative intelligent driving data to obtain the second processing result, and the second processing result needs to be transmitted to the current vehicle. Therefore, after the current vehicle uploads the second sensing data collected in time period 1 (time A1 to time A2, where time A1 is before time A2) to the remote device at time A2, the current vehicle can only receive the second processing result at time B2; where time B2 is after time A2. The current vehicle's sensors continuously collect sensing data, and the current vehicle's intelligent driving system also continuously processes the collected sensing data; therefore, at time B2, the current vehicle's perception module processes the sensing data collected in time period 2 (time B1 to time B2, where time B1 is before time B2) (i.e., the third sensing data). In other words, the second processing result is obtained by processing the sensor data collected in time period 1, and the third processing result is obtained by processing the sensor data collected in time period 2. Therefore, the prediction module can align the detection information of the first object and the detection information of the second object in time and space to obtain the target processing result.
[0300] For example, the vehicle's travel distance can be determined based on the difference between time period 1 and time period 2; then, the detection information of the first object can be compensated based on the vehicle's travel distance; subsequently, the compensated detection information of the first object and the detection information of the second object can be synthesized to obtain the target processing result. The compensated detection information of the first object and the detection information of the second object belong to the same time and space.
[0301] For example, the second sensor data was collected during the period from 20:00:04 to 20:00:06 (i.e., time period 1), and the third sensor data was collected during the period from 20:00:07 to 20:00:09 (i.e., time period 2); the detection information of the first object is: Object1: Bbox: [100,100,200,100], Class:
[0001] ; the current parking speed of the vehicle is 2km / h. Among them, the difference between time period 1 and time period 2 is 3s, and the reversing distance of the vehicle is 1.6m; then, according to the distance between the camera and the object, and the relationship between the size of the object in the image, the detection information of the first object is adjusted, and the adjusted detection information of the first object is: Object1: Bbox: [80,80,150,75], Class:
[0001] (this is only an example and does not represent the actual adjusted detection information of the first object).
[0302] For example, the prediction module can select the detection information of the moving object from the target processing result, predict the motion information of the moving object, and output it to the fusion module.
[0303] Subsequently, the fusion module can perform environmental modeling based on the motion information of moving objects output by the prediction module and the processing results (excluding the detection information of moving objects) output by the perception module, obtaining image data of the current vehicle's scene (which can be 2D or 3D image data, including the position of static objects and the motion information of dynamic objects; it can also be understood as the image describing the changes of objects over a future period of time) and output it to the planning and control module. The planning and control module can then perform path planning and control for the current vehicle based on the image data of the current vehicle's scene.
[0304] Optionally, the prediction module can also output the target processing results to the fusion module; in this way, the fusion module can perform environmental modeling based on the motion information of the moving object output by the prediction module and the target processing results, as well as the processing results output by the perception module other than the detection information of the moving object, to obtain image data of the current vehicle scene.
[0305] Optionally, the above-mentioned S2042 can also be executed by the fusion module. After obtaining the target processing result, the fusion module can output the target processing result to the prediction module. In this way, the prediction module can select the detection information of the moving object from the target processing result, predict the motion information of the moving object, and output it to the fusion module.
[0306] In one possible approach, S2041 is executed by the perception module and the prediction module, S2042 is executed by the fusion module, and S2043 is executed by the fusion module and the planning and control module.
[0307] In this scenario, the third processing result is the output of both the prediction module and the perception module. Specifically, the second sensing data can be input into the perception module, which processes the data and outputs the processing result to both the prediction and fusion modules. The prediction module can select the detection information of the moving object from the processing result output by the perception module; then, based on this detection information, it predicts the motion information of the moving object and outputs it to the fusion module.
[0308] Subsequently, the fusion module can execute S2042 to obtain the target processing result. S2043 can be implemented in the following ways: the fusion module performs environmental modeling based on the target processing result and the motion information of the moving object output by the prediction module to obtain image data of the scene where the vehicle is currently located; then, the planning and control module can perform path planning and control of the current vehicle based on the image data of the scene where the vehicle is currently located.
[0309] For example, when the second processing result includes the attribute information of the first object, a target parking space can be selected based on the attribute information of the first object; a target path can be generated based on the attribute information of the first object and the target parking space; and the current vehicle can be controlled based on the target path.
[0310] For example, since the first object may be in an empty parking space, a target parking space can be selected from multiple parking spaces based on the attribute information of the first object (at this time, the attribute information of the first object can be used to determine the parking space availability). For example, a parking space without the first object or whose attribute information is a soft obstacle can be identified as a target parking space. In addition, the first object may also be on the road leading from the current vehicle to the target parking space. Therefore, a target path can be generated based on the attribute information of the first object and the target parking space to ensure that there are no first objects with attributes of people, animals, or hard obstacles on the target path. This can further improve the safety and reliability of intelligent parking and increase the probability of successful parking. In addition, while generating the target path, the speed and acceleration of the current vehicle can also be generated; then, vehicle control signals (such as throttle control signals, steering wheel control signals, gear control signals, etc.) can be generated based on the target path, speed, and acceleration of the current vehicle; then, the current vehicle is controlled according to the vehicle control signals so that the current vehicle travels to the target location (such as the target parking space) according to the target path.
[0311] It's important to note that during the current route planning process, a virtual parking space can be determined based on the target processing results and parking space indication information. This virtual parking space is then designated as the target parking space, and route planning is performed on the current vehicle based on the target processing results and the target parking space. This allows route planning even when a real target parking space cannot be identified based on sensor data. Consequently, the user doesn't need to drive the vehicle to an empty parking space; the intelligent driving system can still achieve automatic / assisted parking. This is particularly suitable for scenarios where finding a target parking space is difficult in parking lots.
[0312] In other words, this application first determines whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving during each intelligent driving process. If the triggering conditions are met, multi-terminal collaborative intelligent driving is then implemented. If the triggering conditions are not met, the vehicle independently performs intelligent driving. Compared to existing technologies where the vehicle performs multi-terminal collaborative intelligent driving during each intelligent driving process, this application reduces the number of interactions between the current vehicle and remote devices, as well as the number of vehicles interacting with remote devices, thereby alleviating the computational burden on remote devices. Furthermore, it can alleviate network congestion during peak traffic periods, reducing the latency of the vehicle receiving data from remote devices. This allows the vehicle's intelligent driving system to promptly plan and control the vehicle's path based on the data sent by the remote devices, thereby improving the safety and reliability of intelligent driving.
[0313] Secondly, the multi-terminal collaborative intelligent driving data can also include a fourth processing result. The remote device can process the fourth processing result and the second sensor data to generate a second processing result. Compared with the prior art where the remote device generates control signals based solely on sensor data, the data relied upon by this application to generate the second processing result is richer, and therefore the second processing result is more accurate. This can improve the accuracy of the control signals generated by the vehicle, thereby improving the safety and reliability of intelligent driving and the user experience.
[0314] Furthermore, in the process of multi-terminal collaborative intelligent driving, remote devices can dynamically acquire necessary information, i.e., the fourth processing result, based on the complexity of the road scene. In other words, the multi-terminal collaborative intelligent driving data sent by the vehicle to the remote device each time is different, which can reduce the transmission of invalid data and alleviate network pressure.
[0315] For example, before the current vehicle and the remote device cooperate in intelligent driving, that is, before the current vehicle executes S201 to S204, the current vehicle needs to enable the multi-terminal collaborative processing function and the intelligent driving function. The methods for enabling the multi-terminal collaborative processing function and the intelligent driving function can be various, and this application does not limit them. Several methods for enabling the multi-terminal collaborative processing function and the intelligent driving function are described below.
[0316] Figure 3 is a schematic diagram illustrating an exemplary intelligent driving process 300. Process 300 illustrates a method for enabling multi-device collaborative processing and intelligent driving functions.
[0317] S301, in response to the first user's operation, enables multi-device collaborative processing.
[0318] Figures 4A to 4C are schematic diagrams of the interfaces of the vehicle-mounted terminal as examples.
[0319] Referring to FIG4A, by way of example, the main interface 401 of the vehicle terminal may include one or more controls, including but not limited to: intelligent driving options (such as automatic cruise option, assisted cruise option, automatic parking option and assisted parking option, etc.), application icons of various applications (such as application icon of map application, application icon of settings application, application image of commonly used applications and application icon 402 of settings application, etc.), audio playback options, battery icon, network icon, time information, seat setting options, air conditioning setting options, etc., and this application does not limit them.
[0320] Referring again to Figure 4A, exemplarily, a user can click on the application image 402 of the settings application. The in-vehicle terminal can respond to the user's action and display the settings interface 403, as shown in Figure 4B. The settings interface 403 may include one or more controls, including but not limited to: Settings Option 1, Cloud Assistant Option 404, ..., Settings Option n (n is a positive integer), etc. The user can click on the switch control in Cloud Assistant Option 404 (i.e., perform the first user operation). The in-vehicle terminal can respond to the user's action (i.e., respond to the first user operation) by setting the state of the switch control in Cloud Assistant Option 404 to the enabled state, as shown in Figure 4C, and enabling the multi-terminal collaborative processing function.
[0321] S302, in response to a second user's operation, activates the intelligent driving function.
[0322] Referring again to Figure 4A, for example, during the current vehicle operation, the user can click the automatic cruise option (i.e., perform a second user operation), and the vehicle terminal can respond to the user's operation (i.e., respond to the second user operation) and enable the automatic cruise function.
[0323] Referring to Figure 4A, for example, during the current vehicle operation, the user can click the cruise assist option (i.e., perform a second user operation), and the vehicle terminal can respond to the user's operation (i.e., respond to the second user operation) and enable the cruise assist function.
[0324] Referring to Figure 4A, for example, during the current vehicle parking process, the user can click the automatic parking option (i.e., perform a second user operation), and the vehicle terminal can respond to the user's operation (i.e., respond to the second user operation) and enable the automatic parking function.
[0325] Referring to Figure 4A, for example, during the current vehicle parking process, the user can click the assisted parking option (i.e., perform a second user operation), and the vehicle terminal can respond to the user's operation (i.e., respond to the second user operation) and enable the assisted parking function.
[0326] It should be noted that intelligent driving functions may include at least one of automatic cruise control, assisted cruise control, automatic parking, or assisted parking.
[0327] It should be noted that users can also interact with the in-vehicle terminal via voice and gestures to instruct the in-vehicle terminal to enable multi-terminal collaborative processing and intelligent driving functions, and this application does not impose any restrictions on this.
[0328] S303, determine whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein, the current situation includes at least one of the following: the scene in which the vehicle is currently located, the network parameters between the current vehicle and the remote device, the reliability of the first processing result obtained by the current vehicle in processing the first sensor data, or the system state of the current vehicle.
[0329] S304, when it is determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, multi-terminal collaborative intelligent driving data is sent to the remote device; wherein, the multi-terminal collaborative intelligent driving data includes second sensor data.
[0330] S305, receive the second processing result sent by the remote device. The second processing result is obtained by the remote device processing multi-terminal collaborative intelligent driving data.
[0331] S306, based on the second processing result, performs path planning and control for the current vehicle.
[0332] For example, S303 to S306 can be described with reference to the above description of S201 to S204, and will not be repeated here.
[0333] Figure 5 is a schematic diagram illustrating an exemplary intelligent driving process 500. Process 500 demonstrates a method for enabling multi-device collaborative processing and intelligent driving functions. In this process 500, intelligent driving refers to automatic parking or assisted parking.
[0334] S501, detects whether the current vehicle has failed to park.
[0335] For example, it can be determined whether the vehicle has failed to park based on the second sensor data.
[0336] For example, the positional relationship between the target parking space and the current vehicle can be determined based on image data captured by a camera to determine whether the vehicle has failed to park. As another example, speech recognition can be performed on voice data captured by a microphone; the speech recognition result can then be used to determine whether the vehicle has failed to park. It should be understood that this application includes multiple methods for determining whether the vehicle has failed to park based on second sensor data, and this application does not limit these methods.
[0337] For example, the system can detect whether the current vehicle has failed to park according to a preset cycle. If the current vehicle fails to park in the current cycle, steps S502 and S503 can be executed. If the current vehicle fails to park in the current cycle, step S501 can be executed again to detect whether the current vehicle has failed to park in the next cycle.
[0338] S502: When parking failure of the current vehicle is detected, the intelligent driving function is activated.
[0339] In one possible approach, when parking failure is detected, the onboard terminal can directly activate the intelligent driving function.
[0340] In one possible approach, when parking failure is detected, the in-vehicle terminal displays an intelligent driving function activation prompt interface 601, as shown in Figure 6A. The intelligent driving function activation prompt interface 601 may include a confirmation option 602, a denial option 603, and the activation prompt message "Activate automatic parking function?" (or, the prompt message may be "Activate assisted parking function?"). When the user clicks the confirmation option 602, the in-vehicle terminal can respond to the user's action by closing the intelligent driving function activation prompt interface 601 and activating the intelligent driving function. When the user clicks the denial option 603, the in-vehicle terminal can respond to the user's action by closing the intelligent driving function activation prompt interface 601; subsequent parking requires manual operation by the user.
[0341] In one possible approach, when parking failure is detected, the onboard terminal can also display an automatic parking option, as shown in Figure 6B. The user can click the automatic parking option, and the onboard terminal can respond to the user's action by activating the intelligent driving function.
[0342] It should be noted that when parking failure is detected, this application does not restrict what kind of control is displayed on the in-vehicle terminal to enable the intelligent driving function, or the corresponding user's operation to trigger the intelligent driving function.
[0343] It should be understood that users can also interact with the in-vehicle terminal via voice or gestures, select a confirmation option, and instruct the in-vehicle terminal to activate the intelligent driving function; this application does not impose any restrictions on this.
[0344] In one possible approach, when parking failure is detected, the in-vehicle terminal can play a prompt voice message, such as, "Enable intelligent driving function?" The user can then control whether to enable the intelligent driving function via voice. For example, if the user says "Enable," the in-vehicle terminal will receive the user's voice and, upon recognizing the text "Enable," will activate the intelligent driving function. Conversely, if the user says "Disable," the in-vehicle terminal will receive the user's voice and, upon recognizing the text "Disable," will deactivate the intelligent driving function.
[0345] S503: When parking failure of the current vehicle is detected, determine whether the network parameters of the current vehicle meet the preset network conditions.
[0346] For example, when parking failure is detected, the vehicle terminal can also obtain the network parameters of the current vehicle and determine whether the network parameters meet preset network conditions. If the network parameters of the current vehicle meet the preset network conditions, it indicates that the network status between the current vehicle and the remote device is good, and S504 can be executed. If the network parameters of the current vehicle do not meet the preset network status, it indicates that the network status between the current vehicle and the remote device is poor. In this case, the vehicle terminal can call the intelligent driving system to independently perform intelligent driving, and S510 can be executed.
[0347] S504 If the network parameters of the current vehicle meet the preset network conditions, display the switch control for the multi-terminal collaborative processing function.
[0348] In one possible approach, the switch control for the multi-device collaborative processing function can be an enable prompt interface 605 for the multi-device collaborative processing function, as shown in Figure 6C. For example, the enable prompt interface 605 for the multi-device collaborative processing function may include a confirmation option 606, a negative option 607, and the enable prompt message "Enable multi-device collaborative processing function?".
[0349] In one possible approach, the switch control for the multi-terminal collaborative processing function can be the cloud assistant button 608, as shown in Figure 6D.
[0350] S505 enables multi-device collaborative processing in response to third-user operations on switch controls.
[0351] Referring to Figure 6C, exemplarily, a third user operation on the switch control could be clicking the confirmation option 606 in the multi-terminal collaborative processing function activation prompt interface 605. When the user clicks the confirmation option 606, the vehicle terminal can respond to the user's action by closing the multi-terminal collaborative processing function activation prompt interface 605 and enabling the multi-terminal collaborative processing function. When the user clicks the denial option 607, the vehicle terminal can respond to the user's action by closing the multi-terminal collaborative processing function activation prompt interface 605; subsequently, the vehicle terminal can invoke the intelligent driving system to independently perform intelligent driving, and can execute S510.
[0352] Referring to Figure 6D, an exemplary third-user operation on the switch control could be clicking the cloud assistant button 608. After the user clicks the cloud assistant button 608, the vehicle terminal can respond to the user's operation and enable the multi-terminal collaborative processing function.
[0353] It should be understood that users can also interact with the in-vehicle terminal via voice or gestures to instruct the in-vehicle terminal to enable multi-terminal collaborative processing functions, and this application does not impose any restrictions on this.
[0354] In one possible approach, when parking failure is detected, the in-vehicle terminal can play a prompt voice message, such as, "Enable multi-device collaborative processing function?" The user can control whether to enable the multi-device collaborative processing function via voice. For example, if the user says "Enable," the in-vehicle terminal, after receiving the user's voice, can enable the multi-device collaborative processing function by recognizing the text "Enable." Conversely, if the user says "Disable," the in-vehicle terminal, after receiving the user's voice, can disable the multi-device collaborative processing function by recognizing the text "Disable."
[0355] S506, determine whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein, the current situation includes at least one of the following: the scene in which the vehicle is currently located, the network parameters between the current vehicle and the remote device, the reliability of the first processing result obtained by the current vehicle in processing the first sensor data, or the system state of the current vehicle.
[0356] S507, when it is determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, multi-terminal collaborative intelligent driving data is sent to the remote device; wherein, the multi-terminal collaborative intelligent driving data includes second sensor data.
[0357] S508 receives the second processing result sent by the remote device. The second processing result is obtained by the remote device processing multi-terminal collaborative intelligent driving data.
[0358] S509, based on the second processing result, performs path planning and control for the current vehicle.
[0359] For example, S506 to S509 can be described with reference to the above description of S201 to S204, and will not be repeated here.
[0360] S510 processes the third sensor data to generate a third processing result.
[0361] For example, the third sensing data can be processed by the sensing module, and the processing result can be called the third processing result.
[0362] S511, based on the third processing result, performs path planning and control for the current vehicle.
[0363] Next, the third processing result can be output to the prediction module to obtain the trajectory of the moving object; and the third processing result can also be output to the fusion module; then the fusion module can generate image data of the scene where the vehicle is currently located based on the third processing result and the motion information of the moving object. Afterwards, the planning and control module can perform path planning and control of the current vehicle based on the image data of the scene where the vehicle is currently located.
[0364] Figure 7 is a schematic diagram illustrating an exemplary intelligent driving process 700. Process 700 demonstrates a method for enabling multi-device collaborative processing and intelligent driving functions. In this process 700, intelligent driving refers to automatic parking or assisted parking.
[0365] S701, detects whether the vehicle is currently parked.
[0366] For example, it can be determined whether the vehicle is currently parked based on the second sensor data. For instance, scene recognition can be performed on images captured by a camera; then, the scene recognition result can be used to determine whether the vehicle is currently parked. Another example is speech recognition can be performed on voice data captured by a microphone; then, the speech recognition result can be used to determine whether the vehicle is currently parked. It should be understood that this application includes various methods for determining whether the vehicle is currently parked based on the second sensor data, and this application does not limit these methods.
[0367] For example, the system can detect whether the vehicle is currently in a parked state according to a preset cycle; if the vehicle is detected to be in a parked state in the current cycle, S702 and S703 can be executed; if the vehicle is not detected to be in a parked state in the current cycle, S701 can be executed again, that is, the system can detect whether the vehicle is currently in a parked state in the next cycle.
[0368] S702: When the vehicle is detected to be parked, the intelligent driving function is activated.
[0369] For example, the method of enabling intelligent driving function in S702 can be referred to the description of S502 above, and will not be repeated here.
[0370] S703, when the current vehicle is detected to be in a parked state, determines whether the network parameters of the current vehicle meet the preset network conditions.
[0371] For example, S703 can be described with reference to the above description of S503, and will not be repeated here.
[0372] S704 If the network parameters of the current vehicle meet the preset network conditions, display the switch control for the multi-terminal collaborative processing function.
[0373] For example, the way the switch control for displaying the multi-terminal collaborative processing function is shown in S704 can be referred to the description in S504 above, and will not be repeated here.
[0374] The S705 enables multi-device collaborative processing in response to third-user operations on the switch control.
[0375] For example, the method of enabling multi-terminal collaborative processing function in S705 can be referred to the description of S505 above, and will not be repeated here.
[0376] S706, determine whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein, the current situation includes at least one of the following: the scene in which the current vehicle is located, the network parameters between the current vehicle and the remote device, the reliability of the first processing result obtained by the current vehicle in processing the first sensor data, or the system state of the current vehicle.
[0377] S707, when it is determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, multi-terminal collaborative intelligent driving data is sent to the remote device; wherein, the multi-terminal collaborative intelligent driving data includes second sensor data.
[0378] S708 receives the second processing result sent by the remote device. The second processing result is obtained by the remote device processing multi-terminal collaborative intelligent driving data.
[0379] S709, based on the second processing result, performs path planning and control for the current vehicle.
[0380] S710 processes the third sensor data to generate a third processing result.
[0381] S711, based on the third processing result, performs path planning and control for the current vehicle.
[0382] For example, S706 to S711 can be referred to the description of S506 to S511 above, and will not be repeated here.
[0383] Figures 8A to 8G are schematic diagrams of the interfaces of the vehicle-mounted terminal as examples.
[0384] Referring to Figures 8A and 8B, exemplarily, during assisted parking, the on-board terminal can display parking images and connection status icons between the vehicle and the remote device (as shown in 801 and 802). Here, 801 indicates that the vehicle is currently sending data to the remote device, and 802 indicates that the vehicle is currently receiving data from the remote device.
[0385] For example, the current vehicle can generate and display a prompt message based on the text description information of the current vehicle's current scene in the second processing result.
[0386] For example, a warning message can be generated based on the text description of the obstacle, such as "There is a tree stump protruding from the grass in the image, which may be rotten and its height may cause it to scrape against the vehicle chassis," as shown in Figure 8C.
[0387] For example, a warning message can be generated based on the text description of the obstacle, "There is a rock ahead of the road, and its height may cause it to scrape against the vehicle chassis," such as the text "Please be aware of the rock ahead of the road!" and an obstacle icon, as shown in Figure 8D.
[0388] For example, a prompt message can be generated based on the text description information of the intersection, "A vehicle may be entering from the right side of the intersection ahead," such as the text "Please note that a vehicle may be entering from the right side of the intersection ahead!" and a vehicle icon, as shown in Figure 8E.
[0389] For example, prompts can be generated based on the text description of the parking space, as shown in Figure 8F. In Figure 8F, prompt 1 indicates that there is one empty parking space on the left side of the image, and prompt 2 indicates that there are three empty parking spaces on the right side of the image.
[0390] It should be understood that other prompts can also be generated based on the text description information of the current vehicle scene in the second processing result; this application does not limit this.
[0391] Referring to Figure 8G, the in-vehicle terminal can also display parking progress prompts, as shown in Figure 8G, displaying the vehicle identification and the text "Cloud Assistant completes parking".
[0392] It should be understood that the vehicle-mounted terminal may also display other information, and this application does not impose any restrictions on this.
[0393] For example, after completing intelligent driving, the in-vehicle terminal can display buttons such as "Parking Experience Sharing." When the user clicks this button, the in-vehicle terminal responds to the user's action by uploading sensor data from the intelligent driving process, as well as the processing results from the perception module, prediction module, fusion module, and control module (which can be referred to as historical data). For example, this historical data can be uploaded when the vehicle is currently idle and the network connection is good. Subsequently, the large model of the remote device can combine the historical data to process the second sensor data to obtain a second processing result.
[0394] The following describes the processing procedure of remote devices in multi-device collaborative intelligent driving.
[0395] Figure 9 is a schematic diagram illustrating an exemplary intelligent driving process 900. Process 900 describes the processing procedure of a remote device.
[0396] S901 receives multi-terminal collaborative intelligent driving data sent by the current vehicle; wherein, the multi-terminal collaborative intelligent driving data includes the trigger condition identifier of multi-terminal collaborative intelligent driving that is satisfied in the current situation and the second sensor data.
[0397] S902 determines the target large model from multiple large models based on the trigger condition identifiers of multi-terminal collaborative intelligent driving that are met in the current situation.
[0398] For example, one or more models applied to the data processing can be pre-set for each of the various triggering conditions for multi-terminal collaborative intelligent driving.
[0399] For example, for the trigger condition of multi-terminal collaborative intelligent driving: "The current scene of the vehicle is a grass scene", the large model used to process the data is set as a visual language large model.
[0400] For example, regarding the triggering condition for multi-terminal collaborative intelligent driving: "the confidence level of the processing result output by the perception module is lower than the confidence level threshold", the large model applied to the data processing is set as the visual basic large model.
[0401] For example, regarding the triggering condition for multi-terminal collaborative intelligent driving: "The current vehicle is in a scenario where it is close to obstacles or curbs", the large models used to process the data are set as a visual language large model and a visual basic large model.
[0402] Subsequently, a mapping relationship can be established between each trigger condition in the various trigger conditions for multi-terminal collaborative intelligent driving and one or more corresponding large models. For example, a mapping relationship can be established between the trigger condition identifier of each trigger condition in the various trigger conditions for multi-terminal collaborative intelligent driving and the model identifier of one or more corresponding large models. The model identifier is used to uniquely identify a model; for example, the model identifier for the visual language large model is M1, the model identifier for the visual foundation large model is M2, and so on.
[0403] In this way, after receiving multi-terminal collaborative intelligent driving data, the remote device can obtain the trigger condition identifier of multi-terminal collaborative intelligent driving that is satisfied in the current situation from the multi-terminal collaborative intelligent driving data; then, it can look up the mapping relationship based on the trigger condition identifier of multi-terminal collaborative intelligent driving that is satisfied in the current situation, so as to determine one or more target large models from multiple large models.
[0404] It should be understood that the above-mentioned large model for data application set for each triggering condition of multi-terminal collaborative intelligent driving is only an example. This application does not limit which large model is set for data processing for each triggering condition of multi-terminal collaborative intelligent driving.
[0405] It should be understood that the large model used for processing data can be the same or different for different triggering conditions of multi-terminal collaborative intelligent driving; this application does not impose any restrictions on this.
[0406] S903, based on the target large model and the second sensor data, determines the second processing result.
[0407] For example, when there is only one target large model, the second sensing data can be input into the target large model, and the target large model can process the second sensing data to obtain the second processing result.
[0408] For example, when there are multiple target large models, the second sensing data can be input into each target large model to obtain the second processing result output by each target large model.
[0409] For example, when the target large model is a visual language large model, the prompt words and image data from the second sensor data can be input into the visual language large model to obtain the text description information of the scene where the vehicle is currently located and the detection information of the first object.
[0410] For example, the prompts for the visual language big model can be set as needed, and this application does not impose any restrictions on this.
[0411] For example, you can set prompts to understand the scene and identify available parking spaces and risks. For example, "You are a parking expert and now need to find an available parking space [Available Parking Space Request]. Please identify the scene in the image, parking space information, obstacles near the parking space, and risks. Please provide parking suggestions [Parking Suggestion Requirements]. Please use concise language to describe the situation [Text Requirements]."
[0412] For example, prompts can be set for scene understanding or task processing in a specified area, such as "Identify targets in the specified area of the image that affect vehicle parking / exit, and provide relevant location and category information".
[0413] For example, a prompt could be set to confirm whether there are inconsistencies in the previous output of the target large model based on the fourth processing result. For example, "Please reconfirm the previous processing result of the visual language large model based on the fourth processing result {such as obstacles / parking space areas}, such as the previous output image showing an underground parking lot with a horizontal railing on the road ahead} to see if it contains parking space information, obstacles near the parking space, and risks."
[0414] For example, image data from various scenes can be collected. Then, for each image, multiple sets of text descriptions and key-value pairs of objects in the image (such as Bbox-Class, Polygon-Class, Edge-Class, etc.) can be generated; and multiple sets of prompts can be set. A single image, its corresponding set of text descriptions, a set of key-value pairs (which may include key-value pairs of multiple objects), and a set of prompts can constitute a set of training data. Then, the image data and prompts from the training data can be input into a large-scale visual language model to obtain a set of predicted text descriptions and a set of predicted key-value pairs output by the large-scale visual language model. The predicted text descriptions and key-value pairs can be compared with the set of text descriptions and key-value pairs from the training data to perform backpropagation on the large-scale visual language model.
[0415] For example, when the target large model is a visual basic large model, the second sensing data can be input into the visual basic large model, processed by the visual basic large model, and the detection information of the first object can be output.
[0416] For example, a large-scale visual model can be trained in a supervised manner using image data and multi-task annotations; for details, please refer to the descriptions in the prior art, which will not be repeated here.
[0417] For example, when the target large model is a dedicated task large model, the image data / laser data in the second sensing data can be input into the dedicated task large model, and the dedicated task large model can perform a dedicated task to obtain the detection information of the first object.
[0418] For example, a large model for a specific task can be trained in a supervised manner using image data and single-task annotations; for details, please refer to the description in the prior art, which will not be repeated here.
[0419] For example, when the target large model is a multimodal large model, the second sensing data (including sensing data of multiple modalities) can be input into the multimodal large model to output the detection information of the first object.
[0420] For example, based on the training data of the visual language large model, other sensor data corresponding to the image data can be expanded, such as LiDAR point cloud data, millimeter-wave radar point cloud data, ultrasonic radar time-series signals, etc., to realize the construction of multimodal text pairs, and then the multimodal model can be trained in a self-supervised manner; for details, please refer to the description in the prior art, which will not be repeated here.
[0421] For example, when the target large model is a visual language large model or a multimodal model, one way to determine the second processing result based on the target large model and the second sensing data can be as follows: S9031 to S9032:
[0422] S9031, the level of acquiring the second sensor data and the level of the fourth processing result.
[0423] For example, the fourth processing result includes N sets of data, each corresponding to one of the N levels. For instance, the fourth processing result includes four sets of data: the processing result from the perception module, the processing result from the prediction module, the processing result from the fusion module, and the processing result from the planning and control module.
[0424] For example, the levels of the second sensing data and the N sets of data included in the fourth processed data can be preset. For instance, the second sensing data corresponds to level 0, the processing result of the sensing module corresponds to level 1, the processing result of the prediction module corresponds to level 2, the processing result of the fusion module corresponds to level 3, and the processing result of the control module corresponds to level 4; this application does not limit this.
[0425] S9032, based on the level of the second sensing data and the level of the fourth processing result, the target large model is used to process the second sensing data and the fourth processing result to obtain the second processing result.
[0426] Figure 10 is a schematic diagram of the structure of an exemplary intelligent driving system. Figure 10 is based on Figure 1C. Compared with Figure 1C, the large model service in Figure 10 includes a data-level processing function: that is, the large model service can call the target large model and process the multi-terminal collaborative intelligent driving data according to the data level to obtain a second processing result.
[0427] Specifically, the following steps can be repeated until i equals N, with the initial value of i being 1;
[0428] The i-th group of data, the (i-1)-th group of intermediate results, and the second sensor data from the fourth processing result are input into the target large model to obtain the i-th group of intermediate results; where the i-th group of data corresponds to level i, and the 0-th group of intermediate results is obtained by inputting the second sensor data into the target large model;
[0429] Increment i by 1;
[0430] When i equals N, the second processing result is determined based on the N intermediate results from the first intermediate result to the Nth intermediate result.
[0431] In other words, the second sensor data can be input into the target large model first to obtain the 0th intermediate result. Then, the 1st data from the fourth processing result, the 0th intermediate result, and the second sensor data are input into the target large model to obtain the 1st intermediate result. Afterwards, the 2nd data from the fourth processing result, the 1st intermediate result, and the second sensor data are input into the target large model to obtain the 2nd intermediate result; and so on. For example, any one or more intermediate results from the 1st to the Nth intermediate results can be determined as the second processing result.
[0432] For example:
[0433] For example, the second sensing data includes image data from one forward-looking camera and image data from four surround-view cameras, totaling five images. The processing result of the perception module is the detection bounding boxes for objects such as vehicles, pedestrians, pillars, and water-filled barriers in the above five images, as well as the segmentation results of drivable road areas and parking spaces in the above five images. The processing result of the prediction module is the motion information such as the movement trajectory and speed of moving (or motion-carrying) objects such as vehicles and pedestrians obtained by the perception module over a future period (e.g., 3 seconds). The fusion module merges the processing results of the perception module and the prediction module to obtain the changes in road information such as moving objects such as vehicles and pedestrians, stationary objects such as pillars and water-filled barriers, drivable road areas, and parking spaces over a future period. The processing result of the planning and control module is the target path, speed, and acceleration of the current vehicle over a future period; as well as control signals such as accelerator, brake, steering wheel angle, and gear position.
[0434] First, the five images mentioned above, along with the prompt "You are a parking expert and need to find an empty parking space. Please identify the scene, parking space information, obstacles near the parking space, and risks in the image. Please provide parking suggestions in concise language." are input into the visual language model, which outputs the processing result (i.e., the 0th intermediate result). When the 0th intermediate result output by the visual language model (for example, the 0th intermediate result is "This image shows an underground parking garage scene. Regarding parking spaces: There are no empty parking spaces in the current image; regarding obstacles: There are no vehicles, pedestrians, or objects blocking the road; risks: There are no risks at present; parking suggestions: Please continue driving along the road to observe the availability of parking spaces") lacks key information (such as obstacles, parking space, risk information, etc.), the visual language model can be called again to process the processing result of the perception module at level 1, the 0th intermediate result, and the second sensor data at level 0 to obtain the 1st intermediate result.
[0435] For example, the processing results of the perception module, "detection boxes of objects such as vehicles, pedestrians, pillars, and water barriers, and segmentation results of drivable areas and parking spaces on the road," and the 0th intermediate result output by the visual language model can be filled into the prompt word. The resulting prompt word is like this: "You are a parking expert and now need to find an empty parking space. According to the current vehicle-side perception results, the information is as follows: There is a parking space in the [top left corner (10,10), length 100, width 50; corresponding to a rectangular area] region of the [first; corresponding to a certain image] image (segmentation results of drivable areas and parking spaces on the road). Regarding your previous judgment, regarding parking spaces: there are no empty parking spaces in the current image; regarding obstacles: there are no vehicles, pedestrians, or objects blocking the road; risk: there is no risk at present (input the 0th intermediate result). Please confirm again, identify parking space information, obstacles near the parking space, and risks, and provide parking suggestions." Afterwards, this prompt word and the above 5 images can be input into the visual language model to obtain the 1st intermediate result. If the first set of intermediate results (e.g., "Regarding parking spaces: There is a parking space in the indicated area, but it is not available; Regarding obstacles: There are vehicles in the indicated area, but no pedestrians; Risk: There is no risk at present; Parking advice: Please continue driving along the road to observe the availability of parking spaces") still does not contain key information, the visual language big model can be called again to process the processing results of the prediction module at level 2, the first set of intermediate results, and the second sensor data at level 0 to obtain the first set of intermediate results.
[0436] For example, the processing result of the prediction module, "motion information of moving objects such as vehicles and pedestrians in the future (e.g., 3 seconds)," and the first set of intermediate results can be filled into the prompt word. The resulting prompt word is like this: "You are a parking expert and now need to find an empty parking space. According to the current vehicle-side prediction result, the information is as follows: In the first image, the vehicle in the area of the top left corner (20,20), length 90, width 50; corresponding to a rectangular area" may be shown as the red line in the image in the next 3 seconds (project the trajectory onto the image and represent it with a red line; the color is not limited to the corresponding prompt word). Regarding your previous judgment result, regarding the parking space: there is a parking space in the indicated area, but it is not vacant; regarding obstacles: there are vehicles in the indicated area, but no pedestrians; risk: there is no risk at present (input the first set of intermediate results), confirm again, identify obstacles and risks near the parking space, and provide parking suggestions." After that, the above 5 images and the prompt word can be input into the visual language large model to obtain the second set of intermediate results. If the second set of intermediate results (e.g., "Regarding obstacles: There are vehicles in the indicated area. The vehicles may move, but their positions are far away, there is no risk of collision, and there are no pedestrians; Risk: No risk at present; Parking suggestion: Please continue driving along the road and observe the availability of parking spaces") still does not contain key information, the visual language big model can be called again to process the processing results of the fusion module at level 3, the second set of intermediate results, and the second sensor data at level 0 to obtain the third set of intermediate results.
[0437] For example, the processing results of the fusion module, "motion information of moving objects such as vehicles and pedestrians, the position of stationary objects such as pillars and water barriers, and changes in the occupancy of drivable areas and parking spaces within a future period of time (e.g., 3 seconds)," and the second set of intermediate results can be filled into the prompt words. The resulting prompt words are such as "You are a parking expert and now need to find an empty parking space. Please refer to the current vehicle-side fusion results. The information is as follows: In the first image [corresponding to a certain image], the vehicles in the [top left corner (20,20), length 90, width 50; corresponding to a certain rectangular area] region may be […] within the next 3 seconds." The red line in the image (projecting the trajectory onto the image and representing it with a red line; the color is not limited to the corresponding prompt word) indicates that the vehicle in the area of the top left corner (50,100), length 70, width 30; corresponding to a rectangular area in the second image may be shown as the green line in the image within the next 3 seconds. Regarding your previous judgment: Regarding obstacles: there are vehicles in the indicated area; the vehicles may move, but their positions are far away, with no collision risk and no pedestrians; Risk: currently no risk. Reconfirm, identify obstacles and risks near the parking space, and provide parking suggestions.” Then, input the prompt word and the above 5 images into the visual language model to obtain the third set of intermediate results. If the third set of intermediate results (e.g., "Regarding obstacles: The first image indicates that there is a vehicle in the area, which may move but there is no risk of collision; the second image indicates that there is a vehicle in the area, which may move but there is no risk of collision; Risk: No risk at present; Parking suggestion: Please continue driving along the road and observe the availability of parking spaces") still does not contain key information, the visual language big model can be called again to process the processing results of the level 4 control module, the third set of intermediate results, and the second sensor data of level 0 to obtain the fourth set of intermediate results.
[0438] For example, the processing results of the traffic control module, "the target vehicle's driving trajectory and speed over a future period, and signals such as vehicle acceleration, steering wheel angle, and gear position," along with the third set of intermediate results, can be filled into a prompt message. The resulting prompt message might be: "You are a parking expert and now need to find an empty parking space. Please refer to the current vehicle-side traffic control results. The information is as follows: The vehicle's driving trajectory over the next 3 seconds is shown by the blue line in the first image, with a speed of 5 km / s. Currently, the vehicle is being controlled to move straight, with an acceleration of -0.1 km / s." 2The steering wheel angle is 0 degrees. Regarding your previous judgment about obstacles: the first image indicates a vehicle exists in the area; the vehicle may move, but there is no risk of collision. The second image also indicates a vehicle exists in the area; the vehicle may move, but there is no risk of collision. Risk: There is currently no risk. Reconfirm, identify obstacles and risks near the parking space, and provide parking suggestions.” Afterwards, this prompt and the aforementioned 5 images can be input into the visual language model to obtain the 4th set of intermediate results.
[0439] In one possible scenario, if the intermediate result in group 0 (e.g., "This image shows an underground parking garage scene. Regarding parking spaces: There are empty parking spaces in the current image; Regarding obstacles: There is broken glass near the parking space, which may puncture tires; Risk: There is broken glass on the road surface, indicating a possible accident occurred here previously. Please be aware of the current road conditions; Parking advice: Please contact the cleaning staff to remove the broken glass before parking in an empty space, or please drive away and find a new empty parking space.") contains key information (such as information indicating risks), then the intermediate result in group 0 will be determined as the fourth processing result.
[0440] In one possible scenario, if the first set of intermediate results (e.g., "Regarding parking spaces: the indicated area is an empty parking space; regarding obstacles: no vehicles or pedestrians; risk: no risk at present; parking suggestion: please drive to the vicinity of the empty parking space to park") contains key information (such as information indicating an empty parking space), the first set of intermediate results will be determined as the fourth processing result.
[0441] In one possible scenario, if the second set of intermediate results (e.g., "Regarding obstacles: There are vehicles in the indicated area. The vehicles may move and are in close proximity, which could lead to a collision. There are no pedestrians. Risk: The moving vehicles are in close proximity. Please control your speed, maintain a safe distance, and avoid a collision. Parking advice: Slow down and give way. Park in the empty space after the moving vehicles have left") contains key information (such as information about collision risks and empty parking spaces), then the second set of intermediate results will be determined as the fourth processing result.
[0442] One possible scenario is that when the third set of intermediate results (e.g., "Regarding obstacles: The first image indicates the presence of a vehicle in the area, which may move and is relatively close, potentially posing a collision risk; the second image indicates the presence of a vehicle in the area, which may move but poses no collision risk; Risk: The moving vehicle is relatively close, please control your speed, maintain a safe distance, and avoid a collision; Parking advice: Slow down and yield, and park in the empty parking space after the moving vehicle has left") contains key information (such as information about collision risks and empty parking spaces), the third set of intermediate results will be determined as the fourth processing result.
[0443] In one possible scenario, if the fourth intermediate result (e.g., "Regarding obstacles: There is a tilted barrier near the vehicle's trajectory indicated by the image; Risk: The vehicle is close to the barrier and may easily run over it; Parking advice: Please avoid the tilted barrier, or stop and remove the barrier before continuing to drive along the road to observe available parking spaces") contains key information (such as information indicating risks), the fourth intermediate result will be determined as the fourth processing result.
[0444] In one possible scenario, if none of the first four intermediate results contain crucial information, the remote device cannot generate a second processing result. In this case, the remote device may return a failure message or not return any message to the vehicle. If the vehicle receives a failure message from the remote device within a preset time period, or does not receive a message from the remote device, it can be determined that the remote device collaboration has failed, and the vehicle can proceed with independent intelligent driving.
[0445] For example, another implementation of S903 may include the following S9033 to S9034:
[0446] S9033, integrates the second sensor data and the fourth processing result to obtain the intermediate processing result.
[0447] For example, the second sensor data is the image captured by the vehicle's camera, and the fourth processing result is the parking trajectory output by the planning and control module. The image captured by the vehicle's camera and the parking trajectory output by the planning and control module can be fused. For example, the area covered by the vehicle's parking can be superimposed on the image, and the intermediate processing result obtained is an image containing a semi-transparent color area (or a curved frame of the area the vehicle is expected to pass through).
[0448] S9034 uses a target large model to process the intermediate processing results to obtain the second processing result.
[0449] The intermediate processing results can then be input into the target large model, which will process these results to obtain the second processing result. For example, an image containing a semi-transparent colored area (or a curved frame of the vehicle's expected passage area) can be input into the target large model. The target large model can then determine whether the parking lock within the semi-transparent colored area (or the curved frame of the vehicle's expected passage area) is open or closed. The output second processing result can include: parking lock open or parking lock closed. Compared to simply inputting the second sensor data into the target large model, this approach utilizes more comprehensive and richer information to generate a more accurate second processing result, thereby improving the reliability and safety of autonomous driving.
[0450] It should be noted that in practical applications, when implementing S903, at least one of the implementation methods of S9031 to S9032 or S9033 to S9034 can be used.
[0451] S904, send the second processing result to the current vehicle.
[0452] In one possible approach, when only one large model is deployed in the remote device, it is not necessary to perform the operation of "determining the target large model from multiple large models based on the trigger condition identifier of multi-terminal collaborative intelligent driving that is satisfied in the current situation" and directly call the large model to process the second sensing data.
[0453] In one possible approach, when multiple large models are deployed on a remote device, their priority can be prioritized; for example, models with language processing capabilities have higher priority than those without. Furthermore, the visual language large model / multimodal large model can be called first to process the second sensor data, followed by the visual basic large model / dedicated task large model.
[0454] In one possible approach, when multiple large models are deployed on the remote device, the visual language large model can be called first to process the second sensor data. Then, based on the processing results of the visual language large model, a large model is selected to process both the second sensor data and the visual language large model's processing results. For example, if an image is identified by the visual language large model as a scene of parking on a curb, then the basic visual model is called to process the image, obtaining the curb segmentation result, thereby achieving a more refined and accurate perception result and improving the vehicle's perception performance.
[0455] For example, the remote device can also receive sensor data sent by the field device; based on the target large model and the second sensor data, determine the second processing result, including: processing the second sensor data and the sensor data sent by the field device using the target large model to obtain the second processing result. In this way, more comprehensive and richer information can be used to generate a more accurate second processing result; thereby improving the reliability and safety of intelligent driving.
[0456] For example, the second processing result is obtained by using the target large model to process the second sensing data sent by the current vehicle and the sensing data sent by the field device. This can be understood as: the sensing data sent by the field device is regarded as auxiliary information, and the target large model combines the auxiliary information to process the second sensing data to obtain the second processing result.
[0457] For example, the remote device can also use a target large model to process the second sensor data and historical data to obtain a second processing result; wherein, the historical data includes historical sensor data and / or historical processing results. In this way, the information used to generate the second processing result is more comprehensive and richer, which can improve the accuracy of the second processing result, thereby enhancing the reliability and safety of intelligent driving.
[0458] In addition, the method of "using the target large model to process the second sensor data and historical data to obtain the second processing result" can introduce historical data into the processing of the target large model. In this way, there is no need to use these historical data to retrain the large model, which can save training costs.
[0459] For example, the second processing result is obtained by using a target large model to process the second sensing data and historical data. This can be understood as treating the historical data as auxiliary information, and the target large model combines this auxiliary information to process the second sensing data to obtain the second processing result.
[0460] In one example, FIG11 shows a schematic block diagram of an apparatus 1100 according to an embodiment of the present application. The apparatus 1100 may include a processor 1101 and a transceiver / transceiver pin 1102, and optionally, a memory 1103.
[0461] The various components of device 1100 are coupled together via bus 1104, which includes a data bus, a power bus, a control bus, and a status signal bus. However, for clarity, all buses are referred to as bus 1104 in the figure.
[0462] Optionally, the memory 1103 can be used to store instructions from the aforementioned method embodiments. The processor 1101 can be used to execute the instructions in the memory 1103, control the receive pin to receive signals, and control the transmit pin to transmit signals.
[0463] Device 1100 may be an electronic device or a chip of an electronic device in the above method embodiments.
[0464] For example, the electronic device may be a remote device or a vehicle-mounted terminal.
[0465] Figure 12 shows an exemplary structural diagram of electronic device 120. Electronic device 120 may be a vehicle-mounted terminal.
[0466] According to Figure 12, the electronic device 120 includes components such as an application processor 1201, a memory 1202, a wireless communication module 1203, a graphics processing unit (GPU) 1204, and an input / output (I / O) device 1205. Those skilled in the art will understand that the hardware structure shown in Figure 12 does not constitute a limitation on the electronic device 120, and the electronic device 120 may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0467] The following is a detailed description of each component of the electronic device 120 with reference to Figure 12:
[0468] The application processor 1201 is the control center of the electronic device 120, connecting various components of the electronic device 120 via various interfaces and buses. In some embodiments, the application processor 1201 may include one or more processing modules.
[0469] Memory 1202 stores computer programs, such as operating system 1222 and application program 1221 shown in Figure 12. Application processor 1201 is configured to execute the computer program in memory 1202 to implement the functions defined by the computer program. For example, application processor 1201 executes operating system 1222 to implement various functions of the operating system on electronic device 120. Memory 1202 also stores other data besides computer programs, such as data generated during the operation of operating system 1222 and application program 1221. Memory 1202 is a non-volatile storage medium, generally including main memory and secondary storage. Main memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), or cache. Secondary storage includes, but is not limited to, flash memory, hard disk, optical disk, universal serial bus (USB) disk, etc. Computer programs are usually stored on secondary storage, and the processor loads the program from secondary storage into main memory before executing the computer program.
[0470] The memory 1202 can be independent and connected to the application processor 1201 via a bus; the memory 1202 can also be integrated with the application processor 1201 into a chip subsystem.
[0471] The wireless communication module 1203 provides solutions for wireless communication applications on the electronic device 120, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 1203 can be one or more devices integrating at least one communication processing module. The wireless communication module 1203 receives electromagnetic waves via an antenna, demodulates and filters the electromagnetic wave signals, and sends the processed signals to the application processor 1201. The wireless communication module 1203 can also receive signals to be transmitted from the application processor 1201, frequency modulate and amplify them, and then convert them into electromagnetic waves for radiation via the antenna.
[0472] GPU 1204: Performs drawing and rendering calculations on image data to generate the image to be displayed. Also known as a display core or visual processor, it is a microprocessor that performs image processing tasks and may include 2D (dimensional) and / or 3D processing capabilities. Electronic device 120 may include one or more GPUs that execute program instructions to generate or modify display information.
[0473] Input / output devices 1205 include, but are not limited to: display 1251, touch screen 1253, and audio circuit 1255.
[0474] The touchscreen 1253 can collect touch events from the user of the electronic device 120 on or near the touchscreen 1253 (such as user actions on or near the touchscreen 1253 using a finger, stylus, or any suitable object), and send the collected touch events to other devices (such as the application processor 1201). The user's actions near the touchscreen 1253 can be called hover touch; through hover touch, the user can select, move, or drag targets (such as icons) without directly touching the touchscreen 1253. Furthermore, the touchscreen 1253 can be implemented using various types of touchscreens, including resistive, capacitive, infrared, and surface acoustic wave.
[0475] The display (also called a screen) 1251 is used to display information input by the user or information shown to the user. The display can be configured using a liquid crystal display (LCD), organic light emitting diode (OLED), or similar methods. The touchscreen 1253 can cover the display 1251. When the touchscreen 1253 detects a touch event, it transmits the information to the application processor 1201 to determine the type of touch event. The application processor 1201 then provides corresponding visual output on the display 1251 based on the type of touch event. Although in Figure 12, the touchscreen 1253 and the display 1251 are shown as two separate components implementing the input and output functions of the electronic device 120, in some embodiments, the touchscreen 1253 and the display 1251 can be integrated to achieve the input and output functions of the electronic device 120. Furthermore, the touchscreen 1253 and the display 1251 can be configured as a full-panel display on the front of the electronic device 120 to achieve a borderless structure. For example, the display can be used to display the display interfaces of Figures 4A to 4C, 6A to 6D, and 8A to 8G.
[0476] The audio circuit 1255, speaker 1256, and microphone 1257 provide an audio interface between the user and the electronic device 120. The audio circuit 1255 converts received audio data into electrical signals, transmits them to the speaker 1256, and the speaker 1256 converts them into sound signals for output. On the other hand, the microphone 1257 converts collected sound signals into electrical signals, which are received by the audio circuit 1255, converted into audio data, and then transmitted to, for example, another electronic device via a modem processor and radio frequency module, or output to the memory 1202 for further processing.
[0477] Optionally, the electronic device 120 may further include a microcontroller unit (MCU) 1206, which is a coprocessor for acquiring and processing data from the sensor 1261. The MCU 1206 has lower processing power and power consumption than the application processor 1201, but features an "always-on" characteristic, allowing it to continuously collect and process sensor data while the application processor 1201 is in sleep mode, ensuring normal sensor operation with extremely low power consumption. In one embodiment, the MCU 1206 may be a sensor hub chip. The sensor 1261 may include a light sensor and a gyroscope sensor. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display 1251 according to the ambient light level, and the proximity sensor can turn off the power to the display when the electronic device 120 is moved to the ear. The gyroscope sensor can be used to determine the motion posture of the electronic device 120. In some embodiments, the angular velocity of the electronic device 120 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor. The angular velocity information obtained by the gyroscope sensor can be converted into a rotation matrix that describes the object's rotation, which can then be used for image alignment. The gyroscope sensor can also be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor detects the angle of vibration of the electronic device 120, calculates the distance the lens module needs to compensate based on the angle, and allows the lens to counteract the vibration of the electronic device 120 through reverse movement, thus achieving image stabilization. The gyroscope sensor can also be used in navigation and motion-sensing game scenarios. Sensor 1261 can also include other sensors such as accelerometers, barometers, hygrometers, thermometers, and infrared sensors, which will not be elaborated here. The MCU 1206 and sensor 1261 can be integrated onto the same chip or are separate components connected via a bus.
[0478] Optionally, the electronic device 120 may also include an image signal processor (ISP) 1207. The ISP 1207 interfaces with the camera 1271 to acquire images and perform image processing (such as exposure control, white balance, color calibration, or noise removal) to generate image data. It may include a processor core that performs the necessary software processing or be implemented purely in hardware.
[0479] Optionally, the electronic device 120 may further include a mobile communication module 1208. The mobile communication module 1208 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on the electronic device 120. The mobile communication module 1208 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 1208 can receive electromagnetic waves via an antenna, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 1208 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna. In some embodiments, at least some functional modules of the mobile communication module 1208 may be housed in the application processor 1201. In some embodiments, at least some functional modules of the mobile communication module 1208 and at least some modules of the application processor 1201 may be housed in the same device.
[0480] Furthermore, the operating system 1222 mounted on the electronic device 120 can provide... This application does not impose any restrictions on other operating systems, including those used in the embodiments of the present application.
[0481] Those skilled in the art will understand that electronic device 120 may include fewer or more components than those shown in FIG12, which only includes components more relevant to the various implementations disclosed in the embodiments of this application.
[0482] In one possible implementation, the intelligent driving method of this application embodiment can be executed by an electronic device 120. For example, when a user interacts with an in-vehicle terminal during driving to activate the intelligent driving function of the in-vehicle terminal, the intelligent driving method provided in this application embodiment can be used for intelligent driving.
[0483] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0484] This application also provides a chip, including one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the above-described related method steps that implement the method in the above embodiments are executed. The interface circuit is a transceiver / transceiver pin 1102.
[0485] This embodiment also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the methods described in the above embodiments.
[0486] This embodiment also provides a computer program product containing computer instructions that, when executed by a computer or processor, cause the computer to perform the aforementioned related steps to implement the methods described in the above embodiments.
[0487] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component, or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the methods in the above-described method embodiments.
[0488] In this embodiment, the electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0489] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0490] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0491] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0492] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0493] Any content in the various embodiments of this application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.
[0494] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0495] The steps of the methods or algorithms described in conjunction with the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium well known in the art. One exemplary embodiment couples a storage medium to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.
[0496] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer-readable storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0497] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An intelligent driving method, characterized in that, The method includes: Determine whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein, the current situation includes at least one of the following: the scene in which the current vehicle is located, the network parameters between the current vehicle and the remote device, the reliability of the first processing result obtained by the current vehicle in processing the first sensor data, or the system state of the current vehicle; When it is determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, multi-terminal collaborative intelligent driving data is sent to the remote device; wherein, the multi-terminal collaborative intelligent driving data includes second sensor data, which is collected by the current vehicle after collecting the first sensor data or the second sensor data is the first sensor data; Receive a second processing result sent by the remote device, wherein the second processing result is obtained by the remote device processing the multi-terminal collaborative intelligent driving data; Based on the second processing result, path planning and control are performed on the current vehicle.
2. The method according to claim 1, characterized in that, The triggering conditions for the multi-terminal collaborative intelligent driving include at least one of the following: The current scenario in which the vehicle is located is a preset scenario; The network parameters between the current vehicle and the remote device meet the preset network conditions; The reliability of the first processing result meets the preset reliability condition; The current system state of the vehicle meets the preset state.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Based on the second sensor data, the scene in which the current vehicle is located is identified.
4. The method according to any one of claims 1 to 3, characterized in that, The preset reliability conditions include at least one of the following: The probability of the first processing result is greater than the probability threshold; The confidence level of the first processing result is less than the confidence level threshold; The confidence variance of the first processing result is greater than the variance threshold.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Before determining whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, the multi-terminal collaborative processing function is enabled in response to the first user's operation. In response to a second user's action, the intelligent driving function is activated.
6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Before determining whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, when the current vehicle is detected to have failed to park, the switch control for enabling intelligent driving function and displaying multi-terminal collaborative processing function is activated. In response to a third user operation on the switch control, the multi-terminal collaborative processing function is enabled.
7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Before determining whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, when it is detected that the current vehicle is in a parking state, the switch control for enabling intelligent driving function and displaying multi-terminal collaborative processing function is activated. In response to a fourth user operation on the switch control, the multi-terminal collaborative processing function is enabled.
8. The method according to any one of claims 1 to 7, characterized in that, The step of performing path planning and control on the current vehicle based on the second processing result includes: The third sensor data is processed to obtain a third processing result; wherein the third sensor data is collected by the current vehicle after the second sensor data is collected; By combining the second processing result and the third processing result, the target processing result is obtained; Based on the target processing results, path planning and control are performed on the current vehicle.
9. The method according to claim 8, characterized in that, The second processing result includes detection information of the first object, and the third processing result includes detection information of the second object; fusing the second processing result and the third processing result to obtain the target processing result includes: The detection information of the first object and the detection information of the second object are spatiotemporally aligned to obtain the target processing result.
10. The method according to claim 8 or 9, characterized in that, The intelligent driving is intelligent parking, and the second processing result also includes parking space indication information. The step of planning a path for the current vehicle based on the target processing result includes: Based on the target processing result and the parking space indication information, determine the virtual parking space; The virtual parking space is identified as the target parking space, and a path is planned for the current vehicle based on the target processing result and the target parking space.
11. The method according to any one of claims 1 to 10, characterized in that, The second processing result includes textual description information of the current vehicle's location scene, and the method further includes: Based on the text description of the current vehicle's location, a prompt message is generated and displayed.
12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: This displays the current network connection status between the current vehicle and the remote device.
13. The method according to any one of claims 1 to 12, characterized in that, The multi-terminal collaborative intelligent driving data also includes at least one of the trigger condition identifier of multi-terminal collaborative intelligent driving that is satisfied in the current situation or the fourth processing result, wherein the fourth processing result is obtained by the current vehicle processing the second sensor data.
14. The method according to any one of claims 1 to 13, characterized in that, The intelligent driving is intelligent parking, and the second processing result includes the attribute information of the first object; the step of performing path planning and control on the current vehicle based on the second processing result includes: Select the target parking space based on the attribute information of the first object; Based on the attribute information of the first object and the target parking space, generate the target path for the current vehicle; The current vehicle is controlled according to its target path.
15. An intelligent driving method, characterized in that, The method includes: Receive multi-terminal collaborative intelligent driving data sent by the current vehicle; wherein, the multi-terminal collaborative intelligent driving data includes the trigger condition identifier of multi-terminal collaborative intelligent driving that is satisfied in the current situation and the second sensor data; Based on the trigger condition identifiers of multi-terminal collaborative intelligent driving satisfied by the current situation, the target large model is determined from multiple large models; Based on the target large model and the second sensing data, the second processing result is determined; The second processing result is sent to the current vehicle, and the second processing result is used to perform path planning and control on the current vehicle.
16. The method according to claim 15, characterized in that, The step of determining the target large model from multiple large models based on the triggering conditions of multi-terminal collaborative intelligent driving satisfied by the current situation includes: Based on the triggering conditions of multi-terminal collaborative intelligent driving satisfied by the current situation, a mapping relationship is found to determine one or more target large models from multiple large models; The mapping relationship includes the relationship between each trigger condition in the multi-terminal collaborative intelligent driving trigger conditions and one or more corresponding large models.
17. The method according to claim 15 or 16, characterized in that, The multi-terminal collaborative intelligent driving data also includes a fourth processing result, which is obtained by the current vehicle processing the second sensor data; The determination of the second processing result based on the target large model and the second sensing data includes: By fusing the second sensing data and the fourth processing result, an intermediate processing result is obtained; The intermediate processing results are processed using the target large model to obtain the second processing result.
18. The method according to claim 15 or 16, characterized in that, The multi-terminal collaborative intelligent driving data also includes a fourth processing result, which is obtained by the current vehicle processing the second sensor data; The determination of the second processing result based on the target large model and the second sensing data includes: Obtain the level of the second sensing data and the level of the fourth processing result; Based on the level of the second sensing data and the level of the fourth processing result, the target large model is used to process the second sensing data and the fourth processing result to obtain the second processing result.
19. The method according to claim 18, characterized in that, The fourth processing result includes N sets of data, which correspond one-to-one with N levels. The second sensing data corresponds to one level, and N is a positive integer. The step of processing the second sensing data and the fourth processing result using the target large model based on the level of the second sensing data and the level of the fourth processing result to obtain the second processing result includes: Repeat the following steps until i equals N, with i initially set to 1; The i-th group of data, the (i-1)-th group of intermediate results, and the second sensing data from the fourth processing result are input into the target large model to obtain the i-th group of intermediate results; wherein, the i-th group of data corresponds to level i, and the 0-th group of intermediate results is obtained by inputting the second sensing data into the target large model; Increment i by 1; When i equals N, the second processing result is determined based on the N intermediate results from the first intermediate result to the Nth intermediate result.
20. The method according to any one of claims 15 to 19, characterized in that, The method further includes: Receive sensor data sent by the field-end equipment; The determination of the second processing result based on the target large model and the second sensing data includes: The second sensing data and the sensing data sent by the field terminal device are processed using the target large model to obtain the second processing result.
21. The method according to any one of claims 15 to 19, characterized in that, The determination of the second processing result based on the target large model and the second sensing data includes: The target large model is used to process the second sensing data and historical data to obtain the second processing result; The historical data includes historical sensor data and / or historical processing results.
22. The method according to any one of claims 15 to 21, characterized in that, The second processing result includes: detection information of the first object, text description information of the scene where the current vehicle is located, and attribute information of the first object.
23. A multi-terminal collaborative intelligent driving system, characterized in that, The multi-terminal collaborative intelligent driving system includes a vehicle and remote devices, wherein: The vehicle is used to determine whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein, the current situation includes at least one of the following: the scene in which the vehicle is located, the network parameters between the vehicle and the remote device, the reliability of the first processing result obtained by the vehicle in processing the first sensor data, or the system state of the vehicle. The vehicle is configured to send multi-terminal collaborative intelligent driving data to the remote device when it is determined that the current situation meets the triggering conditions of the multi-terminal collaborative intelligent driving; wherein, the multi-terminal collaborative intelligent driving data includes second sensor data, which is collected by the vehicle after collecting the first sensor data or the second sensor data is the first sensor data; The remote device is used to process the second sensor data, obtain a second processing result, and send the second processing result to the vehicle. The vehicle is also used to perform path planning and control on the vehicle based on the second processing result.
24. The system according to claim 23, characterized in that, The multi-terminal collaborative intelligent driving data also includes the trigger condition identifiers for multi-terminal collaborative intelligent driving that are met in the current situation; The remote device is further configured to determine a target large model from multiple large models based on the trigger condition identifier of multi-terminal collaborative intelligent driving satisfied by the current situation; and to determine a second processing result based on the target large model and the second sensing data.
25. A vehicle, characterized in that, The vehicle is used for: Determine whether the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving; wherein, the current situation includes at least one of the following: the scene in which the vehicle is located, the network parameters between the vehicle and the remote device, the reliability of the first processing result obtained by the vehicle in processing the first sensor data, or the system state of the vehicle. When it is determined that the current situation meets the triggering conditions for multi-terminal collaborative intelligent driving, multi-terminal collaborative intelligent driving data is sent to the remote device; wherein, the multi-terminal collaborative intelligent driving data includes second sensor data, which is collected by the vehicle after collecting the first sensor data or the second sensor data is the first sensor data; Receive a second processing result sent by the remote device, wherein the second processing result is obtained by the remote device processing the multi-terminal collaborative intelligent driving data; Based on the second processing result, the vehicle is subjected to path planning and control.
26. A remote device, characterized in that, The remote device is used for: Receive multi-terminal collaborative intelligent driving data sent by the current vehicle; wherein, the multi-terminal collaborative intelligent driving data includes the trigger condition identifier of multi-terminal collaborative intelligent driving that is satisfied in the current situation and the second sensor data; Based on the trigger condition identifiers of multi-terminal collaborative intelligent driving satisfied by the current situation, the target large model is determined from multiple large models; Based on the target large model and the second sensing data, the second processing result is determined; The second processing result is sent to the current vehicle, and the second processing result is used to perform path planning and control on the current vehicle.
27. A vehicle-mounted terminal, characterized in that, include: A memory and a processor, wherein the memory is coupled to the processor; The memory stores program instructions that, when executed by the processor, cause the vehicle terminal to perform the method as described in any one of claims 1 to 14.
28. A remote device, characterized in that, include: A memory and a processor, wherein the memory is coupled to the processor; The memory stores program instructions that, when executed by the processor, cause the remote device to perform the method as described in any one of claims 15 to 22.
29. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed on a computer or processor, causes the computer or processor to perform the method as described in any one of claims 1 to 22.
30. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a computer or processor, cause the steps of the method as described in any one of claims 1 to 22 to be performed.
Citation Information
Patent Citations
Methods And Systems For Remote Parking Assistance
CN107957724A
Cooperative vehicle path generation
CN116547495A
Vehicle parking processing method and device and computer readable storage medium
CN117775013A
Remote automatic parking assist system and control method thereof
KR1020170025206A