A processing method and device of an interactive interface in autonomous driving, and a vehicle-mounted device

By presenting inference data after sensor data processing on the display screen of autonomous vehicles, the problem of users having difficulty understanding the autonomous driving process is solved, thus improving the user experience.

CN122211409APending Publication Date: 2026-06-16SHANGHAI LIXIANG AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI LIXIANG AUTOMOBILE CO LTD
Filing Date
2025-12-16
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

During autonomous driving, users find it difficult to intuitively understand the vehicle's driving reasoning process, resulting in a poor user experience.

Method used

Data is collected by sensing devices, processed by intelligent models in the intelligent driving system to generate autonomous driving inference data, and presented on the vehicle's display screen as trajectory prediction, obstacle recognition, scene description and question-and-answer data, forming an inference interface and an environmental perception interface, thereby improving the user's understanding of the autonomous driving process.

Benefits of technology

Users can intuitively understand the reasoning process in autonomous driving, which improves the autonomous driving experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122211409A_ABST
    Figure CN122211409A_ABST
Patent Text Reader

Abstract

The application discloses a processing method and device of an interactive interface in automatic driving and a vehicle-mounted equipment, and relates to the field of automobiles. The method comprises the following steps: determining automatic driving inference data according to sensing data collected by at least one sensing device arranged on a target vehicle; and presenting the automatic driving inference data on a display screen of the target vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 2024118630150, filed on December 16, 2024, entitled "A method, apparatus and vehicle-mounted device for processing an interactive interface in autonomous driving", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of intelligent driving technology, and in particular to a method, device and vehicle-mounted equipment for processing interactive interfaces in autonomous driving. Background Technology

[0003] In autonomous driving scenarios, the intelligent driving system deployed in the vehicle can output an environmental perception interface on the in-vehicle screen, which can provide environmental information and navigation information about the current driving environment. Summary of the Invention

[0004] In view of the above problems, this application provides a method, apparatus, and in-vehicle equipment for processing the interactive interface in autonomous driving, so as to improve the user experience of autonomous driving. The specific solution is as follows:

[0005] The first aspect of this application provides a method for processing an interactive interface in autonomous driving, comprising:

[0006] Autonomous driving inference data is determined based on sensor data collected by at least one sensor device deployed on the target vehicle.

[0007] The autonomous driving inference data is presented on the display screen of the target vehicle.

[0008] In one possible implementation, the autonomous driving inference data includes at least one of the following:

[0009] The trajectory prediction data of the target vehicle obtained by processing the sensor data;

[0010] Obstacle recognition data obtained by processing the aforementioned sensor data;

[0011] Scene description data of the environment in which the target vehicle is located, obtained by processing the sensor data;

[0012] Scene decision data obtained based on the scene description data;

[0013] Question and answer data in the question and answer operation based on the sensor data.

[0014] In one possible implementation, the trajectory prediction data includes: at least one predicted trajectory line; the predicted trajectory line corresponds to a predicted probability value; at least one of the predicted trajectory lines is presented on the display screen in the form of a trajectory image; wherein, the trajectory image corresponding to the predicted trajectory line with the largest predicted probability value has a target identifier, and the trajectory image also includes at least vehicle identifiers corresponding to other vehicles in the environment where the target vehicle is located.

[0015] In one possible implementation, the presentation position of the trajectory image with the target identifier on the display screen changes in response to a change in a switching condition that is met.

[0016] In one possible implementation, the obstacle recognition data includes: attention scores corresponding to obstacles in the environment where the target vehicle is located; the attention scores are output in the form of an attention layer in the driving scene image presented on the display screen; the driving scene image is an image obtained by using an image acquisition device to acquire images of the environment where the target vehicle is located; wherein the hue distribution of the attention layer matches the attention scores.

[0017] In one possible implementation, the scene description data and the scene decision data are presented on the display screen in the form of target text, and the display screen also presents a driving scene image corresponding to the target text; the driving scene image is an image obtained by using an image acquisition device to capture images of the environment where the target vehicle is located;

[0018] In the case of multiple driving scene images, the display screen presents the driving scene images and their corresponding target text in the order of their acquisition time.

[0019] In one possible implementation, the driving scene image and its corresponding target text are presented on the display screen in the form of a question-and-answer format.

[0020] The driving scene image is the question in the image-text Q&A, and the target text is the answer in the image-text Q&A.

[0021] In one possible implementation, the display screen includes a first display screen and a second display screen, wherein the deployment position of the second display screen matches the passenger seat position of the target vehicle, and the deployment position of the first display screen matches the driver seat position of the target vehicle.

[0022] The method further includes:

[0023] After the autonomous driving inference data is presented on the first display screen, in response to a push command for the autonomous driving inference data, the autonomous driving inference data is presented on the second display screen.

[0024] A second aspect of this application provides a processing apparatus for an interactive interface in autonomous driving, the apparatus comprising:

[0025] The data determination unit is used to determine autonomous driving reasoning data based on sensing data collected by at least one sensing device deployed on the target vehicle.

[0026] The reasoning presentation unit is used to present the autonomous driving reasoning data on the display screen of the target vehicle.

[0027] A third aspect of this application provides a computer program product, including computer-readable instructions, which, when executed on an in-vehicle device, cause the in-vehicle device to implement the method for processing the interactive interface in autonomous driving as described in the first aspect or any implementation thereof.

[0028] A fourth aspect of this application provides an in-vehicle device, comprising:

[0029] At least one sensing device, said sensing device being used to collect sensing data;

[0030] At least one processor is configured to determine autonomous driving inference data based on the sensing data and to present the autonomous driving inference data on a display screen of the target vehicle.

[0031] The fifth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an in-vehicle device, enable the in-vehicle device to process the interactive interface in autonomous driving as described in the first aspect or any implementation thereof.

[0032] By means of the above technical solution, the processing method, device and vehicle equipment for the interactive interface in autonomous driving provided in this application present autonomous driving inference data determined by the sensing data collected by the sensing device on the display screen of the target vehicle. In this way, during the autonomous driving process, not only can other information of the target vehicle driving process be displayed, but also the autonomous driving inference data that realizes the vehicle driving control can be presented, so that users can intuitively understand the inference process in the autonomous driving process through these autonomous driving inference data, thereby improving the user's autonomous driving experience. Attached Figure Description

[0033] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0034] Figure 1 A flowchart illustrating a method for processing an interactive interface in autonomous driving, as provided in an embodiment of this application;

[0035] Figure 2 This is an architecture diagram of the target vehicle to which the embodiments of this application apply;

[0036] Figure 3 This is an example diagram showing the output of the inference interface and the environment perception interface on the first display screen in an embodiment of this application;

[0037] Figure 4 This is another example diagram showing the output of the inference interface and the environment perception interface on the first display screen in this application embodiment;

[0038] Figure 5 This is an example diagram showing the output of the inference interface in an embodiment of this application;

[0039] Figure 6 This is an example diagram illustrating the triggering of the reasoning interface exit in an embodiment of this application;

[0040] Figure 7 This is an example diagram illustrating the presentation of autonomous driving reasoning data in the reasoning interface in an embodiment of this application;

[0041] Figure 8 This is an example diagram showing the presentation of predicted trajectory data in the inference interface in an embodiment of this application;

[0042] Figure 9 This is an example diagram illustrating the switching of the presentation position of the trajectory image in the inference interface in an embodiment of this application;

[0043] Figure 10 This is an example diagram showing the presentation of the attention layer in the inference interface in an embodiment of this application;

[0044] Figure 11 This is another example diagram illustrating the presentation of the attention layer in the inference interface in this application embodiment;

[0045] Figure 12 This is an example diagram showing the presentation of target text in the inference interface in an embodiment of this application;

[0046] Figure 13 This is an example diagram illustrating the pushing of the inference interface between the first and second displays in an embodiment of this application.

[0047] Figure 14This is an example diagram illustrating the presentation of question-and-answer data in the reasoning interface in this embodiment of the application;

[0048] Figure 15 A schematic diagram of the structure of a processing device for an interactive interface in autonomous driving provided in an embodiment of this application;

[0049] Figure 16 This is a schematic diagram of the structure of a vehicle-mounted device provided in an embodiment of this application;

[0050] Figure 17 This is a diagram of the intelligent driving system architecture for presenting autonomous driving reasoning data in this application embodiment. Detailed Implementation

[0051] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0052] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0053] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0054] Reference Figure 1 This is a flowchart illustrating a method for processing an interactive interface in autonomous driving, provided by an embodiment of this application. This method can be applied to in-vehicle devices with processors deploying intelligent driving systems. Figure 2 As shown, the target vehicle is equipped with at least one sensing device and a processor, such as an on-board chip. Based on the processor on the target vehicle, an intelligent driving system can be deployed. The intelligent driving system can include one or more intelligent models, such as an end-to-end model and a Vision-Language Model (VLM). The technical solution in this embodiment is mainly used to improve the user experience of intelligent driving.

[0055] Specifically, the method in this embodiment may include the following steps:

[0056] Step 101: Determine autonomous driving inference data based on the sensor data collected by at least one sensor device deployed on the target vehicle.

[0057] The sensing devices may include cameras, radar, inertial measurement units (IMUs), positioning devices, wheel sensors, etc. Therefore, the sensing data obtained in this embodiment may include: multimedia data such as images or video streams collected by each camera, scanning data collected by various types of radar, acceleration and angular velocity data collected by the IMU, coordinate data collected by the positioning device, wheel rotation speed data collected by the wheel sensors, and so on.

[0058] In one implementation, this embodiment can process the sensor data using an intelligent model in the autonomous driving system to obtain autonomous driving reasoning data.

[0059] Specifically, autonomous driving inference data can be used to generate driving control data for the target vehicle. The target vehicle can then respond to the driving control data by moving or stopping along the corresponding driving trajectory.

[0060] It should be noted that the intelligent model can include both the end-to-end model deployed on the target vehicle and the VLM deployed on the target vehicle. Based on this, the end-to-end model can process sensor data and other data, such as map data, to obtain corresponding autonomous driving inference data. Simultaneously, the VLM can process sensor data and other data, such as map data and navigation data, to obtain corresponding autonomous driving inference data. Based on this, the end-to-end model, according to its own autonomous driving inference data or further combined with the VLM's autonomous driving inference data, performs trajectory filtering, trajectory optimization, and other processing according to relevant rules such as traffic laws and safety rules to generate driving control data. This data may include the driving trajectory, control commands for moving or stopping, and driving parameters such as speed along the driving trajectory. Based on this, the target vehicle responds to the driving control data by moving or stopping along the driving trajectory according to the corresponding driving parameters.

[0061] In a specific implementation, this embodiment can read autonomous driving inference data obtained by intelligent models such as end-to-end models and / or VLMs in the data storage area corresponding to the intelligent model.

[0062] For example, the end-to-end model processes multimedia data, radar scan data, positioning coordinate data, and map data collected by cameras into bird's-eye view (BEV) feature stitching, attention recognition, and trajectory planning to obtain autonomous driving inference data, such as predicted and planned trajectory lines, identified obstacles, and corresponding attention scores. This embodiment acquires this autonomous driving inference data.

[0063] For example, after using the camera, the VLM performs scene understanding, description, and decision-making processes based on the multimedia data, radar scan data, positioning coordinate data, map data, and navigation data collected by the camera to obtain autonomous driving inference data, such as scene description text and decision description text of the target vehicle's environment. This embodiment acquires this autonomous driving inference data.

[0064] Step 102: Present the autonomous driving inference data on the target vehicle's display screen.

[0065] In one implementation, the display screen outputs a reasoning interface used to present autonomous driving reasoning data. Furthermore, the display screen also outputs an environmental perception interface used to present environmental information and navigation information about the target vehicle's location.

[0066] It should be noted that the target vehicle's displays include a first display and a second display. The first display is positioned to match the driver's seat of the target vehicle; for example, it could be the display corresponding to the driver's seat, also known as the central control screen. The second display is positioned to match the passenger seat of the target vehicle; for example, it could be the passenger-side screen corresponding to the passenger seat. The inference interface and the environmental perception interface can be output to the first display. The environmental perception interface can be called the Environmental Intelligence Detection (EID) interface. The environmental information in the environmental perception interface can be obtained by the end-to-end model through BEV feature stitching based on sensor data. Navigation information can be determined based on the target vehicle's starting point and destination. Based on this, the end-to-end model fuses the environmental information and navigation information to obtain the environmental perception interface.

[0067] For example, such as Figure 3As shown, the environmental perception interface is output in the left area of ​​the first display screen. This interface can present the target vehicle's environment with other vehicles, pedestrians, roadside obstacles, the target vehicle's current speed and speed limit information at its current location, as well as lane information and navigation information. The inference interface is output in the right area of ​​the first display screen.

[0068] It should be noted that when there is only one type of autonomous driving inference data, the inference interface is filled with autonomous driving inference data of that type; when there are multiple types of autonomous driving inference data, the inference interface can be divided into multiple inference areas, each used to present the corresponding type of autonomous driving inference data, such as... Figure 4 As shown in the image.

[0069] In one implementation, the inference interface in this embodiment can be output in response to a trigger control (which can be deployed on the environment-aware interface or the navigation interface) corresponding to the inference interface on the first display screen. For example, Figure 5 As shown, a trigger control such as "AI Inference" is deployed on the first display screen. Users can click on the trigger control. In this embodiment, in response to the user's click on the trigger control, the inference interface is output on the first display screen. The inference interface displays autonomous driving inference data. At this time, the original environmental perception interface on the first display screen switches from full-screen output to left-side area output. Correspondingly, the inference interface occupies the right-side area output.

[0070] Furthermore, in this embodiment, the inference interface can be closed in response to an exit operation on the first display screen. This exit operation can be a custom operation, such as a three-finger swipe down on the inference interface. For example, ... Figure 6 As shown, the first display screen outputs an environmental awareness interface and an inference interface. When the user performs a three-finger swipe-down exit operation on the display area where the inference interface is located, the inference interface is closed in response to the user's three-finger swipe-down exit operation, and the environmental awareness interface is displayed in full screen on the first display screen.

[0071] By means of the above technical solution, in the method for processing the interactive interface in autonomous driving provided in the embodiments of this application, the autonomous driving inference data determined by the sensing data collected by the sensing device is presented on the display screen of the target vehicle. In this way, not only can other information in the driving process of the target vehicle be displayed, but also the autonomous driving inference data that realizes the driving control of the vehicle can be presented, so that the user can intuitively understand the inference process in the autonomous driving process through these autonomous driving inference data, thereby improving the user's autonomous driving experience.

[0072] In one implementation, taking an intelligent model that includes at least one of an end-to-end model and a visual language model as an example, the autonomous driving reasoning data obtained in this embodiment may include at least one of the following:

[0073] The trajectory prediction data for the target vehicle is obtained by processing sensor data. Specifically, this can be achieved through an end-to-end model. For example, the trajectory prediction data may include predictions for multiple trajectory lines planned by the end-to-end model. Therefore, presenting the trajectory prediction data in the inference interface allows users to understand the processing operations performed by the end-to-end model to plan the optimal route for them.

[0074] Obstacle recognition data is obtained by processing sensor data. Specifically, obstacle recognition data can be obtained through an end-to-end model that processes sensor data. For example, obstacle recognition data can be understood as the data obtained by the end-to-end model in identifying obstacles such as other vehicles, pedestrians, and curbs in the target vehicle's environment. Therefore, presenting obstacle recognition data in the inference interface allows users to understand the obstacles identified by the end-to-end model to prevent the target vehicle from being endangered, thus enabling users to understand the processing operations performed by the end-to-end model to achieve safe driving.

[0075] The process involves processing the scene description data of the target vehicle's environment obtained from sensor data. Specifically, this scene description data can be processed using a visual language model. For example, scene description data can be understood as the scene text in which the VLM (Visual Language Model) understands and describes the target vehicle's environment in natural language. Therefore, presenting this scene description data in the inference interface allows users to understand the VLM's understanding of the current driving scenario.

[0076] Scene decision data is derived from scene description data. Specifically, it can be obtained through a visual language model based on scene description data. For example, scene decision data can be understood as the decision text obtained by the Visual Language Model (VLM) from analyzing the environment in which the target vehicle is located and making decisions. Therefore, presenting this scene decision data in the inference interface allows users to understand the decisions made by the VLM based on scene understanding.

[0077] The question-and-answer data is generated from question-and-answer operations based on sensor data. Specifically, this can be achieved through question-and-answer data generated by a visual language model (VLM) based on sensor data. For example, to obtain scene description data and scene decision data, the VLM can perform multiple question-and-answer operations. In this embodiment, the question-and-answer data generated by the VLM during these operations is acquired and used as inference data for autonomous driving. Presenting this question-and-answer data in the inference interface allows users to understand the processing operations performed by the VLM during scene understanding and decision-making.

[0078] For example, such as Figure 7 As shown, the inference interface is divided into three inference areas: the inference area in the upper left corner is used to present trajectory prediction data; the inference area in the lower left corner is used to present obstacle recognition data; and the inference area on the right is used to present scene description data and scene decision data.

[0079] As can be seen, this embodiment can present various types or contents of autonomous driving reasoning data on the reasoning interface, enabling users to understand a richer and more detailed reasoning process of the intelligent model, thereby improving the user's experience of autonomous driving.

[0080] In one implementation, the trajectory prediction data may include: at least one predicted trajectory line; each predicted trajectory line corresponds to a prediction probability value, and the prediction probability value corresponding to at least one predicted trajectory line can characterize the confidence level of the corresponding predicted trajectory line, that is, the confidence level generated by the end-to-end model in planning the predicted trajectory line, indicating that the predicted trajectory line is the optimal trajectory. The predicted trajectory line starts from the vehicle position of the target vehicle.

[0081] It should be noted that each predicted trajectory line is presented as a trajectory image on the display screen, as shown in the inference interface of the first display screen. The trajectory image corresponding to each predicted trajectory line occupies a portion of the image area in the inference interface, and the trajectory images are arranged in rows and columns within the inference interface. For example, as shown... Figure 8 As shown, there are 10 predicted trajectory lines and 10 corresponding trajectory images, which are arranged in two rows and five columns in the inference interface.

[0082] Among them, the trajectory image corresponding to the predicted trajectory line with the highest predicted probability value has a target identifier, which can be at least one of color identifiers, line identifiers, and character identifiers. For example, Figure 8 As shown, the target is identified by the outline of the trajectory image, which is highlighted in a color such as blue or yellow to indicate to the user that the predicted trajectory in the trajectory image is the optimal trajectory; the prediction probability value is presented as a percentage.

[0083] Furthermore, the trajectory image also includes at least vehicle identifiers corresponding to other vehicles in the target vehicle's environment. Additionally, it may include obstacle identifiers such as pedestrians or curbs, and lane markings such as lane lines. Specifically, in this embodiment, a trajectory image can be generated based on the predicted trajectory line and environmental perception data. The end-to-end model can generate environmental perception data based on sensor data, which includes data on obstacles and lane lines in the target vehicle's environment. Therefore, the trajectory image generated in this embodiment may include obstacle graphics and lane lines in the target vehicle's environment. For example, as... Figure 8 As shown, the trajectory image includes not only the vehicle identifier of the vehicle (i.e. the target vehicle) and the predicted trajectory line starting from the position of the vehicle, but also obstacle identifiers such as other vehicles, pedestrians, lanes, and curbs. In addition, there can be multiple lane lines in the trajectory image.

[0084] The obstacle marker can be a graphic symbol composed of preset outlines corresponding to the obstacle. For example, pedestrians, other vehicles, and curbs are each pre-set with corresponding graphic symbols, which represent the outline or shape of the obstacle. The vehicle marker for the target vehicle can be a graphic symbol composed of the outlines of the target vehicle. For example, ... Figure 8 As shown, the vehicle identifiers for both the vehicle itself and other vehicles are graphic representations that match the outline of the vehicle type (such as a truck, car, van, etc.).

[0085] As can be seen, in this embodiment, at least two predicted trajectory lines are presented in the inference interface, and the predicted trajectory line with the highest predicted probability value is marked with a conspicuous target identifier. This allows users to more intuitively perceive the different trajectory lines planned by the end-to-end model, thereby letting users understand the processing operations performed by the end-to-end model to provide users with a better driving experience.

[0086] It should be noted that in this embodiment, the trajectory image can be updated in real time in the inference interface according to the target frame rate, such as 10 to 20 frames per second.

[0087] Based on the above implementation, the trajectory image can display the predicted trajectory line, other vehicles, and lanes from a top-down perspective based on the target vehicle. Furthermore, in this embodiment, to improve the viewing experience of the trajectory image, the predicted trajectory line in the trajectory image can be optimized, as follows:

[0088] In one implementation, the predicted trajectory line can be smoothed to make it conform to the natural laws of driving trajectory.

[0089] In another implementation, the shape of at least one predicted trajectory line can be adjusted in this embodiment so that the similarity of the line shapes between different predicted trajectory lines is less than or equal to a similarity threshold. The similarity threshold can be set according to business needs. Based on the setting of the similarity threshold, the shape differences between the predicted trajectory lines presented in the inference interface in this embodiment are relatively large, so that users can more intuitively feel the different trajectory lines planned by the end-to-end model.

[0090] Specifically, in this embodiment, N predicted trajectory lines can be selected from the multiple predicted trajectory lines initially planned by the end-to-end model. For example, the N predicted trajectory lines with predicted probability values ​​sorted from largest to smallest can be selected. Then, for each selected predicted trajectory line, the trajectory line can be changed according to the left or right change direction based on at least one change coefficient. This can result in multiple new predicted trajectory lines, thereby improving the difference between the predicted trajectory lines and other trajectory lines.

[0091] The variation coefficient represents the degree of lane change of the trajectory line. The larger the variation coefficient, the greater the magnitude of the lane change to the left or right of the predicted trajectory line. Correspondingly, the difference between the new predicted trajectory line and other trajectories is greater. Therefore, presenting these significantly different predicted trajectories in the inference interface allows users to more intuitively perceive the different trajectories planned by the end-to-end model, thereby enabling users to understand the work done by the end-to-end model to provide a better driving experience.

[0092] In another implementation, the length of at least one predicted trajectory line can be adjusted in this embodiment to make the line lengths different between different predicted trajectory lines. The line length can represent the predicted driving distance of the target vehicle. Based on this, this embodiment presents the line lengths of the predicted trajectory lines according to their respective prediction probability values. For example, the higher the prediction probability value, the longer the line length of the predicted trajectory line, and correspondingly, the longer the driving prediction distance planned by the end-to-end model presented to the user through the inference interface.

[0093] Where N is a positive integer greater than or equal to 1, such as 5 or 6. For example, in this embodiment, the predicted trajectory lines with the last 7 predicted probability values ​​are selected from the 10 predicted trajectory lines. For each selected trajectory line, the first 3 are processed according to the left lane change direction, based on three different change coefficients, thus obtaining 3 new predicted trajectory lines after the left lane change; the last 3 are processed according to the right lane change direction, based on three different change coefficients, thus obtaining 3 new predicted trajectory lines after the right lane change; the length of the middle predicted trajectory line is adjusted according to the magnitude of the predicted probability value. Thus, in this embodiment, these 7 predicted trajectory lines are varied to different degrees and dimensions, resulting in significant differences between the predicted trajectory lines presented on the inference interface, allowing users to more intuitively perceive the different trajectory lines planned by the end-to-end model.

[0094] Based on the above implementation, in this embodiment, the presentation position of the trajectory image with the target identifier on the display screen changes in response to the fulfillment of switching conditions. Specifically, in this embodiment, for the trajectory image in the inference interface, the presentation position of the trajectory image with the target identifier in the inference interface can be adjusted in response to the fulfillment of switching conditions.

[0095] The switching condition can be: the time interval since the last adjustment of the presentation position of the trajectory image with the target identifier reaches the target time.

[0096] Alternatively, the switching condition can be: the duration of the trajectory image with the target identifier at its current display position reaches the target duration. The target duration can be set according to business needs, such as 4 seconds or 6 seconds.

[0097] In one implementation, this embodiment allows the selection of M target locations from the candidate locations used to output the trajectory image in the inference interface. These target locations have a switching order. Based on this, in this embodiment, each time a switching condition is detected to be met, the identified trajectory image can be switched from its current presentation position in the inference interface to the next target position. That is, the identified trajectory images are presented sequentially at the corresponding target positions in the inference interface according to the switching order between the target positions.

[0098] M can be a positive integer greater than or equal to 2. The value of M can be set according to business requirements.

[0099] In another implementation, this embodiment can randomly select a target location from other candidate locations for outputting the trajectory image in the inference interface each time the switching condition is detected to be met, and then switch the trajectory image with the identifier from the current display position in the inference interface to the randomly selected target location.

[0100] For example, such as Figure 9 As shown, the inference interface has two rows of five columns of candidate positions, such as position 1 to position 10. These candidate positions are used to present the trajectory image. In this embodiment, the timing starts from the trajectory image of the blue box (i.e., the trajectory image of the predicted trajectory line with the highest predicted probability value) at the current position, such as position 2. When the timer reaches 5 seconds (i.e., the target duration), the trajectory image of the blue box is switched from position 2 to a randomly selected target position or the next target position, such as position 5, and the timing restarts. When the timer reaches 5 seconds, the trajectory image with the blue box (which may be the trajectory image at position 5, or a trajectory image at another position marked with a blue box because the predicted probability value has changed to the maximum) is switched from the current position, such as position 5 (or another position), to a randomly selected target position or the next target position, such as position 8, and the timing restarts. When the timer reaches 5 seconds, the trajectory image of the blue box (which may be the trajectory image at position 8, or a trajectory image at another position marked with a blue box because the predicted probability value has changed to the maximum) is switched from the current position, such as position 8 (or another position), to a randomly selected target position or the next target position, such as position 2, and so on. Thus, the trajectory image of the blue box is constantly switched at different positions, demonstrating that the end-to-end model plans the optimal trajectory line in real time.

[0101] As can be seen, by switching the position of the trajectory image with the highest predicted probability value on the inference interface in this embodiment, users can experience the end-to-end model continuously performing real-time calculations and continuously obtaining the optimal trajectory. This allows users to feel that the end-to-end model is running in real time, thereby improving the user's experience with autonomous driving.

[0102] In one implementation, obstacle recognition data may include attention scores corresponding to obstacles in the environment where the target vehicle is located. The attention scores are output as an attention layer in the driving scene image displayed on the screen. The driving scene image is an image obtained by capturing images of the target vehicle's environment using an image acquisition device. For example, the driving scene image may be an image captured by a 120-degree forward-facing camera, or it may be an image stitched together from images captured by multiple cameras.

[0103] The hue distribution of the attention layer is matched with the attention score. Each obstacle in the environment where the target vehicle is located corresponds to an image region in the driving scene image; the attention layer is overlaid on the image region corresponding to the corresponding obstacle.

[0104] In this implementation, the attention score is pre-divided into corresponding score levels, and a hue parameter is set for each score range. For example, the attention score can be divided into 5 score levels, and these 5 score levels are each assigned a different hue parameter according to the level of the highest level (i.e., the size of the attention score). Specifically, the score levels from low to high correspond to: blue, green, yellow, orange, and red. Thus, the higher the attention score, the lower the hue value of the attention layer, and the more eye-catching the attention layer appears to the user with the corresponding color.

[0105] Based on the above implementation, in this embodiment, when rendering the attention layer, the entire attention layer can be rendered using the hue parameters corresponding to the attention score. For example, as... Figure 10 As shown in the image, in the driving scene, the vehicle in front has the highest attention score, and a red attention layer is superimposed and rendered on the image area where the vehicle in front is located. The attention scores for the pedestrian on the left front and the guardrail on the right front are lower. Based on this, a green attention layer is superimposed and rendered on the image area where the pedestrian on the left front is located, and a green attention layer is superimposed and rendered on the image area where the guardrail on the right front is located.

[0106] Alternatively, in this embodiment, when rendering the attention layer, the central area of ​​the attention layer (i.e., the central area of ​​the obstacle) can be rendered as the hue parameter corresponding to the attention score. Based on this, according to the hue gradient rule, along the direction from the central area to the edge area, a hue is changed every target distance, thereby making the attention layer form a hue gradient effect from the center to the edge.

[0107] The hue gradient rule can be a gradient from red, orange, yellow, green to blue. For example, as... Figure 11 As shown in the image, in the driving scene, the vehicle in front has the highest attention score. An attention layer centered on red and gradually turning green towards the edge is overlaid on the image area where the vehicle in front is located. The pedestrian on the left front and the railing on the right front have lower attention scores. Based on this, an attention layer centered on green and gradually turning blue towards the edge is overlaid on the image area where the pedestrian on the left front is located, and an attention layer centered on green and gradually turning blue towards the edge is overlaid on the image area where the railing on the right front is located.

[0108] As can be seen, in this embodiment, different hues can be used to render corresponding layers for the user in the image area where the obstacle is located in the driving scene image, so as to distinguish the degree of attention required for the obstacle identified by the end-to-end model. This allows the user to more intuitively understand which obstacles the end-to-end model pays more attention to in order to avoid obstacles and plan the trajectory, thereby improving the user's experience of autonomous driving.

[0109] It should be noted that, in this embodiment, the area of ​​the attention layer rendered in the driving scene image can be less than or equal to the area of ​​the obstacle's image region in the driving scene image, thus ensuring that the attention layer does not completely obscure the obstacle. Furthermore, the shape of the attention layer rendered in this embodiment can match the outline of the obstacle in the driving scene image. Therefore, this embodiment can present the degree to which the end-to-end model notices the obstacle using different hues while not completely obscuring it, thereby further improving the user experience of autonomous driving.

[0110] In this embodiment, the attention score obtained is based on the vehicle coordinate system corresponding to the end-to-end model. Based on this, before rendering the attention layer, the obstacle coordinates corresponding to the attention score can be converted to the camera coordinate system corresponding to the driving scene image. Then, according to the converted obstacle coordinates, the corresponding attention layer is rendered on the image area where the obstacle is located in the driving scene image.

[0111] In one implementation, scene description data and scene decision data are presented on the display screen in the form of target text, which includes scene text corresponding to the scene description data and decision text corresponding to the scene decision data.

[0112] Specifically, scene text can be generated according to a scene description template, and decision text can also be generated according to a corresponding decision description template. These scene and decision description templates can be obtained from third-party applications or the cloud. These templates are more aligned with human descriptive habits. Based on this, the target text can be obtained by concatenating the scene and decision texts.

[0113] The display screen also shows the driving scene image corresponding to the target text. This can be understood as follows: the target text is the text generated by the visual language model through scene understanding and decision-making based on the driving scene image. The driving scene image can be an image obtained by using an image acquisition device to capture images of the environment in which the target vehicle is located.

[0114] For example, the driving scene image can be an image captured by a 30-degree forward-facing camera.

[0115] In the specific implementation, when there are multiple driving scene images, there are also multiple target texts, and the driving scene images and target texts have a one-to-one mapping relationship. In this case, the driving scene images and their corresponding target texts can be displayed on the screen in chronological order of their acquisition time. For example, the later the acquisition time (i.e., the closer to the current time), the closer the driving scene image and the corresponding target text will be to the top or bottom of the corresponding inference area in the inference interface of the display screen; conversely, the earlier the acquisition time (i.e., the further away from the current time), the further away the driving scene image and the corresponding target text will be from the top or bottom of the corresponding inference area in the inference interface of the display screen.

[0116] It should be noted that in this embodiment, the driving scene image can be regarded as the question content for the visual language model, and the target text is the answer content of the visual language model. Based on this, a one-to-one image-text question and answer can be formed between the driving scene image and the target text. In this embodiment, the driving scene image and its corresponding target text can be presented on the display screen in the form of image-text question and answer. At this time, the driving scene image is the question in the image-text question and answer, and the target text is the answer in the image-text question and answer. Based on the different acquisition time of the driving scene image, the image-text question and answer has a corresponding timestamp. In the corresponding reasoning area of ​​the reasoning interface, the corresponding image-text question and answer can be presented in order from early to late according to the timestamp. Thus, users can view more image-text question and answer during the reasoning process of the visual language model as needed, thereby improving the user experience of autonomous driving.

[0117] For example, such as Figure 12 As shown, in the inference interface, the driving scene image is output as the question content first, and the corresponding target text is output as the VLM's answer content later. When the VLM processes multiple driving scene images, the driving scene images and their corresponding target texts are displayed in a combined manner, scrolling from top to bottom (or bottom to top) in the inference interface. That is, a question and answer is presented as a group, with the driving scene image and its corresponding target text collected later appearing above the driving scene image and its corresponding target text collected earlier. If the user wants to view the driving scene image and its corresponding target text collected earlier, they can swipe up on the inference interface to view the historical question and answer.

[0118] As can be seen, in this embodiment, the scene description data and scene decision data obtained by the VLM in real time can be presented to the user through the inference interface. This allows the user to experience how the VLM continuously understands and makes decisions about the real-time changing driving environment, and then continuously updates the scene description data and scene decision data to the end-to-end model. The end-to-end model then continuously combines the latest scene description data and scene decision data to generate more accurate driving control data. This allows the user to feel that the visual language model and the end-to-end model are performing corresponding processing for autonomous driving in real time, thereby improving the user's experience with autonomous driving.

[0119] In one implementation, the target vehicle further includes a second display screen positioned to match the passenger side of the target vehicle, while the first display screen is positioned to match the driver's side of the target vehicle. For example, as... Figure 13 As shown, the first display screen can be the central control screen corresponding to the driver's position, and the second display screen can be the passenger screen corresponding to the passenger's position.

[0120] Based on this, in this embodiment, after the autonomous driving inference data is presented on the first display screen, in response to a push command for the autonomous driving inference data, the autonomous driving inference data can also be presented on the second display screen.

[0121] The first display screen can output push controls, such as... Figure 13 As shown in the figure, "Push to passenger screen" means that after the user clicks the push control, in this embodiment, a push command is generated in response to the user's click operation. Then, in response to the push command, the autonomous driving inference data can be presented to the inference interface output by the second display screen in the passenger seat.

[0122] It should be noted that, in this embodiment, after the autonomous driving inference data is presented to the inference interface output by the second display screen, the inference interface output by the first display screen can be turned off. The second display screen may or may not output the environmental perception interface.

[0123] For example, such as Figure 13As shown, the target vehicle has a central control screen in the driver's seat and a passenger screen in the passenger's seat. Initially, the central control screen displays the environmental perception interface, while the passenger screen displays the entertainment interface. After the user clicks a trigger control on the central control screen, it can display both the environmental perception interface and an inference interface. This provides the user with environmental and navigation information, as well as inference data for autonomous driving, which is necessary for the intelligent model to achieve autonomous driving. Later, after the user clicks a push control on the central control screen, the inference interface can be deactivated, and instead, it can be displayed on the passenger screen (the entertainment interface exits or is minimized). Furthermore, the passenger screen can display either the environmental perception interface or not while displaying the inference interface.

[0124] Additionally, an exit control can be displayed on both the first and second displays. Figure 13 The "Exit Inference Mode" control shown in the image closes the inference interface regardless of whether it is displayed on the first or second screen.

[0125] Furthermore, when the inference interface is output on the second display screen, the second display screen can have push-back controls, such as... Figure 13 The "Project to Central Control Screen" control shown in the image closes the inference interface displayed on the second screen and displays the inference interface on the first screen when the user clicks the push-back control. At this time, the autonomous driving inference data is presented in the inference interface displayed on the first screen.

[0126] Therefore, the autonomous driving inference data in this embodiment can not only be presented to the user in the driver's seat, but also pushed to the passenger seat for the passenger to view, thereby improving the user's experience of autonomous driving.

[0127] In one implementation, question-and-answer data can be presented in the reasoning interface as a combination of text. This data includes the question content corresponding to the visual language model and the answer content corresponding to the visual language model. The question and answer content form a text combination and are displayed on the reasoning interface of the screen.

[0128] For example, such as Figure 14As shown, the question can be a prompt word input into the VLM, and the corresponding answer can be text generated by the VLM in response to the prompt. For example, if the first prompt is "Continue driving forward, what should I pay attention to?", the first VLM response text is "Ahead is a T-junction, the lighting is dim, there are many road users, and cyclists may cross from the side." If the second prompt is "Then what should I do?", the second VLM response text is "It is recommended that you slow down to observe road users on both sides and slow down in time to give way." Furthermore, the inference interface can also present corresponding decision data based on these question-and-answer data, such as "Turn right ahead, wait at the right-turn red light, and then slow down."

[0129] The above describes a method for processing an interactive interface in autonomous driving provided by an embodiment of this application. The following will describe the apparatus for executing the above-described method for processing an interactive interface in autonomous driving.

[0130] refer to Figure 15 This is a schematic diagram of the structure of a processing device for an interactive interface in autonomous driving, provided in an embodiment of this application. This device can be configured in an in-vehicle device with a processor-based intelligent driving system. Figure 2 As shown, the target vehicle is equipped with at least one sensing device and a processor, such as a vehicle-mounted chip. Based on the processor on the target vehicle, an intelligent driving system can be deployed. The intelligent driving system can include one or more intelligent models, such as an end-to-end model and a Vehicle Modeling Library (VLM). The technical solution in this embodiment is mainly used to improve the user experience of intelligent driving.

[0131] Specifically, the apparatus in this embodiment may include the following units:

[0132] The data determination unit 1501 is used to determine autonomous driving reasoning data based on sensing data collected by at least one sensing device deployed on the target vehicle.

[0133] The reasoning presentation unit 1502 is used to present the autonomous driving reasoning data on the display screen of the target vehicle;

[0134] The display screen can also output an environmental perception interface, which is used to present environmental information and navigation information of the target vehicle's location.

[0135] As can be seen from the above technical solution, in the processing device for the interactive interface in autonomous driving provided in this application embodiment, the autonomous driving inference data determined by the sensing data collected by the sensing device is presented on the display screen of the target vehicle. In this way, during the autonomous driving process, not only can other information of the target vehicle driving process be displayed, but also the autonomous driving inference data that realizes vehicle driving control can be presented, so that users can intuitively understand the inference process in the autonomous driving process through these autonomous driving inference data, thereby improving the user's autonomous driving experience.

[0136] In one implementation, the intelligent model includes at least one of an end-to-end model and a visual language model;

[0137] The autonomous driving inference data includes at least one of the following:

[0138] The trajectory prediction data of the target vehicle obtained by processing the sensor data;

[0139] Obstacle recognition data obtained by processing the aforementioned sensor data;

[0140] Scene description data of the environment in which the target vehicle is located, obtained by processing the sensor data;

[0141] Scene decision data obtained based on the scene description data;

[0142] Question and answer data in the question and answer operation based on the sensor data.

[0143] In one implementation, the trajectory prediction data includes: at least one predicted trajectory line; the predicted trajectory line corresponds to a predicted probability value; at least one of the predicted trajectory lines is presented on the display screen in the form of a trajectory image; wherein, the trajectory image corresponding to the predicted trajectory line with the largest predicted probability value has a target identifier, and the trajectory image also includes at least vehicle identifiers corresponding to other vehicles in the environment where the target vehicle is located.

[0144] In one implementation, the presentation position of the trajectory image with the target identifier on the display screen changes in response to a satisfied switching condition. Specifically, the inference presentation unit 1502 is further configured to: adjust the presentation position of the trajectory image with the target identifier on the inference interface in response to a satisfied switching condition.

[0145] In one implementation, the obstacle recognition data includes: attention scores corresponding to obstacles in the environment where the target vehicle is located; the attention scores are output in the form of an attention layer in the driving scene image presented on the display screen; the driving scene image is an image obtained by using an image acquisition device to acquire images of the environment where the target vehicle is located; wherein the hue distribution of the attention layer matches the attention scores.

[0146] In one implementation, the scene description data and the scene decision data are presented on the display screen in the form of target text, and the display screen also displays a driving scene image corresponding to the target text; the driving scene image is an image obtained by using an image acquisition device to capture images of the environment where the target vehicle is located; wherein, when there are multiple driving scene images, the display screen presents the driving scene images and their corresponding target text in the order of their acquisition time.

[0147] The driving scene image and its corresponding target text are presented on the display screen in the form of a question and answer; wherein the driving scene image is the question in the question and answer, and the target text is the answer.

[0148] In one implementation, the display screen includes a first display screen and a second display screen, wherein the deployment position of the second display screen matches the passenger seat position of the target vehicle, and the deployment position of the first display screen matches the driver seat position of the target vehicle; wherein, the inference presentation unit 1502 is further configured to: after presenting the autonomous driving inference data on the first display screen, in response to a push command for the autonomous driving inference data, present the autonomous driving inference data to the inference interface output on the second display screen.

[0149] It should be noted that the specific implementation of each unit in this embodiment can be referred to the corresponding content above, and will not be described in detail here.

[0150] refer to Figure 16 This is a schematic diagram of the structure of a vehicle-mounted device provided in an embodiment of this application. The vehicle-mounted device can be deployed on a target vehicle. Specifically, the vehicle-mounted device may include the following structure:

[0151] At least one sensing device 1601, the sensing device 1601 being used to collect sensing data;

[0152] At least one processor 1602 is configured to determine autonomous driving inference data based on the sensing data and to present the autonomous driving inference data on a display screen of the target vehicle.

[0153] In addition, the in-vehicle equipment may also include memory, input / output (I / O) interfaces, and a bus. The memory stores the data required for the operation of the intelligent driving system and the visual language model, as well as the data generated by their respective operations. The processor 1502 and the memory can be connected to each other via the bus. The input / output (I / O) interfaces are also connected to the bus.

[0154] The in-vehicle equipment may also include the following devices, all of which can be connected to the input / output interface:

[0155] This includes input devices such as touchscreens, touchpads, cameras, microphones, accelerometers, and gyroscopes; output devices such as liquid crystal displays (LCDs), speakers, and vibrators; storage devices such as memory cards and hard drives; and communication devices. The communication devices allow the in-vehicle equipment to communicate wirelessly or wiredly with other devices to exchange data.

[0156] Although Figure 16 The vehicle-mounted equipment shown includes various devices; however, it should be understood that implementation or possession of all shown devices is not required. More or fewer devices may be implemented or possessed alternatively. Figure 16 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0157] By means of the above technical solution, the in-vehicle device provided in this application displays autonomous driving inference data determined by sensing data collected by sensing devices on the display screen of the target vehicle. In this way, during the autonomous driving process, not only can other information of the target vehicle driving process be displayed, but also the autonomous driving inference data that realizes vehicle driving control can be presented, so that users can intuitively understand the inference process in the autonomous driving process through these autonomous driving inference data, thereby improving the user's autonomous driving experience.

[0158] This application also provides a computer program product including computer-readable instructions, which, when executed on an in-vehicle device, cause the electronic device to implement any of the interactive interface processing methods provided in this application embodiment for autonomous driving.

[0159] This application also provides a computer storage medium that carries one or more computer programs. When the one or more computer programs are executed by an in-vehicle device, the in-vehicle device can implement any of the interactive interface processing methods provided in this application embodiment for autonomous driving.

[0160] In specific implementation, such as Figure 17The diagram shown illustrates the architecture of the intelligent driving system in this embodiment for presenting autonomous driving inference data. The intelligent driving system in the target vehicle may include the following functional modules: a fully-self-driving (FSD) functional module (such as an end-to-end model and VLM) and an in-vehicle infotainment system (HU). The FSD deploys an interface connected to the HU, namely a human user interface (HUI). This embodiment primarily includes the presentation of autonomous driving inference data across three dimensions, as follows:

[0161] The end-to-end model presents the trajectory lines: FSD selects 10 predicted trajectory lines from the generated trajectory lines, and HUI processes these 10 predicted trajectory lines, such as smoothing and lane changing according to left and right lanes. Then, HU renders the inference interface on the first or second display screen, which contains the trajectory images corresponding to these predicted trajectory lines.

[0162] VLM presents scene descriptions and scene decisions: VLM provides scene text and decision text to HUI based on sensor data (System 1 data) or combined with other data (System 2 data, such as scene description templates and decision description templates). HUI concatenates the scene text and decision text and presents the resulting target text to HU. HU renders the inference interface on the first or second display screen, which contains these target texts.

[0163] Presentation of attention layers in the physical world: FSD provides attention scores to HUI, and HUI determines the corresponding hue, such as red or green, according to the attention scores. HU then renders the inference interface on the first or second display screen. In the inference interface, the attention layer corresponding to the attention score is superimposed on the image area where the obstacle is located.

[0164] As can be seen, in this embodiment, the intelligent driving system can present users with various autonomous driving reasoning data involved in autonomous driving, thereby improving the user's experience with autonomous driving.

[0165] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0166] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0167] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0168] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A method for processing the interactive interface in autonomous driving, characterized in that, The method includes: Autonomous driving inference data is determined based on sensor data collected by at least one sensor device deployed on the target vehicle. The autonomous driving inference data is presented on the display screen of the target vehicle.

2. The method according to claim 1, characterized in that, The autonomous driving inference data includes at least one of the following: The trajectory prediction data of the target vehicle obtained by processing the sensor data; Obstacle recognition data obtained by processing the aforementioned sensor data; Scene description data of the environment in which the target vehicle is located, obtained by processing the sensor data; Scene decision data obtained based on the scene description data; Question and answer data in the question and answer operation based on the sensor data.

3. The method according to claim 2, characterized in that, The trajectory prediction data includes: at least one predicted trajectory line; each predicted trajectory line corresponds to a predicted probability value; and at least one predicted trajectory line is presented on the display screen in the form of a trajectory image. The trajectory image corresponding to the predicted trajectory line with the highest predicted probability value has a target identifier, and the trajectory image also contains at least the vehicle identifiers of other vehicles in the environment where the target vehicle is located.

4. The method according to claim 3, characterized in that, The position of the trajectory image with the target identifier displayed on the screen changes in response to a change in the switching condition that is met.

5. The method according to claim 2, 3 or 4, characterized in that, The obstacle recognition data includes: attention scores corresponding to obstacles in the environment where the target vehicle is located; the attention scores are output in the form of attention layers in the driving scene image presented on the display screen; the driving scene image is an image obtained by using an image acquisition device to acquire images of the environment where the target vehicle is located; The hue distribution of the attention layer is matched with the attention score.

6. The method according to claim 2, 3 or 4, characterized in that, The scene description data and the scene decision data are presented on the display screen in the form of target text, and the display screen also displays the driving scene image corresponding to the target text; the driving scene image is an image obtained by using an image acquisition device to capture images of the environment where the target vehicle is located; In the case of multiple driving scene images, the display screen presents the driving scene images and their corresponding target text in the order of their acquisition time.

7. The method according to claim 6, characterized in that, The driving scene image and its corresponding target text are presented on the display screen in the form of a question and answer with images. The driving scene image is the question in the image-text Q&A, and the target text is the answer in the image-text Q&A.

8. The method according to claim 1, 2, 3, 4, 5, 6 or 7, characterized in that, The display screen includes a first display screen and a second display screen, the second display screen is deployed at a position matching the passenger seat of the target vehicle, and the first display screen is deployed at a position matching the driver seat of the target vehicle. The method further includes: After the autonomous driving inference data is presented on the first display screen, in response to a push command for the autonomous driving inference data, the autonomous driving inference data is presented on the second display screen.

9. A processing device for an interactive interface in autonomous driving, characterized in that, The device includes: The data determination unit is used to determine autonomous driving reasoning data based on sensing data collected by at least one sensing device deployed on the target vehicle. The reasoning presentation unit is used to present the autonomous driving reasoning data on the display screen of the target vehicle.

10. A vehicle-mounted device, characterized in that, include: At least one sensing device, said sensing device being used to collect sensing data; At least one processor is configured to determine autonomous driving inference data based on the sensing data and to present the autonomous driving inference data on a display screen of the target vehicle.