Interaction method, apparatus and mobile device
By acquiring and processing images of target objects in the driver assistance system, generating focused visual images and displaying target prompts, the problem of users not being able to understand the motivations behind vehicle behavior is solved, thus improving the human-machine interaction experience and collaborative efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU XIAOPENG CONNECTIVITY TECH CO LTD
- Filing Date
- 2026-04-17
- Publication Date
- 2026-06-26
AI Technical Summary
Existing driver assistance systems are unable to effectively identify obstacles and traffic signs and intuitively display the basis for decision-making in complex road conditions, resulting in users being unable to understand the motivations behind vehicle behavior, poor human-machine interaction experience, and low collaborative efficiency.
By acquiring the first environmental image containing the target object, only visual information directly related to vehicle control behavior is extracted, and target prompt information is displayed on the target page of the display. Redundant information in the full-scene environmental image is discarded, and image processing techniques such as image cropping, annotation and scaling are used to generate a focused visual image to ensure the high focus and semantic clarity of the information.
It enables the accurate transmission of vehicle decision-making basis within a limited display area, enhances the user's understanding of the motivation behind the driver assistance system's behavior, and improves the user's driving experience and interaction efficiency.
Smart Images

Figure CN122275588A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent driving technology, and more specifically, to an interaction method, device, and mobile device. Background Technology
[0002] With the development of intelligent driving technology, driver assistance systems (ADAS) need to identify obstacles and traffic signs in real time and make decisions in complex road conditions. While ADAS can recognize traffic signs through visual perception and optical character recognition (OCR), and identify obstacles using object detection algorithms, thus triggering vehicle control behaviors such as deceleration and avoidance, their human-machine interface only presents raw images or abstract symbols. Especially in scenarios with diverse traffic sign semantics, redundant recognition text, and opaque vehicle control intentions, users cannot understand the vehicle's behavioral motivations, resulting in a poor human-machine interaction experience and low collaborative efficiency.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides an interaction method, apparatus, and mobile device to at least solve the technical problem of poor human-computer interaction experience and low collaborative efficiency caused by users' inability to understand the motivation of vehicle behavior in the driver assistance system.
[0005] According to one aspect of the embodiments of this application, an interaction method is provided, comprising: acquiring a first environmental image containing a target object, wherein the target object is one of the factors affecting the driving state of a mobile device; and displaying target prompt information for the target object on a target page of a display corresponding to the mobile device, wherein the target prompt information is generated based on the first environmental image.
[0006] Optionally, acquiring a first environmental image containing the target object includes: acquiring a second environmental image using the image acquisition sensor of the mobile device, wherein the second environmental image is used to reflect the driving environment of the mobile device; determining whether the second environmental image contains the target object; if the second environmental image contains the target object, determining a first environmental image based on the second environmental image and the target object, wherein the first environmental image is an image region in the second environmental image that at least includes the target object.
[0007] Optionally, determining whether the second environment image contains a target object includes: performing object recognition on the second environment image to obtain at least one candidate object; and determining whether a target object exists among the at least one candidate object based on the category information and feature information corresponding to the at least one candidate object.
[0008] Optionally, a target object is determined to exist among at least one candidate object if at least one of the following conditions is met: if the category information of the candidate object is a static obstacle, and the feature information of the candidate object satisfies a first feature condition; wherein the first feature condition includes: the location of the object corresponding to the candidate object is within the path coverage area of the current driving path of the mobile device, and the size of the object corresponding to the candidate object is greater than a preset size threshold; or, if the category information of the candidate object is a dynamic obstacle, and the feature information of the candidate object satisfies a second feature condition; wherein the second feature condition includes one of the following: the object moving speed of the candidate object is greater than a preset speed threshold, and the distance between the candidate object and the mobile device decreases in the future time period; the predicted moving trajectory of the object corresponding to the candidate object intersects with the current driving path of the mobile device; or, if the category information of the candidate object is a traffic sign, and the feature information of the candidate object satisfies a third feature condition; wherein the third feature condition includes: the object content and object location corresponding to the candidate object are associated with the driving state of the mobile device.
[0009] Optionally, determining the first environment image based on the second environment image and the target object includes: performing an image processing operation on the second environment image based on the target object in the second environment image to obtain the first environment image, wherein the image processing operation includes at least one of the following: performing an image cropping operation on the target object in the second environment image, performing an image annotation operation on the target object in the second environment image, and performing image scaling processing on the target object in the second environment image.
[0010] Optionally, the target prompt information includes: a first environmental image. Displaying the target prompt information for the target object on the target page of the display corresponding to the mobile device includes: responding to the second environmental image including multiple target objects, determining the target object to be displayed based on the priority information of the multiple target objects, and displaying the first environmental image corresponding to the target object to be displayed on the target page, wherein the priority information is determined based on the degree of influence of the multiple target objects on the driving state of the mobile device, or the priority information is determined based on user-preset settings.
[0011] Optionally, the display is a head-up display or a mobile device screen; the target page is a scene-rendered map or a two-dimensional vector map.
[0012] Optionally, the target page may also include at least one of the following driving prompts: driving status information of the mobile device, decision prompts corresponding to the target object, and driving mode switching information. The driving status information is used to determine the activation status of the mobile device's assisted driving function, and the decision prompts corresponding to the target object include at least one of the following: environmental label information associated with the target object, driving decision information, and decision reasoning process information.
[0013] Optionally, at least one of the following methods may be used to display driving prompt information: displaying interface cards, displaying text pop-ups, or providing voice prompts; or highlighting or annotating the target object in the first environmental image.
[0014] Optionally, the method for obtaining the decision prompt information corresponding to the target object includes: in response to the target object being a traffic sign and the target object containing text content, clustering analysis is performed on the text content using the sign scene attribute information to obtain the decision prompt information corresponding to the target object. The sign scene attribute information is used to record the mapping relationship between multiple attribute fields, including: scene sign information, text sign type, sign priority, and prompt information corresponding to the target object.
[0015] Optionally, before performing cluster analysis on the text content using the identified scene attribute information, the method further includes: in response to the mobile device meeting preset state conditions, obtaining an updated configuration file from the server, wherein the preset state conditions are used to indicate that the mobile device is in a powered-on state and in the parking gear; and in response to the updated configuration file passing cyclic redundancy check, updating the identified scene attribute information based on the updated configuration file.
[0016] Optionally, the target page is a scene rendering map, and the target prompt information includes: an identifier element or a first environment image. The target prompt information for the target object is displayed on the target page of the corresponding display of the mobile device, including: rendering the identifier element corresponding to the target object in the scene rendering map; or, displaying the first environment image in the first layer of the target page and displaying the scene rendering map in the second layer of the target page, wherein the first layer is the layer above the second layer.
[0017] Optionally, the target page may also display a virtual navigation guide light carpet corresponding to the mobile device, wherein the virtual navigation guide light carpet is a driving plan channel for the mobile device in the future time period after the target object is identified.
[0018] Optionally, the mobile device includes a first controller and a second controller. Displaying the first environment image on the target page of the mobile device includes: using the first controller to acquire a first environment image including a target object, and sending the first environment image to the second controller; using the second controller to render the target page to display the first environment image on the target page of the mobile device.
[0019] According to another aspect of the embodiments of this application, an interactive device is also provided, including: a first acquisition module, configured to acquire a first environmental image containing a target object, wherein the target object is one of the factors affecting the driving state of the mobile device; and a display module, configured to display target prompt information for the target object on a target page of a display corresponding to the mobile device, wherein the target prompt information is generated based on the first environmental image.
[0020] Optionally, the first acquisition module is further configured to: acquire a second environmental image using the image acquisition sensor of the mobile device, wherein the second environmental image is used to reflect the driving environment of the mobile device; determine whether the second environmental image contains a target object; if the second environmental image contains a target object, determine a first environmental image based on the second environmental image and the target object, wherein the first environmental image is an image region in the second environmental image that at least includes the target object.
[0021] Optionally, the first acquisition module is further configured to: perform object recognition on the second environmental image to obtain at least one candidate object; and determine whether a target object exists among the at least one candidate object based on the category information and feature information corresponding to the at least one candidate object.
[0022] Optionally, the first acquisition module is further configured to: if the category information of the candidate object is a static obstacle category, and the feature information of the candidate object satisfies a first feature condition; wherein the first feature condition includes: the location of the object corresponding to the candidate object is within the path coverage area of the current driving path of the mobile device, and the size of the object corresponding to the candidate object is greater than a preset size threshold; or, if the category information of the candidate object is a dynamic obstacle category, and the feature information of the candidate object satisfies a second feature condition; wherein the second feature condition includes one of the following: the object moving speed of the candidate object is greater than a preset speed threshold, and the distance between the candidate object and the mobile device decreases in the future time period; the predicted moving trajectory of the object corresponding to the candidate object intersects with the current driving path of the mobile device; or, if the category information of the candidate object is a traffic sign category, and the feature information of the candidate object satisfies a third feature condition; wherein the third feature condition includes: the object content and object location corresponding to the candidate object are associated with the driving state of the mobile device.
[0023] Optionally, the first acquisition module is further configured to: perform image processing operations on the second environment image based on the target object in the second environment image to obtain a first environment image, wherein the image processing operations include at least one of the following: performing an image cropping operation on the target object in the second environment image, performing an image annotation operation on the target object in the second environment image, and performing image scaling processing on the target object in the second environment image.
[0024] Optionally, the target prompt information includes: a first environmental image. The display module is further configured to: respond to the second environmental image including multiple target objects, determine the target object to be displayed based on the priority information of the multiple target objects, and display the first environmental image corresponding to the target object to be displayed on the target page, wherein the priority information is determined based on the degree of influence of the multiple target objects on the driving state of the mobile device, or the priority information is determined based on user-preset settings.
[0025] Optionally, the display is a head-up display or a mobile device screen; the target page is a scene-rendered map or a two-dimensional vector map.
[0026] Optionally, the target page may also include at least one of the following driving prompts: driving status information of the mobile device, decision prompts corresponding to the target object, and driving mode switching information. The driving status information is used to determine the activation status of the mobile device's assisted driving function, and the decision prompts corresponding to the target object include at least one of the following: environmental label information associated with the target object, driving decision information, and decision reasoning process information.
[0027] Optionally, at least one of the following methods may be used to display driving prompt information: displaying interface cards, displaying text pop-ups, or providing voice prompts; or highlighting or annotating the target object in the first environmental image.
[0028] Optionally, the interactive device further includes: a second acquisition module, used to perform cluster analysis on the text content using the sign scene attribute information in response to the target object being a traffic sign and the target object containing text content, to obtain decision prompt information corresponding to the target object, wherein the sign scene attribute information is used to record the mapping relationship between multiple attribute fields, and the multiple attribute fields include: scene sign information, text sign type, sign priority, and prompt information corresponding to the target object.
[0029] Optionally, the interactive device further includes: an update module, configured to: obtain an update configuration file from the server in response to the mobile device meeting preset state conditions, wherein the preset state conditions are used to indicate that the mobile device is in a powered-on state and in the parking gear; and update the identification scene attribute information based on the update configuration file in response to the update configuration file passing cyclic redundancy check.
[0030] Optionally, the target page is a scene rendering map, and the target prompt information includes: an identifier element or a first environment image. The display module is also used to: render the identifier element corresponding to the target object in the scene rendering map; or, display the first environment image in the first layer of the target page and display the scene rendering map in the second layer of the target page, wherein the first layer is the layer above the second layer.
[0031] Optionally, the target page may also display a virtual navigation guide light carpet corresponding to the mobile device, wherein the virtual navigation guide light carpet is a driving plan channel for the mobile device in the future time period after the target object is identified.
[0032] According to another aspect of the embodiments of this application, a mobile device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0033] In this embodiment, a focused visual image is generated by extracting only the target object directly related to the vehicle control behavior. By acquiring a first environmental image containing the target object, which is one of the factors affecting the driving state of the mobile device, target prompt information for the target object is displayed on the target page of the mobile device's corresponding display. The target prompt information is generated based on the first environmental image, discarding redundant information from the full-scene environmental image and retaining only the complete visual representation of the target object, presenting it as an independent visual unit on the human-computer interaction interface. This achieves the goal of accurately conveying the vehicle decision-making basis within a limited display area, thereby enhancing the user's understanding of the behavioral motivation of the assisted driving system and improving the user's driving experience and interaction efficiency. This solves the technical problem of poor human-computer interaction experience and low collaborative efficiency caused by the user's inability to understand the vehicle's behavioral motivation in the assisted driving system. Attached Figure Description
[0034] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0035] Figure 1 This is a flowchart of an optional interaction method according to an embodiment of this application;
[0036] Figure 2 This is a schematic diagram of an optional image cropping operation according to an embodiment of this application;
[0037] Figure 3 This is a schematic diagram of an optional image annotation operation according to an embodiment of this application;
[0038] Figure 4 This is a schematic diagram of an optional image scaling operation according to an embodiment of this application;
[0039] Figure 5 This is a schematic diagram of an optional target page prompt content according to an embodiment of this application;
[0040] Figure 6This is a schematic diagram of an optional interaction method according to an embodiment of this application;
[0041] Figure 7 This is a flowchart of an optional update configuration file according to an embodiment of this application;
[0042] Figure 8 This is a schematic diagram of another optional interaction method according to an embodiment of this application;
[0043] Figure 9 This is a schematic diagram of an optional mobile device according to an embodiment of this application;
[0044] Figure 10 This is a structural block diagram of an optional interactive device according to an embodiment of this application. Detailed Implementation
[0045] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0046] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0047] According to an embodiment of this application, an embodiment of an interactive method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0048] With the increasing prevalence of driver assistance systems, vehicles need to rely on visual perception and OCR recognition technologies to detect traffic signs and obstacles affecting driving in real time, and trigger vehicle control behaviors such as deceleration and avoidance accordingly. However, in related technologies, the human-machine interface usually only presents the perception results as abstract symbols, global environment reconstruction, or raw camera images, lacking a precise mapping of key decision triggers. Moreover, users cannot intuitively know whether the vehicle deceleration is due to "construction ahead" or "crosswind warning," nor can they see the correlation between the original sign image that triggered the decision and the semantic text. Especially in scenarios where OCR recognition text is redundant, semantically repetitive, and multi-source heterogeneous, the interface is information overloaded and the key points are blurred, leading to user misunderstanding of the vehicle's intentions, causing a crisis of trust and unnecessary takeover. At the same time, limited by the display space and rendering performance of the in-vehicle screen, existing interfaces often present the perception results in the form of raw images, multiple labels stacked, or blurred text, resulting in information redundancy and blurred key points, making it difficult for users to efficiently obtain the core intentions in a short time. Moreover, there are many types and forms of signs in traffic scenarios, and new signs are constantly being added, such as "Caution for electric vehicles" and "Slippery road surface". Traditional solutions rely on software firmware upgrades to support the display of new signs, which has a long development cycle and high iteration costs, making it difficult to meet the practical needs of products to quickly respond to user feedback and test data.
[0049] In summary, the failure of relevant driver assistance systems to intuitively display the visual and semantic information that underpins decision-making processes makes it difficult for users to understand vehicle behavior, severely restricting the reliability, user experience, and efficiency of human-machine collaboration.
[0050] This application provides an interaction method. The interaction method can be used to provide perception decision-making, data processing, and visualization functions for autonomous driving in preset application scenarios. These preset application scenarios may include the following scenarios in the vehicle field: autonomous driving for commuting, artificial intelligence (AI) assisted driving scenarios for family cars, automatic parking assistance (APA) scenarios (such as memory parking for self-owned parking spaces in garages, intelligent parking for designated parking spaces in parking lots, etc.), and navigation-guided pilot (NGP) scenarios in urban or highway areas.
[0051] When the aforementioned preset application scenario falls within a field other than vehicles, those skilled in the art should understand that the vehicle in the above interaction method can be replaced with other objects, such as autonomous commercial vehicles, unmanned delivery vehicles, low-speed robots, drones, unmanned boats, or wearable devices equipped with augmented reality display functions. Correspondingly, the target page can be replaced with display pages related to other objects, such as augmented reality (AR) glasses display interfaces, vehicle head-up displays (HUDs), smartphone application interfaces, roadside unit (RSU) electronic screens, home smart terminal screens, pedestrian assistance device display interfaces, or digital twin screens of cloud-based traffic management platforms. Based on this, this application embodiment uses the vehicle field as an example to exemplify the specific implementation of the above interaction method.
[0052] Figure 1 This is a flowchart of an optional interaction method according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0053] Step S102: Obtain a first environmental image containing the target object, wherein the target object is one of the factors affecting the driving state of the mobile device;
[0054] Step S104: Display target prompt information for the target object on the target page of the corresponding display of the mobile device, wherein the target prompt information is generated based on the first environment image.
[0055] The aforementioned mobile devices refer to mobile intelligent platforms equipped with driver assistance systems, environmental perception sensors, onboard computing units, and human-machine interface displays. These can be intelligent vehicles, autonomous commercial vehicles, unmanned delivery vehicles, low-speed robots, drones, unmanned boats, or wearable devices equipped with augmented reality (AR) display capabilities. These mobile devices possess the ability to acquire external environmental images, perform perception and decision-making processes, generate vehicle control commands, and output human-machine interaction information, and at least include a perception and decision-making module and an AR display rendering module.
[0056] The aforementioned target objects refer to external entities identified by the environmental perception system in the vehicle's driving environment that can trigger the driver assistance system to adjust vehicle control behavior. Specific types of target objects include, but are not limited to, static environmental elements such as traffic signs, construction fences, streetlights, and guardrails, or dynamic environmental elements such as pedestrians, bicycles, and animals.
[0057] The aforementioned first environmental image represents a sub-image extracted from the original environmental image, with the target object as the core region, and after spatial cropping and other standardization processes. Its size can be 464×262 pixels. Other standardization operations include distortion correction, color normalization, and time synchronization. Both the original environmental image and the first environmental image are acquired by environmental perception sensors in the perception decision module. The content of the first environmental image includes the complete visual range of the target object in the original image, used to focus on presenting specific environmental cues that trigger vehicle behavior. Furthermore, when acquiring the first environmental image containing the target object, triggering can be initiated when the target object is identified as having a causal relationship with the vehicle's driving state, ensuring a clear semantic correspondence between the content of the first environmental image and the vehicle's control behavior.
[0058] The aforementioned target page refers to any display terminal interface with graphics rendering capabilities. This can be a Scene Reconstruction Display (SR) interface for vehicles, an Augmented Reality (AR) glasses display interface, a Head-Up Display (HUD) interface, a smartphone application interface, a Roadside Unit (RSU) electronic screen, a home smart terminal screen, a pedestrian assistance device display interface, or a digital twin screen for a cloud-based traffic management platform. The target page is displayed in a fixed position within the preset visual area of the central control screen or head-up display device, used to intuitively present the mobile device's understanding of the surrounding environment and its current behavioral intentions to the driver or user.
[0059] The aforementioned target prompt information represents the semantic information of the target object in the first environmental image, that is, the semantic visual prompt content for the user. Its manifestation includes heat-rendered areas, edge highlighting, text descriptions, and virtual identifier layers aligned with the target object space.
[0060] When displaying the first environmental image on the target page of a mobile device, the first environmental image can be treated as an independent visual unit. After being enlarged, cropped, or rendered by the SR display rendering module, it can be fully presented in the designated display area of the target page to ensure that the driver can directly identify the environmental factors that trigger the vehicle's behavior through visual focus.
[0061] Based on steps S102 to S104, this embodiment of the application adopts a method of extracting only the target object directly related to the vehicle control behavior and generating a focused visual image. By acquiring a first environmental image containing the target object, where the target object is one of the factors affecting the driving state of the mobile device, target prompt information for the target object is displayed on the target page of the corresponding display of the mobile device. The target prompt information is generated based on the first environmental image, discarding redundant information from the full-scene environmental image, retaining only the complete visual representation of the target object and presenting it as an independent visual unit on the human-computer interaction interface. This achieves the purpose of accurately conveying the vehicle decision-making basis within a limited display area, thereby enhancing the user's understanding of the behavioral motivation of the assisted driving system and improving the user's driving experience and interaction efficiency. This solves the technical problem of poor human-computer interaction experience and low collaborative efficiency caused by the user's inability to understand the vehicle behavioral motivation in the assisted driving system.
[0062] The interaction methods in the embodiments of this application will be further described below.
[0063] Optionally, acquiring a first environmental image containing the target object includes: acquiring a second environmental image using the image acquisition sensor of the mobile device, wherein the second environmental image is used to reflect the driving environment of the mobile device; determining whether the second environmental image contains the target object; if the second environmental image contains the target object, determining a first environmental image based on the second environmental image and the target object, wherein the first environmental image is an image region in the second environmental image that at least includes the target object.
[0064] The aforementioned second environmental image represents the original environmental image that is collected in real time by the image acquisition sensor and fully reflects the current driving environment of the mobile device. Its content includes all visual information of roads, vehicles, buildings, traffic signs and other environmental elements, and is the original input source for the perception and decision-making module to perform target detection and category recognition.
[0065] Before determining the first environment image based on the second environment image and the target object, it is first necessary to determine whether the target object is contained in the second environment image. Specifically, when identifying whether the target object is contained in the second environment image, the perception decision module can perform semantic segmentation and target recognition on the second environment image based on the target detection model. When there is an object matching a predefined category in the detection result, and its confidence level is higher than a preset threshold, it is determined that the target object exists in the second environment image.
[0066] After confirming that the target object is contained within the second environment image, the first environment image is then determined based on the second environment image and the target object. Specifically, during the generation of the first environment, the pixel data of the region enclosed by the two-dimensional (2D) bounding box coordinates of the target object in the second environment image can be used as a reference. Cropping, padding, and size normalization operations are then performed to form an image subset containing only the complete visual content of the target object. During the cropping process, the target object is ensured not to be truncated, and the image center is aligned with the geometric center of the target object to maintain the stability of the visual focus.
[0067] Based on the above optional embodiments, the embodiments of this application realize the accurate extraction of visual cues related only to vehicle control behavior from the full-scene environment image, and eliminate irrelevant background interference, so that the image content presented in the subsequent display stage has high focus and semantic clarity, thereby improving the driver's recognition speed and understanding accuracy of vehicle behavior causes without increasing the display area occupancy.
[0068] Optionally, determining whether the second environment image contains a target object includes: performing object recognition on the second environment image to obtain at least one candidate object; and determining whether a target object exists among the at least one candidate object based on the category information and feature information corresponding to the at least one candidate object.
[0069] The aforementioned candidate objects refer to visual regions in the second environmental image that are initially identified by the object detection model and may belong to the target object category. They possess two-dimensional bounding box coordinates, confidence scores, category labels, and local texture features, but have not yet been finally determined through semantic and behavioral correlation. The object detection model can be the YOLO (You Only LookOnce) series of models or the Single Shot MultiBox Detector (SSD) series of models.
[0070] When performing object recognition on the second environmental image, the perception and decision module can call a target detection algorithm based on a deep neural network to perform pixel-by-pixel analysis on the second environmental image and output multiple candidate regions with bounding boxes, class probabilities, and feature vectors. The recognition process does not rely on manual rules or preset templates, but automatically identifies all objects in the image that may affect the vehicle control behavior adjustment of the assisted driving system through end-to-end training of the model. The output result is a set of unfiltered candidate objects.
[0071] The category information mentioned above represents the semantic category label to which the target object belongs, and is a predefined, structured classification label. In this application embodiment, the category information of the target object includes, but is not limited to: static obstacle category, dynamic obstacle category, and traffic sign category.
[0072] The aforementioned feature information represents a high-dimensional numerical vector extracted by the target detection model from the local image region of the candidate object. This vector is used to characterize the visual structure, texture distribution, motion characteristics, and environmental semantic consistency of the candidate object. It is used to provide a deep semantic-based discrimination criterion when there is ambiguity, noise, or semantic ambiguity in the category label, so as to distinguish the real target object from the interference, such as reflections, shadows, false labels, and non-target objects.
[0073] This application embodiment determines the presence of a target object among at least one candidate object based on the category information and feature information corresponding to at least one candidate object. When determining whether a target object exists among the candidate objects, each object in the aforementioned candidate object set is compared to see if its category label belongs to a predefined target object type list. Simultaneously, its OCR-recognized text semantic content and spatial location features are combined to perform dual semantic and behavioral correlation verification. Only when the category matches and the semantics and features meet the vehicle control triggering conditions is the candidate object confirmed as the real target object. For example, spatial location features include whether it is located directly in front of the lane or whether it intersects with the vehicle's trajectory. The above-described process of determining whether a target object exists among the candidate objects based on the category information and feature information of the candidate objects achieves precise screening from "visual detection" to "semantic decision-making," reducing interference from irrelevant signs, incomplete signs, or other environmental objects.
[0074] Based on the above optional embodiments, this application embodiment obtains candidate objects by performing object recognition in the second environmental image, and determines whether a target object exists based on the category information and feature information of the candidate objects. The first environmental image is generated only when the existence of the target object is confirmed. This achieves the accurate screening of specific marker objects that are strongly associated with vehicle control behavior from massive environmental visual data, avoiding redundant processing and misdisplay of irrelevant visual information. Thus, without relying on complex text descriptions, it ensures that the image content presented by the human-computer interaction interface has high semantic focus and behavioral orientation, effectively improving the driver's immediate recognition efficiency and cognitive accuracy of vehicle behavior causes.
[0075] Optionally, a target object is determined to exist among at least one candidate object if at least one of the following conditions is met: if the category information of the candidate object is a static obstacle, and the feature information of the candidate object satisfies a first feature condition; wherein the first feature condition includes: the location of the object corresponding to the candidate object is within the path coverage area of the current driving path of the mobile device, and the size of the object corresponding to the candidate object is greater than a preset size threshold; or, if the category information of the candidate object is a dynamic obstacle, and the feature information of the candidate object satisfies a second feature condition; wherein the second feature condition includes one of the following: the object moving speed of the candidate object is greater than a preset speed threshold, and the distance between the candidate object and the mobile device decreases in the future time period; the predicted moving trajectory of the object corresponding to the candidate object intersects with the current driving path of the mobile device; or, if the category information of the candidate object is a traffic sign, and the feature information of the candidate object satisfies a third feature condition; wherein the third feature condition includes: the object content and object location corresponding to the candidate object are associated with the driving state of the mobile device.
[0076] The aforementioned candidate objects refer to potential environmental entities output by the perception and decision-making module based on the second environmental image through an object detection algorithm. These entities possess two-dimensional bounding box coordinates, category labels, confidence scores, and local visual features. Their categories cover three types: static obstacles, dynamic obstacles, and traffic signs.
[0077] The above-mentioned static obstacle category refers to objects whose positions are basically fixed in the driving environment, such as street lamp poles, traffic cones, construction fences, and guardrails. Static obstacles do not have significant displacement in the image sequence.
[0078] The above categories of dynamic obstacles refer to objects that are in relative motion in the driving environment, such as vehicles, pedestrians, cyclists, etc. The position of dynamic obstacles changes continuously over time.
[0079] The aforementioned traffic sign categories represent road signs with standardized graphic or textual semantics used to convey regulations or warning information to drivers, such as signs for "crosswinds," "school zone," and "construction ahead." Their recognition relies on the combined results of optical character recognition and graphic template matching.
[0080] The first feature condition mentioned above represents the static obstacle screening rule used to determine whether a candidate object is a valid target object. It includes two necessary conditions: First, the object position of the candidate object is within the path coverage area of the current driving path of the mobile device, that is, after the center point of the candidate object's two-dimensional bounding box is projected onto the vehicle's trajectory plane, it is located within the lane centerline buffer area of a preset width. Second, the object size of the candidate object is greater than a preset size threshold, that is, the product of the pixel width and height of the candidate object in the image exceeds the minimum visually effective area preset by the system, ensuring that it has sufficient visual salience.
[0081] The second characteristic condition mentioned above represents one of two mutually exclusive screening conditions used to determine whether a dynamic obstacle constitutes a basis for vehicle control decisions: First, the candidate object's moving speed is greater than a preset speed threshold, meaning its displacement speed between consecutive frames exceeds the system's set low-speed stationary reference. The preset speed threshold can be 5 km / h. Alternatively, the distance between the candidate object and the mobile device may decrease in the future time period, meaning the relative distance calculated based on the motion prediction model shows a convergence trend within the next 1.5 seconds. Second, the predicted trajectory of the candidate object intersects with the current driving path of the mobile device, meaning its expected motion direction vector and the vehicle's heading vector form a non-parallel angle in space, posing a potential collision risk.
[0082] The third characteristic condition mentioned above represents the semantic-spatial correlation criterion used to determine whether a traffic sign triggers vehicle control behavior. It requires that the content of the candidate object, i.e., the text or graphic semantics obtained through OCR recognition, has a clear causal relationship with the driving state of the mobile device. For example, if the text "crosswind" is recognized and the vehicle is currently decelerating, or if the text "school zone" is recognized and the vehicle has entered a speed limit zone, the object's position must be directly in front of the lane and within the effective perception range of the longitudinal distance from the vehicle.
[0083] When determining whether a candidate object is the target object, there are three possible scenarios:
[0084] First, when the candidate object's category information is a static obstacle category, and the candidate object's feature information meets the first feature condition, the candidate object is determined to be a target object. During the judgment process, after detecting that the candidate object is a static object, the perception and decision-making module further verifies whether it is located in the vehicle's upcoming driving path and has sufficient visual scale. Only when both conditions are met is it confirmed as a target object that can trigger behavioral decisions. This judgment process eliminates irrelevant roadside signs, small distant markings, or distractions not directly in front of the lane, improving the accuracy of target selection.
[0085] Second, when the candidate object's category information is a dynamic obstacle category and the candidate object's feature information meets the second feature condition, the candidate object is determined to be a target object. During the judgment process, when a candidate object is identified as a dynamic entity such as a vehicle or pedestrian ahead, the system does not rely on a single speed or distance judgment. Instead, it predicts whether its movement trend constitutes a potential threat and only identifies it as a target object when either abnormal speed or trajectory intersection is met, thus avoiding false triggering of slow-moving vehicles at a distance or parallel objects.
[0086] Third, when the candidate object's category information is a traffic sign category, and the candidate object's feature information meets the third feature condition, the candidate object is determined to be a target object. During the judgment process, the system not only recognizes the visual form of the traffic sign, but also requires that its semantic content be directly related to the current vehicle behavior before confirming the sign as a target object. For example, the sign is only confirmed as a target object when the text "Construction Ahead" is recognized and the vehicle has already initiated a detour maneuver. This judgment mechanism achieves a semantic closed-loop verification of "perception-intent-behavior," avoiding the display of isolated signs without a response from the uncontrolled vehicle.
[0087] Based on the above optional embodiments, this application embodiment sets independent screening conditions for three types of targets—static obstacles, dynamic obstacles, and traffic signs—according to the category and feature information of candidate objects. This achieves multi-dimensional semantic filtering of the perception results. Only when a candidate object fully meets the preset three types of feature conditions in terms of category, location, size, movement trend, or semantic relevance is it confirmed as a target object. The above steps significantly reduce the interference of irrelevant environmental elements on the human-computer interaction interface, ensure that the generation of the first environmental image has a clear behavioral trigger basis, and make the displayed content always maintain a strong consistency with the actual vehicle control intention, thereby improving the driver's understanding and trust in the system's decisions.
[0088] Optionally, determining the first environment image based on the second environment image and the target object includes: performing an image processing operation on the second environment image based on the target object in the second environment image to obtain the first environment image, wherein the image processing operation includes at least one of the following: performing an image cropping operation on the target object in the second environment image, performing an image annotation operation on the target object in the second environment image, and performing image scaling processing on the target object in the second environment image.
[0089] The image cropping operation described above involves extracting all pixel regions enclosed by the two-dimensional bounding box coordinates of the target object in the second environmental image, and then cropping out the complete visual content of the target object to ensure it is not truncated by edges, thus achieving precise positioning of the visual focus. For example, Figure 2 This is a schematic diagram of an optional image cropping operation according to an embodiment of this application, such as... Figure 2 As shown. Specifically, based on the object detection model, a rectangular sub-image is extracted from the second environment image. This sub-image completely contains the target object and its surrounding local environment, thus obtaining the image before cropping. The rectangular sub-image is then cropped to a certain resolution, such as 464. A 262-pixel image containing the target object yields the cropped, standardized image, also known as the first environment image.
[0090] In one optional embodiment, in response to the target object's display area in the second environmental image being smaller than a preset area, an image cropping operation is performed on the second environmental image based on the target object's corner coordinate information to obtain a first environmental image, wherein the corner coordinate information is used to determine the boundary of the target object.
[0091] The aforementioned display area represents the size of the pixel region occupied by the target object in the second environmental image, used to measure the physical proportion of the target object in the image. The aforementioned preset area represents a pre-set area threshold, used to determine whether the target object appears as a visually too small area in the image due to excessive distance or a small viewing angle.
[0092] The corner coordinate information mentioned above represents the pixel coordinates of the four corner points of the target object's two-dimensional bounding box in the second environment image, used to accurately describe the spatial extent and geometric boundaries of the target object. The boundary of the target object represents its contour edge in the image, defined by the corner coordinate information, and used to limit the area of the image cropping operation.
[0093] When the display area of the target object in the second environmental image is smaller than a preset area, an image cropping operation is performed on the second environmental image based on the corner coordinate information of the target object to obtain the first environmental image. Specifically, when the perception system detects that the pixel area occupied by the target object in the originally acquired second environmental image is lower than a preset area threshold, the system obtains the precise pixel coordinates of the four corner points corresponding to the two-dimensional bounding box of the target object in the second environmental image. Based on the rectangular area enclosed by the above four corner coordinates, a local image sub-region containing only the target object and its immediate surrounding environment is cropped from the second environmental image to obtain the first environmental image.
[0094] The image annotation operations described above involve visual enhancement processing of the edge regions of the target object, such as applying soft blur, highlighting outlines, or gradient color shading, to highlight its salience in the image and improve the human eye's ability to quickly identify key information. For example, Figure 3 This is a schematic diagram of an optional image annotation operation according to an embodiment of this application, such as... Figure 3 As shown, the image on the right displays the result of annotating the first environmental image, i.e., highlighting and outlining the edge regions, and the image contains the target object.
[0095] In one optional embodiment, in response to the second environment image including multiple target objects, an image annotation operation is performed on the second environment image based on the pixel position information of the target objects to obtain a first environment image, wherein the pixel position information is used to determine the position of the multiple target objects in the second environment image.
[0096] The pixel location information mentioned above represents the set of two-dimensional coordinates of each pixel of the target object in the second environmental image, which is used to determine the spatial distribution and relative position of the target object in the image plane.
[0097] When the second environmental image includes multiple target objects, image annotation is performed on the second environmental image based on the pixel position information of the target objects to obtain the first environmental image. Specifically, when the vehicle perception system simultaneously identifies two or more target objects in the originally acquired second environmental image, the system acquires the two-dimensional spatial coordinate data of each target object in the second environmental image, including but not limited to the center point coordinates, corner point coordinates, or pixel range of the covered area of each target object's bounding box. Then, based on the pixel position information of each target object, the system draws corresponding visual identifier elements for each target object on the second environmental image, including but not limited to: adding a highlighted border to the boundary area of each target object, generating a thermal rendering area, overlaying transparent color blocks or dynamic pulse effects, and associating corresponding semantic text at the corresponding positions, so that multiple target objects are clearly distinguished and highlighted in the image. The image generated after the annotation operation is the first environmental image, which serves as the rendering input for the human-computer interaction interface. It retains the original environmental background and overlays visual annotations at the precise pixel positions of each target object to form an enhanced visual image.
[0098] The image scaling process described above involves uniformly enlarging or reducing the cropped target area by a certain factor. This allows the text and graphic details of the target object to achieve higher pixel density within a limited display space, enhancing readability. The magnification or reduction factor can be adaptively adjusted based on real-world human-machine interface experience, such as 1.2. For example, Figure 4 This is a schematic diagram of an optional image scaling operation according to an embodiment of this application, such as... Figure 4 As shown, the image on the right displays the result of scaling the first environment image, specifically reducing it by a factor of 1.2, and the image contains the target object.
[0099] In one optional embodiment, in response to the target object having an image resolution in the second environment image that is less than a preset resolution, an image scaling operation is performed on the second environment image based on the pixel position information of the target object to obtain a first environment image.
[0100] The image resolution mentioned above represents the number of pixels contained within a unit length of the target object in the second environmental image, reflecting the clarity and richness of detail. The preset resolution mentioned above represents a pre-set resolution threshold used to determine whether the target object is too small or too far away to meet the requirements for clear recognition and display.
[0101] When the resolution of the target object in the second environmental image is less than a preset resolution, an image scaling operation is performed on the second environmental image based on the pixel position information of the target object to obtain the first environmental image. Specifically, when the perception system detects that the number of pixels per unit length of the target object in the originally acquired second environmental image is lower than a preset resolution threshold, the system acquires the precise spatial coordinate data of the target object in the second environmental image, including the corner coordinates of the target object's bounding box or the pixel range of the covered area. Then, based on the pixel data information of the target object, the system extracts a local image region containing the target object from the second environmental image and performs a magnification processing based on an interpolation algorithm on the local region to increase the pixel density of the target object in the image and increase the clarity of texture, edge, and text information, thereby generating a locally enhanced image after image scaling, i.e., the first environmental image.
[0102] In the process of obtaining the first environment image by performing image processing operations on the second environment image based on the target object in the second environment image, after confirming the existence of the target object, the system can choose to perform at least one of the following operations: image cropping, image annotation, and image scaling. The execution order and combination are dynamically determined by the resource load and visual priority strategy of the display rendering module. Among them, image cropping ensures content focus, image annotation enhances visual guidance, and image scaling improves semantic clarity. The three work together to form a visual condensation process with minimal redundancy and maximum information density of the original environment image.
[0103] Based on the above optional embodiments, this application embodiment generates a first environment image by performing at least one image processing operation, such as image cropping, image annotation, or image scaling, on the second environment image based on the target object in the second environment image. This achieves targeted optimization and semantic focus of the original environmental visual data, so that the final image content retains only high-value visual cues directly related to vehicle control behavior, while eliminating background interference and redundant information. Thus, without increasing the display area, it significantly improves the visual recognition, semantic clarity, and human-computer interaction understandability of the target object, ensuring that the driver can accurately perceive the environmental factors of vehicle behavior in a moment of attention.
[0104] Optionally, the target prompt information includes: a first environmental image. Displaying the target prompt information for the target object on the target page of the display corresponding to the mobile device includes: responding to the second environmental image including multiple target objects, determining the target object to be displayed based on the priority information of the multiple target objects, and displaying the first environmental image corresponding to the target object to be displayed on the target page, wherein the priority information is determined based on the degree of influence of the multiple target objects on the driving state of the mobile device, or the priority information is determined based on user-preset settings.
[0105] The aforementioned multiple target objects refer to multiple objects that are simultaneously identified and confirmed by the perception and decision-making module in the second environmental image of the current driving environment as objects that can trigger vehicle control behavior. These objects can be one or more of static obstacles, dynamic obstacles, and traffic signs.
[0106] The priority information mentioned above represents a quantitative ranking criterion used to measure the impact of multiple target objects on vehicle behavior in the current driving context. It can be determined based on the degree of influence of multiple target objects on the driving state of the mobile device, i.e., grading the vehicle control behavior triggered by the semantics of the sign according to the urgency and mandatory nature of the action. For example, "Construction Ahead" triggering a detour has a higher priority than "Crosswind" triggering a deceleration action. Priority information can also be determined based on user-preset settings. Specifically, if the driver assigns a higher display priority to specific sign types, such as "Schoolway," in the vehicle settings menu, the system will adjust the default priority ranking accordingly.
[0107] When the second environment image includes multiple target objects, the target object to be displayed is determined based on the priority information of these multiple target objects. Specifically, when the system detects the simultaneous presence of multiple target objects, the priority information serves as the sole sorting criterion. The system sorts the target objects from highest to lowest priority and selects the highest priority target object as the object to be displayed. This method of determining the target object to be displayed based on the priority information of multiple target objects ensures that only one target object's first environment image is projected onto the target page at any given time, avoiding visual information overload.
[0108] When displaying the first environment image corresponding to the target object on the target page, the system renders and outputs the first environment image corresponding to the selected target object according to the preset display coordinate position, such as the upper left corner or upper center of the screen. The image maintains a standardized size and enhances the edges to ensure that it occupies the visual focus in the limited display area and does not overlap or interfere with other interface elements.
[0109] Based on the above optional embodiments, this application embodiment, in response to the inclusion of multiple target objects in the second environmental image, determines the target object to be displayed based on the priority information of the multiple target objects, and displays the first environmental image corresponding to the target object to be displayed on the target page. This achieves strict selective delivery of display content in multi-target concurrent scenarios, ensuring that the driver only receives the highest priority environmental behavior inducement information at any given time. It avoids the increase in cognitive load and distraction caused by the simultaneous presentation of multiple sign images, thereby maintaining the clarity, uniqueness, and decision relevance of information presentation within a limited display space, and significantly improving the readability of the human-computer interaction interface and driving safety.
[0110] Optionally, the display is a head-up display or a mobile device screen; the target page is a scene-rendered map or a two-dimensional vector map.
[0111] The aforementioned head-up display page refers to the display interface of an optical imaging system that projects information onto the area of the vehicle's windshield. The display position of this head-up display page is at the level of the driver's line of sight, allowing the driver's gaze to remain on the road ahead.
[0112] The aforementioned mobile device display page refers to the physical screen display interface located on the center console or dashboard, and the mobile device display page is a traditional flat display device with high resolution and multi-layer rendering capabilities.
[0113] The aforementioned scene rendering map represents a dynamic visualization environment model generated based on multi-sensor fusion data including image acquisition sensors. It includes real-world three-dimensional geometric structures and semantic labels, reconstructing the spatial location and appearance features of elements such as roads, vehicles, pedestrians, and traffic signs with pixel-level precision, presenting a visual effect close to the real scene.
[0114] The aforementioned two-dimensional vector map represents a simplified spatial model of road structure, lane boundaries, traffic sign locations, and semantic attributes expressed in mathematical vector form. It does not rely on image pixels but describes environmental elements through coordinate points, line segments, polygons, and attribute fields. It has a small data volume, efficient rendering, and lossless scaling, making it suitable for lightweight human-computer interaction scenarios.
[0115] The display can be a head-up display page or a mobile device display page. That is, the display architecture of this application embodiment is compatible with two mainstream vehicle display terminals. The head-up display page is suitable for high-safety scenarios that require keeping the line of sight forward, while the mobile device display page is suitable for scenarios with high information complexity and requiring in-depth reading. The above compatibility design can be adapted to different vehicle platforms and does not depend on specific hardware configurations.
[0116] The target page can be a scene-rendered map or a two-dimensional vector map. That is, the target page supported in this embodiment includes two rendering modes: scene-rendered map mode and two-dimensional vector map mode. In scene-rendered map mode, target prompts are overlaid on the reconstructed real-world image as high-fidelity visual elements. In two-dimensional vector map mode, target prompts are presented as simplified graphic symbols and text labels, associated with the corresponding road element locations in the vector coordinate system. Both modes satisfy the spatial alignment and clear expression requirements of the target object's semantic information.
[0117] Based on the above optional embodiments, this application embodiment achieves unified semantic expression and spatial alignment of target prompt information under different hardware configurations and user interaction scenarios by being compatible with two display terminal forms and two map rendering modes. This ensures that users can obtain consistent, clear, and understandable vehicle behavior intention feedback regardless of the display device and map view they use, thereby improving the adaptability, scalability, and consistency of the human-computer interaction system.
[0118] Optionally, the target page may also include at least one of the following driving prompts: driving status information of the mobile device, decision prompts corresponding to the target object, and driving mode switching information. The driving status information is used to determine the activation status of the mobile device's assisted driving function, and the decision prompts corresponding to the target object include at least one of the following: environmental label information associated with the target object, driving decision information, and decision reasoning process information.
[0119] The aforementioned driving status information can represent text or icon prompts that clearly indicate the current activation status of the mobile device's driver assistance function, including semantically clear status indicators such as "Navigation-assisted driving is activated" and "Adaptive cruise is running." These originate from the system status module of the autonomous driving domain controller and are used to communicate to the driver whether the vehicle is currently in driver assistance decision-making mode.
[0120] The decision information corresponding to the aforementioned target object represents semantic auxiliary explanatory content associated with the first identified and displayed environmental image. This includes at least one of the following: environmental label information associated with the target object, i.e., standardized naming of the environmental type to which the target object belongs, such as "crosswind section," "school area," "construction ahead," etc.; driving decision information, i.e., a concise description of the control actions taken by the vehicle towards the target object, such as "vehicle is decelerating" or "about to detour," derived from the vehicle control behavior decision flag of the perception decision module; and decision reasoning process information, i.e., a concise logical explanation of why the vehicle made this decision, such as "due to the detection of the 'crosswind' sign and the wind speed prediction exceeding the standard, deceleration intervention is initiated." Based on the perception results and vehicle control decisions, the system dynamically generates and overlays one or more of these three types of semantic prompts. Environmental label information establishes semantic anchors, driving decision information clarifies behavioral intent, and decision reasoning process information enhances system transparency. Together, these three constitute a complete behavioral explanation chain.
[0121] The driving mode switching information mentioned above indicates the state transition information prompted when the assisted driving function changes mode due to the target object or user intervention, such as "Switched to low-speed following mode" or "Assisted driving will exit", which is used to inform the driver of the dynamic adjustment of the system's operating level.
[0122] The target page also includes at least one driving prompt, indicating that the content displayed on the target page is not limited to the first environmental image, but also includes at least one of the following: driving status information, decision information corresponding to the target object, or driving mode switching information. The aforementioned target page display prompts break through the single information dimension of traditional SR interfaces that only display visual images. It structurally integrates the semantic information of system status, behavioral intent, and decision logic with visual images, forming a composite prompt mechanism of images and text.
[0123] Figure 5 This is a schematic diagram of an optional target page prompt content according to an embodiment of this application, such as... Figure 5 As shown, the top left corner displays the first environmental image corresponding to the target object. Here, "NGP Assisted Driving" represents the mobile device's driving status information, "Crosswind Section, Vehicle is Decelerating" represents the decision information corresponding to the target object, and "Lane Centering Control (LCC)" represents the driving mode switching information. Specifically, in the decision information corresponding to the target object, "Crosswind Section" is the environmental label information, "Vehicle is Decelerating" is the driving decision information, and "Crosswind Section, Vehicle is Decelerating" is the decision reasoning process information.
[0124] Based on the above optional embodiments, this application embodiment achieves a composite information presentation mode that overlays structured semantic prompts on visual images by presenting at least one of driving status information, decision information corresponding to the target object, or driving mode switching information on the target page. This enables the driver to simultaneously obtain the vehicle system operating status, behavioral intent interpretation, and mode change notification, significantly improving the interpretability of autonomous driving decisions and the transparency of human-machine collaboration. Thus, without increasing visual complexity, it alleviates panic caused by unpredictable behavior and enhances the driver's trust in the system and readiness to take over.
[0125] Optionally, at least one of the following methods may be used to display driving prompt information: displaying interface cards, displaying text pop-ups, or providing voice prompts; or highlighting or annotating the target object in the first environmental image.
[0126] The aforementioned driving prompt information refers to non-image semantic information presented on the target page to assist the driver in understanding the vehicle's behavioral intentions, including driving status information, prompt information corresponding to the target object, or driving mode switching information. Its content comes from the structured output of the system's perception and decision-making module and has clear semantic boundaries.
[0127] The aforementioned interface card display indicates that the prompt content is embedded in a fixed area of the target page in the form of a rectangular frame, with borders, background color, icons and text layout. The visual structure is independent of the first environment image to ensure information readability and layout stability.
[0128] The aforementioned text pop-up display indicates that the prompt content appears in a dynamically floating, lightweight text window in the unobstructed area of the target page. The appearance of the text pop-up is synchronized with the recognition of the target object, and its disappearance is triggered by the system timeout mechanism or new information, thus possessing both temporary and focused characteristics.
[0129] The aforementioned voice prompts are delivered via the vehicle's audio system in natural language, such as "There is a crosswind sign ahead, the vehicle will slow down soon." The semantics are consistent with the visual prompts, providing a multimodal perception channel.
[0130] The aforementioned highlighting indicates that visual enhancement processing, such as edge softening, gradient color shading, or brightness enhancement, is applied to the boundary area of the target object in the first environmental image to strengthen its visual salience.
[0131] The above annotations indicate that graphic symbols or text labels are superimposed on the first environmental image, such as adding the word "crosswind" or arrows to indicate direction below traffic signs, to clarify the semantics of the objects.
[0132] At least one method of displaying driving prompt information includes: displaying interface cards, displaying text pop-ups, and providing voice prompts. This indicates that the presentation format of the prompt content in this embodiment has multimodal selectability, that is, the three interaction channels of interface cards, pop-ups, and voice are structurally combined, allowing the system to dynamically select the optimal combination method based on environmental noise, driver attention status, and system load, thereby achieving redundant backup and scene adaptation for information transmission. Figure 5 In this context, the prompts are displayed as interface cards.
[0133] The target object can be highlighted or labeled in the first environmental image. That is, in this embodiment of the application, the target object is visually enhanced within the first environmental image. Highlighting highlights the outline of the object through edge soft light, while labeling clarifies its identity by adding semantic tags. The two work together to ensure that even in strong light or low contrast environments, the driver can still quickly identify the semantic attributes of the target object and avoid information loss due to image blurring or small size.
[0134] Based on the above optional embodiments, this application embodiment presents at least one driving prompt information in at least one of the following ways: displaying interface cards, displaying text pop-ups, or providing voice prompts. At the same time, it highlights or marks the target object in the first environmental image and supports the target page as a head-up display page or a mobile device display page. This achieves multi-channel, multi-form, and multi-terminal collaborative output of prompt information, ensuring that the driver can obtain key behavioral intentions through the most suitable perception mode in different driving scenarios and attention states. This significantly improves the reliability, adaptability, and robustness of information transmission and human-computer interaction.
[0135] Optionally, the method for obtaining the decision prompt information corresponding to the target object includes: in response to the target object being a traffic sign and the target object containing text content, clustering analysis is performed on the text content using the sign scene attribute information to obtain the decision information corresponding to the target object. The sign scene attribute information is used to record the mapping relationship between multiple attribute fields, including: scene sign information, text sign type, sign priority, and decision information corresponding to the target object.
[0136] The above text content represents the recognizable text sequence contained on the surface of the traffic sign, output after the second environmental image captured by the vehicle-mounted camera is processed by the OCR model. Its original form is an unprocessed string, which may contain redundant or non-standard expressions.
[0137] The aforementioned scene attribute information represents structured configuration data stored in the vehicle-side scene library attribute list. This data is used to establish mapping relationships between multiple predefined attribute fields. The fields of the scene attribute information include: scene identifier information, text identifier type, identifier priority, and decision information corresponding to the target object. Specifically, scene identifier information is a unique numerical identifier assigned to each type of traffic sign, such as "1" representing a crosswind sign and "2" representing a school zone sign; text identifier type is the semantic classification label for the OCR-recognized text, such as "contains 'Caution Crosswind' text" or "contains 'Construction Ahead' text"; identifier priority is the ranking weight set according to the importance of traffic regulations and the urgency of vehicle control behavior, such as "school zone" having a priority of 1 and "crosswind" having a priority of 2; and decision information corresponding to the target object is a standardized semantic description automatically generated by the system based on the input text, directed towards the driver, such as "crosswind zone, vehicle is slowing down."
[0138] The clustering analysis described above represents the semantic merging and standardization transformation of the original OCR text content based on preset text matching rules in the identified scene attribute information. Its essence is a string matching and template mapping process based on keyword matching. The table below shows the text matching rules used for clustering analysis in the embodiments of this application.
[0139]
[0140] When the target object is a traffic sign and contains text content, the system uses the sign scene attribute information to perform cluster analysis on the text content to obtain the decision information corresponding to the target object. Specifically, when the system identifies that the target object belongs to the traffic sign category in the output of the perception and decision module and confirms that it contains readable text information, it compares the original text output by OCR with the "text sign type" field in the sign scene attribute information line by line. When the system detects that the text contains predefined keywords, such as "school ahead" or "school zone", it automatically retrieves the corresponding "scene sign information", "sign priority", and "decision information corresponding to the target object" and outputs the latter as standardized prompt content. For example, variations such as "school ahead" and "caution school zone" are uniformly mapped to "school zone, vehicle is slowing down", achieving semantic convergence from heterogeneous input to unified output.
[0141] Figure 6 This is a schematic diagram of an optional interaction method according to an embodiment of this application, such as... Figure 6 As shown. First, the perception and decision module outputs the original image of the traffic sign, the coordinates of the 2D bounding box, and the original text content recognized by OCR. Then, the system receives the above data, performs cluster analysis on OCR text with similar semantics based on the preset scene library attribute list, and selects the highest priority text identifier type and clustered text prompt information based on preset priority rules. At the same time, the original image is cropped, labeled, and scaled based on the 2D bounding box. Only when the "vehicle control behavior decision flag" output by the perception and decision module is in a valid state, that is, the system has initiated response behaviors such as deceleration and detour, will the system send the scene identifier information, text identifier type, processed image, and clustered text prompt information of the traffic sign to the SR display rendering module. Finally, the information is displayed in a fixed or configurable position on the environment reconstruction interface (SR) of the in-vehicle screen, outputting scene identifier information, text identifier type, prompt information, and the first environment image, realizing a closed-loop interaction of "behavior triggering, accurate display, and dynamic updating".
[0142] Based on the above optional embodiments, this application embodiment, in response to the target object being a traffic sign containing text content, uses the sign scene attribute information to perform cluster analysis on the text content to obtain the decision information corresponding to the target object. This achieves semantic standardization and structural transformation of the original OCR recognition results, uniformly mapping diverse and non-standardized text expressions into fixed semantic prompt statements. Thus, without relying on large-scale model retraining, it supports the rapid expansion of new sign types through cloud-based hot updates of scene library attribute lists, significantly improving the accuracy, consistency, and iteration efficiency of prompt information generation, and ensuring reliable output of human-computer interaction content under multiple scenarios and semantic variations.
[0143] Optionally, before performing cluster analysis on the text content using the identified scene attribute information, the method further includes: in response to the mobile device meeting preset state conditions, obtaining an updated configuration file from the server, wherein the preset state conditions are used to indicate that the mobile device is in a powered-on state and in the parking gear; and in response to the updated configuration file passing cyclic redundancy check, updating the identified scene attribute information based on the updated configuration file.
[0144] The aforementioned preset state conditions represent two necessary operational constraints for the system to determine whether to allow cloud configuration file updates. The first is that the mobile device is powered on, meaning that both the vehicle's high-voltage and low-voltage power supplies are connected, and the onboard computing system has started up and entered a stable operating mode. The second is that the mobile device is in the parking position, meaning that the gear sensor indicates "P," indicating that the vehicle is not in motion and has safe conditions for background updates, preventing system anomalies caused by configuration switching while driving.
[0145] The aforementioned server represents a remote configuration management center operated by the vehicle manufacturer. It is equipped with a standardized scenario library attribute list configuration file, which is encoded in comma-separated values (CSV) format. The file contains fields such as scenario identification information, text identification type, identification priority, and prompt information corresponding to the target object, and is accompanied by version number, generation date, and cyclic redundancy check value.
[0146] The aforementioned update configuration file represents a configuration data file containing the latest scene library attribute information pushed from the server to the mobile device. Its content is used to replace or incrementally update the scene library attribute list on the vehicle's local device.
[0147] The Cyclic Redundancy Check (CRC) mentioned above means that a mathematical operation is performed on all bytes of data in the updated configuration file using a preset polynomial algorithm to generate a unique check value. This value is used to verify whether the file has been corrupted, tampered with, or incomplete during transmission. The check result is a binary judgment, either pass or fail.
[0148] When a mobile device meets preset conditions, it retrieves an updated configuration file from the server. Specifically, when the system detects that the mobile device is simultaneously powered on and in the park position, it actively establishes a communication link with the server and initiates a version comparison request. If the server-side configuration file version number is higher than the current version on the vehicle, a download process is triggered, and the updated configuration file is completely transferred to the vehicle's temporary storage space. The above process of retrieving the updated configuration file is only performed when the vehicle is stationary and the system power supply is stable, ensuring that the update operation does not interfere with the driving control logic.
[0149] When the updated configuration file passes the cyclic redundancy check, the system updates the identification scene attribute information based on the updated configuration file. Specifically, after the file is downloaded, the system recalculates the file checksum according to a preset CRC algorithm and compares it with the original checksum carried in the file header. If they match, the file is deemed complete and valid. The system then writes all attribute fields from the updated configuration file into the identification scene attribute information file in the specified directory on the vehicle, replacing the original mapping relationship and completing the dynamic upgrade of the local knowledge base. The update process does not involve a system restart and can be completed seamlessly in the background.
[0150] Figure 7 This is a flowchart of an optional update configuration file according to an embodiment of this application, such as... Figure 7 As shown, the vehicle-side system periodically polls the cloud server to obtain the version number of the latest scene library configuration file. If a version update is detected, it securely downloads a CSV configuration file containing scene identification information, text identification type, identification priority, and prompt information corresponding to the target object to a designated directory on the vehicle. The application then calculates a CRC-32 checksum and compares it with the embedded checksum in the file to ensure the file is intact and unaltered. Once the checksum is successful, the system atomically writes the new configuration to the official configuration directory, replacing the old version, and applies this configuration file. It also retains historical versions for rollback support. Subsequently, it notifies the SR display rendering module to dynamically load the new strategy. This achieves cloud-based hot updates and seamless activation of mobile object display rules without requiring a vehicle restart or firmware upgrade, significantly improving system scalability, security, and maintenance efficiency.
[0151] Based on the above optional embodiments, this application embodiment obtains an updated configuration file from the server in response to the mobile device meeting preset state conditions, and updates the identification scene attribute information based on the updated configuration file after the updated configuration file passes cyclic redundancy check. This realizes the remote hot update capability of traffic sign semantic mapping rules, enabling this application embodiment to dynamically expand the recognition and prompting logic of new types of traffic signs in the cloud without relying on vehicle software version upgrades. This significantly improves the system's adaptability and iteration efficiency to new road signs. At the same time, the dual state constraints of power-on state and parking gear and the dual verification mechanism of cyclic redundancy check ensure the security, integrity and reliability of the configuration update process.
[0152] Optionally, the target page is a scene rendering map, and the target prompt information includes: an identifier element or a first environment image. The target prompt information for the target object is displayed on the target page of the corresponding display of the mobile device, including: rendering the identifier element corresponding to the target object in the scene rendering map; or, displaying the first environment image in the first layer of the target page and displaying the scene rendering map in the second layer of the target page, wherein the first layer is the layer above the second layer.
[0153] The aforementioned identifier elements represent enhanced visual objects with clear semantic expression that are not original images generated in the scene rendering map. These include, but are not limited to: thermal regions of the target object, highlighted borders, dynamic pulse halos, virtual overlay layers aligned with the space of real signs, and clustered text descriptions, used to highlight the semantic attributes of the target object in the scene rendering map. Figure 8 This is a schematic diagram illustrating another optional interaction method according to an embodiment of this application, such as... Figure 8 As shown, if the target object is a "crosswind" sign, the rendered sign will be displayed in the scene rendering map.
[0154] The first layer mentioned above represents the uppermost image rendering layer on the target page. Its content is displayed before the lower layers and is used to present original visual information or overlay prompts, ensuring that it visually covers the lower content and enhancing the user's perception of the priority of key information.
[0155] The second layer mentioned above represents the lower image rendering layer in the target page. Its content is a scene rendering map, which serves as the environmental background and provides lane structure and road semantics aligned with the real world, providing a spatial reference system for the upper layers.
[0156] The target page used in this application uses a scene rendering map as the underlying environment display basis. Its essence is a three-dimensional realistic environment model reconstructed based on sensor data, rather than a two-dimensional vector map or static image, to ensure that all prompts are established in a semantic coordinate system consistent with the real road space, avoiding cognitive misalignment caused by coordinate offset.
[0157] The target prompt information in this application embodiment includes either an identifier element or the first environment image. That is, this application embodiment provides two independent but complementary ways of expressing target prompt information: the first is to generate standardized identifier elements to abstractly express semantics, and the second is to directly overlay the original first environment image to retain perceptual details. Both can be used as implementation forms of target prompt information to meet the differentiated needs for clarity and realism in different scenarios.
[0158] When rendering the identifier elements corresponding to the target object in the scene rendering map, the system generates identifier elements containing heat regions, highlighted edges and clustered text based on the recognition results of the target object, and maps their spatial coordinates to the three-dimensional coordinate system of the scene rendering map. The SR display rendering module then performs graphic overlay rendering at the corresponding physical location on the scene rendering map to achieve accurate fusion of semantic information and environmental background.
[0159] When the first environmental image is displayed on the first layer of the target page, and the scene rendering map is displayed on the second layer of the target page, the SR display rendering module renders the original first environmental image as the upper layer on the target page after processing its transparency, and at the same time uses the scene rendering map as the lower layer as the base background. By controlling the layer stacking order, the original image content is covered on the reconstructed environment, which not only preserves the original visual characteristics of the sign, but also ensures its accurate positioning in the lane-level spatial structure.
[0160] Based on the above optional embodiments, the embodiments of this application realize two high-precision, spatially aligned, and visually enhanced ways of expressing the semantic information of the target object, which not only ensures the clarity and readability of the prompt information, but also retains the authenticity and credibility of the original perception, thereby significantly improving the user's intuitive understanding and expectation consistency of autonomous driving behavior intentions in complex driving scenarios.
[0161] Optionally, the target page may also display a virtual navigation guide light carpet corresponding to the mobile device, wherein the virtual navigation guide light carpet is a driving plan channel for the mobile device in the future time period after the target object is identified.
[0162] The aforementioned virtual navigation guide light carpet represents a dynamically rendered, semi-transparent light strip on the target page, visually guiding the mobile device's future trajectory after recognizing the target object. The geometry of the virtual navigation guide light carpet is output by the path planning module of the autonomous driving domain controller. Its width corresponds to the lateral space required for safe vehicle avoidance, and its length covers the longitudinal time range from the current moment until the completion of deceleration, detour, or lane change. Its visual characteristics are a soft, gradually changing, low-brightness halo that does not obscure the road structure, serving only as a spatiotemporal extension of the behavioral intent.
[0163] The aforementioned future time period represents the time window from the current moment until the mobile device completes the vehicle control behavior triggered by the target object, such as decelerating to a safe speed or completing lane departure. Its duration is dynamically calculated by the vehicle dynamics model and environmental constraints, and is usually a continuous trajectory interval of 3 to 10 seconds.
[0164] The target page can display a virtual navigation guide light carpet corresponding to the mobile device. That is, in this embodiment of the application, the target page can further overlay a virtual navigation guide light carpet on top of the first environmental image and semantic prompt information. The aforementioned virtual navigation guide light carpet allows the driver to intuitively predict the trajectory changes that the vehicle is about to execute, especially in complex road conditions, providing a clear expectation of the direction of travel and avoiding misjudgment and takeover anxiety caused by sudden behavior.
[0165] A virtual navigation guide light carpet represents the planned driving path of a mobile device over a future time period after identifying a target object. Specifically, after the system senses the target object and generates a vehicle control decision, the path planning module outputs a sequence of future trajectory points. The SR display rendering module then maps these trajectory points into a continuous light strip with a specific width. The edges of the light carpet are softened, and its brightness decreases over time to indicate the dynamic evolution of the trajectory. This virtual navigation guide light carpet is always aligned with the vehicle's centerline and its curvature is adjusted in real time according to the steering angle to ensure geometric consistency with the actual driving path.
[0166] Based on the above optional embodiments, this application embodiment displays a virtual navigation guide light carpet corresponding to the mobile device on the target page. This virtual navigation guide light carpet is the driving planning channel of the mobile device in the future time period after recognizing the target object. This realizes the direct projection of the path planning intention onto the environment reconstruction interface in the form of a visual light domain, enabling the driver to perceive the spatiotemporal behavior trajectory of the vehicle after recognizing traffic signs in advance. This significantly enhances the intuitiveness and predictability of human-machine alignment, thereby effectively alleviating the psychological uncertainty caused by the invisibility of autonomous driving behavior without increasing cognitive load.
[0167] Optionally, the mobile device includes a first controller and a second controller. Displaying the first environment image on the target page of the mobile device includes: using the first controller to obtain a first environment image containing a target object, and sending the first environment image to the second controller; using the second controller to render the target page to display the first environment image on the target page of the mobile device.
[0168] The aforementioned first controller refers to the autonomous driving domain controller, which is the core computing unit in the mobile device responsible for environmental perception, target recognition, and behavior decision-making. It has a built-in image acquisition interface, OCR recognition model, and vehicle control decision-making logic, and includes a perception and decision-making module. The input source of the first controller is the original second environmental image captured by the vehicle-mounted camera, and the output includes the 2D bounding box coordinates of the target object, OCR-recognized text, vehicle control behavior decision flags, and the first environmental image after cropping and magnification.
[0169] The aforementioned second controller refers to the display control unit in the vehicle system. It is a computing module in the mobile device responsible for rendering the human-computer interaction interface and compositing multiple layers. It includes an SR display rendering module, which has graphics processing capabilities and an SR rendering engine. Its input is standardized image data and semantic information sent by the first controller, and its output is the visual content finally presented on the target page.
[0170] After acquiring a first environmental image containing the target object, the first controller sends the first environmental image to the second controller. That is, after recognizing the target object, the first controller crops and magnifies the original image by 1.2 times based on its 2D bounding box coordinates to generate a first environmental image that conforms to the display specifications. The first controller then transmits the image data packet and its attribute information, such as scene identification information, text identification type, and identification priority, to the second controller through an Ethernet communication link, realizing standardized data docking between perception output and display input.
[0171] The second controller renders the target page and displays the first environmental image on the target page of the mobile device. That is, after receiving the first environmental image, the second controller combines it with text prompts, virtual navigation guide light carpets and other SR layers, and performs layer overlay, color correction and frame synchronization processing through the graphics rendering engine, and finally outputs it to the target page display device, completing a complete closed loop from perception data to human-computer interaction visual presentation.
[0172] Figure 9 This is a schematic diagram of an optional mobile device according to an embodiment of this application, such as... Figure 9 The diagram illustrates a mobile device whose target object is a traffic sign. The first controller acquires a first environmental image containing the target object through a perception and decision-making module. The data processing module then performs cluster analysis to obtain corresponding text prompts. The first environmental image and text prompts are then sent to the second controller via Ethernet. The SR display rendering module in the second controller renders and displays the first environmental image, i.e., the target page, to achieve a complete human-computer interactive visual presentation.
[0173] Based on the above optional embodiments, this application embodiment uses a mobile device including a first controller and a second controller. The first controller acquires a first environmental image containing the target object and sends it to the second controller, which then renders the target page to display the first environmental image. This achieves physical separation and collaborative processing of perception computing and display rendering, ensuring high-quality, low-latency rendering of the reconstructed environmental image. At the same time, it avoids decision delays caused by graphics processing load in the perception decision module, ensuring stable operation of the assisted driving system and smooth human-machine interaction response in high-load scenarios.
[0174] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0175] Figure 10 This is a structural block diagram of an optional interactive device according to an embodiment of this application, such as... Figure 10 As shown, it should be noted that this device can be used to execute the above-described interaction method. The device includes: a first acquisition module 1001, used to acquire a first environmental image containing a target object, wherein the target object is one of the factors affecting the driving state of the mobile device; and a display module 1002, used to display target prompt information for the target object on a target page of the mobile device's corresponding display, wherein the target prompt information is generated based on the first environmental image.
[0176] Optionally, the first acquisition module 1001 is further configured to: acquire a second environmental image using the image acquisition sensor of the mobile device, wherein the second environmental image is used to reflect the driving environment of the mobile device; determine whether the second environmental image contains a target object; if the second environmental image contains a target object, determine a first environmental image based on the second environmental image and the target object, wherein the first environmental image is an image region in the second environmental image that at least includes the target object.
[0177] Optionally, the first acquisition module 1001 is further configured to: perform object recognition on the second environmental image to obtain at least one candidate object; and determine whether a target object exists among the at least one candidate object based on the category information and feature information corresponding to the at least one candidate object.
[0178] Optionally, if at least one of the following conditions is met, the first acquisition module 1001 is further configured to: if the category information of the candidate object is a static obstacle category, and the feature information of the candidate object satisfies a first feature condition; wherein the first feature condition includes: the object location corresponding to the candidate object is within the path coverage area of the current driving path of the mobile device, and the object size corresponding to the candidate object is greater than a preset size threshold; or, if the category information of the candidate object is a dynamic obstacle category, and the feature information of the candidate object satisfies a second feature condition; wherein the second feature condition includes one of the following: the object movement speed of the candidate object is greater than a preset speed threshold, and the distance between the candidate object and the mobile device decreases in the future time period; the predicted movement trajectory of the object corresponding to the candidate object intersects with the current driving path of the mobile device; or, if the category information of the candidate object is a traffic sign category, and the feature information of the candidate object satisfies a third feature condition; wherein the third feature condition includes: the object content and object location corresponding to the candidate object are associated with the driving state of the mobile device.
[0179] Optionally, the first acquisition module 1001 is further configured to: perform image processing operations on the second environment image based on the target object in the second environment image to obtain a first environment image, wherein the image processing operations include at least one of the following: performing image cropping operations on the target object in the second environment image, performing image annotation operations on the target object in the second environment image, and performing image scaling operations on the target object in the second environment image.
[0180] Optionally, the target prompt information includes: a first environmental image. The display module 1002 is further configured to: respond to the second environmental image including multiple target objects, determine the target object to be displayed based on the priority information of the multiple target objects, and display the first environmental image corresponding to the target object to be displayed on the target page, wherein the priority information is determined based on the degree of influence of the multiple target objects on the driving state of the mobile device, or the priority information is determined based on user-preset settings.
[0181] Optionally, the display is a head-up display or a mobile device screen; the target page is a scene-rendered map or a two-dimensional vector map.
[0182] Optionally, the target page may also include at least one of the following driving prompts: driving status information of the mobile device, decision prompts corresponding to the target object, and driving mode switching information. The driving status information is used to determine the activation status of the mobile device's assisted driving function, and the decision prompts corresponding to the target object include at least one of the following: environmental label information associated with the target object, driving decision information, and decision reasoning process information.
[0183] Optionally, at least one of the following methods may be used to display driving prompt information: displaying interface cards, displaying text pop-ups, or providing voice prompts; or highlighting or annotating the target object in the first environmental image.
[0184] Optionally, the interactive device further includes: a second acquisition module 1003, used to perform cluster analysis on the text content using the sign scene attribute information in response to the target object being a traffic sign and the target object containing text content, to obtain decision prompt information corresponding to the target object, wherein the sign scene attribute information is used to record the mapping relationship between multiple attribute fields, and the multiple attribute fields include: scene sign information, text sign type, sign priority, and prompt information corresponding to the target object.
[0185] Optionally, the interactive device further includes: an update module 1004, configured to: obtain an update configuration file from the server in response to the mobile device meeting preset state conditions, wherein the preset state conditions are used to indicate that the mobile device is in a powered-on state and in the parking gear; and update the identification scene attribute information based on the update configuration file in response to the update configuration file passing cyclic redundancy check.
[0186] Optionally, the target page is a scene rendering map, and the target prompt information includes: an identifier element or a first environment image. The display module 1002 is also used to: render the identifier element corresponding to the target object in the scene rendering map; or, display the first environment image in the first layer of the target page and display the scene rendering map in the second layer of the target page, wherein the first layer is the layer above the second layer.
[0187] Optionally, the target page may also display a virtual navigation guide light carpet corresponding to the mobile device, wherein the virtual navigation guide light carpet is a driving plan channel for the mobile device in the future time period after the target object is identified.
[0188] Embodiments of this application also provide a mobile device, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application during runtime.
[0189] Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0190] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0191] Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.
[0192] Embodiments of this application also provide a computer program that, when executed by a processor, implements the methods described in the various embodiments of this application.
[0193] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0194] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0195] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0196] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0197] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0198] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An interaction method, characterized in that, The method is applied to a mobile device, and the method includes: Acquire a first environmental image containing a target object, wherein the target object is one of the factors affecting the driving state of the mobile device; The target prompt information for the target object is displayed on the target page of the display corresponding to the mobile device, wherein the target prompt information is generated based on the first environmental image.
2. The interaction method according to claim 1, characterized in that, The step of acquiring the first environmental image containing the target object includes: A second environmental image is acquired using the image acquisition sensor of the mobile device, wherein the second environmental image is used to reflect the driving environment of the mobile device; Determine whether the target object is contained within the second environmental image; If the second environment image contains the target object, the first environment image is determined based on the second environment image and the target object, wherein the first environment image is an image region in the second environment image that includes at least the target object.
3. The interaction method according to claim 2, characterized in that, The step of determining whether the target object is contained within the second environmental image includes: Perform object recognition on the second environmental image to obtain at least one candidate object; Based on the category information and feature information corresponding to the at least one candidate object, determine whether the target object exists among the at least one candidate object.
4. The interaction method according to claim 3, characterized in that, The target object is determined to exist among the at least one candidate object if at least one of the following conditions is met: If the category information of the candidate object is a static obstacle, and the feature information of the candidate object satisfies a first feature condition; wherein the first feature condition includes: the location of the object corresponding to the candidate object is within the path coverage area of the current driving path of the mobile device, and the size of the object corresponding to the candidate object is greater than a preset size threshold; or... If the category information of the candidate object is a dynamic obstacle category, and the feature information of the candidate object satisfies a second feature condition; wherein the second feature condition includes one of the following: the object moving speed of the candidate object is greater than a preset speed threshold, and the distance between the candidate object and the mobile device decreases in the future time period; the predicted movement trajectory of the object corresponding to the candidate object intersects with the current driving path of the mobile device; or, If the category information of the candidate object is a traffic sign category, and the feature information of the candidate object satisfies the third feature condition; wherein, the third feature condition includes: the object content and object location corresponding to the candidate object are associated with the driving status of the mobile device.
5. The interaction method according to claim 2, characterized in that, Determining the first environment image based on the second environment image and the target object includes: Based on the target object in the second environmental image, an image processing operation is performed on the second environmental image to obtain the first environmental image, wherein the image processing operation includes at least one of the following: performing an image cropping operation on the target object in the second environmental image, performing an image annotation operation on the target object in the second environmental image, and performing image scaling processing on the target object in the second environmental image.
6. The interaction method according to claim 2, characterized in that, The target prompt information includes: the first environmental image, and the display of the target prompt information for the target object on the target page of the display corresponding to the mobile device includes: In response to the second environmental image including multiple target objects, a target object to be displayed is determined based on the priority information of the multiple target objects, and the first environmental image corresponding to the target object to be displayed is displayed on the target page, wherein the priority information is determined based on the degree of influence of the multiple target objects on the driving state of the mobile device, or the priority information is determined based on user-preset settings.
7. The interaction method according to claim 1, characterized in that, The display is a head-up display or a mobile device screen; the target page is a scene rendering map or a two-dimensional vector map.
8. The interaction method according to claim 1, characterized in that, The target page also includes at least one of the following driving prompt information: driving status information of the mobile device, decision prompt information corresponding to the target object, and driving mode switching information. The driving status information is used to determine the activation status of the mobile device regarding the assisted driving function. The decision prompt information corresponding to the target object includes at least one of the following: environmental label information associated with the target object, driving decision information, and decision reasoning process information.
9. The interaction method according to claim 8, characterized in that, The display method of the at least one driving prompt information includes at least one of the following: display of interface card, display of text pop-up window, and voice prompt; the target object is highlighted or marked in the first environmental image.
10. The interaction method according to claim 8, characterized in that, The methods for obtaining the decision-making prompt information corresponding to the target object include: In response to the target object being a traffic sign and the target object containing text content, cluster analysis is performed on the text content using the sign scene attribute information to obtain decision prompt information corresponding to the target object. The sign scene attribute information is used to record the mapping relationship between multiple attribute fields, including: scene sign information, text sign type, sign priority, and prompt information corresponding to the target object.
11. The interaction method according to claim 10, characterized in that, Before performing cluster analysis on the text content using the identifier scene attribute information, the method further includes: In response to the mobile device meeting preset state conditions, an updated configuration file is obtained from the server, wherein the preset state conditions are used to indicate that the mobile device is in a powered-on state and in the parking gear position; In response to the updated configuration file passing cyclic redundancy check, the identification scene attribute information is updated based on the updated configuration file.
12. The interaction method according to any one of claims 1 to 10, characterized in that, The target page is a scene-rendered map, and the target prompt information includes: an identifier element or the first environment image. Displaying the target prompt information for the target object on the target page of the mobile device's corresponding display includes: Render the identifier element corresponding to the target object in the scene rendering map; or... The first environment image is displayed in the first layer of the target page, and the scene rendering map is displayed in the second layer of the target page, wherein the first layer is the layer above the second layer.
13. The interaction method according to any one of claims 1 to 10, characterized in that, The target page is also used to display a virtual navigation guide light carpet corresponding to the mobile device, wherein the virtual navigation guide light carpet is a driving plan channel for the mobile device in the future time period after the target object is identified.
14. The interaction method according to claim 1, characterized in that, The mobile device includes: a first controller and a second controller, and displaying the first environmental image on the target page of the mobile device includes: The first controller is used to acquire the first environmental image including the target object, and the first environmental image is sent to the second controller. The second controller is used to render the target page to display the first environment image on the target page of the mobile device.
15. An interactive device, characterized in that, The device is used in a mobile device, and the device includes: An acquisition module is used to acquire a first environmental image containing a target object, wherein the target object is one of the factors affecting the driving state of the mobile device; The display module is used to display target prompt information for the target object on the target page of the display corresponding to the mobile device, wherein the target prompt information is generated based on the first environmental image.
16. A mobile device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 14.