Vehicle window intelligent display method, system and device, electronic equipment and storage medium
By collecting the driver's voice commands, converting them into image elements, and displaying the outlines of targets around the vehicle on the window, the problem of the driver's line of sight deviating when searching for targets outside the vehicle is solved, improving driving safety and target search efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA FAW CO LTD
- Filing Date
- 2023-06-15
- Publication Date
- 2026-04-24
AI Technical Summary
When drivers are looking for targets outside the vehicle, their eyes are easily deviated from the road ahead, affecting driving safety. Distraction also leads to a decrease in reaction ability, making it difficult to maintain attention to road conditions at the same time.
By collecting the driver's voice commands, analyzing the target and feature information, converting it into image elements, using the vehicle's onboard camera to capture images of the vehicle's surroundings, and displaying image outlines that match the features on the car window, the driver can be assisted in identifying targets outside the vehicle.
While maintaining driving safety, it improves the efficiency of finding targets and reduces the driver's burden. Through the display method that corresponds to both virtual and real objects, it makes it easier for the driver to control the driving trajectory and pay attention to objects outside the vehicle.
Smart Images

Figure CN116834663B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of displays, and more particularly to intelligent display methods, intelligent display systems, intelligent display devices, electronic devices, and storage media for vehicle windows. Background Technology
[0002] When driving, users often encounter situations where they need to look for people, shops, office buildings, road signs, and other targets outside the vehicle while driving. This leads to: their gaze frequently being deviated from the road ahead for extended periods, causing them to ignore road conditions and increasing the risk of traffic accidents; when using their eyes to identify and locate targets outside the vehicle, the brain allocates a significant amount of energy to identifying and locking onto the targets, neglecting driving operations and reducing the ability to react to emergencies, seriously affecting driving safety; and because the gaze needs to continuously search for targets outside the vehicle, the vehicle speed inevitably decreases for a longer period, easily causing congestion for vehicles behind.
[0003] Therefore, there is a need for a solution that can automatically assist drivers in searching for things around the vehicle based on voice commands, so that drivers can keep their attention on the road conditions ahead. Summary of the Invention
[0004] The purpose of this invention is to provide a method, system, device, electronic equipment and storage medium for intelligent display of vehicle windows, thereby solving at least one of the aforementioned technical problems.
[0005] This invention provides the following solution:
[0006] According to one aspect of the present invention, a method for intelligent display of vehicle windows is provided, the method comprising:
[0007] Collect speech information and analyze the information of the target to be searched and its features in the speech.
[0008] The pre-search target and its feature information are parsed into image element information;
[0009] The vehicle-mounted camera captures images of the area around the vehicle;
[0010] The captured images of the area around the vehicle are displayed on the car window from the driver's perspective.
[0011] The acquired images around the vehicle are segmented to generate multiple images of the objects to be identified;
[0012] Based on the pre-search target and the image elements of the pre-search target's features, scan the real-time acquired images around the vehicle;
[0013] If an image of a target object matching the image elements of the pre-searched target is scanned, a label is added to the corresponding image of the target object displayed on the car window, indicating that it is the pre-searched target.
[0014] Furthermore, the process of collecting voice information and parsing the information about the target and its features in the voice includes:
[0015] Analyze speech sentences containing feature information and target information;
[0016] Based on the parsed features pointing to the parsed target, determine the target to be searched and the features of the target to be searched.
[0017] Furthermore, the step of parsing the pre-search target and its feature information into image element information includes:
[0018] Based on the pre-search target described in the voice, select the example image corresponding to the pre-search target from the preset database;
[0019] Based on the features of the target to be searched in the speech description, the image features to be compared are marked on the example image;
[0020] Based on example images and labeled pre-comparison image features, determine the degree of similarity between the image to be identified and the example image;
[0021] If the objects in the image to be identified and the objects in the example image are classified as the same type of objects, then the image features of the image to be identified are extracted based on the pre-comparison image features.
[0022] If the image features of the object to be identified are more similar to the image features of the example image than a preset threshold, then the image to be identified is labeled and displayed as a target to be searched.
[0023] Furthermore, displaying the captured images of the vehicle's surroundings, corresponding to the driver's perspective, on the vehicle window includes:
[0024] Based on the segmented images of the area surrounding the vehicle, multiple images of the objects to be identified are generated, and the images are then processed in layers.
[0025] The layering process includes displaying the outline of the object image to be identified on the car window;
[0026] Add annotations to the image of the object to be identified displayed on the car window, and display the outline of the image of the object to be identified in a preset display style.
[0027] Furthermore, determining the similarity between the image to be identified and the example image based on the example image and the labeled pre-comparison image features includes:
[0028] Assign values to image features based on their identifiability;
[0029] Based on the assigned values of the pre-lookup target features, the similarity between the image of the object to be identified and the example image is determined.
[0030] Furthermore, if the objects in the image to be identified and the example image are classified as the same type of object, then the extraction of image features from the image to be identified based on the pre-comparison image features includes:
[0031] Based on the preset first-level image features, classify the types of objects in the image to be identified and the example image;
[0032] Based on the preset second-level image features, determine the degree of similarity between the image of the object to be identified and the example image;
[0033] Among them, based on the pre-compared image features, a second level of image features is preset.
[0034] According to a second aspect of the present invention, a vehicle window intelligent display system is provided, the vehicle window intelligent display system comprising: a target recognition module and a target display module;
[0035] The target recognition module is used to identify similar images of the target to be identified in the images around the vehicle based on the pre-search target parsed from the speech.
[0036] The target display module is used to display images of the vehicle's surroundings on the car window corresponding to the driver's perspective, including adding annotations similar to the pre-searched target to the image of the object to be identified on the car window display according to the pre-searched target parsed in the voice.
[0037] According to three aspects of the present invention, a smart display device for vehicle windows is provided, the smart display device for vehicle windows comprising:
[0038] The voice acquisition module is used to acquire voice information and parse the information of the target to be searched and its features in the voice.
[0039] The image conversion module is used to parse the pre-search target and its feature information into image element information;
[0040] Image acquisition module, used by the vehicle-mounted camera to acquire images of the area around the vehicle;
[0041] The image display module is used to display the captured images of the vehicle's surroundings on the windshield, corresponding to the driver's perspective.
[0042] The image segmentation module is used to segment the acquired images around the vehicle and generate multiple images of objects to be identified.
[0043] The image scanning module is used to scan real-time images of the area around the vehicle based on the image elements of the pre-search target and its features.
[0044] The image annotation module is used to add annotations to the corresponding image of the object to be identified displayed on the car window if an image of the object to be identified that matches the image elements of the pre-search target is scanned, indicating that it is the pre-search target.
[0045] According to four aspects of the present invention, an electronic device is provided, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0046] The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the intelligent display method for vehicle windows.
[0047] According to five aspects of the present invention, a computer-readable storage medium is provided, comprising: storing a computer program executable by an electronic device, wherein when the computer program is run on the electronic device, the electronic device performs the steps of the intelligent display method for vehicle windows.
[0048] The above solution achieves the following beneficial technical effects:
[0049] This application collects the driver's voice, analyzes the target the driver is looking for, and displays the target on the car window. This allows the driver to stay focused on the road conditions while reducing the burden of searching for the target, thus improving the efficiency of finding the target while maintaining vehicle safety.
[0050] This application displays objects around the vehicle on the car window from the driver's perspective, so that the objects outside the car observed by the driver are in a "virtual" correspondence with the real-time captured images. This makes it easier for the driver to control the driving trajectory and pay attention to the objects outside the car, and presents a positional change relationship corresponding to the driver's sense of movement, reducing the burden of interpreting image data for the driver.
[0051] This application displays the outline of an image on the car window while retaining most of the transparent portion of the window, thus maintaining the driver's first-person perspective while driving. By aligning the outline with the actual observed object, the influence of the image displayed on the car window on the driver is minimized, maintaining absolute control over the driver's line of sight. Attached Figure Description
[0052] Figure 1 This is a flowchart of a smart display method for vehicle windows provided by one or more embodiments of the present invention.
[0053] Figure 2 This is a structural diagram of a smart display device for vehicle windows provided in one or more embodiments of the present invention.
[0054] Figure 3This is a structural diagram of a vehicle window intelligent display system provided in one or more embodiments of the present invention.
[0055] Figure 4 This is a schematic diagram of the main process of displaying a target on a car window based on voice according to a specific embodiment of the present invention.
[0056] Figure 5 This is a schematic diagram of a system for displaying a target on a car window based on voice according to a specific embodiment of the present invention.
[0057] Figure 6 This is a schematic diagram of the target detection process according to a specific embodiment of the present invention.
[0058] Figure 7 This is a schematic diagram of a target detection process using a multimodal model according to a specific embodiment of the present invention.
[0059] Figure 8 This is a schematic diagram of the main process change using a multimodal model in a specific embodiment of the present invention.
[0060] Figure 9 This is a block diagram of an electronic device structure for a smart display method for vehicle windows provided in one or more embodiments of the present invention. Detailed Implementation
[0061] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] Figure 1 This is a flowchart of a smart display method for vehicle windows provided by one or more embodiments of the present invention.
[0063] like Figure 1 As shown, the intelligent display method for vehicle windows includes:
[0064] Step S1: Collect speech information and analyze the information of the target to be searched and its features in the speech.
[0065] Step S2: Parse the target to be searched and its feature information into image element information;
[0066] Step S3: The vehicle-mounted camera captures images of the area around the vehicle;
[0067] Step S4: Display the captured images of the vehicle's surroundings, corresponding to the driver's perspective, on the car window;
[0068] Step S5: Segment the acquired images around the vehicle to generate multiple images of the objects to be identified;
[0069] Step S6: Scan the real-time acquired images around the vehicle based on the image elements of the target to be searched and the features of the target to be searched.
[0070] Step S7: If an image of the object to be identified that matches the image elements of the pre-search target is scanned, then a label is added to the corresponding image of the object to be identified displayed on the car window, indicating that it is the pre-search target.
[0071] The above solution achieves the following beneficial technical effects:
[0072] This application collects the driver's voice, analyzes the target the driver is looking for, and displays the target on the car window. This allows the driver to stay focused on the road conditions while reducing the burden of searching for the target, thus improving the efficiency of finding the target while maintaining vehicle safety.
[0073] This application displays objects around the vehicle on the car window from the driver's perspective, so that the objects outside the car observed by the driver are in a "virtual" correspondence with the real-time captured images. This makes it easier for the driver to control the driving trajectory and pay attention to the objects outside the car, and presents a positional change relationship corresponding to the driver's sense of movement, reducing the burden of interpreting image data for the driver.
[0074] This application displays the outline of an image on the car window while retaining most of the transparent portion of the window, thus maintaining the driver's first-person perspective while driving. By aligning the outline with the actual observed object, the influence of the image displayed on the car window on the driver is minimized, maintaining absolute control over the driver's line of sight.
[0075] Specifically, while driving, a driver may need to identify objects outside the vehicle, such as locating passengers to pick up or a restaurant to stop at. People may stand on the roadside to avoid the vehicle, and restaurants may be located on either side of the road, not in the same direction as the driver's line of sight. Diverting the driver's attention to directions other than the road ahead is a dangerous driving behavior.
[0076] In-vehicle infotainment systems typically include a voice recognition module to recognize voice commands such as playing songs or directing navigation. Similarly, drivers can use the system's voice recognition function to analyze speech and identify the target they are looking for. However, unlike traditional voice recognition, the target a driver is looking for is an object with a "shape." Therefore, the voice description of the "target" needs to be converted into image data. Image comparison is then used to assist the driver in finding the target. During voice analysis, in addition to indicating the target, sufficiently rich features need to be added to the target so that the image analysis has enough features for comparison.
[0077] You can use a car camera to record objects outside the car and generate video files. Then, you can compare the video frames to identify the target the driver is looking for.
[0078] The real-time video feed recorded by the vehicle's camera needs to be displayed on the windshield to help the driver maintain their focus on the road ahead. However, the entire video feed cannot be displayed on the windshield, as this would obstruct too much of the driver's view and compromise driving safety.
[0079] It can display the outline of an object on the car window, keeping most of the windows from displaying image information, ensuring that the driver can see outside through the windows.
[0080] For example, existing car window display systems, such as HUD (Head-up Display) systems, can project only the outlines of objects in the image onto the car window.
[0081] For example, by recognizing nouns and adjectives in a speech sentence, the noun is used as the pre-search target, and the adjectives describing the noun are used as the pre-search target features. If the pre-search target and its features match successfully, the pre-search target can be converted into image data. For instance, an image corresponding to the pre-search target is found in an image database and used as an example. Then, using the features of the pre-search target, the pre-comparison image features are annotated on the example image.
[0082] The system can segment real-time images of the vehicle's surroundings into individual image units, each representing an image of an object to be identified. These images are then used as the target for identification, determining whether the driver's desired object exists within them.
[0083] The image of the object to be identified can be labeled according to the pre-comparison image features, and then the pre-search target image in the image of the object to be identified, i.e. the pre-search target in the voice command, can be filtered according to the image features.
[0084] When displayed on the car window, images of objects to be identified that are similar to the target image to be searched are labeled, for example, by displaying different colors on the outline.
[0085] In this embodiment, collecting voice information and parsing the information of the target to be searched and its features in the voice includes:
[0086] Analyze speech sentences containing feature information and target information;
[0087] Based on the parsed features pointing to the parsed target, determine the target to be searched and the features of the target to be searched.
[0088] Specifically, for example, the user's voice input, or unstructured request text, is converted into a structured semantic representation. This structured semantic representation consists of three parts: domain, intent, and slots. The domain is the scope of the user's request, the intent is the type of the user's request, and the slots are entities describing the user's request. For instance, a user requesting "Help me find a tall female student with a ponytail and a red top" has the domain "find target" (i.e., "help me" is the triggering instruction), the intent is "person" (i.e., the target to be searched), and the slots are "hairstyle = ponytail, top = red top, height = tall, gender = female, occupation = student" (i.e., the characteristics of the target to be searched).
[0089] In this embodiment, parsing the pre-search target and its feature information into image element information includes:
[0090] Based on the pre-search target described in the voice, select the example image corresponding to the pre-search target from the preset database;
[0091] Based on the features of the target to be searched in the speech description, the image features to be compared are marked on the example image;
[0092] Based on example images and labeled pre-comparison image features, determine the degree of similarity between the image to be identified and the example image;
[0093] If the objects in the image to be identified and the objects in the example image are classified as the same type of objects, then the image features of the image to be identified are extracted based on the pre-comparison image features.
[0094] If the image features of the object to be identified are more similar to the image features of the example image than a preset threshold, then the image to be identified is labeled and displayed as a target to be searched.
[0095] Specifically, while speech parsing into text provides readability, it cannot be directly used for image feature comparison. Instead, a sample image from an image database can be used to represent the target. For example, if the query is "Find me a tall female student with a ponytail and a red top," and the target is a "person," then an image of a "person" in the image database can be used as a sample. Furthermore, if the target possesses features such as "hairstyle = ponytail, top = red top, height = tall, gender = female, occupation = student," then the images to be identified can be filtered according to these pre-compare image features.
[0096] To reduce the computational load of recognition, feature comparison is performed in a tiered manner. First, objects are classified according to general object types. Then, the images of the objects to be recognized are compared against the pre-comparison image features. While the defined recognition features differ across systems during object type classification, their impact on the results is relatively small, and it's not necessary to strictly adhere to the pre-comparison image features when comparing the images of the objects to be recognized. For example, overly stringent feature comparisons are not required for images of "people." However, features specifically emphasized in the user's voice commands, such as "hairstyle = ponytail, top = red top, height = tall, gender = female, occupation = student," need to be compared against the images of the objects to be recognized one by one.
[0097] It's possible that the driver misspoke one or two features in their voice commands, resulting in a mismatch between the actual image of the target object and the actual image. It's also possible the user's memory is uncertain or ambiguous. To address this, the target features could be pre-assigned values, and the display on the car window could be adjusted based on whether the total score of multiple features exceeds a preset threshold.
[0098] In this embodiment, displaying the captured images of the vehicle's surroundings, corresponding to the driver's perspective, on the vehicle window includes:
[0099] Based on the segmented images of the area surrounding the vehicle, multiple images of the objects to be identified are generated, and the images are then processed in layers.
[0100] Layered processing includes displaying the outline of the object image to be identified on the car window;
[0101] Add annotations to the image of the object to be identified displayed on the car window, and display the outline of the image of the object to be identified in a preset display style.
[0102] Specifically, finding a target in an image surrounding a vehicle requires first segmenting that image. Since vehicle cameras record frames, each frame may contain multiple images of the target object. Directly labeling these frames with features would make it impossible to distinguish which part each feature belongs to during later comparisons. Therefore, the image surrounding the vehicle needs to be segmented into individual images of the target object to facilitate labeling with pre-comparison features. Maintaining a one-to-one feature comparison between the target image and the image of the target object prevents interference from other image features, resulting in a more accurate recognition process.
[0103] The images of the vehicle's surroundings, collected from the driver's perspective, are displayed on the car window. The images of objects to be identified are marked with outlines. In practice, this means that outlines are marked on objects outside the vehicle that the driver is actually observing.
[0104] The driver can recognize the person they are looking for by observing the outline outside the car. If the outline were displayed on a regular screen, the driver would need to switch from their current perspective to the camera's viewpoint, requiring a strong ability to mentally adapt to different perspectives, thus increasing the driver's mental workload. Therefore, displaying the outline on the car window is more beneficial for the driver's user experience.
[0105] In this embodiment, determining the similarity between the image to be identified and the example image based on the example image and the labeled pre-comparison image features includes:
[0106] Assign values to image features based on their identifiability;
[0107] Based on the assigned values of the pre-lookup target features, the similarity between the image of the object to be identified and the example image is determined.
[0108] Specifically, it's possible that the features described in the driver's speech may not perfectly match the features of the actual image of the object to be identified. For example, the driver might want to pick someone up at the station but doesn't know the person's current appearance or attire. Therefore, the features provided in the speech description contain considerable uncertainty. For instance, features like "hairstyle = ponytail, top = red top, height = tall, gender = female, occupation = student" are less reliable as it's difficult to determine whether the hairstyle or clothing has changed. Therefore, a higher weight could be assigned to features like "height = tall, gender = female, occupation = student," while a lower weight could be assigned to features like "hairstyle = ponytail, top = red top." A comprehensive scoring system could then be used to determine the degree of similarity.
[0109] Features can also be assigned values based on how easily they can be visualized for the target search. For example, the feature "occupation = student" is difficult to represent in the form of image data, but the features "hairstyle = ponytail, top = red top, height = tall, gender = female" are easy to represent in the form of image data and have slightly higher reliability. Therefore, a lower weight can be assigned to the feature "occupation = student", while a higher weight can be assigned to the features "hairstyle = ponytail, top = red top, height = tall, gender = female".
[0110] In this embodiment, if the object to be identified and the objects in the example image are classified as the same type of object, then the extraction of image features of the object to be identified based on the pre-comparison image features includes:
[0111] Based on the preset first-level image features, classify the types of objects in the image to be identified and the example image;
[0112] Based on the preset second-level image features, determine the degree of similarity between the image of the object to be identified and the example image;
[0113] Among them, based on the pre-compared image features, a second level of image features is preset.
[0114] Specifically, the first-level image feature setting is used to roughly classify the image of the object to be identified and the example image, preventing overly detailed feature information from obscuring basic features. For example, regarding the word "person," the example image might be male, while the image of the object to be identified might be female, but this does not prevent both from being classified as "person" for subsequent detailed comparison. However, if the basic feature of "person" is buried under detailed features such as "clothing, hairstyle, and color," the comparison results may be biased, making it difficult to distinguish the key features.
[0115] You can start by identifying relatively easy and basic types such as "people" or "women" to perform first-level image feature recognition, narrowing down the range of contrasting features, and then perform second-level image feature comparison.
[0116] Second-level image features can directly utilize pre-compare image features. This reduces computational load and increases recognition accuracy while narrowing the range of contrasting features using first-level image features. For example, using "object = human, hairstyle = ponytail, top = red top, height = tall, gender = female, occupation = student" directly as second-level image features can lead to errors if "object = human" is incorrectly identified. Conversely, using "object = human" as first-level image features and "hairstyle = ponytail, top = red top, height = tall, gender = female, occupation = student" as second-level image features ensures that even if some second-level image features are missing, it won't affect the recognition of "human." Using "object = human" as first-level image features results in a very low error rate.
[0117] Of course, "people" are not the only primary image feature; the scope of primary image features can be adjusted temporarily. For example, on a sunny day, the accuracy and reliability of clothing color recognition are high, so clothing color can be used as a primary image feature. However, features with too much detail, such as clothing patterns, are not suitable as primary image features.
[0118] Figure 2 This is a structural diagram of a smart display device for vehicle windows provided in one or more embodiments of the present invention.
[0119] like Figure 2 As shown, the intelligent display device for car windows includes: a voice acquisition module, an image conversion module, an image acquisition module, an image display module, an image segmentation module, an image scanning module, and an image annotation module;
[0120] The voice acquisition module is used to acquire voice information and parse the information of the target to be searched and its features in the voice.
[0121] The image conversion module is used to parse the target to be searched and its feature information into image element information;
[0122] Image acquisition module, used by the vehicle-mounted camera to acquire images of the area around the vehicle;
[0123] The image display module is used to display the captured images of the vehicle's surroundings on the windshield, corresponding to the driver's perspective.
[0124] The image segmentation module is used to segment the acquired images around the vehicle and generate multiple images of objects to be identified.
[0125] The image scanning module is used to scan real-time images of the area around the vehicle based on the image elements of the target to be searched and the features of the target.
[0126] The image annotation module is used to add annotations to the corresponding image of the object to be identified displayed on the car window if an image of the object to be identified that matches the image elements of the pre-search target is scanned, indicating that it is the pre-search target.
[0127] It is worth noting that although this system only discloses the voice acquisition module, image conversion module, image acquisition module, image display module, image segmentation module, image scanning module, and image annotation module, it does not mean that this device is limited to the above-mentioned basic functional modules. On the contrary, what this invention intends to express is that, based on the above-mentioned basic functional modules, those skilled in the art can arbitrarily add one or more functional modules in combination with existing technology to form an infinite number of embodiments or technical solutions. That is to say, this system is open rather than closed. It should not be assumed that the scope of protection of the claims of this invention is limited to the above-disclosed basic functional modules just because this embodiment only discloses a few basic functional modules.
[0128] Figure 3 This is a structural diagram of a vehicle window intelligent display system provided in one or more embodiments of the present invention.
[0129] like Figure 3 As shown, the intelligent display system for vehicle windows includes: a target recognition module and a target display module;
[0130] The target recognition module is used to identify similar images of the target to be identified in the images around the vehicle based on the pre-search target parsed from the speech.
[0131] The target display module is used to display images of the vehicle's surroundings on the car window corresponding to the driver's perspective, including adding annotations similar to the pre-searched target to the image of the object to be identified on the car window display according to the pre-searched target parsed in the voice.
[0132] Specifically, the pre-search target and its features for speech recognition are text-based data that need to be converted into image-based data before they can be used to compare and identify images of the vehicle's surroundings.
[0133] In terms of display method, only by displaying on the car window can the effect that other screen displays cannot achieve be achieved. That is, the objects outside the car currently observed by the driver are in a "virtual" corresponding relationship with the real-time collected images, which makes it easier for the driver to control the driving trajectory and pay attention to the objects outside the car in a positional relationship corresponding to the body sense, reducing the burden on the driver to interpret the image data.
[0134] Of course, the car windows are not limited to the windshield. Depending on the camera's field of view, the image of the object to be identified can also be displayed on the side windows. Although displaying the image of the object to be identified on the side windows is not conducive to the driver's forward observation of road conditions, the vehicle's automated image recognition process does not actually require the driver to constantly look at the side windows. The driver only needs to confirm the image when the image of the object to be identified is the target image to be searched.
[0135] Of course, displaying the image of the object to be identified on the vehicle's side window requires a corresponding camera perspective. For example, when passing through an alley, the image of the object to be identified in the surrounding area is captured as the target to be searched. At this moment, the image of the object to be identified is annotated and displayed on the vehicle's side window as appropriate.
[0136] It can issue an audible warning based on the scanned image of the object to be identified that matches the image elements of the pre-searched target, further reducing the visual burden and preventing the vehicle from missing the information that is marked with an outline on the windshield for reminder due to the dynamic movement of the vehicle.
[0137] Figure 4 This is a schematic diagram of the main process of displaying a target on a car window based on voice according to a specific embodiment of the present invention.
[0138] Figure 5 This is a schematic diagram of a system for displaying a target on a car window based on voice according to a specific embodiment of the present invention.
[0139] Figure 6 This is a schematic diagram of the target detection process according to a specific embodiment of the present invention.
[0140] Figure 7 This is a schematic diagram of a target detection process using a multimodal model according to a specific embodiment of the present invention.
[0141] Figure 8 This is a schematic diagram of the main process change using a multimodal model in a specific embodiment of the present invention.
[0142] like Figure 4 , 5As shown, a process is performed on the system. The driver speaks a voice command to find a target, such as "Find me a tall female student with a ponytail and wearing a red top" or "Find me a good Sichuan restaurant." The voice acquisition unit inputs the above voice command into the voice recognition module (which includes a fully trained voice recognition model and an intent understanding model). The voice recognition module determines whether the driver intends to find a target outside the vehicle. If so, it outputs keywords describing the entity requested by the user. Keywords describing the driver's request entity are input into the target matching module (including the matching algorithm). Simultaneously, the external image acquisition unit (including cameras, radar, etc.) begins real-time acquisition of video images within a certain range in front of the vehicle and inputs them into the target detection module. The target detection module (including a well-trained semantic segmentation model, image classification model, and feature-to-text model) extracts video frames from the acquired video images for semantic segmentation. The image classification algorithm labels all targets extracted by semantic segmentation, selects targets (multiple of which may belong to the same category as the driver's search intent), and further classifies the targets (the classification method is consistent with the slot classification of the speech intent model). The target image classification results (i.e., image features) are processed by the feature-to-text model to output all feature keywords of the target, and then the feature keywords are also input into the target matching module. The target matching module compares the characteristic keywords of the target with the keywords describing the entity requested by the driver and calculates a weighted score. When the score reaches a set threshold, the target is considered to match the driver's search intent. At the same time, all target categories output by the target detection module are input into the 3D model library, and 3D models of each target category are retrieved. The display module (including 3D rendering software) uses the 3D models selected from the model library and the three-dimensional position information of each target relative to the vehicle obtained by the external image acquisition unit for real-time 3D environment reconstruction of the images acquired by the image acquisition unit, and displays it on the windshield in front of the driver (which can be a HUD solution or the windshield itself as a screen). It also uses the relative position information and category information of the targets that match the search intent to determine the position of the targets in the 3D model and marks the 3D models of the targets that match the search intent with a specific color. The driver determines the location and distance of the target by comparing the relative positions of the target model marked with color on the car window with the model of the vehicle itself and other reference models in the surrounding area. Then, the driver looks at the location of the target in the real environment and makes a final manual judgment to determine whether the target has appeared. If it is not the target, the driver continues to wait for the next target marking or speaks a voice command containing more details about the target to start a new round of searching.
[0143] like Figure 6As shown, the target detection module extracts video frames from the acquired video images and performs semantic segmentation; the image classification model labels all targets extracted by semantic segmentation; targets of the same category as the driver's search intent (which can be multiple) are selected and further classified; the target image classification results (image features) are processed by the feature-to-text model to output all feature keywords of the target.
[0144] In speech recognition models, intent understanding transforms unstructured user-inputted request text into a structured semantic representation. This structured semantic representation comprises three parts: domain, intent, and slots. The domain defines the scope of the user's request, the intent defines the type of the request, and the slots describe the entities described in the request. For example, a user requesting "Find me a tall female student with a ponytail and a red top" has the domain "find target," the intent "person," and the slots "hairstyle = ponytail, top = red top, height = tall, gender = female, occupation = student." In this system, the "domain" only has two categories: "find target" and "other." Therefore, the main tasks of intent understanding are intent classification and entity recognition. Common methods for intent understanding models include, but are not limited to, rule-based methods and deep learning-based methods. Based on common scenarios of searching for targets outside a vehicle while driving, intents can be classified as people, cars, animals, shops, restaurants, tourist attractions, traffic signs, brand logos, etc. Different intents correspond to different slots. For example, the intent "person" corresponds to slots such as hairstyle, top, bottom, shoes, hat, height, body type, gender, age, occupation, glasses, accessories, posture, and carried items. The intent understanding model ultimately needs to extract slots to output keywords describing the entity requested by the user.
[0145] In the target matching module, a matching algorithm is used for quantification. The speech recognition module outputs keywords describing the entity requested by the user, such as: hairstyle = ponytail, top = red top, height = tall, gender = female, occupation = student. The target detection module outputs feature keywords for all targets in the image. Assuming there are two targets in the image classified as "person", the feature keywords for the first target are: hairstyle = ponytail, top = red top, height = tall, gender = female, occupation = unknown; and the feature keywords for the second target are: hairstyle = ponytail, top = pink top, height = tall, gender = female, occupation = student. For the category "person", the feature weights are defined as follows: "gender" 10%, "hairstyle" 10%, top 10%, height 5%, occupation 5%, bottom 10%, shoes 5%, hat 10%, body type 5%, age 10%, glasses 5%, posture 5%, accessories 5%, and carried items 5%. The keywords describing the user's requested entity output by the speech recognition module are compared with the feature keywords of each target in the image output by the object detection module. A match is scored as 1 point, and a non-match as 0 points. Then, the matching score of each keyword is multiplied by its weight, and finally, the scores are added together to calculate the final score. Because the keywords describing the user's requested entity may not necessarily cover all slots corresponding to the intent, the score also needs to be divided by the percentage of the sum of the weights of the covered keywords to 100% of the total weight. The comparison results for the first objective are as follows: Hairstyle matching scores 1 point multiplied by the hairstyle weight of 10%; Top matching scores 1 point multiplied by the top weight of 10%; Height matching scores 1 point multiplied by the height weight of 5%; Gender matching scores 1 point multiplied by the gender weight of 10%; Occupation mismatch scores 0 points. Adding these values together gives 0.35. Dividing this by the weight percentage of the covered keywords (0.4), the final score is 0.875. The comparison results for the second objective are as follows: Hairstyle matching scores 1 point multiplied by the hairstyle weight of 10%; Top mismatch scores 0 points; Height matching scores 1 point multiplied by the height weight of 5%; Gender matching scores 1 point multiplied by the gender weight of 10%; Occupation matching scores 1 point multiplied by the occupation weight of 5%. Adding these values together gives 0.30. Dividing this by the weight percentage of the covered keywords (0.4), the final score is 0.75. If the threshold is set to 0.8, the score of the first target (0.875 > 0.8) is considered a match for the search intent, while the score of the second target (0.75 < 0.8) is considered a mismatch for the search intent.
[0146] like Figure 7 As shown, the target matching module can be used not only with matching algorithms but also with multimodal models. If a multimodal model is used, the target detection module does not need to output feature keywords. The target detection module extracts video frames from the acquired video images for semantic segmentation; the image classification module labels all targets identified by semantic segmentation.
[0147] The keywords describing the user's requested entity output by the speech recognition module and all target images output by the object detection module are input into the object matching module (including the multimodal model) to calculate similarity and obtain a matching score. Targets with scores reaching a set threshold are considered to match the search intent. The multimodal model performs semantic similarity rating on the image-text, with the scoring criterion based on semantic text similarity. If a multimodal model is used, then... Figure 6 Adjustments were made to steps 4 and 5, as can be seen. Figure 8 .
[0148] In another embodiment, a practical application scenario is described. For example, Xiao Zhang is going to pick up a client from out of town, and they agree to meet on a certain road. Before setting off, he asks the client what he is wearing today, and the client says he is wearing a dark blue suit, a red tie, and carrying a brown leather bag. Because the agreed-upon route is in a relatively busy area, as the vehicle approaches the target route, Xiao Zhang issues a voice command to the vehicle: "Find me a man wearing a dark blue suit, a red tie, and carrying a brown leather bag." The intent understanding model converts the above command into a structured semantic representation, with the domain being "find target," the intent being "person," and the slots being "shirt = dark blue suit, gender = male, age = adult, items carried = brown leather bag, accessories = red tie." The vehicle immediately entered intelligent target recognition mode. When the vehicle found a target matching the search intent on the right front side of the road, the 3D model diagram on the driver's windshield showed that the target was 20 meters away on the right front side of the road the vehicle was traveling on. According to the positional relationship of each reference object in the 3D diagram, Xiao Zhang turned his gaze to the right side of the windshield and quickly found the customer he was looking for. Xiao Zhang safely and smoothly parked the car next to the customer.
[0149] In another embodiment, a real-world application scenario is described. For example, Xiao Zhang is going to have dinner with a blind date he's meeting for the first time. The restaurant the girl chose is on Beiyuan Road and is called "Haoweidao Sichuan Restaurant." Xiao Zhang has never been to this restaurant before. After setting off, Xiao Zhang issues the command "Find Haoweidao Sichuan Restaurant for me" to the vehicle. The intent understanding model converts the above command into a structured semantic representation, with the domain being "Find target," the intent being "restaurant," and the slot being "Name = Haoweidao Sichuan Restaurant." After driving for a short while, the 3D environment model on the car window marks the target location. Sure enough, it is Haoweidao Sichuan Restaurant, but it's quite far from Beiyuan Road. He thinks, "Oh, this restaurant is a chain. There's one so close to my home, and I've never been there before." After the vehicle reached Beiyuan Road, the target location was marked again on the car window. This should be the restaurant that the girl had booked. Although the navigation indicated that it had reached the vicinity of the destination and automatically ended the navigation, the restaurant was actually located on the side of the main road and was not easy to find. Fortunately, the intelligent search system helped him keep an eye on the outside of the car at all times. Otherwise, if he had missed it and had to drive to a place far away to turn around, he might have missed the agreed time.
[0150] Figure 9 This is a block diagram of an electronic device structure for a smart display method for vehicle windows provided in one or more embodiments of the present invention.
[0151] like Figure 9 As shown, this application provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0152] The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of a smart display method for vehicle windows.
[0153] This application also provides a computer-readable storage medium storing a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of a smart window display method.
[0154] This application also provides a vehicle, including:
[0155] Electronic devices, steps for implementing intelligent display methods for vehicle windows;
[0156] The processor runs a program, and when the program runs, it executes the steps of the intelligent display method for vehicle windows based on data output from electronic devices.
[0157] Storage medium for storing programs that, when running, execute steps of the intelligent window display method based on data output from electronic devices.
[0158] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0159] The electronic device comprises a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory. The operating system can be any one or more computer operating systems that control the electronic device through processes, such as Linux, Unix, Android, iOS, or Windows. Furthermore, in this embodiment of the invention, the electronic device can be a smartphone, tablet computer, or other handheld device, or a desktop computer, portable computer, or other electronic device; there is no particular limitation in this embodiment.
[0160] In this embodiment of the invention, the executing entity for electronic device control can be an electronic device itself, or a functional module within an electronic device capable of calling and executing a program. The electronic device can obtain the firmware corresponding to the storage medium. This firmware is provided by the supplier, and different storage media may have the same or different firmware; no limitation is made here. After obtaining the firmware corresponding to the storage medium, the electronic device can write this firmware into the storage medium; specifically, it burns the firmware corresponding to the storage medium into the storage medium. The process of burning the firmware into the storage medium can be implemented using existing technology, and will not be elaborated upon in this embodiment of the invention.
[0161] Electronic devices can also obtain reset commands corresponding to storage media. These reset commands are provided by the supplier, and the reset commands for different storage media can be the same or different, which is not limited here.
[0162] At this time, the storage medium of the electronic device is a storage medium on which the corresponding firmware has been written. The electronic device can respond to the reset command corresponding to the storage medium on which the corresponding firmware has been written, thereby resetting the storage medium on which the corresponding firmware has been written according to the reset command. The process of resetting the storage medium according to the reset command can be implemented by existing technology and will not be described in detail in this embodiment of the invention.
[0163] For ease of description, the above devices are described separately by function as various units and modules. Of course, in implementing this application, the functions of each unit and module can be implemented in one or more software and / or hardware.
[0164] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined.
[0165] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0166] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent display of vehicle windows, characterized in that, The intelligent display method for vehicle windows includes: Collect speech information and analyze the information of the target to be searched and its features in the speech. Among them, analyzing speech sentences containing feature information and target information; Based on the parsed features pointing to the parsed target, determine the target to be searched and the features of the target to be searched. The pre-search target and its feature information are parsed into image element information; The vehicle-mounted camera captures images of the area around the vehicle; The captured images of the area around the vehicle are displayed on the car window from the driver's perspective. The acquired images around the vehicle are segmented to generate multiple images of the objects to be identified; In this process, multiple images of objects to be identified are generated based on the segmented images of the area surrounding the vehicle, and the images are then processed in layers. The layering process includes displaying the outline of the object image to be identified on the car window; Add annotations to the image of the object to be identified displayed on the car window, and display the outline of the image of the object to be identified in the preset display style; Based on the pre-search target and the image elements of the pre-search target's features, scan the real-time acquired images around the vehicle; Among them, based on the pre-search target described by voice, an example image corresponding to the pre-search target is selected from a preset database; Based on the features of the target to be searched in the speech description, the image features to be compared are marked on the example image; Based on example images and labeled pre-comparison image features, determine the degree of similarity between the image to be identified and the example image; If the objects in the image to be identified and the objects in the example image are classified as the same type of objects, then the image features of the image to be identified are extracted based on the pre-comparison image features. If the image features of the object to be identified and the image features of the example image have a similarity exceeding a preset threshold, then the image of the object to be identified is labeled and displayed as the target to be searched. If an image of a target object matching the features of the pre-searched target is found, a label is added to the corresponding image of the target object displayed on the car window, indicating that it is the pre-searched target.
2. The intelligent display method for vehicle windows according to claim 1, characterized in that, The step of determining the similarity between the image to be identified and the example image based on the example image and the labeled pre-comparison image features includes: Assign values to image features based on their identifiability; Based on the assigned values of the pre-lookup target features, the similarity between the image of the object to be identified and the example image is determined.
3. The intelligent display method for vehicle windows according to claim 2, characterized in that, If the objects in the image to be identified and the example image are classified as the same type of object, then the image features of the image to be identified extracted based on the pre-comparison image features include: Based on the preset first-level image features, classify the types of objects in the image to be identified and the example image; Based on the preset second-level image features, determine the degree of similarity between the image of the object to be identified and the example image; Among them, based on the pre-compared image features, a second level of image features is preset.
4. A vehicle window intelligent display system, characterized in that, Based on the intelligent display method for vehicle windows according to any one of claims 1 to 3, the intelligent display system for vehicle windows includes: a target recognition module and a target display module; The target recognition module is used to identify similar images of the target to be identified in the images around the vehicle based on the pre-search target parsed from the speech. The target display module is used to display images of the vehicle's surroundings on the car window corresponding to the driver's perspective, including adding annotations similar to the pre-searched target to the image of the object to be identified on the car window display according to the pre-searched target parsed in the voice.
5. A smart display device for vehicle windows, characterized in that, Based on the intelligent display method for vehicle windows according to any one of claims 1 to 3, the intelligent display device for vehicle windows includes: The voice acquisition module is used to acquire voice information and parse the information of the target to be searched and its features in the voice. The image conversion module is used to parse the pre-search target and its feature information into image element information; Image acquisition module, used by the vehicle-mounted camera to acquire images of the area around the vehicle; The image display module is used to display the captured images of the vehicle's surroundings on the windshield, corresponding to the driver's perspective. The image segmentation module is used to segment the acquired images around the vehicle and generate multiple images of objects to be identified. The image scanning module is used to scan real-time images of the area around the vehicle based on the image elements of the pre-search target and its features. The image annotation module is used to add annotations to the corresponding image of the object to be identified displayed on the car window if an image of the object to be identified that matches the image elements of the pre-search target is scanned, indicating that it is the pre-search target.
6. An electronic device, characterized in that, include: The processor, communication interface, memory, and communication bus are connected, with the processor, communication interface, and memory communicating with each other via the communication bus. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the intelligent display method for vehicle windows according to any one of claims 1 to 3.
7. A computer-readable storage medium, characterized in that, include: It stores a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of the intelligent display method for vehicle windows as described in any one of claims 1 to 3.
Citation Information
Patent Citations
A head-up display system based on a reflex panoramic camera
CN107656374A
Vehicle-mounted image feature automatic identification method and system based on cloud computing
CN110647658A
Semantic association character recognition method and device
CN113221904A
Voice-based image generation method and device, equipment and storage medium
CN113256751A