Target object query method and device, medium and equipment

By using a camera in the vehicle to collect images of the window area, determine the object projection profile and the target position of the user's line of sight, and lock the target object that the user intends to query, solving the accuracy of target object query in the vehicle, achieving efficient user experience improvement.

CN120067466APending Publication Date: 2025-05-30XG TECHNOLOGIES PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510061786.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

During the vehicle driving, how to effectively lock the target object that the user intends to query and provide the user with corresponding query result information has become a technical problem that needs to be solved urgently.

Method used

By acquiring the first image in the preset window area collected by the first camera in the vehicle, the projection profile information of each object in the image on the preset window area is determined, the target position information of the user's line of sight on the preset window area is detected, and the target object corresponding to the user's line of sight is determined based on the projection profile information and the target position information, and the corresponding query result information is output.

Benefits of technology

It realizes accurate locking of the target object that the user's line of sight is focused on, ensures the accuracy of the target object, and thus provides users with relevant information query results and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067466A_ABST
    Figure CN120067466A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a target object query method and device, a medium and equipment, and the method comprises the steps: obtaining a first image, collected by a first camera in a vehicle, in a preset vehicle window region; determining projection contour information of each object in the first image on a preset vehicle window area; detecting target position information corresponding to the sight of the user in the vehicle on the preset vehicle window area; determining a target object corresponding to the sight line of the user based on the projection contour information and the target position information of the object; and outputting query result information corresponding to the target object. According to the embodiment of the invention, the target object watched by the sight line of the user can be effectively locked, the accuracy of the target object is ensured, the user can inquire the related information of the object through the sight line watching, and the user experience is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to computer vision technology, and in particular, to a method, apparatus, medium, and device for querying target objects. Background Art

[0002] During the driving of a vehicle, a user in the vehicle (such as a driver, a co-pilot user, etc.) can query information about surrounding objects of interest. How to effectively lock the target object that the user intends to query to provide the corresponding query result information for the user has become a technical problem that urgently needs to be solved. Summary of the Invention

[0003] Embodiments of the present disclosure provide a method, apparatus, medium, and device for querying target objects to effectively lock the target object that the user intends to query based on the user's line of sight and improve the accuracy of the locked target object.

[0004] In a first aspect of the embodiments of the present disclosure, a method for querying a target object is provided, including: obtaining a first image in a preset window area collected by a first camera in the vehicle; determining projection contour information of each object in the first image on the preset window area; detecting target position information corresponding to the line of sight of a user in the vehicle on the preset window area; determining the target object corresponding to the user's line of sight based on the projection contour information of the object and the target position information; and outputting query result information corresponding to the target object.

[0005] In a second aspect of the embodiments of the present disclosure, a device for querying a target object is provided, including: a first processing module for obtaining a first image in a preset window area collected by a first camera in the vehicle; a second processing module for determining projection contour information of each object in the first image on the preset window area; a third processing module for detecting target position information corresponding to the line of sight of a user in the vehicle on the preset window area; a fourth processing module for determining the target object corresponding to the user's line of sight based on the projection contour information of the object and the target position information; and an output module for outputting query result information corresponding to the target object.

[0006] In a third aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The storage medium stores a computer program, and the computer program is used to execute the method for querying a target object according to any one of the above embodiments of the present disclosure.

[0007] In a fourth aspect of the embodiments of the present disclosure, an electronic device is provided, the electronic device includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the query method for the target object described in any one of the above embodiments of the present disclosure.

[0008] In a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, when the instructions in the computer program product are executed by a processor, the query method for the target object provided in any one of the above embodiments of the present disclosure is executed.

[0009] Based on the query method, device, medium, and device for the target object provided in the above embodiments of the present disclosure, during the driving of the vehicle, since the user's line of sight can be projected onto the preset window area, the objects outside the window can be viewed through the preset window area, and the objects outside the window can also be projected onto the preset window area. Thus, a first image within the preset window area is collected by a first camera, and the objects outside the window will be presented in the first image. By projecting each object onto the preset window area through the first image, combining the relative position relationship between the projection position of the user's line of sight on the preset window area and the projection contour of the object on the preset window area, the target object being stared at by the user's line of sight can be effectively locked, ensuring the accuracy of the target object. Furthermore, the query result information corresponding to the target object is output, enabling the user to query the relevant information of the object by staring with the line of sight, effectively improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is an exemplary application scenario of the query method for the target object provided by the present disclosure;

[0011] Figure 2 is a schematic flowchart of the query method for the target object provided by an exemplary embodiment of the present disclosure;

[0012] Figure 3 is a schematic flowchart of the query method for the target object provided by another exemplary embodiment of the present disclosure;

[0013] Figure 4 is a schematic flowchart of the query method for the target object provided by still another exemplary embodiment of the present disclosure;

[0014] Figure 5 is a schematic flowchart of the query method for the target object provided by yet another exemplary embodiment of the present disclosure;

[0015] Figure 6 is a schematic flowchart of the query method for the target object provided by still another exemplary embodiment of the present disclosure;

[0016] Figure 7It is a flowchart of a method for querying a target object provided by an exemplary embodiment of the present disclosure;

[0017] Figure 8 It is a schematic structural diagram of a query device for a target object provided by an exemplary embodiment of the present disclosure;

[0018] Figure 9 It is a schematic structural diagram of a query device for a target object provided by another exemplary embodiment of the present disclosure;

[0019] Figure 10 It is a schematic structural diagram of a query device for a target object provided by still another exemplary embodiment of the present disclosure;

[0020] Figure 11 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0021] To explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments. It should be understood that the present disclosure is not limited by the exemplary embodiments.

[0022] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.

[0023] Overview of the present disclosure

[0024] In the process of implementing the present disclosure, the inventors found that during the driving of a vehicle, users in the vehicle (such as the driver, the co-pilot user, etc.) can query information about surrounding objects of interest. How to effectively lock the target object that the user intends to query in order to provide the user with corresponding query result information has become a technical problem that urgently needs to be solved.

[0025] Exemplary overview

[0026] Figure 1 It is an exemplary application scenario of the method for querying a target object provided by the present disclosure. As Figure 1As shown, user 11 is inside vehicle 12. A first camera 13 is provided inside vehicle 12 near the eye position of user 11. For example, the first camera 13 can be set above the head of user 11. The field of view angle of the first camera 13 covers a preset window area 14 of vehicle 11 and can collect images within the preset window area 14. Since the preset window area 14 is a transparent area, the images are essentially images of the physical space outside the vehicle. That is, objects outside the vehicle (such as object A, object B, and object C in the figure) will appear in the images. Using the method for querying target objects in the embodiments of the present disclosure, a first image within the preset window area 14 collected by the first camera 13 inside vehicle 12 can be obtained; then, the projection contour information of each object in the first image on the preset window area 14 can be determined. For example, the projection contour b of object B in the figure on the preset window area 14; the target position information corresponding to the line of sight of user 11 inside vehicle 12 on the preset window area can be detected, such as the target position P in the figure. Based on the projection contour information of the object and the target position information, the target object corresponding to the line of sight of user 11 can be determined. In the figure, the target object is object B; then, the query result information corresponding to the target object can be output. The query result information can, for example, include the name, type, and other descriptive information corresponding to the target object. For example, if the target object is a building, the query result information can include the name, function, construction history, etc. of the building. For example, the building is a certain convention center and is mainly used for holding a certain type of meeting. Again, for example, if the target object is a vehicle, the query result information can include the brand, price, displacement, etc. of the vehicle. The specific query result information is not limited. In the method for querying target objects in the embodiments of the present disclosure, since the user's line of sight can be projected onto the preset window area to view objects outside the vehicle through the preset window area, and the objects outside the vehicle can also be projected onto the preset window area, the first image within the preset window area can be collected by the first camera. The objects outside the vehicle will appear in the first image. By projecting each object onto the preset window area through the first image and combining the relative position relationship between the projection position of the user's line of sight on the preset window area and the projection contour of the object on the preset window area, the target object being stared at by the user's line of sight can be effectively locked, ensuring the accuracy of the target object. Then, the query result information corresponding to the target object is output, enabling the user to query relevant information of the object by staring with the line of sight and effectively improving the user experience.

[0027] Exemplary method

[0028] Figure 2 It is a flowchart of a method for querying target objects provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, specifically, for example, on a vehicle computing platform (or vehicle terminal), such as Figure 2 As shown, the method of the embodiments of the present disclosure can include the following steps:

[0029] Step 210: Obtain a first image within a preset window area captured by a first camera inside the vehicle.

[0030] Among them, the first camera can be set at a position inside the vehicle close to the user's eyes, such as above the user's head, and the specific position is not limited. The field of view angle of the first camera covers the preset window area of the vehicle. The preset window area can be the front window area of the vehicle. Through the first camera, a first image within the preset window area can be captured. Since the preset window area is a transparent area, the first camera can capture the environmental image outside the vehicle (such as in front of the vehicle) through the preset window area, so that the objects outside the vehicle can be presented in the first image.

[0031] In some optional embodiments, the first camera can communicate with the in-vehicle terminal and send the captured first image to the in-vehicle terminal in real time, and the in-vehicle terminal can then obtain the first image within the preset window area captured by the first camera.

[0032] Step 220: Determine the projection contour information of each object in the first image on the preset window area.

[0033] Among them, each object in the first image is an object outside the vehicle. For example, during the driving of the vehicle on the road, the objects on the road in front of the vehicle (outside the front window) and on both sides of the road, such as Figure 1 where object A and object B can be buildings on both sides of the road, and object C can be other vehicles on the road. The projection contour information of the object on the preset window area is the contour information obtained by projecting the object contour in the first image onto the plane where the preset window area is located according to a certain projection conversion relationship. The projection contour information can be expressed as a projection contour box or a set of projection contour points, and the specific form is not limited.

[0034] In some optional embodiments, the projection conversion relationship between the first image and the plane of the preset window area can be determined according to the position association relationship between the first camera and the user's eyes, so as to simulate the projection of the object on the preset window area when the first camera imitates the user's eyes looking through the preset window area at the objects outside the vehicle.

[0035] Step 230: Detect the target position information corresponding to the user's line of sight in the vehicle on the preset window area.

[0036] Among them, the user inside the vehicle can be the driver, the co-pilot user, etc. The target position information corresponding to the user's line of sight on the preset window area refers to the position information where the user's line of sight projects onto the preset window area.

[0037] In some alternative embodiments, the line-of-sight information of the user can be detected by any line-of-sight detection method. The line-of-sight information may include, for example, the starting point of the user's line of sight (the position of the eyes or pupils) and the line-of-sight direction. According to the relative relationship between the user's line-of-sight information and a preset window area, the target position information corresponding to the user's line of sight on the preset window area is determined. For example, the user's line-of-sight information and the plane where the preset window area is located can be unified into the same coordinate system in the vehicle interior space, and the target position information is determined according to the intersection relationship between the line and the plane.

[0038] In some alternative embodiments, the line-of-sight information of the user can be obtained through one or more of a Driver Monitor System (DMS for short), an Occupancy Monitoring System (OMS for short), etc.

[0039] In some alternative embodiments, a second camera can be arranged inside the vehicle. The field of view angle of the second camera covers the user's face area. A second image including the user's eyes is collected through the second camera, and then the line-of-sight information of the user inside the vehicle is detected based on the second image. According to the line-of-sight information of the user, the target position information corresponding to the user's line of sight on the preset window area is determined.

[0040] It should be noted that the operations of determining the projection contour information in steps 210 to 220 and the operation of determining the target position information in step 230 are not executed in a specific order.

[0041] Step 240: Based on the projection contour information of the object and the target position information, determine the target object corresponding to the user's line of sight.

[0042] Among them, the projection contour information of the object represents the projection contour of the object on the preset window area when the first camera simulates the user's eyes looking at the object outside the vehicle. The target position information represents the projection position of the user's line of sight on the preset window area. Therefore, the target object corresponding to the user's line of sight can be determined based on the projection contour information and the target position information of each object.

[0043] In some alternative embodiments, the target object corresponding to the user's line of sight can be determined based on the relative position relationship between the target position information and the projection contour information of each object. For example, the target position information can be matched with the projection contour information of each object to determine whether the target position corresponding to the target position information is within the object projection contour. If the target position is within the projection contour, it means that the object being looked at by the user's line of sight is this object, and then this object is the target object corresponding to the user's line of sight. If the target position is not within the projection contour, it means that this object is not the target object being looked at by the user.

[0044] Step 250, output the query result information corresponding to the target object.

[0045] Among them, the query result information corresponding to the target object is the description information describing the target object obtained according to the target object gazed at by the user. Different objects can correspond to different description information. For example, the query result information can include the name, type, and other description information corresponding to the target object. Taking the target object as a building as an example, the query result information can include relevant information such as the name, function, and construction history of the building. The specific content of the query result information of the target object is not limited.

[0046] In some alternative embodiments, the query result information corresponding to the target object can be obtained through a search engine provided on the vehicle, and then the query result information corresponding to the target object is output.

[0047] In some alternative embodiments, the output manner of the query result information can be any manner. For example, the query result information can be output through voice, or output through voice + screen display, and so on.

[0048] In some alternative embodiments, after determining the target object gazed at by the user, further interaction with the user can be performed to confirm whether the locked target object is the object that the user wants to query. After the user confirms, the query result information corresponding to the target object is output. Interaction with the user can also be performed to obtain which aspects of information about the target object the user intends to query, so as to output the query result information required by the user according to the user's intention.

[0049] In the query method of the target object provided in this embodiment, during the driving of the vehicle, since the user's line of sight can be projected onto the preset window area, the objects outside the vehicle window can be viewed through the preset window area, and the objects outside the vehicle window can also be projected onto the preset window area. Thus, the first image in the preset window area is collected by the first camera, and the objects outside the vehicle window will appear in the first image. The objects are projected onto the preset window area through the first image to simulate the scene where the user's eyes observe the objects outside the vehicle window. Combining the relative position relationship between the projection position of the user's line of sight on the preset window area and the projection contour of the object on the preset window area, the target object gazed at by the user can be effectively locked, ensuring the accuracy of the target object. Furthermore, the query result information corresponding to the target object is output, enabling the user to query the relevant information of the object by gazing, effectively improving the user experience.

[0050] Figure 3 It is a flowchart of the query method of the target object provided by another exemplary embodiment of the present disclosure.

[0051] In some alternative embodiments, on the basis of the above Figure 2 shown embodiment, asFigure 3 As shown in Figure 3 , determining the projected contour information of each object in the first image on the preset window area in step 220 may include:

[0052] Step 2210, based on the first image, determine the object contour information corresponding to each object in the first image.

[0053] Among them, the object contour information is the contour information of the object in the first image, that is, the contour information of the object in the image coordinate system.

[0054] In some optional embodiments, the object contour information may be represented as the detection box of the object obtained by performing object detection on the first image, or may be represented as the bounding box of the object pixel point set obtained by performing semantic segmentation on the first image, or the object contour information may be represented as the target box obtained by expanding the detection box or the bounding box outward by a certain proportion. The specific representation method of the object contour information is not limited.

[0055] In some optional embodiments, the object contour information may be represented as a set of contour points of the object in the first image.

[0056] In some optional embodiments, semantic segmentation may be performed on the first image to obtain the pixel point sets corresponding to at least one object in the first image; for any object among the at least one object, based on the pixel point set corresponding to the object, determine the object contour information of the object in the first image. Optionally, a pre-trained semantic segmentation model may be used to perform semantic segmentation on the first image to obtain the pixel point sets corresponding to each object. The specific network structure of the semantic segmentation model is not limited.

[0057] In some optional embodiments, object detection may be performed on the first image to obtain the detection boxes corresponding to each object in the first image, and based on the detection boxes of each object, determine the object contour information corresponding to each object. Optionally, a pre-trained object detection model may be used to perform object detection on the first image to obtain the detection boxes corresponding to each object. The specific network structure of the object detection model is not limited.

[0058] Step 2220, project the object contour information corresponding to each object onto the preset window area to obtain the projected contour information corresponding to each object.

[0059] In some optional embodiments, the object contour information corresponding to each object may be projected onto the preset window area based on the projection conversion relationship between the first image and the preset window area to obtain the projected contour information corresponding to each object.

[0060] In some alternative embodiments, the projection conversion relationship may be determined according to the relative positional relationship between the first camera and the preset window area in the vehicle interior space. For example, based on the position and orientation of the image coordinate system of the first camera in the vehicle interior space and the position and orientation of the plane where the preset window area is located in the vehicle interior space, the projection conversion relationship between the first image and the preset window area is determined.

[0061] In some alternative embodiments, the projection conversion relationship may be determined according to the relative positional relationship between the first camera, the user's eyes, and the preset window area in the vehicle interior space. For example, there is a certain deviation between the viewing angle of the first camera and the viewing angle of the user's eyes. The projection conversion relationship may be corrected by combining the relative positional relationship between the first camera and the user's eyes in the vehicle interior space, so as to reduce the deviation of each object observed by the first camera and the user's eyes, and further improve the accuracy of locking the target object. Specifically, a virtual camera coordinate system (or virtual camera coordinate system) may be established at the position of the user's eyes. According to the conversion relationship between the first camera coordinate system and the virtual camera coordinate system, and the projection relationship between the image coordinate system of the virtual camera and the preset window area, the projection conversion relationship between the first image and the preset window area is determined to ensure that the projected contour information of the object is consistent with or has a small deviation from the contour of the object observed by the user's eyes in the preset window area. For example, the first camera and the virtual camera may be used as a binocular camera, and the target contour corresponding to the object contour in the first image in the image coordinate system of the virtual camera is determined through parallax, and then the target contour is projected onto the preset window area in the virtual camera coordinate system.

[0062] In some alternative embodiments, for different positional relationships between the user's eyes and the first camera, the corresponding projection conversion relationships may be pre-calibrated. After the vehicle is started, the positional relationship between the current user's eyes and the first camera may be determined first, and the projection conversion relationship corresponding to the current user is obtained. During the driving of the vehicle, according to the projection conversion relationship, the objects in the first image may be projected onto the preset window area in real time to obtain the projected contour information of each object.

[0063] In the embodiments of the present disclosure, by determining the object contour information corresponding to each object in the first image and then projecting it onto the preset window area, the projected contour information corresponding to each object is obtained, providing effective object projected contour information for locking the target object gazed by the user.

[0064] In some alternative embodiments, projecting the object contour information corresponding to each object onto the preset window area in step 2220 to obtain the projected contour information corresponding to each object may include:

[0065] For any object, based on a pre-configured first mapping relationship from a first image to a preset window area, project the object contour information corresponding to the object onto the preset window area to obtain the projected contour information corresponding to the object.

[0066] Among them, the first mapping relationship from the first image to the preset window area is the above-mentioned projection conversion relationship between the first image and the preset window area. For details, reference can be made to the foregoing embodiments and will not be elaborated herein.

[0067] In the embodiments of the present disclosure, through the pre-configured first mapping relationship from the first image to the preset window area, effective projection of the object contour information in the image to the preset window area is realized, providing effective projected contour information for locking the target object gazed at by the user's line of sight.

[0068] In some optional embodiments, the step of projecting the object contour information corresponding to each object onto the preset window area to obtain the projected contour information corresponding to each object in step 2220 may include:

[0069] For the object contour information corresponding to any object, determine the rays emitted by the optical center of the first camera to the multiple contour points corresponding to the object contour information; based on the intersection points of the rays corresponding to each contour point with the preset window area, determine the projected contour information corresponding to the object.

[0070] Among them, the multiple contour points corresponding to the object contour information are pixel points on the first image, and the rays emitted by the optical center of the first camera to the multiple contour points corresponding to the object contour information refer to the rays emitted from the optical center of the first camera to the pixels corresponding to each contour point in the image coordinate system. The intersection points of the rays corresponding to each contour point with the preset window area are equivalent to taking the preset window area as the imaging plane and the contour points of the object on this imaging plane. Therefore, based on the intersection points of the rays corresponding to each contour point with the preset window area, the projected contour information of the object in the preset window area can be obtained.

[0071] In some optional embodiments, based on the intersection points of the rays corresponding to each contour point with the preset window area, combined with the virtual camera corresponding to the position of the user's eyes and the first camera, the projected contour information corresponding to the object can be comprehensively determined to further improve the consistency between the projected contour information and the contour of the object observed by the user's eyes. That is, considering the parallax between observing the object from the user's eyes and from the first camera, the projected contour information is made closer to the contour of the object observed from the user's eyes.

[0072] In the embodiments of the present disclosure, through the rays emitted by the optical center of the first camera to the object contour points, effective projection of the object contour information in the image to the preset window area is realized, providing effective projected contour information for locking the target object gazed at by the user's line of sight.

[0073] Figure 4 It is a schematic flowchart of a method for querying a target object provided by another exemplary embodiment of the present disclosure.

[0074] In some optional embodiments, based on any of the above embodiments, as Figure 4 shown, the detecting of the target position information corresponding to the line of sight of the user in the vehicle on the preset window area in step 230 may include:

[0075] Step 2310, determining a second image including the user's eyes collected by a second camera in the vehicle.

[0076] Wherein, the second camera may be a camera arranged in the vehicle for monitoring the user, and the viewing direction of the second camera faces the user to cover the user's face area. For example, the second camera may be an OMS camera in the vehicle.

[0077] Step 2320, based on the second image, determining the position of the user's eyes and the line of sight direction of the eyes looking at the preset window area.

[0078] Wherein, the line of sight of the user may be detected based on the second image to obtain the position of the user's eyes and the line of sight direction of the eyes looking at the preset window area.

[0079] In some optional embodiments, the position of the user's eyes and the line of sight direction may be the position and direction in the second camera coordinate system, or may be the position and direction after being converted to a pre-calibrated coordinate system in the vehicle space, which is not specifically limited. The pre-calibrated coordinate system may be a coordinate system with a preset position in the vehicle as the origin. And the plane position and orientation of the preset window area in the pre-calibrated coordinate system may be pre-calibrated. The position and line of sight direction in the second camera coordinate system may be converted to the pre-calibrated coordinate system according to the conversion relationship between the second camera coordinate system and the pre-calibrated coordinate system.

[0080] In some optional embodiments, the pre-calibrated coordinate system may be the camera coordinate system of the first camera or the second camera, which is not specifically limited.

[0081] In some optional embodiments, the pre-calibrated coordinate system may be the camera coordinate system of the first camera or the second camera, which is not specifically limited.

[0082] Step 2330, based on the position of the user's eyes and the line of sight direction, determining the target position information corresponding to the user's line of sight on the preset window area.

[0083] Among them, based on the position and line-of-sight direction of the user's eyes, the user's line-of-sight ray can be determined, and according to the intersection point of the line-of-sight ray and the preset window area, the target position information corresponding to the user's line of sight on the preset window area can be determined. Specifically, the position of the user's eyes can be used as the starting point of the line-of-sight ray, and a ray is emitted along the line-of-sight direction as the line-of-sight ray.

[0084] In the embodiments of the present disclosure, the position and line-of-sight direction of the user's eyes are detected by the second camera. The position and line-of-sight direction of the user's eyes can represent the line-of-sight ray emitted from the user's eyes, and the intersection point of the ray and the preset window area can represent the target position of the preset window area that the user is looking at, so as to effectively obtain the target position information corresponding to the user's line of sight on the preset window area, and provide an effective gaze position reference for locking the target object that the user is looking at.

[0085] Figure 5 It is a schematic flowchart of a method for querying a target object provided by another exemplary embodiment of the present disclosure.

[0086] In some optional embodiments, on the basis of any of the above embodiments, as Figure 5 shown, determining the target object corresponding to the user's line of sight based on the projection contour information and target position information in step 240 may include:

[0087] Step 2410, determining candidate objects from each object based on the projection contour information and target position information of each object.

[0088] Among them, the target position information can be matched with the projection contour information of each object to determine the matching relationship between the target position information and the projection contour information of each object, and candidate objects can be determined from each object according to the matching relationship.

[0089] In some optional embodiments, for the projection contour information of each object, the coordinate range of the projection contour can be determined according to the projection contour information, and it can be determined whether the position coordinate corresponding to the target position information is within the coordinate range of the projection contour. If the target position information is within the coordinate range of the projection contour, it is determined that the matching relationship is a match, that is, the object is a candidate object. If the position coordinate corresponding to the target position information is not within the coordinate range of the projection contour, it is determined that the matching relationship is a mismatch, and the object is not a candidate object.

[0090] In some optional embodiments, for the projection contour information of each object, the center point coordinate and size of the projection contour can be calculated, the distance between the position coordinate corresponding to the target position information and the center point coordinate of the projection contour can be calculated, and the matching relationship can be determined in combination with the distance and the size of the projection contour. The specific manner of determining candidate objects is not limited.

[0091] Step 2420: Based on the projection contour information corresponding to the candidate object, display a projection frame corresponding to the candidate object on the preset window area.

[0092] The projection frame is used to lock the candidate object outside the vehicle corresponding to the user's line of sight.

[0093] In some alternative embodiments, the projection frame corresponding to the candidate object can be displayed on the preset window area through a Head-up Display (HUD) device. The projection frame can refer to Figure 1 the projection contour b in

[0094] In some alternative embodiments, the display parameters of the HUD device can be generated based on the projection contour information corresponding to the candidate object, and then the HUD is controlled according to the display parameters to project and display the projection frame corresponding to the candidate object on the preset window area. The specific control principle of the HUD device will not be elaborated.

[0095] Step 2430: In response to obtaining the query intention information of the user for the candidate object, determine that the candidate object is the target object corresponding to the user's line of sight.

[0096] The query intention information of the user for the candidate object is the intention information of the query type indicated by the user. For example, the query intention information can include what brand this car is, what the displacement of this car is, what building this building is, and so on.

[0097] In some alternative embodiments, the query intention information of the user for the candidate object can be obtained by voice detection. For example, after the user sees the projection frame displayed on the preset window area, according to the overlapping situation between the projection frame and the object actually gazed at by the user, it is determined whether the candidate object locked by the projection frame is the target object to be queried. If it is confirmed to be the target object, the user can output by voice "what brand this car is", "what building this building is", etc. The in-vehicle terminal collects the user's voice data through a voice collection device, and obtains the query intention information of the user for the candidate object through voice recognition and semantic understanding.

[0098] In the embodiments of the present disclosure, by displaying the projection frame of the candidate object on the preset window area, it is convenient for the user to confirm whether the candidate object is the target object to be queried, and then accurate and effective query result information can be provided for the user according to the user's query intention, improving the accuracy and effectiveness of the query result.

[0099] In some alternative embodiments, the step of displaying a projection frame corresponding to the candidate object on the preset window area based on the projection contour information corresponding to the candidate object in step 2420 may include:

[0100] Determine the display parameters of the head-up display device according to the projection contour information corresponding to the candidate object; based on the display parameters, control the head-up display device to project the projection frame corresponding to the projection contour information onto a preset window area.

[0101] Among them, the display parameters are the parameters required for the head-up display device to display the projection frame at the corresponding position in the preset window area. For example, the display parameters may include the projection direction, projection content, etc. of the head-up display device to ensure that the corresponding projection frame can be displayed in the contour area corresponding to the projection contour information.

[0102] In the embodiments of the present disclosure, the projection frame is projected and displayed on the preset window area through the head-up display device. Since the preset window area is a transparent area, through the display of the head-up display device, the locked candidate object can be displayed for the user without affecting the user's line of sight, facilitating the user to confirm whether the locked candidate object is the target object to be queried on the basis of ensuring driving safety, and further improving the user experience.

[0103] Figure 6 It is a schematic flowchart of a method for querying a target object provided by another exemplary embodiment of the present disclosure.

[0104] In some optional embodiments, after displaying the projection frame corresponding to the candidate object on the preset window area in step 2420 based on the projection contour information corresponding to the candidate object, the method of the embodiments of the present disclosure may further include:

[0105] Step 310, obtain the voice information of the user.

[0106] Among them, the voice information (or voice data) of the user can be obtained through a voice collection device on the vehicle. The voice collection device may include, for example, a microphone, a microphone array, etc.

[0107] Step 320, based on the voice information, determine whether the query intention information of the user for the candidate object is obtained.

[0108] In some optional embodiments, voice recognition and semantic understanding may be performed on the voice information to determine whether the query intention information of the user for the candidate object is obtained. The query intention information may refer to the foregoing embodiments.

[0109] In some optional embodiments, the voice information may be recognized based on a pre-trained voice recognition model to obtain the text information corresponding to the voice information. Then, semantic understanding is performed on the text information to determine whether the text information contains the query intention information of the user for the candidate object. Optionally, semantic understanding may be performed based on pre-configured semantic understanding rules or based on a pre-trained language model (such as a large language model). The specific network structures of the voice recognition model and the language model are not limited.

[0110] In some alternative embodiments, speech information can be recognized and understood based on a pre-trained large language model, that is, the speech information is used as the input data of the large language model, and the large language model outputs the user's intent information. According to the type of the intent information, it is determined whether query intent information is obtained.

[0111] In the embodiments of the present disclosure, the query intent information of the user for the candidate object is determined through the user's speech information, so that after the user sees the projection frame displayed in the preset window area, the user can express his query intent through voice interaction, so as to provide the user with more accurate and effective query result information and further improve the user experience.

[0112] In some alternative embodiments, step 320 of determining whether query intent information of the user for the candidate object is obtained based on the speech information may include:

[0113] Perform semantic parsing on the speech information to determine the user's intent information; determine the user's intent type according to the intent information; in response to the intent type being the query type, determine that the query intent information of the user for the candidate object is obtained.

[0114] Among them, the intent type may include a query type, a correction type, etc. The user's intent information may include at least one of intent information of the query type (i.e., query intent information), intent information of the correction type (or correction intent information), and other types of intent information. The intent information of the correction type refers to the intent information of the user to correct the candidate object when the user sees that the object seen through the projection frame area is not the target object he wants to query. For example, by voice outputting "the car on the left", "the building on the right", etc.

[0115] In some alternative embodiments, semantic parsing of the speech information can be performed through a speech recognition model and semantic understanding rules to determine the user's intent information. Alternatively, the large language model can be used to perform semantic parsing on the speech information to obtain the user's intent information.

[0116] In some alternative embodiments, intent keywords corresponding to the query type and the correction type can be preset respectively, and the user's intent type is determined by matching the intent information with the intent keywords of each type. If the user's intent type is the query type, it is determined that the query intent information of the user for the candidate object is obtained.

[0117] In the embodiments of the present disclosure, the user's intent information is obtained through semantic parsing. Since the intent information contains the keywords of the user's intent, it is possible to effectively determine whether the intent information is the intent information of the query type according to the intent information, so as to confirm the target object through the user's voice interaction and ensure the accuracy of the target object.

[0118] In some alternative embodiments, determining the target object corresponding to the user's line of sight based on the projection profile information and target position information of the object in step 240 may further include:

[0119] In response to the intent type being a correction type, determine the corrected object according to the user's intent information; adjust the display position of the projection frame in the preset window area according to the projection profile information corresponding to the corrected object, so that the adjusted projection frame locks the corrected object; and determine the corrected object as the target object.

[0120] Among them, if it is determined according to the intent information that the user's intent type is a correction type, the corrected object can be determined according to the intent information of the user's correction type. For example, based on the projection frame of the candidate object currently displayed, if the user's intent information includes correction type intent information such as "the car on the left" and "the building on the right", the corrected object can be determined from other objects according to the user's intent information and the relative position relationship between other objects (that is, objects other than the candidate object) and the candidate object in each object. Then, according to the projection profile information corresponding to the corrected object, adjust the display position of the projection frame in the preset window area. For example, determine the corrected display parameters of the head-up display device according to the projection profile information corresponding to the corrected object, and then project and display the projection frame corresponding to the corrected object at the corresponding position in the preset window area based on the corrected display parameters.

[0121] In some alternative embodiments, after locking the corrected object, it is also possible to further confirm by voice interaction whether the locked projection frame is the target object that the user intends to query, and output the query result information after confirmation, or output the corresponding query result information according to the user's query intent information.

[0122] In the embodiments of the present disclosure, in the case where there is a deviation between the candidate object locked by the projection frame and the user's true intention, the user can correct the target object to be queried by voice interaction, so as to ensure the accuracy of the target object and provide accurate and effective query result information for the user.

[0123] In some alternative embodiments, based on any of the above embodiments, outputting the query result information corresponding to the target object in step 250 may include:

[0124] In response to obtaining the user's query intent information for the candidate object, determine the query result information of the target object according to the query intent information; output the query result information.

[0125] Among them, when the query intention information of the candidate object is obtained, if it is determined that the candidate object is the target object intended to be queried by the user, the query result information corresponding to the target object is searched according to the query intention information of the user, and then the query result information is output.

[0126] In the embodiments of the present disclosure, the query result information of the target object is determined according to the query intention information of the user, so that the query result information can better meet the needs of the user, and the accuracy and effectiveness of the query result information are improved.

[0127] In some alternative embodiments, Figure 7 is a flowchart of a method for querying a target object provided by an exemplary embodiment of the present disclosure. As Figure 7 shown, taking the driver as an example, the preset window area is the front window area, and the method of the embodiments of the present disclosure may include the following steps:

[0128] Step 410, when the first image of the front window area is obtained, object detection and segmentation are performed on the first image to obtain the object contour information of each object in the first image.

[0129] Step 420, project the object contour onto the front window area. That is, determine the projected contour information of each object in the first image on the front window area.

[0130] Step 430, perform line-of-sight recognition based on the driver's image. That is, determine the position of the driver's eyes (or the starting point of the line of sight) and the line-of-sight direction.

[0131] Step 440, determine the line-of-sight landing position. That is, based on the position of the driver's eyes and the line-of-sight direction, determine the target position information corresponding to the user's line of sight on the preset window area.

[0132] Step 450, target object positioning. That is, based on the projected contour information of the object and the target position information, determine the target object corresponding to the user's line of sight.

[0133] Step 460, HUD display, that is, project and display the projection frame corresponding to the target object on the front window area through the HUD.

[0134] Step 470, perform speech recognition based on the driver's voice command (i.e., voice information or voice data) to obtain the text information corresponding to the voice command.

[0135] Step 480, perform semantic understanding on the text information to obtain the user's intention information.

[0136] Step 490, determine whether the locked target object is correct according to the user's intention information. If the target is incorrect, return to step 450, and adjust the position of the projection frame through repositioning. If the target is correct, go to step 4100.

[0137] Step 4100, obtain query result information. That is, obtain the query result information corresponding to the searched target object.

[0138] Step 4110, output the query result information.

[0139] For the specific operations of the above steps 410 to 4110, reference can be made to the foregoing embodiments, and details are not described herein again.

[0140] In the query method of the target object in the embodiments of the present disclosure, the first camera is used to simulate the user's eyes to observe the situation outside the preset vehicle window, and the outline of the target object stared at by the user's line of sight is projected and displayed on the corresponding position area of the preset vehicle window area through the HUD, providing visual feedback to the user to facilitate the user to confirm the locked target object. Then, combined with the user's voice interaction, the target stared at by the line of sight is confirmed or corrected, and the user can observe the correction result in real time to ensure the accuracy of the target object, so as to provide accurate and effective query result information for the user, enabling the user to query the relevant information of the target object outside the vehicle based on simple voice and line of sight, and effectively improving the user experience.

[0141] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs. Moreover, in the technical solution of the present disclosure, the collection and use of the user's personal information involved are carried out with the user's knowledge and authorization, and do not involve the illegal collection and use of the user's personal information.

[0142] The above embodiments of the present disclosure can be implemented separately or combined in any combination without conflict, and can be specifically set according to actual needs, which is not limited in the present disclosure.

[0143] Any query method of the target object provided in the embodiments of the present disclosure can be executed by any suitable electronic device with data processing capabilities, including but not limited to: terminal devices, servers, and other electronic devices. Or, any query method of the target object provided in the embodiments of the present disclosure can be executed by a processor. For example, the processor executes any query method of the target object mentioned in the embodiments of the present disclosure by calling the corresponding instructions stored in the memory. Details are not described herein again.

[0144] Exemplary device

[0145] Figure 8 It is a schematic structural diagram of a query device for a target object provided by an exemplary embodiment of the present disclosure. The device in this embodiment can be used to implement the corresponding method embodiments of the present disclosure, such as Figure 8The device shown may include: a first processing module 51, a second processing module 52, a third processing module 53, a fourth processing module 54, and an output module 55.

[0146] The first processing module 51 is configured to obtain a first image within a preset window area collected by a first camera in the vehicle.

[0147] The second processing module 52 is configured to determine projection contour information of each object in the first image on the preset window area.

[0148] The third processing module 53 is configured to detect target position information corresponding to the line of sight of a user in the vehicle on the preset window area.

[0149] The fourth processing module 54 is configured to determine a target object corresponding to the user's line of sight based on the projection contour information of the object and the target position information.

[0150] The output module 55 is configured to output query result information corresponding to the target object.

[0151] Figure 9 It is a schematic structural diagram of a query device for a target object provided by another exemplary embodiment of the present disclosure.

[0152] In some alternative embodiments, on the basis of the above Figure 8 shown embodiment, as Figure 9 shown, the second processing module 52 may include: a first determination unit 521 and a first processing unit 522.

[0153] The first determination unit 521 is configured to determine object contour information corresponding to each object in the first image based on the first image.

[0154] The first processing unit 522 is configured to project the object contour information corresponding to each object onto the preset window area to obtain the projection contour information corresponding to each object.

[0155] In some alternative embodiments, the first determination unit 521 is specifically configured to: perform semantic segmentation on the first image to obtain a pixel point set corresponding to at least one object in the first image; for any one of the at least one object, based on the pixel point set corresponding to the object, determine the object contour information of the object in the first image.

[0156] In some alternative embodiments, the first processing unit 522 is specifically configured to:

[0157] For any object, project the object contour information corresponding to the object onto the preset window area based on a pre-configured first mapping relationship from the first image to the preset window area to obtain the projection contour information corresponding to the object.

[0158] In some alternative embodiments, the first processing unit 522 is specifically configured to:

[0159] For the object contour information corresponding to any object, determine the rays emitted by the optical center of the first camera to the multiple contour points corresponding to the object contour information; based on the intersection points of the rays respectively corresponding to the contour points with the preset window area, determine the projection contour information corresponding to the object.

[0160] In some alternative embodiments, on the basis of any of the above embodiments, as Figure 9 shown, the third processing module 53 may include: a second determination unit 531, a third determination unit 532, and a fourth determination unit 533.

[0161] The second determination unit 531 is configured to determine a second image including the user's eyes collected by the second camera in the vehicle.

[0162] The third determination unit 532 is configured to determine the position of the user's eyes and the line-of-sight direction of the eyes looking at the preset window area based on the second image.

[0163] The fourth determination unit 533 is configured to determine the target position information corresponding to the user's line of sight on the preset window area based on the position and line-of-sight direction of the user's eyes.

[0164] In some alternative embodiments, on the basis of any of the above embodiments, as Figure 9 shown, the fourth processing module 54 may include: a second processing unit 541, a third processing unit 542, and a fourth processing unit 543.

[0165] The second processing unit 541 is configured to determine candidate objects from the objects based on the projection contour information and target position information of the objects.

[0166] The third processing unit 542 is configured to display a projection frame corresponding to the candidate object on the preset window area based on the projection contour information corresponding to the candidate object.

[0167] Wherein, the projection frame is used to lock the candidate object outside the vehicle corresponding to the user's line of sight.

[0168] The fourth processing unit 543 is configured to determine that the candidate object is the target object corresponding to the user's line of sight in response to obtaining the query intention information of the user for the candidate object.

[0169] In some alternative embodiments, the third processing unit 542 is specifically configured to:

[0170] Determine the display parameters of the head-up display device according to the projection contour information corresponding to the candidate object; based on the display parameters, control the head-up display device to project the projection frame corresponding to the projection contour information onto a preset window area.

[0171] Figure 10 It is a schematic structural diagram of a query device for a target object provided by another exemplary embodiment of the present disclosure.

[0172] In some optional embodiments, such as Figure 10 shown, the fourth processing module 54 may further include:

[0173] An acquisition unit 544, configured to acquire the voice information of the user.

[0174] A fifth processing unit 545, configured to determine whether query intention information of the user for the candidate object is acquired based on the voice information.

[0175] In some optional embodiments, the fifth processing unit 545 is specifically configured to:

[0176] Perform semantic parsing on the voice information to determine the intention information of the user; determine the intention type of the user according to the intention information; in response to the intention type being the query type, determine that the query intention information of the user for the candidate object is acquired.

[0177] In some optional embodiments, the fourth processing unit 543 is further configured to:

[0178] In response to the intention type being the correction type, determine the corrected object according to the intention information of the user; adjust the display position of the projection frame in the preset window area according to the projection contour information corresponding to the corrected object, so that the adjusted projection frame locks the corrected object; determine the corrected object as the target object.

[0179] In some optional embodiments, based on any of the above embodiments, the output module 55 is specifically configured to: in response to acquiring the query intention information of the user for the candidate object, determine the query result information of the target object according to the query intention information; output the query result information.

[0180] For the beneficial technical effects corresponding to the exemplary embodiments of this device, reference may be made to the corresponding beneficial technical effects in the above exemplary method part, which will not be elaborated here.

[0181] Exemplary electronic device

[0182] Figure 11 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure, including at least one processor 91 and a memory 92.

[0183] The processor 91 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 90 to perform desired functions.

[0184] The memory 92 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 91 may run one or more computer program instructions to implement the methods of the various embodiments of the present disclosure above and / or other desired functions.

[0185] In one example, the electronic device 90 may further include: an input device 93 and an output device 94, and these components are interconnected through a bus system and / or other form of connection mechanism (not shown).

[0186] The input device 93 may further include, for example, a touch screen, a microphone, various sensors, etc. The sensors may include, for example, an image sensor (such as a camera, a webcam, etc.), lidar, millimeter-wave radar, ultrasonic radar, a positioning sensor, a pressure sensor, an air quality sensor, a temperature sensor, etc. The image sensor, lidar, millimeter-wave radar, ultrasonic radar, etc. may be used for the perception of the surrounding environment, that is, to detect static and dynamic objects in the surrounding environment. Static and dynamic objects may include, for example, static objects such as lane lines, curbs, arrows, signs, trees, buildings, etc., and dynamic objects such as surrounding vehicles, pedestrians, cyclists, etc. The positioning sensor is used to implement the positioning of the movable device where the electronic device is located (such as a vehicle, a robot, etc.). The positioning sensor may include, for example, an inertial measurement unit (IMU) and a global positioning system (GPS), etc. The pressure sensor may be used to detect the seat pressure. The temperature sensor may be used to detect the temperature inside the vehicle cockpit. The air quality sensor may be used to detect the air quality inside the vehicle cockpit.

[0187] The output device 94 may output various information to the outside, which may include, for example, a display, a speaker, and a communication network and its connected remote output devices, etc.

[0188] Of course, for simplicity, Figure 11Only some of the components of the electronic device 90 related to the present disclosure are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 90 may further include any other appropriate components.

[0189] Exemplary computer program product and computer-readable storage medium

[0190] In addition to the above methods and devices, embodiments of the present disclosure may also provide a computer program product, including computer program instructions, which when run by a processor cause the processor to execute the steps in the methods of various embodiments of the present disclosure described in the above "Exemplary Method" section.

[0191] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0192] In addition, embodiments of the present disclosure may also be a computer-readable storage medium, on which computer program instructions are stored, which when run by a processor cause the processor to execute the steps in the methods of various embodiments of the present disclosure described in the above "Exemplary Method" section.

[0193] The computer-readable storage medium may adopt any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium, for example but not limited to, includes systems, devices or components of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0194] The basic principles of the present disclosure have been described in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that they are essential for each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purposes of illustration and facilitating understanding, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0195] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these changes and modifications.

Claims

1. A method for querying a target object, comprising: Acquire a first image in a preset vehicle window area captured by a first camera in the vehicle; Determine projection contour information of each object in the first image on the preset vehicle window area; Detecting target position information corresponding to the sight line of the user in the vehicle on the preset window area; Determining a target object corresponding to the user's line of sight based on the projection contour information of the object and the target position information; Output the query result information corresponding to the target object.

2. The method according to claim 1, wherein: The determining of projection contour information of each object in the first image on the preset vehicle window area includes: Based on the first image, determining object contour information corresponding to each of the objects in the first image; The object contour information corresponding to each of the objects is projected onto the preset vehicle window area to obtain the projection contour information corresponding to each of the objects.

3. The method according to claim 2, wherein: The step of projecting the object contour information corresponding to each of the objects onto the preset vehicle window area to obtain the projection contour information corresponding to each of the objects includes: For any of the objects, based on a preconfigured first mapping relationship from the first image to the preset window area, projecting the object contour information corresponding to the object to the preset window area to obtain the projection contour information corresponding to the object; and / or, For the object contour information corresponding to any of the objects, determine the rays emitted from the optical center of the first camera to multiple contour points corresponding to the object contour information; based on the intersection of the rays corresponding to each of the contour points and the preset window area, determine the projection contour information corresponding to the object.

4. The method according to claim 1, wherein: The detecting the target position information corresponding to the sight line of the user in the vehicle on the preset window area includes: Determining a second image captured by a second camera in the vehicle that includes the user's eyes; Based on the second image, determining the position of the user's eyes and the direction of the eyes' sight toward the preset vehicle window area; Based on the position of the user's eyes and the line of sight direction, the target position information corresponding to the user's line of sight on the preset vehicle window area is determined.

5. The method according to any one of claims 1 to 4, wherein: The determining the target object corresponding to the user's sight line based on the projection contour information of the object and the target position information includes: Determine a candidate object from among the objects based on the projection profile information and the target position information of each of the objects; Based on the projection contour information corresponding to the candidate object, a projection frame corresponding to the candidate object is displayed on the preset vehicle window area; the projection frame is used to lock the candidate object outside the vehicle corresponding to the user's line of sight; In response to obtaining the query intention information of the user for the candidate object, the candidate object is determined to be the target object corresponding to the user's line of sight.

6. The method according to claim 5, wherein: The step of displaying a projection frame corresponding to the candidate object on the preset vehicle window area based on the projection contour information corresponding to the candidate object includes: Determining display parameters of a head-up display device according to the projection profile information corresponding to the candidate object; Based on the display parameters, the head-up display device is controlled to project the projection frame corresponding to the projection profile information onto the preset vehicle window area.

7. The method according to claim 5, wherein: After displaying the projection frame corresponding to the candidate object on the preset vehicle window area based on the projection contour information corresponding to the candidate object, the method further includes: Acquiring voice information of the user; Based on the voice information, it is determined whether the user's query intention information for the candidate object is obtained.

8. The method according to claim 7, wherein: The determining, based on the voice information, whether the user's query intention information for the candidate object is obtained includes: Performing semantic analysis on the voice information to determine the user's intention information; Determining the user's intention type according to the intention information; In response to the intention type being a query type, it is determined that query intention information of the user for the candidate object is acquired.

9. The method according to claim 8, wherein: Also includes: In response to the intention type being a correction type, determining a corrected object according to the intention information of the user; According to the projection contour information corresponding to the corrected object, adjusting the display position of the projection frame in the preset vehicle window area so that the adjusted projection frame locks the corrected object; The corrected object is determined as the target object.

10. The method according to claim 8, wherein: The outputting the query result information corresponding to the target object includes: In response to obtaining the user's query intention information for the candidate object, determining query result information of the target object according to the query intention information; Output the query result information.

11. A target object query device, comprising: A first processing module, used to obtain a first image in a preset vehicle window area captured by a first camera in the vehicle; A second processing module, used to determine projection contour information of each object in the first image on the preset vehicle window area; A third processing module is used to detect the target position information corresponding to the sight line of the user in the vehicle on the preset window area; A fourth processing module, configured to determine a target object corresponding to the user's line of sight based on the projection contour information of the object and the target position information; The output module is used to output the query result information corresponding to the target object.

12. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the method according to any one of claims 1 to 10.

13. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1-10.