Image processing methods, devices, and storage media based on gesture recognition in vehicles
By detecting user gesture trajectories and using 3D point cloud data reconstruction technology, images of external targets are generated, solving the problem of not being able to obtain external objects in a timely manner while the vehicle is in motion, thus improving the user's driving experience and communication efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AVATR CO LTD
- Filing Date
- 2023-09-07
- Publication Date
- 2026-05-26
Smart Images

Figure CN117058765B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of intelligent driving and artificial intelligence in vehicle technology, and in particular to an image processing method, apparatus, and storage medium based on gesture recognition in a vehicle. Background Technology
[0002] In daily life, when we see someone or something we want to photograph, we use our phones to take pictures and share them. However, for users in a moving car, it's easy to miss the right moment to take a picture with a phone.
[0003] Currently, most vehicles integrate internal and external camera systems, providing users with the hardware for real-time photography. Users can trigger the camera system to capture images of the outside of the vehicle via buttons or commands, and then use these images for communication. However, these images are full-range, all-scenario images of the vehicle's exterior. When communicating based on panoramic images, users cannot immediately identify the person they want to communicate with from other passengers.
[0004] In conclusion, how to obtain information about objects outside the vehicle in a timely manner while the vehicle is in motion is an urgent problem to be solved. Summary of the Invention
[0005] This application provides an image processing method, apparatus, and storage medium based on gesture recognition in a vehicle, to solve the problem of not being able to acquire images of the vehicle's exterior in a timely manner while the vehicle is in motion.
[0006] In a first aspect, this application provides an image processing method based on gesture recognition in a vehicle, the method comprising:
[0007] When a user is detected to meet a preset condition, the trajectory information of the user's gestures within a preset time period is obtained;
[0008] Determine the target area outside the vehicle formed by the trajectory information, with the user's eye position as the observation point for the target area;
[0009] Based on the image of the vehicle's exterior, determine the target image corresponding to the target area;
[0010] Display the target image.
[0011] Optionally, the preset conditions include: the shape of the user's hand outline conforms to a preset shape and / or the user's voiceprint conforms to a preset voiceprint.
[0012] Optionally, determining the target area outside the vehicle formed by the trajectory information includes:
[0013] Determine whether the trajectory information forms a closed shape;
[0014] If the trajectory information forms a closed shape, the target area corresponding to the closed shape is determined according to the first preset rule;
[0015] If the trajectory information does not form a closed shape, a closed shape is generated according to the trajectory information and the second preset rule, and the target area corresponding to the closed shape is determined according to the first preset rule.
[0016] Optionally, determining the target image corresponding to the target area based on the image outside the vehicle includes:
[0017] The target image of the target area is determined based on images of the vehicle's exterior and constructed 3D scene data of the vehicle's interior and exterior.
[0018] Optionally, determining the target image of the target area based on images of the vehicle's exterior and constructed 3D scene data of the vehicle's interior and exterior includes:
[0019] The system acquires the trajectory information of the user's gestures inside the vehicle and the first position of the user's eyes, where the first position represents the three-dimensional coordinates of the eyes within the vehicle.
[0020] 3D point cloud data of the vehicle's interior and exterior is constructed based on real-time data acquired from the vehicle. The 3D point cloud data provides three-dimensional scene information of the vehicle's surrounding environment and interior.
[0021] Based on the 3D point cloud data and the images captured by the camera outside the vehicle, the target image of the target area pointed to by the user's gesture is determined.
[0022] Optionally, the step of constructing 3D point cloud data of the vehicle's interior and exterior based on real-time vehicle data includes:
[0023] Based on the real-time images captured by the in-vehicle camera and the real-time images captured by the out-of-vehicle camera, parallax calculation is performed to obtain the first depth information of the image from the perspective of the in-vehicle camera and the second depth information of the image from the perspective of the out-of-vehicle camera.
[0024] The image captured by the in-vehicle camera is converted into first point cloud data based on the first depth information, and the image captured by the out-of-vehicle camera is converted into second point cloud data based on the second depth information.
[0025] The first point cloud data and the second point cloud data are respectively registered to the same coordinate system to obtain the 3D point cloud data inside and outside the vehicle.
[0026] Optionally, the step of constructing 3D point cloud data of the vehicle's interior and exterior based on real-time vehicle data includes:
[0027] When the user meets the preset conditions, 3D point cloud data of the inside and outside of the vehicle is constructed from the images captured by the in-vehicle camera and the out-of-vehicle camera at the moment when the user begins to draw graphics in space based on the gesture.
[0028] or,
[0029] When the user meets the preset conditions, 3D point cloud data of the inside and outside of the vehicle is constructed based on the images captured by the vehicle's internal and external cameras at the moment when the user finishes drawing graphics in space according to the gesture.
[0030] Optionally, determining the target image of the target area pointed to by the user's gesture based on the 3D point cloud data and the image captured by the camera outside the vehicle includes:
[0031] An extendable view frustum is constructed from the first position to the boundary of the drawn graphic, the extendable view frustum representing the range of the 3D point cloud data visible from the viewpoint.
[0032] Obtain the target 3D point cloud data within the extendable view frustum range from the 3D point cloud data;
[0033] The target image is obtained based on the 3D target point cloud data.
[0034] Optionally, obtaining the target image based on the 3D target point cloud data includes:
[0035] The target 3D point cloud data is projected onto the plane of the graphic to generate a projected two-dimensional image;
[0036] The target image is rendered based on the projected two-dimensional image and the color data of each point in the 3D point cloud data.
[0037] or,
[0038] The target image is obtained from images captured by an external camera of the vehicle using feature descriptors based on the 3D target point cloud data.
[0039] Optionally, displaying the target image on the vehicle's display screen includes:
[0040] Objects in the image of the target region are labeled to obtain the target image.
[0041] Optionally, the step of annotating objects in the image of the target region to obtain the target image includes:
[0042] The image recognition model identifies object information in the image of the target region by means of an image recognition model, which is a deep learning model pre-trained on multiple labeled image data;
[0043] The target image is obtained by annotating the region corresponding to the object information in the image of the target region.
[0044] or,
[0045] At least one target object is identified in the image of the target region using a target detection algorithm;
[0046] The at least one target object is identified in the image of the target area using a drawing tool to obtain the target image.
[0047] Optionally, the step of annotating objects in the image of the target region to obtain the target image includes:
[0048] Obtain the spectral graphic formed in space by the trajectory information of the user's gestures from the first position viewpoint;
[0049] The perspective image is compared with a preset image to obtain the comparison result;
[0050] If the comparison results are consistent, the object information is marked in the image of the target area to obtain the target image;
[0051] If the comparison results are inconsistent, at least one target object is identified in the image of the target area to obtain a target image.
[0052] Secondly, this application also provides an image processing device based on gesture recognition in a vehicle, the device comprising:
[0053] The acquisition module is used to acquire the trajectory information of the user's gestures within a preset time period when the user meets the preset conditions.
[0054] The first determining module is used to determine the target area outside the vehicle formed by the trajectory information, wherein the target area is observed from the user's eye position.
[0055] The second determining module is used to determine the target image corresponding to the target area based on the images of the vehicle's exterior and interior.
[0056] The display module is used to display the target image.
[0057] Thirdly, this application also provides a vehicle, said vehicle comprising:
[0058] The vehicle body, sensors, cameras, processor, and communication interfaces for interacting with other devices, the processor being used to execute the gesture recognition-based image processing method in the vehicle as described in any of the first aspects.
[0059] Fourthly, this application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the gesture-based image processing method in a vehicle as described in any of the first aspects.
[0060] Fifthly, this application also provides a computer program product, including computer program instructions that cause a computer to perform an image processing method based on gesture recognition in a vehicle as described in any of the first aspects.
[0061] This application provides a gesture recognition-based image processing method, apparatus, and storage medium for vehicles. The method includes: acquiring the trajectory information of the user's gestures within a preset time period when a user is detected to meet preset conditions; determining a target area outside the vehicle formed by the trajectory information, wherein the target area is viewed from the user's eye position; determining a target image corresponding to the target area based on images of the vehicle's exterior and interior; and displaying the target image. This method allows users to obtain the desired image simply by drawing a shape on an object outside the vehicle, making the operation convenient and improving the user's driving experience. Attached Figure Description
[0062] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0063] Figure 1 This application provides an illustration of an image processing method based on gesture recognition in a vehicle.
[0064] Figure 2 A flowchart illustrating an embodiment of an image processing method based on gesture recognition in a vehicle provided in this application;
[0065] Figure 3 A flowchart illustrating a second embodiment of an image processing method based on gesture recognition in a vehicle provided in this application;
[0066] Figure 4 A flowchart illustrating a third embodiment of an image processing method based on gesture recognition in a vehicle provided in this application;
[0067] Figure 5 A flowchart illustrating Embodiment 4 of an image processing method based on gesture recognition in a vehicle provided in this application;
[0068] Figure 6 A flowchart illustrating an example of a gesture recognition-based image processing method in a vehicle provided in this application;
[0069] Figure 7 This application provides a schematic diagram of the structure of an image processing device based on gesture recognition in a vehicle, according to one embodiment.
[0070] Figure 8 This is a structural schematic diagram of a vehicle provided in this application.
[0071] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0072] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0073] It should be noted that the image processing method, device, and storage medium based on gesture recognition in vehicles provided in this application can be used in the fields of intelligent driving and artificial intelligence in the field of vehicle technology, and can also be used in any field other than vehicle technology. The application fields of the image processing method, device, and storage medium based on gesture recognition in vehicles in this application are not limited.
[0074] First, let me explain the terms used in this application:
[0075] Parallax refers to the difference in position of an object within an observer's field of vision when the same object is viewed from different perspectives. Parallax provides information about the object's position in three-dimensional space. In this application, there is parallax between the object outside the vehicle as seen from the user's perspective and the object outside the vehicle as captured by the camera, and there is also parallax between images of the object outside the vehicle captured by cameras at different locations.
[0076] Depth information refers to the distance of an object relative to the observer. It reflects the distance between objects and the observer. Depth information can be used to describe the distance relationships between different objects in a scene, thus helping to understand the three-dimensional structure of the scene.
[0077] The visual cone is a cone-shaped region extending from the observer's eye position that defines the objects visible in the camera's field of view.
[0078] Figure 1 This application provides an illustration of an image processing method based on gesture recognition in a vehicle, as shown in the diagram. Figure 1 As shown, if a passenger sees an object outside the vehicle at point A while the vehicle is in motion and wants to talk to the driver, pointing it out to the driver would require the driver to turn around, creating a safety hazard. If the passenger takes a picture with their phone, the vehicle will quickly move to point B, missing the opportune moment to take the picture. Furthermore, regardless of whether the communication is done via phone or vehicle camera, neither party may be aware of the object the other is referring to.
[0079] In view of the above problems, the inventors discovered during their research in this technical field that when a user finds an object of interest outside the vehicle, the user can draw a range around the object's location with their finger inside the vehicle, and the vehicle can quickly respond to the user's gesture. The vehicle uses data collected by internal sensors and cameras, combined with gesture recognition and image processing technologies, to obtain an image of the object the user wants to circle, and displays it to the user. This method allows for the acquisition of images of specified objects outside the vehicle based on user gestures. Based on this, this application proposes an image processing method, apparatus, and storage medium for vehicles based on gesture recognition.
[0080] This application can be applied not only to vehicles, but also to airplanes, trains, and other scenarios requiring real-time capture of external images, such as conference rooms, medical imaging, industrial inspection, and augmented reality. Applying gesture recognition image processing methods in these scenarios requires hardware devices for collecting data on people and their surrounding environment, such as cameras, sensors, and radar.
[0081] The executing entity of this application can be a processor with processing capabilities in a vehicle, an in-vehicle terminal with processing capabilities, or a chip with processing capabilities. In some cases, the executing entity can also be a cloud server.
[0082] The following describes the technical solution of this application and how it solves the aforementioned technical problems, taking a processor in a vehicle as the executing entity. The specific embodiments described below can be combined with each other, and similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0083] Figure 2 A flowchart illustrating an embodiment of an image processing method based on gesture recognition in a vehicle provided in this application is shown below. Figure 2As shown, the method includes the following steps:
[0084] S11. When the user is detected to meet the preset conditions, obtain the trajectory information of the user's gestures within the preset time period.
[0085] After the vehicle is started, sensors or cameras inside the vehicle collect real-time trajectory data of hand gestures inside the vehicle. This data includes the real-time position information of the user's hand, as well as the outline information of the user's hand.
[0086] When the collected user data meets the preset conditions, the real-time position information of the user's hand is obtained from the collected data from the start of the gesture meeting the preset conditions to the preset time period.
[0087] There are at least the following ways for a user to meet the preset conditions:
[0088] In one implementation, it is determined whether the contour information of the user's hand in the collected data meets a preset contour. Specifically, a neural network model trained on a large amount of hand contour data can be used to identify whether the preset contour is met. The real-time collected hand contour information is input into the neural network model, and the model outputs information indicating whether the preset contour is met. The input contour information can be an image of the user's hand captured by a camera or measurement data collected by sensors.
[0089] In another possible implementation, after the user activates the vehicle's voice assistant, the system detects whether the words spoken by the user match a preset vocabulary. In one specific approach, a pre-built voiceprint model is selected, which stores the user's voiceprint information as a template. The collected user voiceprint information is then input into the voiceprint model for comparison and matching, and the result is output.
[0090] In another possible implementation, after detecting that the user has spoken a preset word, it is also necessary to collect data on the outline information of the user's hand that meets the preset outline.
[0091] S12. Determine the target area outside the vehicle formed by the trajectory information, with the user's eye position as the observation point.
[0092] In this step, a target area is determined based on the collected trajectory information of the user's hand (i.e., the real-time position information of the user's hand). This target area is defined by the user's eye position as the observation point and the shape formed by the user's trajectory information as its range. This target area is used to determine the image outside the vehicle as seen from the user's perspective. Therefore, the target area must be a closed region. However, the region formed by the user's trajectory information is not necessarily a closed region. Therefore, it is necessary to determine whether the trajectory information of the user's hand is a closed region. If it is not a closed region, it needs to be supplemented to form a closed region.
[0093] To determine whether the trajectory information within a preset time period forms a closed shape, specifically, the 3D point cloud data or image data of the hand movement collected by the sensor within the preset time period is input into the gesture trajectory recognition algorithm for gesture trajectory recognition, and the shape formed by the user's gesture trajectory is obtained. The gesture trajectory recognition algorithm is a model that can recognize the shape formed by gesture trajectory by using multiple gesture action data in advance and training through a neural network model.
[0094] If the trajectory information over a preset time period has formed a closed pattern, the target area corresponding to the closed pattern is determined according to the first preset rule. The first preset rule determines the area formed by the boundary of the closed pattern from the user's eye position as the viewing angle.
[0095] If the trajectory information within the preset time period does not form a closed shape, a closed shape is generated based on the trajectory information and a second preset rule, and the target area corresponding to the closed shape is determined according to the first preset rule. The second preset rule is a pre-defined rule used to complete the trajectory information and form a closed shape. The second preset rule may include at least connecting the start and end points of the trajectory points, symmetrically expanding the trajectory information within the preset time period, and generating a closed shape that includes all the trajectory information.
[0096] In one implementation, when the distance between the starting point and the ending point of the trajectory information is within a preset distance range, the starting point and the ending point are connected to generate a closed graph.
[0097] In one implementation, when the distance between the starting point and the ending point is less than a preset distance range, a circular closed shape including all trajectory points is generated, using the center or key point that does not form a closed shape as the center and the farthest trajectory point as the radius or the preset distance as the radius. For example, if the user's hand does not move and the user's trajectory information is a single point with a starting and ending point distance of 0, then a circular target area is generated in a plane perpendicular to the user's viewpoint, using the starting point as the center and the preset distance as the radius.
[0098] In one implementation, when the distance between the starting point and the ending point is greater than a preset distance range, the trajectory information is symmetrically expanded according to the principle of symmetry, with the line connecting the starting point and the ending point as the axis of symmetry, to generate a closed figure.
[0099] It should be noted that the aforementioned preset distance range can be a distance range in real-world coordinates, or it can be a distance range used in image or sensor data to represent the distance between the two.
[0100] S13. Based on the images of the vehicle's exterior, determine the target image corresponding to the target area.
[0101] In one implementation, three-dimensional scene data is constructed based on images of the vehicle's exterior and interior, simulating the image seen through the target area from the user's perspective, obtaining image points seen through the target area from the user's perspective, and then determining the corresponding target image for the target area.
[0102] In the above implementation, the raw data for constructing the 3D scene can be obtained not only from images acquired by the vehicle, but also from data acquired by the vehicle's LiDAR. The data acquired by the vehicle for constructing the 3D scene can be obtained through a depth camera or a stereo camera.
[0103] In another implementation, the coordinates of the user's eye position and the target area in the coordinate system of the vehicle's internal camera are determined. Based on the matrix parameters of the internal and external cameras, the user's eye position and the target area are transformed into the coordinate system of the external camera. The image of the target area is then determined based on the image from the external camera. In a specific approach, the image captured by the external camera is converted into 3D scene data, and the image points seen through the target area from the user's perspective are determined, thereby identifying the corresponding target image for the target area.
[0104] S14. Display the target image.
[0105] This embodiment provides an image processing method based on gesture recognition in a vehicle. When a user meets preset conditions, the method acquires the trajectory information of the user's gestures within a preset time period; determines the target area outside the vehicle formed by the trajectory information, with the user's eye position as the observation point; determines the target image corresponding to the target area based on images outside and inside the vehicle; and displays the target image. This method allows users to acquire images of the outside of the vehicle from their perspective using gestures inside the vehicle, avoiding the problem of easily missing objects when taking photos outside the vehicle in existing technologies, thus improving the efficiency of acquiring images of the outside of the vehicle.
[0106] Figure 3 A flowchart illustrating a second embodiment of an image processing method based on gesture recognition in a vehicle provided in this application is shown below. Figure 3 As shown, the method includes the following steps:
[0107] S101. The system uses sensors to acquire in real time the trajectory information of the user's gestures inside the vehicle and the first position of the user's eyes in the vehicle coordinate system.
[0108] In this solution, it is necessary to recognize the closed trajectory drawn by the user's hand in the air. Therefore, after the vehicle starts, sensors inside the vehicle collect data in real time, including the real-time position and contour information of the user's hand and the position information of the eye. The real-time position information of the user's hand constitutes the trajectory information of the gesture. The position information of the hand and the eye are both in the same coordinate system. For example, a vehicle coordinate system with the geometric center of the vehicle as the center point can be used.
[0109] In one implementation, the in-vehicle sensors are infrared sensors that can detect and track the position and movements of the user's hands in real time. Infrared sensors typically sense objects and measure their distance through infrared light or infrared thermal radiation.
[0110] In another implementation, the in-vehicle sensor is a depth camera, which can provide depth information for each pixel and capture the three-dimensional shape and position of the user's hand in real time.
[0111] In the two implementation methods mentioned above, depth cameras and infrared sensors can convert the collected data into 3D point cloud data in real time, and can record and represent the shape and spatial position of objects in real time.
[0112] Alternatively, in another implementation, a common vision camera (RGB camera) can be used as an alternative. Common vision cameras can include monocular and binocular cameras. Using a binocular camera, the depth information of the object can be calculated to generate 3D point cloud data. Using a monocular camera, computer vision techniques (such as optical flow algorithms) can be used to analyze the movement of pixels within a continuous sequence of images to estimate the object's depth, thereby generating a partial 3D point cloud. This approach can reduce the implementation cost of the solution.
[0113] It should be noted that the user's eye position can be the center point of the line connecting the positions of the two eyes, or a certain area determined by the position of the two eyes; this application does not impose any restrictions on this.
[0114] S102. When determining that the gesture forms a closed shape in space based on the trajectory information, construct 3D point cloud data inside and outside the vehicle based on the vehicle coordinate system and the real-time images captured by the in-vehicle camera and the out-of-vehicle camera.
[0115] The determination of whether a closed shape has been formed based on the trajectory information generated from the real-time collected position information of the user's hand can be achieved in the following ways:
[0116] In one implementation, a gesture recognition algorithm is used to determine whether the trajectory information drawn by the user is closed. This gesture recognition algorithm is a neural network model pre-trained with multiple gesture action data. The input is point cloud data of the trajectory information, and the output is the result of whether it is a closed shape.
[0117] In another implementation, the trajectory information drawn by the user is mapped onto the two-dimensional plane where the user's eyes are located, and it is determined whether the distance between the starting point and the current position is within a preset error range. If it is within the preset error range, a closed figure is formed.
[0118] Both of the above methods require a starting point to determine the initial data for the trajectory information. This starting point is determined based on the 3D point cloud data of the user's hand contour. For example, when a user draws a trajectory, they will extend one finger. Therefore, the starting point is determined based on the data at the moment when the user extends one finger. The data from the starting point to the current time point is input into the algorithm model or mapped onto a two-dimensional plane. It should be noted that the hand contour in this solution is not specifically limited to the number of fingers or which finger. In some other possible methods, a shape can be pre-displayed to represent the starting point for drawing.
[0119] In another implementation, it's possible to determine whether a closed shape is formed using trajectory data within a preset time period. For example, a set of data is acquired every 10 seconds to determine if the user's hand trajectory information within those 10 seconds forms a closed shape. However, performing the judgment at preset time intervals might result in the initial data being in the process of trajectory drawing. To reduce the probability of this anomaly, the time is rolled back during the next data acquisition. For example, if a set of data is acquired from 00:00 to 00:10, and a second set from 00:10 to 00:20, the user's drawing time is approximately one second, so the probability of an anomaly is one in ten. By rolling back, if a set of data is acquired from 00:00 to 00:10, and a second set from 00:05 to 00:15, the probability of an anomaly is almost zero.
[0120] Based on the vehicle coordinate system, 3D point cloud data of the vehicle's interior and exterior are constructed using images captured in real-time by external cameras. This 3D point cloud data provides three-dimensional scene information of the vehicle's surrounding environment and interior. Specifically, 3D point cloud data conforming to the vehicle coordinate system is constructed based on images captured by external cameras and images captured by internal cameras. The two point cloud data sets are then registered together to construct a unified 3D point cloud data set.
[0121] It should be noted that the external cameras of the vehicle can also be the aforementioned depth cameras or multiple vision cameras.
[0122] Optionally, the external camera of the vehicle can also be a light field camera. The image captured by the light field camera records all optical information within the focal length range. Therefore, when the vehicle is moving at high speed, there is no need to focus, and a clear image can still be obtained in subsequent processing, which can avoid the problem of blurry images captured by the camera when the vehicle is moving at high speed.
[0123] S103. Based on 3D point cloud data and images captured by cameras outside the vehicle, determine the image of the target area pointed to by the user's gesture. The target area is the area with the first position as the observation point and the closed shape formed by the trajectory information as the range.
[0124] In this step, to determine the image of the vehicle's exterior as defined by the gesture trajectory, the 3D point cloud data of the exterior image is first obtained. Specifically, in the constructed 3D data of the vehicle's interior and exterior, taking the first position as the observation point, an extendable view frustum is constructed from the first position to the boundary of the closed shape, and the 3D point cloud data within the view frustum is acquired.
[0125] After acquiring 3D point cloud data within the view frustum, at least two processing methods are used to obtain an image of the target region:
[0126] The first implementation involves projecting 3D point cloud data within the view frustum onto the plane containing the closed shape, generating a projected 2D image. Then, based on the color and texture data of each point in the 3D point cloud data, and according to pre-set viewpoints and focal lengths, an image of the target area is rendered. This method produces an image that mirrors the user's perspective.
[0127] The second implementation involves extracting features from the 3D point cloud data within the view frustum using feature descriptors, such as Scale-invariant Feature Transform (SIFT) or Speeded-Up Robust Features (SURF), to obtain the first set of features. The original image captured by the vehicle's external camera is then also extracted using these feature descriptors to obtain the second set of features. A feature matching algorithm (e.g., nearest neighbor matching) is used to match the first and second sets of features. Based on the matching result, the image of the target region in the original image is obtained from the 3D point cloud data. For example, the nearest neighbor matching algorithm matches the feature descriptor in the second set of features with the highest similarity or the feature descriptor with the smallest Euclidean distance, based on one feature descriptor from the first set. The image obtained in this method is from the camera's perspective, resulting in high image fidelity, but it differs from the user's perspective.
[0128] Optionally, in the second implementation described above, the image of the target area that is different from the user's perspective is transformed into an image of the target area from the user's perspective by perspective projection based on the position of the original image camera, the coordinate position of the user's eye, and the camera parameters.
[0129] S104. Label the objects in the image of the target region to obtain the target image.
[0130] In this step, after obtaining the image of the target area, it needs to be labeled to facilitate user observation.
[0131] In one implementation, an object detection algorithm is used to detect at least one target object in the image of the target region. The target object is then marked in the image of the target region using a drawing tool. For example, the target object is circled in the image of the target region to obtain the target image.
[0132] In another implementation, the image of the target region is input into an image recognition model to identify all object information (e.g., name, price, etc.) in the image of the target region, and the object information is then labeled in the image of the target region to obtain the target image. The image recognition model is a deep learning model pre-trained on multiple labeled image data sets.
[0133] S105. Display the target image to other users in the vehicle.
[0134] In this step, the target image is displayed on the vehicle's central control screen for the driver's convenience, or it can be displayed on all the screens in the front and rear of the vehicle to facilitate communication between users.
[0135] Optionally, the target image can be pushed to a user-preset terminal device.
[0136] This embodiment provides an image processing method based on gesture recognition in a vehicle. When the user's gesture trajectory information determines that the gesture forms a closed shape in space, 3D point cloud data of the vehicle's interior and exterior is constructed using the vehicle coordinate system and real-time images captured by the in-vehicle and external cameras. Based on the 3D point cloud data and the images captured by the external cameras, the image of the target area pointed to by the user's gesture is determined. The target area is the region defined by the closed shape formed by the trajectory information, with a first position as the observation point. Objects in the target area image are labeled to obtain the target image. The target image is then displayed to other users in the vehicle. This method allows users to obtain the desired image simply by drawing a circle, making the operation convenient and improving the user's driving experience.
[0137] Based on Example 1, the following provides a detailed explanation of how to construct the overall 3D point cloud data of the vehicle's interior and exterior, and how to determine the point cloud data of the target area within the gesture trajectory range in the 3D point cloud data.
[0138] To obtain an image of a closed shape from the user's perspective, the best method is to take a picture at the user's eye level using a camera. However, the camera inside the vehicle cannot be located at the user's position. Therefore, obtaining an image of the target area from the user's perspective becomes a challenge. The following method constructs 3D point cloud data and renders it to obtain an image of the target area within the gesture trajectory range from the user's perspective.
[0139] Figure 4 A flowchart illustrating a third embodiment of an image processing method based on gesture recognition in a vehicle provided in this application is shown below. Figure 4 As shown, the method includes the following steps:
[0140] S201. Obtain the first depth information of the image from the perspective of the in-vehicle camera and the second depth information of the image from the perspective of the out-of-vehicle camera.
[0141] In this step, two images are acquired simultaneously from at least two monocular vision cameras on the same side of the vehicle exterior, capturing video frames in real time. Representative feature points (e.g., corner points, edge points of objects) are extracted from each image; these feature points possess a degree of uniqueness across different viewpoints. These feature points are then matched using feature descriptors (e.g., SIFT, ORB). Based on the matching relationships, the pixel displacements between the images, i.e., disparity, are calculated. Disparity values for other pixels are obtained using interpolation based on the disparity values of the representative feature points, or using a stereo matching algorithm. Based on the disparity values and known camera parameters (baseline, focal length), second depth information for points in the scene is calculated using triangulation or other depth estimation algorithms (e.g., back projection), where depth information represents the distance or relative position of an object point from the camera. Similarly, first depth information is obtained from images captured by the in-vehicle camera in the same manner.
[0142] When calculating the disparity of representative feature points, at least the following two methods can be used:
[0143] The first implementation method is the Disparity Method: For each pair of matched feature points, their pixel displacement in the image is calculated. Assuming the feature point coordinates in the left image are (x1, y1) and the feature point coordinates in the right image are (x2, y2), the disparity value can be calculated by subtracting the left image coordinates from the right image coordinates: Disparity = x2 - x1. This disparity value represents the horizontal displacement of the feature point, i.e., the difference in the object's position between the two viewpoints.
[0144] The second implementation method is the triangulation method: based on the positions of the matched feature points in the image and the geometric relationship with the camera, the disparity is calculated using the principles of triangulation. Assuming the feature point coordinates in the left image are (x1, y1), and the feature point coordinates in the right image are (x2, y2), the camera baseline length is B, and the focal length is f, then the horizontal disparity can be calculated as: disparity = (x1 - x2) * f / B. The resulting disparity value represents the depth difference of the feature points in three-dimensional space.
[0145] It should be noted that in both of the above methods, coordinate alignment needs to be performed in advance to convert the coordinates of the two images to the same camera coordinate system. This requires obtaining the transformation relationship between the two camera coordinate systems in advance through the parameters of the intrinsic matrix of the two cameras. This transformation relationship does not need to be changed if the cameras are not replaced.
[0146] Optionally, in this solution, all cameras inside and outside the vehicle can also use depth cameras to directly acquire the depth information of image pixels.
[0147] S202. Convert the image captured by the in-vehicle camera into first point cloud data based on the first depth information, and convert the image captured by the out-of-vehicle camera into second point cloud data based on the second depth information.
[0148] In this step, the camera's intrinsic and extrinsic parameter matrices are pre-obtained. The intrinsic parameter matrix describes the camera's internal properties, such as focal length, principal point coordinates, and pixel pitch, while the extrinsic parameter matrix describes the camera's pose, including rotation and translation vectors. Based on the camera's intrinsic and extrinsic parameter matrices, the camera's 3D coordinates in the camera coordinate system are calculated through the inverse operation of depth information and the camera's intrinsic parameter matrix. The points in the camera coordinate system are then converted to points (X, Y, Z) in the vehicle coordinate system based on the extrinsic parameter matrix. The converted 3D points (X, Y, Z) are organized into a point cloud data structure. Point cloud data is a 3D data representation of a set of points, each point having position information (x, y, z coordinates) and other attributes (such as color, normals, etc.).
[0149] Optionally, a light sensor can be installed on the exterior of the vehicle to acquire lighting information from outside the vehicle. Based on the lighting information, the normal vector parameters of the 3D point cloud can be set. This method can simulate a more realistic lighting effect.
[0150] S203. Register the first point cloud data and the second point cloud data to the vehicle coordinate system respectively to obtain 3D point cloud data inside and outside the vehicle.
[0151] Optionally, based on the vehicle's 3D model data, it can be displayed in the vehicle coordinate system. During observation, the effect of vehicle occlusion can be simulated, and the resulting image of the target area is more consistent with reality.
[0152] S204. Construct an extendable view frustum from the first position to the boundary of the closed shape. The extendable view frustum represents the range of 3D point cloud data visible from the viewpoint.
[0153] In this step, to determine the image seen from the user's perspective, it is first necessary to determine the 3D point cloud data seen by the user.
[0154] Specifically, taking the user's initial position as the viewpoint, the boundary of the closed shape defines the visible area. Using the viewpoint and the boundary of the closed shape, an initial view frustum is constructed, where the initial view frustum is a 3D closed shape with the closed shape as its base and the initial position as its vertex. The initial view frustum is then extended to include more 3D point cloud data; this process requires iterative operations.
[0155] First, based on the current view frustum, point cloud data located inside the view frustum are selected from the overall 3D point cloud data. Next, based on the selected point cloud data, the boundary of the view frustum is extended to include as much 3D point cloud data as possible. Visibility detection of the point cloud is performed using ray projection to determine which newly added point cloud data is included in the extended view frustum.
[0156] Then, repeat the above two steps. Through iterative operations, the view frustum will gradually extend to approximate the 3D point cloud data visible from the user's perspective. When a specific termination condition is met, the termination condition can be set to stop when the increase in the visible point cloud data is less than a preset value.
[0157] S205. Obtain target 3D point cloud data within an extendable view frustum range from 3D point cloud data.
[0158] S206. Project the target 3D point cloud data onto the plane of the closed graphic to generate a projected two-dimensional image.
[0159] In one implementation, an orthogonal projection matrix is used to directly map the coordinates of 3D point cloud data onto a plane, ignoring viewpoint and distance factors. The projected two-dimensional coordinates are the projected positions of the point cloud data on the plane.
[0160] In another implementation, the coordinates of the 3D point cloud data are converted to perspective projection coordinates using a perspective projection matrix. These coordinates are then mapped onto a plane by dividing them by the homogeneous coordinates of the perspective projection coordinates.
[0161] S207. Render the image of the target area based on the color data of each point in the projected two-dimensional image and the 3D point cloud data.
[0162] In this step, the projected 2D image is not enough to generate an image from the user's perspective. Color information is also needed. The RGB values included in the 3D point cloud data are associated with the points in the 2D image, and the color information is filled into the 2D image to render the image of the target area.
[0163] This embodiment provides an image processing method based on gesture recognition in a vehicle. From complete 3D point cloud data inside and outside the vehicle, an extendable view frustum is constructed from the user's perspective to determine the 3D point cloud data seen by the user, thereby determining the target area image seen by the user. The image obtained through this method is from the user's perspective, which is more accurate than methods such as cropping images from camera captures.
[0164] In the above embodiment three, an image of the target area was obtained, which can be directly used for display. However, to increase the convenience of communication between users, the image of the target area can be further annotated. The following is a detailed description of the further processing procedure after obtaining the image of the target area.
[0165] Figure 5 A flowchart illustrating Embodiment 4 of the image processing method based on gesture recognition in a vehicle provided in this application is shown below. Figure 5 As shown, the method includes the following steps:
[0166] S301. Obtain the viewpoint graphic formed in space by the trajectory information of the user's gestures from the first position viewpoint.
[0167] In this solution, different shapes formed by user gesture trajectories can represent different meanings. For example, drawing a circle indicates acquiring information about an object, while drawing a square indicates acquiring the entire image. Therefore, it is necessary to identify the shapes formed by the user's gesture trajectory information. However, shapes drawn in the air will appear differently when viewed from different angles. For instance, what appears as a circle at the user's position may appear as a line from a 90-degree angle to the user's side. Therefore, it is necessary to determine the spectral shape formed by the user's trajectory information in the air from the user's perspective based on 3D point cloud data.
[0168] S302. Compare the viewpoint graphic with the preset graphic to obtain the comparison result.
[0169] In this step, different preset graphics can represent different subsequent execution instructions. In this embodiment, only one preset graphic is compared, but this solution is not limited to only one.
[0170] If the comparison results are consistent, proceed to steps S303-S304; if the comparison results are inconsistent, proceed to steps S305-S306.
[0171] S303. Identify object information in the image of the target area using an image recognition model.
[0172] In this step, an image recognition model trained on multiple labeled image data sets is prepared in advance. This model can be a deep learning-based model, such as a convolutional neural network (CNN), or other models suitable for image classification tasks. The image of the target region is used as input to the image recognition model for object information identification. The image recognition model outputs at least one recognition result, namely the identified object name, category, and price.
[0173] Optionally, the image of the target area can be used to search for information about the object on the Internet. For example, the image of the target area can be entered into search software or shopping software to obtain information about all objects in the image.
[0174] S304. Mark the corresponding region of the object information in the image of the target region to obtain the target image.
[0175] In this step, the object information obtained in step S303 is displayed next to the corresponding object, or a marker area is provided in the image, and the object information is expanded by touching the marker area.
[0176] Through steps S303-S304, the information of the object is marked in the image of the target area, which makes it convenient for users to obtain the information of the object in a timely manner without having to perform a separate query.
[0177] S305. Identify at least one target object in the image of the target region using a target detection algorithm.
[0178] In this step, a trained object detection model is required. This model can be based on deep learning algorithms such as Faster R-CNN, YOLO, or SSD. These models can locate and identify multiple target objects in an image. The image of the target region is used as input for the object detection model to identify the target object. The model processes the image and outputs the recognition results, including the category of the detected target object, the bounding box location, and the confidence score.
[0179] S306. In the image of the target area, at least one target object is identified by drawing tools to obtain the target image.
[0180] In this step, based on the target detection results in step S305, the category and bounding box position of at least one identified target object are obtained. Using drawing tools, such as rectangle or polygon drawing tools, the corresponding target object is identified in the image of the target region according to the bounding box position of the target object.
[0181] Optional: If needed, text labels can be added around the label box to indicate the category or other relevant information of the target object.
[0182] Through steps S305-S306, at least one target object is identified in the image of the target area using a drawing tool, resulting in a target image. This visually presents the position and shape of the target object in the image, facilitating user observation.
[0183] This embodiment provides an image processing method based on gesture recognition in a vehicle. Different processing is performed on the image within the target area according to different shapes drawn by the user's gesture. The object is selected in the image of the target area for easy observation by the user. The object information is directly displayed on the image so that the user can easily obtain the object information. The above method can improve the user experience.
[0184] The following example illustrates the entire solution in detail: a rear-seat passenger in a moving vehicle uses gestures to identify items on a billboard in the road.
[0185] Figure 6 A flowchart illustrating an example of a gesture recognition-based image processing method in a vehicle provided in this application is shown below. Figure 6 As shown, the method includes the following steps:
[0186] S401. Obtain the facial data of the user in the vehicle and determine whether the user has gesture interaction permission.
[0187] In this step, sensors or cameras inside the vehicle periodically acquire facial images of the users inside the vehicle after the vehicle is turned on.
[0188] Based on the user's facial image and the preset user permission table for gesture interaction, the system compares whether the facial information of the authorized user in the gesture interaction user permission table includes the facial information in the user's facial image to determine whether the user has the permission to perform human-computer interaction through gestures.
[0189] In this way, the awkward experience of triggering gesture interaction when a stranger gets into a vehicle can be avoided.
[0190] S402, The user draws a circle in the air with their index finger over the billboard.
[0191] In this step, after the user authorizes the gesture interaction, when interested in the billboard outside the vehicle, they can draw a circle with their index finger in the air inside the vehicle from the user's perspective.
[0192] S403: The system uses sensors to acquire in real time the trajectory information of the user's gestures inside the vehicle and the first position of the user's eyes in the vehicle coordinate system.
[0193] This step is similar to step S101 and will not be repeated here.
[0194] S404. Input the 3D point cloud data of the hand and the 3D point cloud data of the hand movement through each trajectory point into the gesture recognition algorithm in real time for gesture recognition.
[0195] In this step, a gesture recognition algorithm is used to identify the user's hand contour and the graphic formed by the gesture trajectory in real time. This gesture recognition algorithm is trained using multiple gesture action data in advance, through methods such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs), to recognize the graphic formed by the hand contour and gesture trajectory. During training, the hand contour and trajectory data need to be associated with corresponding graphic labels so that the model can learn the features and patterns of the gesture.
[0196] S405. When the hand outline is determined to be the shape of an index finger, check whether the shape formed by the gesture trajectory is a closed shape.
[0197] In this step, when the hand outline is the preset shape of an extended index finger, it indicates that the user has started to draw a graphic with their hand. Therefore, the graphic formed by the gesture trajectory is obtained in real time based on the 3D point cloud data of each trajectory of the hand, and it is determined whether a closed shape is formed.
[0198] If a closed shape is formed, proceed to step S406. If no closed shape is formed, continue the detection until a closed shape is formed or the user's hand outline does not conform to the preset outline.
[0199] Understandably, closed shapes and closed figures refer to the same thing.
[0200] S406. Determine the time to acquire the target area image.
[0201] In this step, when the user's hand trajectory forms a closed shape, it indicates that the user's operation is complete. We only need to obtain the image from the user's perspective from the image data. However, determining which moment's image data to use is a challenge. Several approaches can be taken:
[0202] The first method involves acquiring image data from both inside and outside the vehicle's cameras when the user-drawn shape is closed. This method is suitable for scenarios where the vehicle is moving slowly or stationary.
[0203] The second method involves acquiring image data from both inside and outside the vehicle's cameras when the user's drawn shape is closed, based on the user's hand outline conforming to a preset outline. This ensures that the acquired image does not deviate from the closed shape due to excessive vehicle speed, thus improving the accuracy of image acquisition.
[0204] S407. Construct 3D point cloud data to obtain an image of the target area.
[0205] This process is similar to that in Example 2, and will not be described in detail here.
[0206] S408. Obtain the image processing method based on the preset circle.
[0207] In this step, the user-drawn circular shape is obtained. Based on the preset graphic instruction table that includes multiple shapes, when the current circular shape is determined, the instruction to be executed on the target image is to obtain object information in the target area image.
[0208] S409. Input the image of the target area into the image recognition model to obtain information about the billboard.
[0209] In this step, the image of the target area containing the billboard is input into the image recognition model to obtain information from the image, including mobile phone information (e.g., model) and spokesperson information (e.g., name, age).
[0210] Optionally, an image of the target area can be input into the shopping app to obtain the phone's price information.
[0211] S410. Annotate the object information in the target area image and display it on the vehicle's central control screen.
[0212] In this step, the information from the mobile phone and the spokesperson is displayed in an image of the target area and sent to the vehicle's central control screen for the driver to view.
[0213] This example provides a gesture recognition-based image processing method for vehicles. While driving, the system captures an image of a billboard by having the user draw a circle in the air with their hand. The object information of the billboard is then extracted and displayed in the image for the driver to observe. In this way, when a user sees an object of interest, its image can be readily and conveniently displayed, enhancing the user's driving experience.
[0214] Figure 7 This application provides a schematic diagram of the structure of an image processing device based on gesture recognition in a vehicle, as shown in Embodiment 1. Figure 7 As shown, the device 500 includes:
[0215] The acquisition module 501 is used to acquire the trajectory information of the user's gestures within a preset time period when the user meets the preset conditions.
[0216] The first determining module 502 is used to determine the target area outside the vehicle formed by the trajectory information, wherein the target area is observed from the user's eye position.
[0217] The second determining module 503 is used to determine the target image corresponding to the target area based on the image of the vehicle exterior;
[0218] Display module 504 is used to display the target image.
[0219] Optionally, the preset conditions include: the shape of the user's hand outline conforms to a preset shape and / or the user's voiceprint conforms to a preset voiceprint.
[0220] Optionally, the first determining module 502 is specifically used for:
[0221] Determine whether the trajectory information forms a closed shape;
[0222] If the trajectory information forms a closed shape, the target area corresponding to the closed shape is determined according to the first preset rule;
[0223] If the trajectory information does not form a closed shape, a closed shape is generated according to the trajectory information and the second preset rule, and the target area corresponding to the closed shape is determined according to the first preset rule.
[0224] Optionally, the second determining module 503 is specifically used for:
[0225] The target image of the target area is determined by constructing three-dimensional scene data based on real-time data acquired by the vehicle.
[0226] Optionally, the second determining module 503 includes an acquisition unit, a 3D construction unit, and an image acquisition unit:
[0227] The acquisition unit is used to acquire the trajectory information of the user's gestures inside the vehicle and the first position of the user's eyes, wherein the first position is used to represent the three-dimensional coordinates of the eyes in the vehicle;
[0228] The three-dimensional construction unit is used to construct 3D point cloud data of the vehicle's interior and exterior based on data acquired in real time. The 3D point cloud data provides three-dimensional scene information of the vehicle's surrounding environment and interior.
[0229] The image acquisition unit is used to determine the target image of the target area pointed to by the user's gesture based on the 3D point cloud data and the image captured by the camera outside the vehicle.
[0230] The three-dimensional building unit is specifically used for:
[0231] Based on the real-time images captured by the in-vehicle camera and the real-time images captured by the out-of-vehicle camera, parallax calculation is performed to obtain the first depth information of the image from the perspective of the in-vehicle camera and the second depth information of the image from the perspective of the out-of-vehicle camera.
[0232] The image captured by the in-vehicle camera is converted into first point cloud data based on the first depth information, and the image captured by the out-of-vehicle camera is converted into second point cloud data based on the second depth information.
[0233] The first point cloud data and the second point cloud data are respectively registered to the same coordinate system to obtain the 3D point cloud data inside and outside the vehicle.
[0234] The three-dimensional building unit is also used for:
[0235] When the user meets the preset conditions, 3D point cloud data of the inside and outside of the vehicle is constructed from the images captured by the in-vehicle camera and the out-of-vehicle camera at the moment when the user begins to draw graphics in space based on the gesture.
[0236] or,
[0237] When the user meets the preset conditions, 3D point cloud data of the inside and outside of the vehicle is constructed based on the images captured by the vehicle's internal and external cameras at the moment when the user finishes drawing graphics in space according to the gesture.
[0238] The image acquisition unit is specifically used for:
[0239] An extendable view frustum is constructed from the first position to the boundary of the target region, the extendable view frustum representing the range of the 3D point cloud data visible from the viewpoint.
[0240] Obtain the target 3D point cloud data within the extendable view frustum range from the 3D point cloud data;
[0241] The target image is obtained based on the 3D target point cloud data.
[0242] The image acquisition unit is also used for:
[0243] The target 3D point cloud data is projected onto the plane of the graphic to generate a projected two-dimensional image;
[0244] The target image is rendered based on the projected two-dimensional image and the color data of each point in the 3D point cloud data.
[0245] or,
[0246] The target image is obtained from images captured by an external camera of the vehicle using feature descriptors based on the 3D target point cloud data.
[0247] The display module is also used for:
[0248] Objects in the image of the target region are labeled to obtain the target image.
[0249] The display module is also used for:
[0250] The image recognition model identifies object information in the image of the target region by means of an image recognition model, which is a deep learning model pre-trained on multiple labeled image data;
[0251] The target image is obtained by annotating the region corresponding to the object information in the image of the target region.
[0252] or,
[0253] At least one target object is identified in the image of the target region using a target detection algorithm;
[0254] The at least one target object is identified in the image of the target area using a drawing tool to obtain the target image.
[0255] The display module is also used for:
[0256] Obtain the spectral graphic formed in space by the trajectory information of the user's gestures from the first position viewpoint;
[0257] The perspective image is compared with a preset image to obtain the comparison result;
[0258] If the comparison results are consistent, the object information is marked in the image of the target area to obtain the target image;
[0259] If the comparison results are inconsistent, at least one target object is identified in the image of the target area to obtain a target image. The image processing device for vehicles based on gesture recognition provided in this embodiment is used to execute the technical solutions of any of the above method embodiments. Its implementation principle and technical effects are similar and will not be elaborated upon here.
[0260] Figure 8 A structural schematic diagram of a vehicle provided in this application, such as Figure 8 As shown, the vehicle 600 includes: a vehicle body, sensors, cameras, a processor, and communication interfaces for interacting with other devices.
[0261] The vehicle body 611, sensor 612, processor 613, camera 614, and communication interface 615 for interacting with other devices, wherein the processor is used to execute the gesture recognition-based image processing method in the vehicle as described in any of the above method embodiments.
[0262] Optionally, the various devices in the vehicle 600 can be connected via a system bus.
[0263] Optionally, the vehicle also includes a memory 616, which stores processor-executed instructions, data acquired by cameras and sensors, gesture recognition algorithm models, image recognition models, etc.
[0264] The memory can be a separate storage unit or a storage unit integrated into the processor 613.
[0265] Optionally, the aforementioned 614 camera can be a depth camera or a vision camera. In some embodiments, the camera can also be a light field camera.
[0266] Optionally, the vehicle may also include a display for showing the processor's processing results and for human-machine interaction. In some embodiments, the display may be a central control screen for the vehicle; in other embodiments, the display may be a flexible display, or even a non-rectangular, irregularly shaped display, i.e., a non-rectangular screen. The display may be made of materials such as liquid crystal display (LCD) or organic light-emitting diode (OLED).
[0267] It should be understood that the processor 613 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0268] The system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. Memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0269] All or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof.
[0270] The vehicle provided in this application embodiment can be used to execute the gesture recognition-based image processing method in any of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.
[0271] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the gesture-based image processing method in a vehicle as described in any of the foregoing method embodiments.
[0272] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0273] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, it can implement the gesture recognition-based image processing method in a vehicle provided in any of the foregoing embodiments.
[0274] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0275] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. An image processing method based on gesture recognition in a vehicle, characterized in that, The method includes: When a user is detected to meet a preset condition, the trajectory information of the user's gestures within a preset time period is obtained; Determine the target area outside the vehicle formed by the trajectory information, with the user's eye position as the observation point for the target area; Based on the image of the vehicle's exterior, determine the target image corresponding to the target area; Display the target image; The step of determining the target image corresponding to the target region based on the image of the vehicle's exterior includes: Obtain the trajectory information and the three-dimensional coordinate position of the user's eyes in the vehicle coordinate system; Based on the vehicle coordinate system, images captured in real time by the vehicle's internal and external cameras, 3D point cloud data of the vehicle's interior and exterior is constructed; the 3D point cloud data provides three-dimensional scene information of the vehicle's surrounding environment and interior. Using the three-dimensional coordinate position of the user's eye as the vertex and the boundary of the closed shape drawn corresponding to the trajectory information as the base, an initial visual cone is constructed; the initial visual cone is extended until a termination condition is met to obtain the extended visual cone; the termination condition is: the increase in 3D point cloud data is less than a preset value; Obtain the target 3D point cloud data within the extended view cone range from the 3D point cloud data inside and outside the vehicle; Obtain the target image based on the target 3D point cloud data.
2. The method according to claim 1, characterized in that, The preset conditions include: the shape of the user's hand outline conforms to a preset shape and / or the user's voiceprint conforms to a preset voiceprint.
3. The method according to claim 1, characterized in that, Determining the target area outside the vehicle formed by the trajectory information includes: Determine whether the trajectory information forms a closed shape; If the trajectory information forms a closed shape, the target area corresponding to the closed shape is determined according to the first preset rule; the first preset rule is the area formed by the boundary of the closed shape with the user's eye position as the viewing angle. If the trajectory information does not form a closed shape, a closed shape is generated according to the trajectory information and the second preset rule, and the target area corresponding to the closed shape is determined according to the first preset rule; the second preset rule is used to complete the trajectory information to form a closed shape.
4. The method according to claim 1, characterized in that, The step of constructing 3D point cloud data of the vehicle's interior and exterior based on real-time images captured by the vehicle coordinate system, the in-vehicle camera, and the external camera includes: Based on the real-time images captured by the in-vehicle camera and the real-time images captured by the out-of-vehicle camera, parallax calculation is performed to obtain the first depth information of the image from the perspective of the in-vehicle camera and the second depth information of the image from the perspective of the out-of-vehicle camera. The image captured by the in-vehicle camera is converted into first point cloud data based on the first depth information, and the image captured by the out-of-vehicle camera is converted into second point cloud data based on the second depth information. The first point cloud data and the second point cloud data are respectively registered to the vehicle coordinate system to obtain the 3D point cloud data inside and outside the vehicle.
5. The method according to claim 4, characterized in that, The step of calculating parallax based on real-time images captured by the in-vehicle camera and real-time images captured by the out-of-vehicle camera includes: When the user meets the preset conditions, the parallax is calculated based on the images captured by the vehicle's internal camera and external camera at the moment when the user begins to draw the graphic in space according to the gesture. or, When the user meets the preset conditions, the parallax is calculated based on the images captured by the vehicle's internal and external cameras at the moment when the user finishes drawing the graphic in space according to the gesture.
6. The method according to claim 1, characterized in that, The step of obtaining the target image based on the target 3D point cloud data includes: The target 3D point cloud data is projected onto the plane of the drawn closed shape to generate a projected two-dimensional image; The target image is rendered based on the projected two-dimensional image and the color data of each point in the 3D point cloud data. or, The target image is obtained from images captured by an external camera of the vehicle using feature descriptors based on the target 3D point cloud data.
7. The method according to claim 6, characterized in that, Displaying the target image on the vehicle's display screen includes: Objects in the image of the target region are labeled to obtain the target image.
8. The method according to claim 7, characterized in that, The step of labeling objects in the image of the target region to obtain the target image includes: The image recognition model identifies object information in the image of the target region by means of an image recognition model, which is a deep learning model trained on multiple labeled image data; The target image is obtained by annotating the region corresponding to the object information in the image of the target region. or, At least one target object is identified in the image of the target region using a target detection algorithm; The at least one target object is identified in the image of the target area using a drawing tool to obtain the target image.
9. The method according to claim 8, characterized in that, The step of labeling objects in the image of the target region to obtain the target image includes: Obtain the spectral graphic formed in space by the trajectory information of the user's gestures from the first position viewpoint; The perspective image is compared with a preset image to obtain the comparison result; If the comparison results are consistent, the object information is marked in the image of the target area to obtain the target image; If the comparison results are inconsistent, at least one target object is identified in the image of the target area to obtain a target image.
10. An image processing device based on gesture recognition in a vehicle, characterized in that, The device includes: The acquisition module is used to acquire the trajectory information of the user's gestures within a preset time period when the user meets the preset conditions. The first determining module is used to determine the target area outside the vehicle formed by the trajectory information, wherein the target area is observed from the user's eye position. The second determining module is used to determine the target image corresponding to the target area based on the image of the vehicle's exterior; Display module, used to display the target image; The second determining module is specifically used to acquire the trajectory information and the three-dimensional coordinate position of the user's eye in the vehicle coordinate system; construct 3D point cloud data inside and outside the vehicle based on the vehicle coordinate system, images captured in real time by the in-vehicle camera and the out-of-vehicle camera; the 3D point cloud data provides three-dimensional scene information of the vehicle's surrounding environment and interior; construct an initial visual cone with the three-dimensional coordinate position of the user's eye as the vertex and the boundary of the closed shape drawn corresponding to the trajectory information as the base; extend the initial visual cone until a termination condition is met to obtain an extended visual cone; the termination condition is: the increase in 3D point cloud data is less than a preset value; acquire target 3D point cloud data within the range of the extended visual cone from the 3D point cloud data inside and outside the vehicle; and acquire a target image based on the target 3D point cloud data.
11. A vehicle, characterized in that, The vehicles include: The vehicle body, sensors, cameras, processor, and communication interfaces for interacting with other devices, the processor being used to execute the gesture recognition-based image processing method in a vehicle as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the gesture-based image processing method in a vehicle as described in any one of claims 1 to 9.