Vehicle control method and device, electronic equipment and storage medium

Through the combined method of image acquisition and gesture detection combined with identity verification, the problem of insufficient establishment of interactive objects in out-of-vehicle scenarios is solved, the accuracy and efficiency of vehicle interactive objects are improved, and security and privacy are ensured.

CN120340002APending Publication Date: 2025-07-18CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510380188.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the off-vehicle scenario, the existing technology lacks effective interactive object establishment and function initialization schemes, resulting in a single external interaction method and unable to fully support user needs.

Method used

The candidate body object is determined through image acquisition, rough and fine gesture action detection is performed, combined with identity verification, and the vehicle is granted interactive permissions and control target components.

Benefits of technology

It improves the accuracy and efficiency of vehicle interactive objects, ensures the safety and privacy of interactive objects, and improves the operation efficiency of vehicle control methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340002A_ABST
    Figure CN120340002A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle control method and device, electronic equipment and a storage medium, and the method comprises the steps: determining a candidate human body object from a collected image, carrying out the motion detection of the candidate human body object, obtaining a motion detection result, and carrying out the control of a vehicle when determining that the candidate human body object meets a preset vehicle interaction authority granting condition according to the motion detection result. According to the method, the human body objects in the collected images are detected twice, then the candidate human body objects are endowed with the authority of interaction with the vehicle, then the target parts of the vehicle are controlled according to the actions of the candidate human body objects in the interaction process with the vehicle, and the accuracy of the finally determined interaction objects is improved; the candidate human body objects are determined firstly, and the objects having the interaction intention with the vehicle are preliminarily screened, so that the subsequent action detection process can be focused on the objects having the interaction intention with the vehicle, action detection on irrelevant human body objects can be avoided, the interaction object establishment efficiency is improved, and the detection accuracy is improved. And the operation efficiency of the vehicle control method is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle control technology, and in particular to a vehicle control method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of automobile intelligence, the application of human-computer interaction technology in the automotive field is becoming increasingly rich. For example, when a user controls certain functions of a vehicle, the vehicle computer needs to establish the target object that is currently interacting with the vehicle in order to prevent interference from other objects. Therefore, how to establish the target object that is currently interacting with the vehicle to control the vehicle has become a technical problem that needs to be solved urgently. Summary of the invention

[0003] Based on this, a vehicle control method, device, electronic device and storage medium are provided to solve the above technical problems.

[0004] In a first aspect, a vehicle control method is provided, comprising:

[0005] Capturing an image of a target area of the vehicle to obtain a captured image;

[0006] In response to the number of human objects in the acquired image being greater than 1, determining a candidate human object from the acquired image;

[0007] Performing motion detection on the candidate human object to obtain a motion detection result;

[0008] When it is determined according to the action detection result that the candidate human object meets a preset vehicle interaction permission granting condition, granting the candidate human object permission to interact with the vehicle;

[0009] A target component of the vehicle is controlled according to the action of the candidate human object in the process of interacting with the vehicle.

[0010] In one embodiment, in response to the number of human objects in the acquired image being greater than 1, determining a candidate human object from the acquired image comprises:

[0011] Determine a human object detection frame from the acquired image;

[0012] In response to the presence of a target scene image region in the acquired image, and the number of the human object detection frames in the target scene image region is greater than 1, determining the human object in the human object detection frame closest to the center point of the target scene image region as the candidate human object; or

[0013] In response to the number of the detected human object bounding boxes being greater than 1, determine the human object in the detected human object bounding box with the largest area as the candidate human object, or determine the human object in the detected human object bounding box closest to the center point of the acquired image as the candidate human object.

[0014] In one embodiment, the determining the candidate human object from the acquired image includes:

[0015] When it is determined through a rough gesture action detection algorithm that the gesture action of a certain human object in the acquired image meets the rough matching condition of the target gesture action, determine the human object as the candidate human object;

[0016] The performing action detection on the candidate human object to obtain an action detection result includes:

[0017] Judge whether the gesture action of the candidate human object meets the fine matching condition of the target gesture action through a fine gesture action detection algorithm to obtain a judgment result; the judgment result is the action detection result.

[0018] In one embodiment, the target gesture action is a raising hand action; the determining that the gesture action of a certain human object in the acquired image meets the rough matching condition of the target gesture action through a rough gesture action detection algorithm includes:

[0019] Obtain the angle between the upper arm and the lower arm of the human object;

[0020] In response to the angle being greater than a preset angle threshold, determine that the gesture action of the human object meets the rough matching condition of the target gesture action; or, in response to the angle being greater than the angle threshold and the hand key point of the human object being above the face key point of the human object, determine that the gesture action of the human object meets the rough matching condition of the target gesture action, where the hand key point and the face key point are obtained by performing human key point detection on the acquired image.

[0021] In one embodiment, the obtaining the angle between the upper arm and the lower arm of the human object includes:

[0022] Perform human key point detection on the acquired image to obtain the elbow key point, the wrist key point and the shoulder key point of the human object;

[0023] Obtain a first vector between the elbow key point and the wrist key point, and a second vector between the elbow key point and the shoulder key point;

[0024] Calculate the vector angle between the first vector and the second vector, where the vector angle is the angle between the upper arm and the lower arm of the human object.

[0025] In one embodiment, the method of determining whether the gesture action of the candidate human object meets the fine matching condition of the target gesture action through the fine gesture action detection algorithm, and obtaining a judgment result includes:

[0026] Extract the gesture action features of the candidate human object from the captured image; the gesture action features include the features of the fine gesture action of the candidate human object;

[0027] Determine the gesture action of the candidate human object according to the gesture action features and the trained gesture action recognition model;

[0028] Judge whether the gesture action is the target gesture action;

[0029] If so, determine that the gesture action meets the fine matching condition of the target gesture action;

[0030] If not, determine that the gesture action does not meet the fine matching condition of the target gesture action.

[0031] In one embodiment, when it is determined according to the action detection result that the candidate human object meets the preset vehicle interaction permission granting condition, granting the candidate human object the permission to interact with the vehicle includes:

[0032] When the judgment result is yes, obtain verification information for verifying the identity of the candidate human object;

[0033] When it is determined that the identity of the candidate human object is legal according to the verification information, grant the candidate human object the permission to interact with the vehicle.

[0034] In one embodiment, the verification information includes at least one of facial image information, voice information, and password information input through the projection interface of the vehicle; determining that the identity of the candidate human object is legal according to the verification information includes:

[0035] In response to at least one of the verification information matching the corresponding preset verification information, determine that the identity of the candidate human object is legal.

[0036] In a second aspect, the present application provides a vehicle control device, and the device includes:

[0037] An acquisition module, configured to perform image acquisition on a target area of the vehicle to obtain a captured image;

[0038] A determination module, configured to determine a candidate human object from the acquired image in response to the number of human objects in the acquired image being greater than 1;

[0039] A detection module, configured to perform action detection on the candidate human object to obtain an action detection result;

[0040] An authorization module, configured to grant the candidate human object the permission to interact with the vehicle when it is determined according to the action detection result that the candidate human object meets a preset vehicle interaction permission granting condition;

[0041] A control module, configured to control a target component of the vehicle according to the actions during the interaction between the candidate human object and the vehicle.

[0042] In a third aspect, the present application provides an electronic device, including a processor and a memory, where a computer program is stored in the memory, and the processor executes the computer program to implement the vehicle control method in the first aspect above.

[0043] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the vehicle control method in the first aspect above is implemented.

[0044] The vehicle control method, device, electronic device, and storage medium provided by the present application determine a candidate human object from the acquired image, perform action detection on the candidate human object to obtain an action detection result, and grant the candidate human object the permission to interact with the vehicle when it is determined according to the action detection result that the candidate human object meets a preset vehicle interaction permission granting condition; in this process, the candidate human object is determined first, and then action detection is performed on the candidate human object, which is equivalent to performing two detections on the human objects in the acquired image, improving the accuracy of the finally established interaction object; in addition, by first determining the candidate human object and preliminarily screening the objects with the intention of interacting with the vehicle, the subsequent action detection process can focus on the objects with the intention of interacting with the vehicle, and compared with directly performing action detection on the human objects in the acquired image, it can avoid performing action detection on irrelevant human objects, improving the efficiency of establishing the interaction object, and further improving the operation efficiency of the vehicle control method.

[0045] Other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0047] Figure 1 It is a schematic flow chart of the vehicle control method in the first embodiment;

[0048] Figure 2 It is a schematic diagram of the first action of the human object in the first embodiment;

[0049] Figure 3 It is a schematic diagram of the second action of the human object in the first embodiment;

[0050] Figure 4 It is a schematic flow chart of obtaining the angle between the upper arm and the lower arm in the first embodiment;

[0051] Figure 5 It is a schematic structural diagram of the vehicle control device in the second embodiment;

[0052] Figure 6 It is a schematic structural diagram of the electronic device in the third embodiment. Specific Embodiments

[0053] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0054] Embodiment 1:

[0055] Currently, the establishment of interaction objects in most automotive intelligent interaction technologies mainly focuses on the interior of the cockpit. For example, by defining specific dynamic gestures, users can achieve interactive operations on the in-vehicle screen, while completing user identity locking and function initialization; for another example, when performing certain cockpit functions, the user identity is confirmed through facial image information. These technologies have played an important role in improving the convenience and intelligence level of in-vehicle operations, providing users with a more efficient and immersive cockpit experience.

[0056] However, compared with the mature interaction technologies in the cockpit, the exploration of the establishment of interaction objects and function initialization in the current out-of-vehicle scenarios is still insufficient, and no effective solutions have been provided for the recognition and dynamic initialization of interaction objects, resulting in overly single out-of-vehicle interaction methods and insufficient support for user needs.

[0057] In view of this, an embodiment of the present application provides a vehicle control method. It should be noted that the vehicle control method in the embodiment of the present application can be applied to an out-of-vehicle scenario. Of course, it can also be applied to an in-vehicle scenario. Please refer to Figure 1 As shown, the method includes:

[0058] S11: Perform image acquisition on a target area of the vehicle to obtain an acquired image.

[0059] S12: In response to the number of human objects in the acquired image being greater than 1, determine candidate human objects from the acquired image.

[0060] S13: Perform action detection on the candidate human objects to obtain an action detection result.

[0061] S14: When it is determined according to the action detection result that the candidate human objects meet the preset vehicle interaction permission granting conditions, grant the candidate human objects the permission to interact with the vehicle.

[0062] S15: Control the target components of the vehicle according to the actions during the interaction between the candidate human objects and the vehicle.

[0063] Next, the above steps will be introduced in detail.

[0064] In the embodiment of the present application, image acquisition can be performed through an image acquisition device on the vehicle to obtain at least one acquired image. The image acquisition area of the image acquisition device is the target area. The target area in the embodiment of the present application is an out-of-vehicle area. In other embodiments, the target area can also be an in-vehicle area.

[0065] In step S12, it is necessary to determine candidate human objects with an interaction intention with the vehicle from the acquired image.

[0066] In one embodiment, a human object with the highest interaction intention with the vehicle can be determined from the acquired image and used as the candidate human object.

[0067] In one embodiment, all human objects with an interaction intention with the vehicle can be screened out as candidate human objects.

[0068] Specifically, for the above step S12, the candidate human objects can be determined from the acquired image in the following manner:

[0069] Method 1: Determine a human object detection frame from the acquired image; in response to the presence of a target scene image area in the acquired image and the number of human object detection frames in the target scene image area being greater than 1, determine the human object in the human object detection frame closest to the center point of the target scene image area as the candidate human object.

[0070] It can be understood that when the number of human object detection boxes in the target scene image area is 1, the human object in the human object detection box is directly determined as the candidate human object.

[0071] The center point of the target scene image area refers to the point located at the center of the target scene image area. However, it should be noted that the target scene image area is not an absolutely regular image area, and the center point of the target scene image area is not the center in an absolute sense for this image area. For example, when the target scene image area is the image area corresponding to a stage, the point in the very middle of the stage can be used as the center point of the target scene image area here.

[0072] In this embodiment, the human object detection box can be determined through the target box detection model; it should be noted that the target box detection model in the embodiments of this application is used to generate, for each human object in the acquired image, a human object detection box that can cover the corresponding human object. The structure of the target box detection model in the embodiments of this application can include but is not limited to at least one of FCOS (Fully Convolutional One-Stage Object Detection), R-CNN (Region-based Convolutional Neural Networks), YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), and deformable convolutional networks.

[0073] In this embodiment, it can be detected through the target scene detection model whether there is a target scene image area in the acquired image. The target scene image area refers to the area in the image where there is a target scene, and the target scene refers to a scene that meets the preset scene conditions, and the preset scene conditions can be flexibly set by developers.

[0074] Exemplarily, the target scene here can be a stage scene. For example, it can be a virtual stage scene projected through virtual imaging technology, or a stage scene that meets the preset lighting effects.

[0075] It can be understood that when the user has an intention to interact with the vehicle, the user usually tends to approach the center point of the target scene. Therefore, in this embodiment, the human object in the human object detection box closest to the center point of the target scene image area is used as the candidate human object, so that the confirmed candidate human object is exactly the user who needs to interact with the vehicle, improving the accuracy of candidate human object confirmation.

[0076] Method 2: Determine the human object detection frame from the captured image. In response to the number of human object detection frames in the captured image being greater than 1, determine the human object in the human object detection frame closest to the center point of the captured image as the candidate human object.

[0077] In practical applications, when a user has an intention to interact with the vehicle, the user usually tends to stand directly in front of the image acquisition device. Therefore, in this embodiment, the human object in the human object detection frame closest to the center point of the captured image is used as the candidate human object to improve the accuracy of candidate human object confirmation.

[0078] Method 3: Determine the human object detection frame from the captured image. In response to the number of human object detection frames being greater than 1, determine the human object in the human object detection frame with the largest area as the candidate human object.

[0079] The human object detection frame with the largest area means that the human object satisfies a sufficient proximity to the image acquisition device, thus realizing the determination of this human object as the candidate human object with the intention of interacting with the vehicle.

[0080] It can be understood that when the number of human object detection frames determined from the captured image is 1, directly determine the human object in this human object detection frame as the candidate human object.

[0081] In one embodiment, the vehicle can independently select the method for determining the candidate human object. For example, when it is detected that the ambient brightness of the vehicle's environment is greater than or equal to the preset ambient brightness threshold, for example, in the daytime environment, the candidate human object is determined by the above Method 2 or Method 3. When it is detected that the ambient brightness of the vehicle's environment is less than the preset ambient brightness threshold, for example, in the nighttime environment, the candidate human object is determined by the above Method 1. At this time, the projection headlights can be used to build a stage scene outside the vehicle, and the image area corresponding to the stage scene in the captured image is used as the target scene image area.

[0082] Method 4: When it is determined through the rough gesture action detection algorithm that the gesture action of a certain human object in the captured image meets the rough matching condition of the target gesture action, determine this human object as the candidate human object.

[0083] In this embodiment, step S13 includes: judging whether the gesture action of the candidate human object meets the fine matching condition of the target gesture action through the fine gesture action detection algorithm to obtain a judgment result; the judgment result is the action detection result.

[0084] The rough gesture action detection algorithm is an algorithm for roughly detecting gesture actions, used to screen out human objects whose gesture actions meet the rough matching conditions. The fine gesture action detection algorithm is a method for finely detecting gesture actions, used to determine whether the gesture actions of candidate human objects meet the fine matching conditions.

[0085] In one embodiment, if the target gesture action is a raising hand action, then determining that the gesture action of a certain human object in the captured image meets the rough matching conditions of the target gesture action through the rough gesture action detection algorithm includes:

[0086] Obtaining the angle between the upper arm and the lower arm of the human object;

[0087] In response to the angle being greater than a preset angle threshold, determining that the gesture action of the human object meets the rough matching conditions of the target gesture action; or, in response to the angle being greater than the angle threshold and the hand key point of the human object being above the face key point of the human object, determining that the gesture action of the human object meets the rough matching conditions of the target gesture action, where the hand key point and the face key point are obtained by performing human key point detection on the captured image.

[0088] The above-mentioned angle threshold can be flexibly set by developers. Exemplarily, the angle threshold is an obtuse angle, such as 95°. The specific positions corresponding to the hand key point and the face key point can also be flexibly set by developers. For example, the recognized wrist key point can be used as the hand key point, and the recognized eye key point can be used as the face key point.

[0089] It should be noted that after detecting each human object through the rough gesture action detection algorithm, for human objects that do not meet the rough matching conditions, there is no need to further detect them through the fine gesture action detection algorithm. Only the candidate human objects that meet the rough matching conditions need to be further detected, so as to improve the accuracy of establishing the target interaction object for the final interaction with the vehicle.

[0090] Through the above rough gesture action detection algorithm, irrelevant human objects can be excluded, and candidate human objects with the intention of interacting with the vehicle can be preliminarily screened out. For example, if the angle threshold is 90°, when a certain human object makes a scratching head action as shown in Figure 2 , since the angle between the upper arm and the lower arm during scratching the head is less than 90°, it means that this human object is not a candidate human object, and the interference of this human object on the establishment of the interaction object can be excluded. When a certain human object makes an action as shown in Figure 3 , since the angle between the upper arm and the lower arm of this human object is greater than 90° and the hand key point of this human object is above the face key point of this human object, it is determined that this human object is a candidate human object.

[0091] Please refer to Figure 4 as shown, the angle between the upper arm and the lower arm of the human object can be determined in the following way:

[0092] S41: Perform human key point detection on the acquired image to obtain the elbow key point, wrist key point, and shoulder key point of the human object.

[0093] S42: Obtain the first vector between the elbow key point and the wrist key point, and the second vector between the elbow key point and the shoulder key point.

[0094] S43: Calculate the vector angle between the first vector and the second vector, and the vector angle is the angle between the upper arm and the lower arm of the human object.

[0095] In one embodiment, the first vector in step S42 includes a left - hand first vector and a right - hand first vector, and the second vector includes a left - hand second vector and a right - hand second vector. The left - hand first vector is the vector from the left - hand elbow key point to the left - hand wrist key point, the right - hand first vector is the vector from the right - hand elbow key point to the right - hand wrist key point, the left - hand second vector is the vector from the left - hand elbow key point to the left - shoulder key point, and the right - hand second vector is the vector from the right - hand elbow key point to the right - shoulder key point.

[0096] Then in step S43, the left - hand vector angle between the left - hand first vector and the left - hand second vector can be calculated, and the right - hand vector angle between the right - hand first vector and the right - hand second vector can be calculated. When determining the candidate human object, it is necessary to ensure that both the left - hand vector angle and the right - hand vector angle of the human object are greater than the preset angle threshold.

[0097] Specifically, the left - hand vector angle can be calculated by the formula where \(v1\) represents the left - hand first vector, \(v2\) represents the left - hand second vector, and \(\theta1\) represents the left - hand vector angle.

[0098] Similarly, the right - hand vector angle can be calculated by the formula where \(v3\) represents the right - hand first vector, \(v4\) represents the right - hand second vector, and \(\theta2\) represents the right - hand vector angle.

[0099] The candidate human object determined by the above - mentioned rough gesture action detection algorithm is not necessarily an object with an intention to interact with the vehicle. For example, in some cases, the human object may be reaching up to grab an item. Therefore, to improve the accuracy of the finally established target object that interacts with the vehicle, after determining the candidate human object through the rough gesture action detection algorithm, it is necessary to further detect the gesture action through the fine gesture action detection algorithm.

[0100] In this embodiment, it is determined whether the gesture action of the candidate human object meets the fine matching condition of the target gesture action through a fine gesture action detection algorithm, and a judgment result is obtained, including:

[0101] Step 1: Extract the gesture action features of the candidate human object from the captured image; the gesture action features include the features of the fine gesture actions of the candidate human object.

[0102] Step 2: Determine the gesture action of the candidate human object according to the gesture action features and the trained gesture action recognition model.

[0103] Step 3: Judge whether the gesture action is the target gesture action.

[0104] Step 4: If so, determine that the gesture action meets the fine matching condition of the target gesture action; if not, determine that the gesture action does not meet the fine matching condition of the target gesture action.

[0105] The gesture action features include at least one of static gesture action features and dynamic gesture action features. A fine gesture action refers to a hand action that requires high precision and control, such as actions related to fingers, wrists, and palms. When a user grabs an item upward, their fingers usually form a fist shape inward. Therefore, in one embodiment, the fine gesture action is a finger action.

[0106] The structure of the gesture action recognition model may include but is not limited to at least one of CNN (Convolutional Neural Network), Transformer (a deep neural network model based on the attention mechanism), LSTM (Long Short-Term Memory networks), and ResNet (Residual Network).

[0107] In the embodiment of the present application, candidate human objects are initially screened through a rough gesture action detection algorithm, and then the fine gesture actions of the candidate human objects are recognized through the trained gesture action recognition model, so that the raising hand action of the human object can be accurately recognized, improving the accuracy of raising hand action judgment.

[0108] Exemplarily, in the above Step 3, if it is determined that the gesture action of the candidate human object is a fist gesture with fingers inward, it is determined that the candidate human object does not meet the fine matching condition; if it is determined that the gesture action of the candidate human object is a gesture with fingers vertically straightened, it is determined that the candidate human object meets the fine matching condition.

[0109] It should be noted that when the number of collected images is greater than or equal to 2, for each frame of the collected images, the corresponding candidate human object can be determined respectively, and the above steps S13 and S14 can be executed until it is determined that a certain candidate human object meets the preset vehicle interaction permission granting condition. In particular, it should be noted that for the above method four, when the number of collected images is greater than or equal to 2, the human object whose gesture action first meets the rough matching condition can be used as the candidate human object. If it is determined that the candidate human object meets the preset vehicle interaction permission granting condition, the candidate human object is granted the permission to interact with the vehicle, and it is used as the target object in the current interaction process. Subsequently, the target component of the vehicle can be controlled according to the actions of the target object during the interaction with the vehicle.

[0110] In one embodiment, step S14 includes:

[0111] When it is determined that the gesture action of the candidate human object meets the fine matching condition of the target gesture action, obtain the verification information for verifying the identity of the candidate human object. When it is determined that the identity of the candidate human object is legal according to the verification information, the candidate human object is granted the permission to interact with the vehicle, that is, the candidate human object is used as the target object, and subsequently, the target object is tracked to control the target component of the vehicle according to the actions of the target object.

[0112] In the embodiment of the present application, the security and privacy are ensured through identity confirmation, preventing the abuse by illegal users. The confirmation of the verification information often requires more computing resources. In the embodiment of the present application, after it is determined that the candidate human object meets the fine matching condition of the target gesture action, further confirming the identity of the candidate human object can save computing resources and improve the computing efficiency.

[0113] In one embodiment, the verification information includes at least one of facial image information, voice information, and password information input through the projection interface of the vehicle; then determining that the identity of the candidate human object is legal according to the verification information includes:

[0114] Responding to at least one verification information matching the corresponding preset verification information, determining that the identity of the candidate human object is legal.

[0115] For example, when it is determined that the facial image information matches the preset facial image information, it can be determined that the identity of the candidate human object is legal; or, when it is determined that the password information matches the preset password information, it can be determined that the identity of the candidate human object is legal.

[0116] Exemplarily, facial image information or sound information may be collected first, and when it is determined that the facial image information or sound information matches the preset facial image information or sound information, the identity of the candidate human object is determined to be legal. If the match fails, the candidate human image is prompted to enter a password through the projection interface. After the password information entered by the candidate human object is obtained, it is matched with the preset password information. If the match is successful, the identity of the candidate human object is determined to be legal. If the match fails, the identity is determined to be illegal. In this example, identity authentication is performed first through facial image information or sound information, which reduces the operations on the user side and can better improve the user experience.

[0117] It should be noted that, in other embodiments, to further ensure security, the legality of the identity of the candidate human object may be determined when at least two of the facial image information, the voice information, and the password information match the corresponding preset verification information.

[0118] Here, the process of identity authentication based on facial image information is exemplified. In this embodiment, the first facial image information of the candidate human object can be obtained, the first facial image information is compressed to obtain the second facial image information, and the facial image information is sent to the cloud, and the third facial image information sent by the cloud is received. The third facial image information includes an enhanced image of the facial area of the candidate human object, and the enhanced image is compared with the legal face images in the pre-stored list to determine whether the identity of the candidate human object is legal. In this example, the cloud performs image enhancement, which solves the problem of limited local computing power, improves the efficiency of identity recognition, and thus improves control efficiency.

[0119] In one embodiment, the vehicle may project the projection interface through an AR (Augmented Reality) device for the candidate human object to input password information.

[0120] Exemplarily, the password information is a pattern password or a digital password, and the pattern password or digital password input by the user through the projection interface can be used as the password information.

[0121] Taking the application scenario outside the vehicle as an example, the image acquisition device captures images of the area outside the vehicle. Since the perspective of the AR device is inward and the perspective of the image acquisition device is outward, in order to ensure the accuracy of projection and interaction, the perspectives of the image acquisition device and the AR device need to be jointly calibrated in advance.

[0122] AR device coordinate system: The AR device coordinate system is usually part of the in-vehicle coordinate system. Let the AR device coordinate system be C in .

[0123] Image acquisition device coordinate system: The image acquisition device coordinate system can be relative to the global coordinate system outside the vehicle or the vehicle body coordinate system. Let the image acquisition device coordinate system be C out .

[0124] The position and orientation between these two coordinate systems can be mapped through a transformation matrix and a displacement vector.

[0125] Specifically, the joint calibration of the AR device coordinate system and the image acquisition device coordinate system can be carried out in the following way:

[0126] Let the rotation matrix of the AR device coordinate system relative to the vehicle body coordinate system be RAR car , and the displacement vector be TAR car , then: PAR = RAR car ·Pcar + TAR car ; Pcar is a point in the vehicle body coordinate system, and PAR is the virtual projection point corresponding to this point in the AR device coordinate system. In this way, any virtual projection point in the AR device coordinate system can be transformed into the vehicle body coordinate system.

[0127] Let the rotation matrix of the image acquisition device relative to the vehicle body coordinate system be R out-car , and the displacement vector be T out-car , then: Pout = R out-car ·Pcar + T out-car . In this way, any spatial point coordinate Pout detected by the image acquisition device can be transformed into the vehicle body coordinate system.

[0128] Obtain the spatial coordinate Pout of the human object calibration point through the image acquisition device, and then obtain the coordinate Pvirtual of its corresponding virtual projection point through the AR device projection.

[0129] By matching the spatial relationship between these calibration points, calculate the target transformation matrix RAR out and the target displacement vector TAR out .

[0130] Pvirtual = RAR out ·Pout + TAR out ; Pvirtual is the coordinate of the AR device projection point.

[0131] Using the least squares method to solve the error between the calibration points, the optimized target transformation matrix and target displacement vector can be obtained. This process is actually adjusting RAR out and TAR out to make the calibration points captured by the image acquisition device match the projection points on the AR device as much as possible.

[0132] After completing the above calibration, in actual applications, the coordinates of the key points of the captured human object can be directly mapped to the AR device coordinate system according to the RAR out and TAR out In this way, the user can operate the virtual projection interface through gestures, such as inputting a pattern password or a numeric password for verification.

[0133] In one embodiment, the above step S11 may be to perform image acquisition on the target area of the vehicle when receiving a target function activation instruction. Then, in step S14, when it is determined that the candidate human object meets the vehicle interaction permission granting condition, the candidate human object is granted the permission to interact with the vehicle, that is, the candidate human object is used as the target object, and the target function is activated. The target function is a function of controlling the target component of the vehicle according to the actions of the target object. In this way, the activation of the target function can be achieved.

[0134] During the process of the user interacting with the vehicle, the interaction object may need to be switched. At this time, it is necessary to re-determine and lock the target object. Therefore, in another embodiment, the above step S11 may be to perform image acquisition on the target area of the vehicle when receiving an interaction object determination instruction.

[0135] The target component in the embodiments of the present application can be flexibly set by developers. For example, it can be a vehicle suspension, left and right mirrors, sunroof or rear hatch. Exemplarily, the target component is a vehicle suspension. During the interaction between the target object and the vehicle, the vehicle controls the vehicle suspension to rise or fall according to the actions of the target object.

[0136] In one embodiment, at least two vehicle interaction permission granting conditions can be preset, including a first vehicle interaction permission granting condition and a second vehicle interaction permission granting condition. Then step S14 includes:

[0137] When it is determined according to the action detection result that the candidate human object meets the first vehicle interaction permission granting condition, the candidate human object is granted the first permission to interact with the vehicle; when it is determined according to the action detection result that the candidate human object meets the second vehicle interaction permission granting condition, the candidate human object is granted the second permission to interact with the vehicle, where the first permission is higher than the second permission.

[0138] Exemplarily, the first vehicle interaction permission granting condition may be: the gesture action of the candidate human body object satisfies the fine matching condition, and at least two of the verification information of the candidate human body object match the corresponding preset verification information. Then the second vehicle interaction permission granting condition may be: the gesture action of the candidate human body object satisfies the fine matching condition, and one verification information of the candidate human body object matches the corresponding preset verification information, and the remaining verification information does not match the corresponding preset verification information.

[0139] Exemplarily, the first vehicle interaction permission granting condition may be: the gesture action of the candidate human body object satisfies the fine matching condition, and the facial image information or voice information of the candidate human body object matches the corresponding preset facial image information or preset voice information. Then the second vehicle interaction permission granting condition may be: the gesture action of the candidate human body object satisfies the fine matching condition, and the facial image information or voice information of the candidate human body object does not match the corresponding preset facial image information or preset voice information, and the password information input by the candidate human body object matches the preset password information.

[0140] In one embodiment, a higher permission means that more functions can be controlled. Therefore, correspondingly, step S15 includes:

[0141] When the candidate human body object is granted the first permission, control the first function of the target component of the vehicle according to the actions during the interaction between the candidate human body object and the vehicle; when the candidate human body object is granted the second permission, control the second function of the target component of the vehicle according to the actions during the interaction between the candidate human body object and the vehicle, where the second function is a part of the first function.

[0142] In one embodiment, a higher permission means that more target components can be controlled. Therefore, correspondingly, step S15 includes:

[0143] When the candidate human body object is granted the first permission, control the first target component of the vehicle according to the actions during the interaction between the candidate human body object and the vehicle; when the candidate human body object is granted the second permission, control the second target component of the vehicle according to the actions during the interaction between the candidate human body object and the vehicle, where the second target component is a part of the first target component.

[0144] Exemplarily, when the candidate human body object is granted the first permission, the suspension, left and right mirrors, and rear hatch can be controlled according to the actions. When the candidate human body object is authorized the second permission, only the suspension can be controlled.

[0145] Exemplarily, the first target component includes interior components of the cockpit and exterior components of the cockpit, and the second target component is an exterior component of the cockpit. When the candidate human object is granted the first permission, the suspension and air conditioner can be controlled according to the actions. When the candidate human object is authorized with the second permission, only the suspension outside the cockpit can be controlled.

[0146] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the other steps or sub-steps or stages of the other steps.

[0147] Embodiment 2:

[0148] Based on the same inventive concept, please refer to Figure 5 As shown, an embodiment of the present application provides a vehicle control device, including:

[0149] An acquisition module 501, configured to acquire an acquisition image by performing image acquisition on a target area of the vehicle;

[0150] A determination module 502, configured to determine a candidate human object from the acquisition image in response to the number of human objects in the acquisition image being greater than 1;

[0151] A detection module 503, configured to perform action detection on the candidate human object to obtain an action detection result;

[0152] An authorization module 504, configured to grant the candidate human object the permission to interact with the vehicle when it is determined that the candidate human object meets the preset vehicle interaction permission granting conditions according to the action detection result;

[0153] A control module 505, configured to control the target components of the vehicle according to the actions during the interaction between the candidate human object and the vehicle.

[0154] In one embodiment, the determination module 502 is configured to determine a human object detection frame from the acquisition image;

[0155] In response to the existence of a target scene image region in the captured image, and the number of human object detection frames in the target scene image region being greater than 1, determine the human object in the human object detection frame closest to the center point of the target scene image region as the candidate human object; or, in response to the number of human object detection frames being greater than 1, determine the human object in the human object detection frame with the largest area as the candidate human object, or determine the human object in the human object detection frame closest to the center point of the captured image as the candidate human object.

[0156] In one embodiment, the determining module 502 is configured to determine the human object as the candidate human object when it is determined that the gesture action of a certain human object in the captured image satisfies the rough matching condition of the target gesture action through a rough gesture action detection algorithm; the detecting module 503 is configured to determine whether the gesture action of the candidate human object satisfies the fine matching condition of the target gesture action through a fine gesture action detection algorithm, and obtain a judgment result; the judgment result is the action detection result.

[0157] In one embodiment, the determining module 502 is configured to obtain the angle between the upper arm and the lower arm of the human object; in response to the angle being greater than a preset angle threshold, determine that the gesture action of the human object satisfies the rough matching condition of the target gesture action; or, in response to the angle being greater than the angle threshold and the hand key points of the human object being above the face key points of the human object, determine that the gesture action of the human object satisfies the rough matching condition of the target gesture action, where the hand key points and the face key points are obtained by performing human key point detection on the captured image.

[0158] In one embodiment, the determining module 502 is configured to perform human key point detection on the captured image to obtain the elbow key point, the wrist key point, and the shoulder key point of the human object; obtain a first vector between the elbow key point and the wrist key point, and a second vector between the elbow key point and the shoulder key point; calculate the vector angle between the first vector and the second vector, and the vector angle is the angle between the upper arm and the lower arm of the human object.

[0159] In one embodiment, the detection module 503 is configured to extract the gesture action features of the candidate human object from the collected image; the gesture action features include the features of the fine gesture actions of the candidate human object; determine the gesture action of the candidate human object according to the gesture action features and the trained gesture action recognition model; determine whether the gesture action is the target gesture action; if so, determine that the gesture action meets the fine matching condition of the target gesture action; if not, determine that the gesture action does not meet the fine matching condition of the target gesture action.

[0160] In one embodiment, when the judgment result is yes, the authorization module 504 is configured to obtain verification information for verifying the identity of the candidate human object; when it is determined that the identity of the candidate human object is legal according to the verification information, grant the candidate human object the permission to interact with the vehicle.

[0161] In one embodiment, the verification information includes at least one of facial image information, voice information, and password information input through the projection interface of the vehicle; the authorization module 504 is configured to determine that the identity of the candidate human object is legal in response to at least one of the verification information matching the corresponding preset verification information.

[0162] It should be understood that, for the sake of brevity of description, the content described in some embodiments will not be repeated in this embodiment.

[0163] Embodiment Three:

[0164] Please refer to Figure 6 As shown, the embodiment of the present application provides an electronic device, including a processor 601 and a memory 602. A computer program is stored in the memory 602, and the processor 601 executes the computer program. The processor executes the computer program to implement the steps of the method introduced above, which will not be repeated here.

[0165] The processor 601 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 601 may be a general-purpose processor, including a CPU (Central Processing Unit), an NP (Network Processor), etc.; it may also be a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0166] The memory 602 may include, but is not limited to, a RAM (Random Access Memory), a ROM (Read Only Memory), a PROM (Programmable Read Only Memory), an EPROM (Erasable Programmable Read-Only Memory), and an EEPROM (Electrically Erasable Programmable Read Only Memory), etc.

[0167] Those skilled in the art can understand that Figure 6 the structure shown in [the figure] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0168] Based on the same inventive concept, the embodiments of the present application also provide a computer-readable storage medium, such as a floppy disk, an optical disc, a hard disk, a flash memory, a USB flash drive, an SD (Secure Digital) card, an MMC (Multi-Media Card), etc. One or more programs for implementing the above various steps are stored in the computer storage medium. These one or more programs can be executed by one or more processors to implement the steps of the various methods in the above embodiments, which will not be elaborated here.

[0169] Based on the same inventive concept, an embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor implements the method described in any one of the above.

[0170] Among them, the program code for executing the computer program product of the present application can be written in any combination of one or more programming languages. The program code can be completely executed on the user device, partially executed on the user device, executed as an independent software package, partially executed on the user device and partially executed on a remote device, or completely executed on a remote device.

[0171] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.

[0172] The present application is described with reference to the flowcharts and / or block diagrams of the method, device (system), and computer-readable storage medium according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0173] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0174] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of user operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocksFigure 1 Steps of functions specified in one or more boxes.

[0175] It should be noted that the illustrations provided in this embodiment only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex. The structures, proportions, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the conditions for the present application to be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present application can produce and the purposes that can be achieved, should still fall within the scope that the technical content disclosed in the present application can cover. At the same time, the terms such as "upper", "lower", "left", "right", "middle", and "one" cited in this specification are only for the convenience of clear narration, rather than used to limit the scope for the present application to be implemented. The change or adjustment of their relative relationships, without substantial change in the technical content, should also be regarded as the scope within which the present application can be implemented.

[0176] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appearing in various positions in the text does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0177] As shown herein, unless the context clearly indicates an exception, the words "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0178] The definitions included herein, as used herein, the terms "having", "may have", "comprising", or "may comprise" indicate the existence of the corresponding functions, operations, elements, etc. herein, and do not limit the existence of one or more other functions, operations, elements, etc. In addition, it should be understood that, as used herein, the terms "including" or "having" indicate the existence of the characteristics, numbers, steps, operations, elements, components, or combinations thereof described in the specification, and do not exclude the existence or addition of one or more other characteristics, numbers, steps, operations, elements, components, or combinations thereof.

[0179] In the embodiments of the present application, prefix words such as "first" and "second" are only used to distinguish different described objects, and have no restrictive effect on the position, order, priority, quantity, content, etc. of the described objects. The use of prefix words such as ordinal numbers for distinguishing described objects in the embodiments of the present application does not constitute a limitation on the described objects. For the statement of the described objects, refer to the description in the claims or the context of the embodiments. It should not be construed as an unnecessary limitation due to the use of such prefix words. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0180] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0181] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A vehicle control method, characterized in that, Including: Performing image acquisition on a target area of a vehicle to obtain an acquired image; In response to the number of human objects in the acquired image being greater than 1, determining candidate human objects from the acquired image; Performing action detection on the candidate human objects to obtain an action detection result; When it is determined according to the action detection result that the candidate human objects meet a preset vehicle interaction permission granting condition, granting the permission to interact with the vehicle to the candidate human objects; Controlling a target component of the vehicle according to the actions during the interaction between the candidate human objects and the vehicle.

2. The vehicle control method according to claim 1, wherein, The step of, in response to the number of human objects in the acquired image being greater than 1, determining candidate human objects from the acquired image includes: Determining a human object detection frame from the acquired image; In response to the presence of a target scene image area in the acquired image and the number of the human object detection frames in the target scene image area being greater than 1, determining the human object in the human object detection frame closest to the center point of the target scene image area as the candidate human object; or, In response to the number of the human object detection frames being greater than 1, determining the human object in the human object detection frame with the largest area as the candidate human object, or determining the human object in the human object detection frame closest to the center point of the acquired image as the candidate human object.

3. The vehicle control method according to claim 1, wherein, The step of determining candidate human objects from the acquired image includes: When it is determined through a rough gesture action detection algorithm that the gesture action of a certain human object in the acquired image meets the rough matching condition of a target gesture action, determining the human object as the candidate human object; The step of performing action detection on the candidate human objects to obtain an action detection result includes: Judging whether the gesture action of the candidate human object meets the fine matching condition of the target gesture action through a fine gesture action detection algorithm to obtain a judgment result; the judgment result is the action detection result.

4. The vehicle control method according to claim 3, wherein The target gesture action is a raising hand action; the step of, when it is determined through a rough gesture action detection algorithm that the gesture action of a certain human object in the acquired image meets the rough matching condition of the target gesture action, includes: Obtaining the angle between the upper arm and the lower arm of the human object; In response to the angle being greater than a preset angle threshold, determining that the gesture action of the human object meets the rough matching condition of the target gesture action; or, in response to the angle being greater than the angle threshold and the hand key point of the human object being above the face key point of the human object, determining that the gesture action of the human object meets the rough matching condition of the target gesture action, where the hand key point and the face key point are obtained by performing human key point detection on the acquired image.

5. The vehicle control method according to claim 4, wherein The step of obtaining the angle between the upper arm and the lower arm of the human object includes: Performing human key point detection on the acquired image to obtain the elbow key point, the wrist key point and the shoulder key point of the human object; Obtaining a first vector between the elbow key point and the wrist key point, and a second vector between the elbow key point and the shoulder key point; Calculate the vector angle between the first vector and the second vector, where the vector angle is the angle between the upper arm and the lower arm of the human object.

6. The vehicle control method according to claim 4, wherein Judging whether the gesture action of the candidate human object meets the fine matching condition of the target gesture action through a fine gesture action detection algorithm, and obtaining a judgment result, including: Extract the gesture action features of the candidate human object from the captured image; the gesture action features include the features of the fine gesture action of the candidate human object; Determine the gesture action of the candidate human object according to the gesture action features and the trained gesture action recognition model; Judge whether the gesture action is the target gesture action; If so, determine that the gesture action meets the fine matching condition of the target gesture action; If not, determine that the gesture action does not meet the fine matching condition of the target gesture action.

7. The vehicle control method according to claim 3, wherein When it is determined according to the action detection result that the candidate human object meets the preset vehicle interaction permission granting condition, granting the candidate human object the permission to interact with the vehicle, including: When the judgment result is yes, obtain verification information for verifying the identity of the candidate human object; When the identity of the candidate human object is determined to be legal according to the verification information, grant the candidate human object the permission to interact with the vehicle.

8. The vehicle control method according to claim 7, wherein The verification information includes at least one of facial image information, voice information, and password information input through the projection interface of the vehicle; determining that the identity of the candidate human object is legal according to the verification information includes: In response to at least one of the verification information matching the corresponding preset verification information, determine that the identity of the candidate human object is legal.

9. A vehicle control device, characterized in that, The device includes: An acquisition module for acquiring a captured image by performing image acquisition on a target area of the vehicle; A determination module for determining a candidate human object from the captured image in response to the number of human objects in the captured image being greater than 1; A detection module for performing action detection on the candidate human object to obtain an action detection result; An authorization module for granting the candidate human object the permission to interact with the vehicle when it is determined according to the action detection result that the candidate human object meets the preset vehicle interaction permission granting condition; A control module for controlling the target component of the vehicle according to the actions during the interaction between the candidate human object and the vehicle.

10. An electronic device, characterized in that, Comprising a processor and a memory, wherein a computer program is stored in the memory, and the processor executes the computer program to implement the method according to any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by at least one processor, the method according to any one of claims 1-8 is implemented.

Citation Information

Cited By

  • Game interaction control method and device, computer equipment and storage medium

    CN120550406A