Vehicle door control method, device, system, vehicle, electronic equipment and storage medium
By installing an image acquisition module in the vehicle for facial recognition and door opening intent analysis, the system automatically controls the unlocking and opening of the doors, solving the inconvenience of users having to manually operate the car key and improving convenience and security.
Patent Information
- Application Number
- CN202210441785.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-22
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2039-10-22
AI Technical Summary
The need for users to control car doors via car keys presents inconvenience, and car keys are prone to damage, malfunction, or loss.
By installing an image acquisition module on the vehicle to capture video streams, perform facial recognition and door opening intent analysis, and automatically control the unlocking and opening of the car doors.
It enables automatic opening of car doors without manual user operation, improving ease of use and safety.
Smart Images

Figure CN114937294B_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese patent application No. 201911006853.5, filed on October 22, 2019, with the title of “Vehicle door control method, device, system, vehicle, electronic device and storage medium”. TECHNICAL FIELD
[0002] The present disclosure relates to the technical field of computer, and particularly relates to a vehicle door control method, device, system, vehicle, electronic device and storage medium. BACKGROUND
[0003] At present, a user needs to control a vehicle door by using a vehicle key (for example, a mechanical key or a remote control key). For the user, especially for the user who likes sports, it is inconvenient to carry the vehicle key. In addition, the vehicle key has the risk of being damaged, invalid or lost. SUMMARY
[0004] The present disclosure provides a vehicle door control technical solution.
[0005] According to an aspect of the present disclosure, a vehicle door control method is provided, comprising:
[0006] controlling an image acquisition module arranged on a vehicle to acquire a video stream;
[0007] performing face recognition based on at least one image in the video stream to obtain a face recognition result;
[0008] determining an intersection-over-union of images of adjacent frames in the video stream, and determining opening door intention information according to the intersection-over-union of the images of the adjacent frames; and / or, determining an area of a human body region in a plurality of newly acquired frames of the video stream, and determining the opening door intention information according to the area of the human body region in the plurality of newly acquired frames;
[0009] determining control information corresponding to at least one vehicle door of the vehicle based on the face recognition result and the opening door intention information;
[0010] if the control information includes control information for opening any vehicle door of the vehicle, obtaining state information of the vehicle door;
[0011] if the state information of the vehicle door is not unlocked, controlling the vehicle door to be unlocked and opened; and / or, if the state information of the vehicle door is already unlocked but not opened, controlling the vehicle door to be opened.
[0012] According to an aspect of the present disclosure, a vehicle door control device is provided, comprising:
[0013] a first control module configured to control an image acquisition module arranged on a vehicle to acquire a video stream;
[0014] a face recognition module, configured to perform face recognition based on at least one image in the video stream to obtain a face recognition result;
[0015] a second determination module, configured to determine an intersection-over-union of images of adjacent frames in the video stream, and determine the door opening intention information according to the intersection-over-union of the images of the adjacent frames, and / or determine an area of a human body region in a plurality of newly collected images in the video stream, and determine the door opening intention information according to the area of the human body region in the plurality of newly collected images;
[0016] a first determination module, configured to determine control information corresponding to at least one door of the vehicle based on the face recognition result and the door opening intention information;
[0017] a first acquisition module, configured to acquire state information of the door if the control information includes control of opening any door of the vehicle;
[0018] a second control module, configured to control the door to be unlocked and opened if the state information of the door is not unlocked, and / or control the door to be opened if the state information of the door is already unlocked but not opened.
[0019] According to an aspect of the present disclosure, a vehicle door control system is provided, comprising a memory, an object detection module, a face recognition module, and an image acquisition module; the face recognition module is connected with the memory, the object detection module, and the image acquisition module respectively, and the object detection module is connected with the image acquisition module; the face recognition module is further provided with a communication interface for connecting with a vehicle door domain controller, and the face recognition module sends control information for unlocking and opening the door to the vehicle door domain controller through the communication interface;
[0020] The image acquisition module is configured to acquire a video stream.
[0021] The face recognition module is configured to perform face recognition based on at least one image in the video stream to obtain a face recognition result; determine an intersection-over-union of images of adjacent frames in the video stream, and determine the door opening intention information according to the intersection-over-union of the images of the adjacent frames, and / or determine an area of a human body region in a plurality of newly collected images in the video stream, and determine the door opening intention information according to the area of the human body region in the plurality of newly collected images; and determine control information corresponding to at least one door of the vehicle based on the face recognition result and the door opening intention information.
[0022] According to an aspect of the present disclosure, a vehicle is provided, comprising the above-mentioned vehicle door control system, and the vehicle door control system is connected with a vehicle door domain controller of the vehicle.
[0023] In an aspect of the present disclosure, an electronic device is provided, comprising:
[0024] a processor;
[0025] a memory for storing processor-executable instructions;
[0026] wherein the processor is configured to perform the above vehicle door control method.
[0027] In an aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon computer program instructions, which when executed by a processor, implement the above vehicle door control method.
[0028] In the embodiments of the present disclosure, a video stream is collected by controlling an image collection module arranged on a vehicle, face recognition is performed based on at least one image in the video stream to obtain a face recognition result, control information corresponding to at least one vehicle door of the vehicle is determined based on the face recognition result, if the control information includes control of opening any vehicle door of the vehicle, state information of the vehicle door is obtained, if the state information of the vehicle door is not unlocked, the vehicle door is controlled to be unlocked and opened, and / or if the state information of the vehicle door is already unlocked but not opened, the vehicle door is controlled to be opened. Thus, the user can be automatically opened the vehicle door based on face recognition, without manually pulling open the vehicle door, thereby improving the convenience of using the vehicle.
[0029] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present disclosure.
[0030] Other features and aspects of the present disclosure will become apparent from the following detailed description of example embodiments with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0031] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and serve to explain the technical solutions of the present disclosure together with the specification.
[0032] Figure 1 A flow chart of a vehicle door control method provided by an embodiment of the present disclosure is shown.
[0033] Figure 2 A schematic view of a B-pillar of a vehicle is shown.
[0034] Figure 3 A schematic view of the installation height of an image collection module and the identifiable height range in the vehicle door control method provided by an embodiment of the present disclosure is shown.
[0035] Figure 4a A schematic view of an image sensor and a depth sensor in the vehicle door control method provided by an embodiment of the present disclosure is shown.
[0036] Figure 4b Another schematic diagram of the image sensor and the depth sensor in the vehicle door control method provided by the embodiments of the present disclosure is shown.
[0037] Figure 5 A schematic diagram of the vehicle door control method provided by the embodiments of the present disclosure is shown.
[0038] Figure 6 Another schematic diagram of the vehicle door control method provided by the embodiments of the present disclosure is shown.
[0039] Figure 7 A schematic diagram of one example of the living body detection method according to the embodiments of the present disclosure is shown.
[0040] Figure 8 A schematic diagram of one example of determining the living body detection result of the face in the first image based on the first image and the second depth map in the living body detection method according to the embodiments of the present disclosure is shown.
[0041] Figure 9 A schematic diagram of the depth prediction neural network in the vehicle door control method provided by the embodiments of the present disclosure is shown.
[0042] Figure 10 A schematic diagram of the correlation detection neural network in the vehicle door control method provided by the embodiments of the present disclosure is shown.
[0043] Figure 11 An exemplary schematic diagram of the depth map updating in the vehicle door control method provided by the embodiments of the present disclosure is shown.
[0044] Figure 12 A schematic diagram of the surrounding pixels in the vehicle door control method provided by the embodiments of the present disclosure is shown.
[0045] Figure 13 Another schematic diagram of the surrounding pixels in the vehicle door control method provided by the embodiments of the present disclosure is shown.
[0046] Figure 14 A block diagram of the vehicle door control apparatus according to the embodiments of the present disclosure is shown.
[0047] Figure 15 A block diagram of the vehicle door control system provided by the embodiments of the present disclosure is shown.
[0048] Figure 16 A schematic diagram of the vehicle door control system according to the embodiments of the present disclosure is shown.
[0049] Figure 17 A schematic diagram of the vehicle provided by the embodiments of the present disclosure is shown.
[0050] Figure 18A block diagram of an electronic device 800 is shown.
[0051] Figure 19 A block diagram of an electronic device 1900 is shown. DETAILED DESCRIPTION
[0052] Various exemplary embodiments, features, and aspects of the present disclosure will be explained in greater detail below with reference to the accompanying drawings. Like reference numerals may be used to refer to like elements throughout. While various aspects of embodiments are illustrated, the embodiments need not be used to scale unless specifically indicated.
[0053] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.
[0054] The term "and / or", used herein only to represent an association relationship of associated objects, means that there can be three relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality, for example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0055] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the specific embodiments below. Those skilled in the art should understand that the present disclosure can be implemented without certain specific details. In some examples, methods, means, elements and circuits well known to those skilled in the art are not described in detail, in order to highlight the main idea of the present disclosure.
[0056] Figure 1 A flowchart of a vehicle door control method is shown. The execution subject of the vehicle door control method can be a vehicle door control device. In some possible implementations, the vehicle door control method can be implemented by a processor calling computer readable instructions stored in a memory. As Figure 1 As shown, the vehicle door control method includes steps S11 to S15.
[0057] In step S11, a video stream is collected by an image collection module arranged on the vehicle.
[0058] In a possible implementation, the control of the image acquisition module arranged on the vehicle to acquire the video stream comprises: controlling the image acquisition module arranged on the outside of the vehicle to acquire the video stream outside the vehicle. In this implementation, the image acquisition module can be installed on the outside of the vehicle, and the video stream outside the vehicle is acquired by controlling the image acquisition module arranged on the outside of the vehicle, so that the boarding intention of the person outside the vehicle can be detected based on the video stream outside the vehicle.
[0059] In a possible implementation, the image acquisition module can be installed on at least one of the following positions: a B-pillar of the vehicle, at least one door of the vehicle, and at least one rearview mirror. The door of the vehicle in the embodiments of the present disclosure can include a door through which a person enters or exits (for example, a left front door, a right front door, a left rear door, or a right rear door), and can also include a trunk door of the vehicle and the like. Figure 2 A schematic diagram of the B-pillar of the vehicle is shown. For example, the image acquisition module can be installed on the B-pillar at a height of 130 cm to 160 cm from the ground, and the horizontal recognition distance of the image acquisition module can be 30 cm to 100 cm, which is not limited herein. Figure 3 A schematic diagram of the installation height of the image acquisition module and the recognizable height range in the door control method provided by the embodiments of the present disclosure is shown. In the example shown, the installation height of the image acquisition module is 160 cm, and the recognizable height range is 140 cm to 190 cm. Figure 3 The installation height of the image acquisition module is 160 cm, and the recognizable height range is 140 cm to 190 cm.
[0060] In one example, the image acquisition module can be installed on two B-pillars and a trunk of the vehicle. Among them, the image acquisition module can be installed on each B-pillar to face the boarding position of the front-row vehicle occupant (driver or front passenger) and the boarding position of the rear-row vehicle occupant.
[0061] In a possible implementation, the control of the image acquisition module arranged on the vehicle to acquire the video stream comprises: controlling the image acquisition module arranged on the inside of the vehicle to acquire the video stream inside the vehicle. In this implementation, the image acquisition module can be installed on the inside of the vehicle, and the video stream inside the vehicle is acquired by controlling the image acquisition module arranged on the inside of the vehicle, so that the alighting intention of the person inside the vehicle can be detected based on the video stream inside the vehicle.
[0062] As an example of this implementation, the control of the image acquisition module arranged on the inside of the vehicle to acquire the video stream inside the vehicle comprises: in the case that the driving speed of the vehicle is 0 and there is a person inside the vehicle, controlling the image acquisition module arranged on the inside of the vehicle to acquire the video stream inside the vehicle. In this example, by controlling the image acquisition module arranged on the inside of the vehicle to acquire the video stream inside the vehicle in the case that the driving speed of the vehicle is 0 and there is a person inside the vehicle, both safety and power consumption can be ensured.
[0063] In step S12, face recognition is performed based on at least one image in the video stream to obtain a face recognition result.
[0064] For example, face recognition can be performed based on a first image in the video stream to obtain a face recognition result. The first image can contain at least part of a human body or a face. The first image can be an image selected from the video stream, and the image can be selected from the video stream in various ways. In one specific example, the first image is an image selected from the video stream that meets a preset quality condition. The preset quality condition can include one or any combination of the following: whether a human body or a face is contained, whether a human body or a face is located in a central area of the image, whether a human body or a face is completely contained in the image, a proportion of a human body or a face in the image, a state of a human body or a face (e.g., a human body orientation or a face angle), image sharpness, image exposure, and the like, without limitation.
[0065] In one possible implementation, the face recognition includes face authentication. The face recognition based on at least one image in the video stream includes face authentication based on a first image in the video stream and a pre-registered face feature. In this implementation, face authentication is used to extract a face feature in a collected image, compare the face feature in the collected image with the pre-registered face feature, and determine whether the face features belong to the same person. For example, it can be determined whether a face feature in a collected image belongs to a face feature of a car owner or a temporary user (e.g., a friend of the car owner or a courier).
[0066] In one possible implementation, the face recognition further includes live detection. The face recognition based on at least one image in the video stream includes collecting, by a depth sensor in the image collection module, a first depth map corresponding to a first image in the video stream, and performing live detection based on the first image and the first depth map. In this implementation, live detection is used to verify whether it is a live body, for example, to verify whether it is a human body.
[0067] In one example, live detection can be performed before face authentication. For example, if the live detection result of a person is a human live body, a face authentication process is triggered; if the live detection result of a person is a human false body, a face authentication process is not triggered.
[0068] In another example, face authentication can be performed before live detection. For example, if face authentication is passed, a live detection process is triggered; if face authentication is not passed, a live detection process is not triggered.
[0069] In another example, live detection and face authentication can be performed simultaneously.
[0070] In the embodiments of the present disclosure, the depth sensor refers to a sensor for collecting depth information. The working principle and working waveband of the depth sensor are not limited in the embodiments of the present disclosure.
[0071] In the embodiments of the present disclosure, the image sensor and the depth sensor of the image collection module can be separately arranged or arranged together. For example, the image sensor of the image collection module adopts an RGB (Red, Green, Blue) sensor or an infrared sensor, and the depth sensor adopts a binocular infrared sensor or a TOF (Time of Flight) sensor; the image sensor and the depth sensor of the image collection module are arranged together, and the image collection module adopts an RGBD (Red, Green, Blue, Deep) sensor to realize the functions of the image sensor and the depth sensor.
[0072] As an example, the image sensor is an RGB sensor. If the image sensor is an RGB sensor, the image collected by the image sensor is an RGB image.
[0073] As another example, the image sensor is an infrared sensor. If the image sensor is an infrared sensor, the image collected by the image sensor is an infrared image. The infrared image can be an infrared image with a light spot or an infrared image without a light spot.
[0074] In other examples, the image sensor can be other types of sensors, which are not limited in the embodiments of the present disclosure.
[0075] As an example, the depth sensor is a three-dimensional sensor. For example, the depth sensor is a binocular infrared sensor, a TOF sensor or a structured light sensor. The binocular infrared sensor includes two infrared cameras. The structured light sensor can be a coded structured light sensor or a speckle structured light sensor. The depth map of a person obtained by the depth sensor can obtain a high-precision depth map. The embodiments of the present disclosure use the depth map containing a face to perform liveness detection, which can fully exploit the depth information of the face, thereby improving the accuracy of liveness detection.
[0076] In an example, the TOF sensor adopts an infrared waveband-based TOF module. In this example, by adopting the infrared waveband-based TOF module, the influence of external light on the depth map shooting can be reduced.
[0077] In the embodiments of the present disclosure, the first depth map and the first image correspond to each other. For example, the first depth map and the first image are collected by a depth sensor and an image sensor for a same scene, respectively, or the first depth map and the first image are collected by the depth sensor and the image sensor for a same target region at a same time, but the embodiments of the present disclosure are not limited thereto.
[0078] Figure 4a A schematic diagram of the image sensor and the depth sensor in the vehicle door control method provided by the embodiments of the present disclosure is shown. In the embodiments of the present disclosure, Figure 4a In the example shown, the image sensor is an RGB sensor, the camera of the image sensor is an RGB camera, the depth sensor is a binocular infrared sensor, and the binocular infrared sensor includes two infrared (IR) cameras, and the two infrared cameras of the binocular infrared sensor are arranged on the two sides of the RGB camera of the image sensor. Among them, the two infrared cameras collect depth information based on the binocular disparity principle.
[0079] In one example, the image acquisition module further includes at least one light supplementing lamp, the at least one light supplementing lamp is arranged between the infrared camera of the binocular infrared sensor and the camera of the image sensor, and the at least one light supplementing lamp includes at least one of a light supplementing lamp for the image sensor and a light supplementing lamp for the depth sensor. For example, if the image sensor is an RGB sensor, the light supplementing lamp for the image sensor can be a white light lamp; if the image sensor is an infrared sensor, the light supplementing lamp for the image sensor can be an infrared lamp; and if the depth sensor is a binocular infrared sensor, the light supplementing lamp for the depth sensor can be an infrared lamp. Figure 4a In the example shown, an infrared lamp is arranged between the infrared camera of the binocular infrared sensor and the camera of the image sensor. For example, the infrared lamp can use 940 nm infrared rays.
[0080] In one example, the light supplementing lamp can be in a constant-on mode. In this example, when the camera of the image acquisition module is in a working state, the light supplementing lamp is in an on state.
[0081] In another example, the light supplementing lamp can be turned on when the light is insufficient. For example, the ambient light intensity can be obtained by an ambient light sensor, and when the ambient light intensity is lower than a light intensity threshold, it is determined that the light is insufficient, and the light supplementing lamp is turned on.
[0082] Figure 4b Another schematic diagram of the image sensor and the depth sensor in the vehicle door control method provided by the embodiments of the present disclosure is shown. In the embodiments of the present disclosure, Figure 4b In the example shown, the image sensor is an RGB sensor, the camera of the image sensor is an RGB camera, and the depth sensor is a TOF sensor.
[0083] In one example, the image acquisition module further includes a laser, which is disposed between the camera of the depth sensor and the camera of the image sensor. For example, the laser is disposed between the camera of the TOF sensor and the camera of the RGB sensor. For example, the laser can be a VCSEL (Vertical Cavity Surface Emitting Laser), and the TOF sensor can acquire a depth map based on the laser emitted by the VCSEL.
[0084] In the embodiments of the present disclosure, the depth sensor is used to acquire a depth map, and the image sensor is used to acquire a two-dimensional image. It should be noted that, although the image sensor is described by taking the RGB sensor and the infrared sensor as examples, and the depth sensor is described by taking the binocular infrared sensor, the TOF sensor and the structured light sensor as examples, those skilled in the art can understand that the embodiments of the present disclosure should not be limited thereto. Those skilled in the art can select the types of the image sensor and the depth sensor according to actual application requirements, as long as the acquisition of the two-dimensional image and the depth map can be realized respectively.
[0085] In a possible implementation, the face recognition further includes permission authentication; and the face recognition based on the at least one image in the video stream includes: obtaining door opening permission information of the person based on a first image in the video stream; and performing permission authentication based on the door opening permission information of the person. According to this implementation, different door opening permission information can be set for different users, so as to improve the safety of the vehicle.
[0086] As an example of this implementation, the door opening permission information of the person includes one or more of the following: information of a door of the vehicle for which the person has door opening permission, a time for which the person has door opening permission, and a number of times of door opening permission corresponding to the person.
[0087] For example, the information of the door of the vehicle for which the person has door opening permission can be all doors or a trunk door. For example, the doors for which the vehicle owner or the family and friends of the vehicle owner have door opening permission can be all doors, and the door for which a courier or a property staff has door opening permission can be the trunk door. The vehicle owner can set the information of the door for which other persons have door opening permission.
[0088] For example, the time when the person has the door opening permission can be all time or can be a preset time period. For example, the time when the car owner or the family of the car owner has the door opening permission can be all time. The car owner can set the time when other people have the door opening permission. For example, in the application scenario that the car owner lends the car to the friend of the car owner, the car owner can set the time when the friend has the door opening permission as two days. For another example, after the express delivery person contacts the car owner, the car owner can set the time when the express delivery person has the door opening permission as 13:00-14:00 on September 29, 2019.
[0089] For example, the number of times of the door opening permission corresponding to the person can be unlimited times or limited times. For example, the number of times of the door opening permission corresponding to the car owner or the family of the car owner, the friend of the car owner can be unlimited times. For another example, the number of times of the door opening permission corresponding to the express delivery person can be limited times, for example, 1 time.
[0090] In step S13, based on the face recognition result, control information corresponding to at least one door of the car is determined.
[0091] In a possible implementation, before the control information corresponding to at least one door of the car is determined based on the face recognition result, the method further includes: determining door opening intention information based on the video stream; and the control information corresponding to at least one door of the car is determined based on the face recognition result and the door opening intention information.
[0092] In a possible implementation, the door opening intention information can be intentional door opening or unintentional door opening. Intentional door opening can be intentional getting on the car, intentional getting off the car, intentional placing an article in the trunk or intentional taking an article out of the trunk. For example, in the case that the video stream is collected by the image collection module on the B-pillar, if the door opening intention information is intentional door opening, it can indicate that the person intentionally gets on the car or intentionally places an article; if the door opening intention information is unintentional door opening, it can indicate that the person unintentionally gets on the car and unintentionally places an article. In the case that the video stream is collected by the image collection module on the trunk door, if the door opening intention information is intentional door opening, it can indicate that the person intentionally places an article (such as luggage) in the trunk; if the door opening intention information is unintentional door opening, it can indicate that the person unintentionally places an article in the trunk.
[0093] In a possible implementation, the door opening intention information can be determined based on multiple frames of images in the video stream, thereby improving the accuracy of the determined door opening intention information.
[0094] As an example of this implementation, determining the door opening intention information based on the video stream includes: determining the Intersection over Union (IoU) ratio of adjacent frames in the video stream; and determining the door opening intention information based on the IoU ratio of the adjacent frames.
[0095] In one example, determining the intersection-over-union ratio (IoU) of adjacent frames in the video stream may include: determining the IoU of the bounding boxes of human bodies in adjacent frames of the video stream as the IoU of the adjacent frames.
[0096] In another example, determining the intersection-union ratio (CUP) of adjacent frames in the video stream may include: determining the CUP of the bounding boxes of faces in adjacent frames of the video stream as the CUP of the adjacent frames.
[0097] In one example, determining the door-opening intention information based on the intersection-over-union (IoU) ratio of the adjacent frames may include: caching the IoU ratios of the most recently acquired N sets of adjacent frames, where N is an integer greater than 1; determining the average value of the cached IoU ratios; and determining that the door-opening intention information is intentional if the average value is greater than a first preset value for a duration of a first preset time. For example, N equals 10, the first preset value equals 0.93, and the first preset time equals 1.5 seconds. Of course, the specific values of N, the first preset value, and the first preset time can be flexibly set according to the actual application scenario requirements. In this example, the cached N IoU ratios are the IoU ratios of the most recently acquired N sets of adjacent frames. When a new image is acquired, the oldest IoU ratio is deleted from the cache, and the IoU ratio of the most recently acquired image and the previously acquired image is stored in the cache.
[0098] For example, if N equals 3, and the four most recently acquired images are image 1, image 2, image 3, and image 4, where image 4 is the most recently acquired image, then the cached intersection-over-union ratio (I) includes the intersection-over-union ratio (I) of image 1 and image 2. 12 The intersection-union ratio (I) of images 2 and 3 23 Intersection over Union (I) of Images 3 and 4 34 At this point, the average crossover / union ratio of the cache is I. 12 I 23 and I 34 The average value of I. 12 I 23 and I 34 If the average value is greater than the first preset value, then image 5 will continue to be acquired through the image acquisition module, and images with intersection ratios of I will be deleted. 12 The intersection-union ratio (I) of cached image 4 and image 5 45 At this point, the average crossover ratio I of the cache is... 23 I34 and I 45 The average value of the buffered intersection over union. If the average value of the buffered intersection over union is greater than the first preset value for a duration of the first preset time length, it is determined that the door opening intention information is intentional door opening, otherwise it can be determined that the door opening intention information is unintentional door opening.
[0099] In another example, the determining of the door opening intention information according to the intersection over union of the images of the adjacent frames can include: if the number of continuous groups of adjacent frames with the intersection over union greater than the first preset value is greater than a second preset value, it is determined that the door opening intention information is intentional door opening.
[0100] In the above examples, by determining the intersection over union of the images of the adjacent frames in the video stream and determining the door opening intention information according to the intersection over union of the images of the adjacent frames, the door opening intention of the person can be accurately determined.
[0101] As another example of this implementation, the determining of the door opening intention information based on the video stream includes: determining the area of the human body region in the latest collected multiple frames of images in the video stream; and determining the door opening intention information according to the area of the human body region in the latest collected multiple frames of images.
[0102] In one example, the determining of the door opening intention information according to the area of the human body region in the latest collected multiple frames of images can include: if the area of the human body region in the latest collected multiple frames of images is greater than a first preset area, it is determined that the door opening intention information is intentional door opening.
[0103] In another example, the determining of the door opening intention information according to the area of the human body region in the latest collected multiple frames of images can include: if the area of the human body region in the latest collected multiple frames of images gradually increases, it is determined that the door opening intention information is intentional door opening. Wherein, the area of the human body region in the latest collected multiple frames of images gradually increasing can mean that the area of the human body region in the image collected at a time close to the current time is greater than the area of the human body region in the image collected at a time far from the current time, or can mean that the area of the human body region in the image collected at a time close to the current time is greater than or equal to the area of the human body region in the image collected at a time far from the current time.
[0104] In the above examples, by determining the area of the human body region in the latest collected multiple frames of images in the video stream and determining the door opening intention information according to the area of the human body region in the latest collected multiple frames of images, the door opening intention of the person can be accurately determined.
[0105] As another example of the implementation, the determining, based on the video stream, the door opening intention information includes: determining areas of face regions in a plurality of newly captured images in the video stream; and determining the door opening intention information according to the areas of the face regions in the plurality of newly captured images.
[0106] In one example, the determining, based on the areas of the face regions in the plurality of newly captured images, the door opening intention information can include: if the areas of the face regions in the plurality of newly captured images are all greater than a second preset area, determining that the door opening intention information is intentional door opening.
[0107] In another example, the determining, based on the areas of the face regions in the plurality of newly captured images, the door opening intention information can include: if the areas of the face regions in the plurality of newly captured images gradually increase, determining that the door opening intention information is intentional door opening. The areas of the face regions in the plurality of newly captured images gradually increasing can mean that an area of a face region in an image captured at a time close to a current time is greater than an area of a face region in an image captured at a time far from the current time, or can mean that an area of a face region in an image captured at a time close to a current time is greater than or equal to an area of a face region in an image captured at a time far from the current time.
[0108] In the above examples, by determining the areas of the face regions in the plurality of newly captured images in the video stream, and determining the door opening intention information according to the areas of the face regions in the plurality of newly captured images, the door opening intention of the person can be accurately determined.
[0109] In the embodiments of the present disclosure, by controlling at least one door of the vehicle based on the door opening intention information, the possibility of opening the door of the vehicle in the case of unintentional door opening of the user can be reduced, and thus the safety of the vehicle can be improved.
[0110] In a possible implementation, the determining, based on the face recognition result and the door opening intention information, the control information corresponding to the at least one door of the vehicle includes: if the face recognition result is face recognition success, and the door opening intention information is intentional door opening, determining that the control information includes control of opening of the at least one door of the vehicle.
[0111] In a possible implementation, before the control information corresponding to the at least one door of the vehicle is determined based on the face recognition result, the method further includes: performing object detection on at least one image in the video stream to determine object carrying information of the person; and the control information corresponding to the at least one door of the vehicle is determined based on the face recognition result and the object carrying information of the person. In this implementation, the door control can be performed based on the face recognition result and the object carrying information of the person, without considering the door opening intention information.
[0112] As an example of this implementation, the control information corresponding to the at least one door of the vehicle is determined based on the face recognition result and the object carrying information of the person, including: if the face recognition result is that the face recognition is successful, and the object carrying information of the person is that the person carries an object, it is determined that the control information includes controlling the at least one door of the vehicle to be opened. According to this example, when the face recognition result is that the face recognition is successful, and the object carrying information of the person is that the person carries an object, the door of the vehicle can be automatically opened for the user, without the user manually opening the door.
[0113] As an example of this implementation, the control information corresponding to the at least one door of the vehicle is determined based on the face recognition result and the object carrying information of the person, including: if the face recognition result is that the face recognition is successful, and the object carrying information of the person is that the person carries an object of a preset category, it is determined that the control information includes controlling the trunk door of the vehicle to be opened. According to this example, when the face recognition result is that the face recognition is successful, and the object carrying information of the person is that the person carries an object of a preset category, the trunk door of the vehicle can be automatically opened for the user, without the user manually opening the trunk door.
[0114] In a possible implementation, before the control information corresponding to the at least one door of the vehicle is determined based on the face recognition result, the method further includes: performing object detection on at least one image in the video stream to determine object carrying information of the person; and the control information corresponding to the at least one door of the vehicle is determined based on the face recognition result and the object carrying information of the person. In this implementation, the door control can be performed based on the face recognition result and the object carrying information of the person, without considering the door opening intention information.
[0115] In this implementation, the object carrying information of the person can represent information of an object carried by the person. For example, the object carrying information of the person can represent whether the person carries an object; or for another example, the object carrying information of the person can represent a category of an object carried by the person.
[0116] According to the implementation, when the user is not convenient to open the door (for example, the user carries a handbag, a shopping bag, a trolley case, an umbrella, or the like), the door (for example, the left front door, the right front door, the left rear door, the right rear door, or the trunk door of the vehicle) is automatically opened for the user, thereby greatly facilitating the user to get into the vehicle and place the articles in the trunk in the scenario of carrying articles by hand or raining. With the implementation, when the user approaches the vehicle, the face recognition process is automatically triggered without deliberately performing an action (such as touching a button or performing a gesture), so that the door of the vehicle is automatically opened without the user freeing a hand to unlock or open the door, thereby improving the experience of the user to get into the vehicle and place the articles in the trunk.
[0117] As an example of the implementation, the control information corresponding to the at least one door of the vehicle is determined based on the face recognition result, the door opening intention information, and the object carrying information of the person, including: if the face recognition result is face recognition success, the door opening intention information is to open the door, and the object carrying information of the person is that the person carries an object, it is determined that the control information includes control of opening the at least one door of the vehicle.
[0118] In this example, if the object carrying information of the person is that the person carries an object, it can be determined that the person is currently not convenient to manually open the door of the vehicle, for example, the person currently carries a heavy object or holds an umbrella, or the like.
[0119] As an example of the implementation, the object detection is performed on the at least one image in the video stream to determine the object carrying information of the person, including: performing object detection on the at least one image in the video stream to obtain an object detection result; and determining the object carrying information of the person based on the object detection result. For example, object detection is performed on the first image in the video stream to obtain an object detection result.
[0120] In this example, the object detection is performed on the at least one image in the video stream to obtain an object detection result, and the object carrying information of the person is determined based on the object detection result, thereby accurately obtaining the object carrying information of the person.
[0121] In this example, the object detection result can be used as the object carrying information of the person. For example, the object detection result includes an umbrella, and the object carrying information of the person includes the umbrella; for another example, the object detection result includes an umbrella and a trolley case, and the object carrying information of the person includes the umbrella and the trolley case; for another example, the object detection result is empty, and the object carrying information of the person can be empty.
[0122] In this example, an object detection network can be employed to perform object detection on at least one image in the video stream, where the object detection network can be based on a deep learning architecture. In this example, the categories of objects that the object detection network can recognize can not be limited, and a person skilled in the art can flexibly set the categories of objects that the object detection network can recognize according to the actual application scenario requirements. For example, the categories of objects that the object detection network can recognize include umbrellas, trolley cases, trolleys, baby strollers, handbags, shopping bags, and the like. By employing the object detection network to perform object detection on at least one image in the video stream, the accuracy and speed of object detection can be improved.
[0123] In this example, the object detection on the at least one image in the video stream to obtain an object detection result can include: detecting a bounding box of a human body in the at least one image in the video stream; and performing object detection on a region corresponding to the bounding box to obtain the object detection result. For example, a bounding box of a human body in a first image of the video stream can be detected, and object detection can be performed on a region corresponding to the bounding box in the first image. The region corresponding to the bounding box can represent a region defined by the bounding box. In this example, by detecting a bounding box of a human body in at least one image in the video stream and performing object detection on a region corresponding to the bounding box, interference of background parts in the images in the video stream on object detection can be avoided, and thus the accuracy of object detection can be improved.
[0124] In this example, determining the object carrying information of the person based on the object detection result can include: if the object detection result is that an object is detected, obtaining a distance between the object and a hand of the person; and determining the object carrying information of the person based on the distance.
[0125] In one example, if the distance is less than a preset distance, the object carrying information of the person can be determined to be that the person carries an object. In this example, when determining the object carrying information of the person, only the distance between the object and the hand of the person can be considered, and the size of the object does not need to be considered.
[0126] In another example, determining the object carrying information of the person based on the object detection result can further include: if the object detection result is that an object is detected, obtaining a size of the object; and determining the object carrying information of the person based on the distance includes: determining the object carrying information of the person based on the distance and the size. In this example, when determining the object carrying information of the person, both the distance between the object and the hand of the person and the size of the object can be considered.
[0127] The determining the object carrying information of the person based on the distance and the size can include: if the distance is less than or equal to a preset distance, and the size is greater than or equal to a preset size, determining the object carrying information of the person as the person carrying the object.
[0128] In this example, the preset distance can be 0, or the preset distance can be set to be greater than 0.
[0129] In this example, the determining the object carrying information of the person based on the object detection result can include: if the object detection result is detecting the object, obtaining the size of the object; and determining the object carrying information of the person based on the size. In this example, when determining the object carrying information of the person, only the size of the object can be considered, and the distance between the object and the hand of the person does not need to be considered. For example, if the size is greater than a preset size, the object carrying information of the person is determined as the person carrying the object.
[0130] As an example of this implementation, the determining the control information corresponding to the at least one door of the vehicle based on the face recognition result, the door opening intention information and the object carrying information of the person can include: if the face recognition result is face recognition success, the door opening intention information is intentional door opening, and the object carrying information of the person is that the person carries an object of a preset category, determining that the control information includes controlling the trunk door of the vehicle to be opened. The preset category can represent the category of the object suitable for being placed in the trunk. For example, the preset category can include a trolley case and the like. Figure 5 A schematic diagram of the vehicle door control method provided by the embodiments of the present disclosure is shown. In Figure 5 In the example shown, if the face recognition result is face recognition success, the door opening intention information is intentional door opening, and the object carrying information of the person is that the person carries an object of a preset category (such as a trolley case), it is determined that the control information includes controlling the trunk door of the vehicle to be opened. In this example, by determining that the control information includes controlling the trunk door of the vehicle to be opened if the face recognition result is face recognition success, the door opening intention information is intentional door opening, and the object carrying information of the person is that the person carries an object of a preset category, the trunk door can be automatically opened for the person when the person carries an object of a preset category, which facilitates the person to place the object in the trunk.
[0131] As an example of the implementation, the determining the control information corresponding to at least one door of the vehicle based on the face recognition result, the door opening intention information, and the object carrying information of the person includes: if the face recognition result is face recognition success and is not a driver, the door opening intention information is intentional door opening, and the object carrying information of the person is carrying an object, determining that the control information includes controlling at least one non-driver seat door of the vehicle to open. In this example, if the face recognition result is face recognition success and is not a driver, the door opening intention information is intentional door opening, and the object carrying information of the person is carrying an object, the control information is determined to include controlling at least one non-driver seat door of the vehicle to open, thereby enabling a non-driver to automatically open the door corresponding to the seat suitable for him / her to ride.
[0132] In a possible implementation, based on the face recognition result and the door opening intention information, the determining the control information corresponding to at least one door of the vehicle can include: based on the face recognition result and the door opening intention information, determining the control information corresponding to the door corresponding to the image acquisition module acquiring the video stream. The door corresponding to the image acquisition module acquiring the video stream can be determined according to the position of the image acquisition module. For example, if the video stream is acquired by the image acquisition module installed on the left B pillar and facing the front-row vehicle-riding personnel boarding position, the door corresponding to the image acquisition module acquiring the video stream can be the left front door, thereby the control information corresponding to the left front door of the vehicle can be determined based on the face recognition result and the door opening intention information; if the video stream is acquired by the image acquisition module installed on the left B pillar and facing the rear-row vehicle-riding personnel boarding position, the door corresponding to the image acquisition module acquiring the video stream can be the left rear door, thereby the control information corresponding to the left rear door of the vehicle can be determined based on the face recognition result and the door opening intention information; if the video stream is acquired by the image acquisition module installed on the right B pillar and facing the front-row vehicle-riding personnel boarding position, the door corresponding to the image acquisition module acquiring the video stream can be the right front door, thereby the control information corresponding to the right front door of the vehicle can be determined based on the face recognition result and the door opening intention information; if the video stream is acquired by the image acquisition module installed on the right B pillar and facing the rear-row vehicle-riding personnel boarding position, the door corresponding to the image acquisition module acquiring the video stream can be the right rear door, thereby the control information corresponding to the right rear door of the vehicle can be determined based on the face recognition result and the door opening intention information; if the video stream is acquired by the image acquisition module installed on the trunk door, the door corresponding to the image acquisition module acquiring the video stream can be the trunk door, thereby the control information corresponding to the trunk door of the vehicle can be determined based on the face recognition result and the door opening intention information.
[0133] In step S14, if the control information includes control of any door of the vehicle to be opened, the state information of the door is acquired.
[0134] In the embodiments of the present disclosure, the state information of the door can be unlocked, unlocked and not opened, or opened.
[0135] In step S15, if the state information of the door is unlocked, the door is controlled to be unlocked and opened; and / or, if the state information of the door is unlocked and not opened, the door is controlled to be opened.
[0136] In the embodiments of the present disclosure, controlling the door to be opened can mean controlling the door to be popped open, so that the user can enter the vehicle through the opened door (for example, the front door or the rear door), or can place objects through the opened door (for example, the trunk door or the rear door). By controlling the door to be opened, the user does not need to manually pull open the door after the door is unlocked.
[0137] In a possible implementation, the door can be controlled to be unlocked and opened by sending the unlocking instruction and the opening instruction corresponding to the door to the door domain controller; and the door can be controlled to be opened by sending the opening instruction corresponding to the door to the door domain controller.
[0138] In an example, the SoC (System on Chip) of the door control device can send the door unlocking instruction, the opening instruction and the closing instruction to the door domain controller to control the door.
[0139] Figure 6 Another schematic diagram of the door control method provided by the embodiments of the present disclosure is shown. In the embodiments of the present disclosure, Figure 6 In the example shown, the video stream can be acquired by the image acquisition module installed on the B-pillar, the face recognition result and the door opening intention information are obtained based on the video stream, and the control information corresponding to at least one door of the vehicle is determined based on the face recognition result and the door opening intention information.
[0140] In a possible implementation, the image acquisition module installed on the B-pillar acquires the video stream, including: the image acquisition module installed on the trunk door of the vehicle acquires the video stream. In this implementation, the image acquisition module can be installed on the trunk door to detect the intention of placing objects into the trunk or taking objects out of the trunk based on the video stream acquired by the image acquisition module on the trunk door.
[0141] In a possible implementation, after the control information is determined to include control of the trunk door of the vehicle to be opened, the method further includes: in a case where the person is determined to leave the interior of the vehicle according to a video stream collected by an image collection module arranged in the interior of the vehicle, or in a case where the opening door intention information of the person is detected to be an intention to get off the vehicle, the trunk door is controlled to be opened. According to this implementation, if the vehicle occupant places an object in the trunk before getting on the vehicle, the trunk door can be automatically opened when the vehicle occupant gets off the vehicle, so that the vehicle occupant does not need to manually pull open the trunk door, and the vehicle occupant can be reminded to take the object in the trunk.
[0142] In a possible implementation, after the vehicle door is controlled to be opened, the method further includes: in a case where an automatic door closing condition is met, the vehicle door is controlled to be closed, or the vehicle door is controlled to be closed and locked. In this implementation, by controlling the vehicle door to be closed, or the vehicle door to be closed and locked in a case where the automatic door closing condition is met, the safety of the vehicle can be improved.
[0143] As an example of this implementation, the automatic door closing condition includes one or more of the following: the opening door intention information for controlling the vehicle door to be opened is an intention to get on the vehicle, and a person with the intention to get on the vehicle is determined to have taken a seat according to a video stream collected by an image collection module in the interior of the vehicle; the opening door intention information for controlling the vehicle door to be opened is an intention to get off the vehicle, and a person with the intention to get off the vehicle is determined to have left the interior of the vehicle according to a video stream collected by an image collection module in the interior of the vehicle; and a time length for which the vehicle door is opened reaches a second preset time length.
[0144] In an example, if the vehicle doors for which the person has the opening door permission include only the trunk door, the trunk door can be controlled to be closed when the time length for which the trunk door is controlled to be opened reaches a second preset time length, for example, the second preset time length can be 3 minutes. For example, if the vehicle doors for which the delivery person has the opening door permission include only the trunk door, the trunk door can be controlled to be closed when the time length for which the trunk door is controlled to be opened reaches a second preset time length, so that the needs of the delivery person to place the express delivery in the trunk can be met, and the safety of the vehicle can be improved.
[0145] In a possible implementation, the method further includes one or both of the following: user registration is performed according to a face image collected by the image collection module; and remote registration is performed according to a face image collected or uploaded by a first terminal, and registration information is sent to the vehicle, where the first terminal is a terminal corresponding to the vehicle owner, and the registration information includes the collected or uploaded face image.
[0146] In one example, the owner registration is performed according to a face image collected by the image collection module, including: when detecting that a registration button on the touch screen is clicked, requesting a user to input a password, after the password verification is passed, starting the RGB camera in the image collection module to acquire a face image, and performing registration according to the acquired face image, and extracting a face feature in the face image as a pre-registered face feature, so as to perform face comparison based on the pre-registered face feature in subsequent face authentication.
[0147] In one example, remote registration is performed according to a face image collected or uploaded by a first terminal, and registration information is sent to the vehicle, wherein the registration information includes the collected or uploaded face image. In this example, a user (for example, the owner) can send a registration request to the TSP (Telematics Service Provider, automotive telematics service provider) cloud through a mobile phone App (Application), wherein the registration request can carry a face image collected or uploaded by the first terminal, for example, the face image collected by the first terminal can be the face image of the user (the owner) himself, and the face image uploaded by the first terminal can be the face image of the user (the owner) himself, the user's friends or a courier, etc.; the TSP cloud sends the registration request to the vehicle-mounted T-Box (Telematics Box) of the vehicle door control device, and the vehicle-mounted T-Box activates the face recognition function according to the registration request, and takes the face feature in the face image carried in the registration request as a pre-registered face feature, so as to perform face comparison based on the pre-registered face feature in subsequent face authentication.
[0148] As an example of this implementation, the face image uploaded by the first terminal includes a face image sent by a second terminal to the first terminal, and the second terminal is a terminal corresponding to a temporary user; and the registration information further includes door opening permission information corresponding to the uploaded face image. For example, the temporary user can be a courier, etc. In this example, the owner can set the door opening permission information for the temporary user such as a courier.
[0149] In one possible implementation, the method further includes: acquiring information of adjusting a seat by a passenger of the vehicle; and generating or updating seat preference information corresponding to the passenger according to the information of adjusting the seat by the passenger. The seat preference information corresponding to the passenger can reflect the preference information of adjusting the seat by the passenger when the passenger rides the vehicle. In this implementation, by generating or updating the seat preference information corresponding to the passenger, the seat can be automatically adjusted according to the seat preference information corresponding to the passenger when the passenger rides the vehicle next time, so as to improve the riding experience of the passenger.
[0150] In a possible implementation, the generating or updating the seat preference information corresponding to the passenger according to the information of the passenger adjusting the seat comprises: generating or updating the seat preference information corresponding to the passenger according to the position information of the seat on which the passenger is seated and the information of the passenger adjusting the seat. In this implementation, the seat preference information corresponding to the passenger can be associated not only with the information of the passenger adjusting the seat, but also with the position information of the seat on which the passenger is seated, that is, the seat preference information corresponding to the seat in different positions can be recorded for the passenger, thereby further improving the driving experience of the user.
[0151] In a possible implementation, the method further comprises: obtaining seat preference information corresponding to the passenger based on the face recognition result; and adjusting the seat on which the passenger is seated according to the seat preference information corresponding to the passenger. In this implementation, the seat information is automatically adjusted for the passenger according to the seat preference information corresponding to the passenger, without manual adjustment by the passenger, thereby improving the driving or riding experience of the passenger.
[0152] In an example, one or more of the height, the front-back, the backrest, and the temperature of the seat can be adjusted.
[0153] As an example of this implementation, the adjusting the seat on which the passenger is seated according to the seat preference information corresponding to the passenger comprises: determining position information of the seat on which the passenger is seated; and adjusting the seat on which the passenger is seated according to the position information of the seat on which the passenger is seated and the seat preference information corresponding to the passenger. In this implementation, the seat information is automatically adjusted for the passenger according to the position information of the seat on which the passenger is seated and the seat preference information corresponding to the passenger, without manual adjustment by the passenger, thereby improving the driving or riding experience of the passenger.
[0154] In other possible implementations, other personalized information, such as light information, temperature information, air conditioner wind force information, music information, and the like, corresponding to the passenger can also be obtained based on the face recognition result, and automatic setting can be performed according to the obtained personalized information.
[0155] In a possible implementation, before the image acquisition module arranged on the vehicle is controlled to acquire the video stream, the method further includes: searching, by a Bluetooth module arranged on the vehicle, for a Bluetooth device with a preset identifier; in response to searching for the Bluetooth device with the preset identifier, establishing a Bluetooth pairing connection between the Bluetooth module and the Bluetooth device with the preset identifier; in response to the Bluetooth pairing connection being successful, waking up a face recognition module arranged on the vehicle; and the controlling the image acquisition module arranged on the vehicle to acquire the video stream includes: controlling the face recognition module that is woken up to acquire the video stream.
[0156] As an example of this implementation, the searching, by the Bluetooth module arranged on the vehicle, for the Bluetooth device with the preset identifier includes: searching, by the Bluetooth module arranged on the vehicle, for the Bluetooth device with the preset identifier when the vehicle is in an engine-off state or in an engine-off and door-locked state. In this example, the Bluetooth module does not need to search for the Bluetooth device with the preset identifier before the vehicle is turned off, or does not need to search for the Bluetooth device with the preset identifier before the vehicle is turned off and when the vehicle is in the engine-off state but the door is not in the locked state, thereby further reducing power consumption.
[0157] As an example of this implementation, the Bluetooth module can be a Bluetooth Low Energy (BLE) module. In this example, when the vehicle is in an engine-off state or in an engine-off and door-locked state, the Bluetooth module can be in a broadcast mode and broadcast a broadcast data packet to the surroundings every certain time (for example, 100 milliseconds). When performing a scanning action, a Bluetooth device in the surroundings sends a scanning request to the Bluetooth module if the broadcast data packet broadcast by the Bluetooth module is received, and the Bluetooth module returns a scanning response data packet to the Bluetooth device sending the scanning request in response to the scanning request. In this implementation, if a scanning request from a Bluetooth device with a preset identifier is received, it is determined that the Bluetooth device with the preset identifier is searched.
[0158] As another example of this implementation, when the vehicle is in an engine-off state or in an engine-off and door-locked state, the Bluetooth module can be in a scanning state, and if a Bluetooth device with a preset identifier is scanned, it is determined that the Bluetooth device with the preset identifier is searched.
[0159] As an example of this implementation, the Bluetooth module and the face recognition module can be integrated in a face recognition system.
[0160] As another example of this implementation, the Bluetooth module can be independent of the face recognition system. That is, the Bluetooth module can be arranged outside the face recognition system.
[0161] The implementation does not limit the maximum search distance of the Bluetooth module. In an example, the maximum search distance can be about 30 m.
[0162] In the implementation, the identification of the Bluetooth device can refer to a unique identifier of the Bluetooth device. For example, the identification of the Bluetooth device can be an ID, a name, or an address of the Bluetooth device, etc.
[0163] In the implementation, the preset identification can be an identification of a device that has been successfully paired with the Bluetooth module of the vehicle in advance based on the Bluetooth security connection technology.
[0164] In the implementation, the number of Bluetooth devices with the preset identification can be one or more. For example, if the identification of the Bluetooth device is an ID of the Bluetooth device, one or more Bluetooth IDs that are authorized to open the door of the vehicle can be preset. For example, if the number of Bluetooth devices with the preset identification is one, the Bluetooth device with the preset identification can be a Bluetooth device of the owner of the vehicle. If the number of Bluetooth devices with the preset identification is more than one, the Bluetooth devices with the preset identification can include a Bluetooth device of the owner of the vehicle and Bluetooth devices of family members, friends, and pre-registered contacts of the owner of the vehicle. The pre-registered contacts can be pre-registered delivery personnel or property staff, etc.
[0165] In the implementation, the Bluetooth device can be any mobile device with Bluetooth function, for example, the Bluetooth device can be a mobile phone, a wearable device, or an electronic key, etc. The wearable device can be a smart bracelet or smart glasses, etc.
[0166] As an example of the implementation, if the number of Bluetooth devices with the preset identification is more than one, a Bluetooth pairing connection between the Bluetooth module and any Bluetooth device with the preset identification is established in response to searching for the Bluetooth device with the preset identification.
[0167] As an example of the implementation, in response to searching for the Bluetooth device with the preset identification, the Bluetooth module performs identity authentication on the Bluetooth device with the preset identification. After the identity authentication is passed, the Bluetooth pairing connection between the Bluetooth module and the Bluetooth device with the preset identification is established. In this way, the security of the Bluetooth pairing connection can be improved.
[0168] In the implementation, when the Bluetooth pairing connection with the Bluetooth device with the preset identifier is not established, the face recognition module can be in a dormant state to keep low-power operation, so as to reduce the operation power consumption of the face recognition opening door method, and the face recognition module can be in an operable state when the user carrying the Bluetooth device with the preset identifier reaches the door, and after the first image is collected by the image collection module, the face image processing can be quickly performed through the awakened face recognition module, so as to improve the face recognition efficiency and improve the user experience. Therefore, the embodiments of the present disclosure can meet the requirements of low-power operation and fast door opening.
[0169] In the implementation, if the Bluetooth device with the preset identifier is searched, it can be indicated to a large extent that the user (for example, the owner) carrying the Bluetooth device with the preset identifier enters the search range of the Bluetooth module. At this time, the Bluetooth pairing connection between the Bluetooth module and the Bluetooth device with the preset identifier is established in response to searching the Bluetooth device with the preset identifier, and the face recognition module is awakened and the image collection module is controlled to collect the video stream in response to the successful Bluetooth pairing connection. Therefore, the face recognition module is awakened based on the successful Bluetooth pairing connection, which can effectively reduce the probability of false awakening of the face recognition module, and even avoid false awakening of the face recognition module, so as to improve the user experience and effectively reduce the power consumption of the face recognition module. In addition, compared with ultrasonic, infrared and other short-distance sensor technologies, the pairing connection based on Bluetooth has the advantages of high security and supporting a larger distance. Practice shows that the time for the user carrying the Bluetooth device with the preset identifier to reach the car through the distance (the distance between the user and the car when the Bluetooth pairing connection is successful) is roughly matched with the time for the car to awaken the face recognition module from the dormant state to the working state. Therefore, when the user reaches the door, the face recognition module can be awakened to perform face recognition to open the door, without the need for the user to wait for the face recognition module to be awakened after reaching the door, so as to improve the face recognition efficiency and improve the user experience. In addition, the user has no perception during the Bluetooth pairing connection, so as to further improve the user experience. Therefore, the implementation provides a solution that can better balance the power consumption saving, user experience and safety of the face recognition module.
[0170] In another possible implementation, the face recognition module can be awakened in response to the user touching the face recognition module. According to the implementation, when the user forgets to carry the mobile phone or other Bluetooth device, the face recognition opening door function can still be used.
[0171] In a possible implementation, after the face recognition module of the vehicle is woken up, the method further includes: if a face image is not captured within a preset time, controlling the face recognition module to enter a dormant state. This implementation can reduce power consumption by controlling the face recognition module to enter a dormant state if a face image is not captured within a preset time after the face recognition module is woken up.
[0172] In a possible implementation, after the face recognition module of the vehicle is woken up, the method further includes: if face recognition fails within a preset time, controlling the face recognition module to enter a dormant state. This implementation can reduce power consumption by controlling the face recognition module to enter a dormant state if face recognition fails within a preset time after the face recognition module is woken up.
[0173] In a possible implementation, after the face recognition module of the vehicle is woken up, the method further includes: if the driving speed of the vehicle is not 0, controlling the face recognition module to enter a dormant state. In this implementation, the safety of face-unlocking can be improved, and power consumption can be reduced, by controlling the face recognition module to enter a dormant state if the driving speed of the vehicle is not 0.
[0174] In another possible implementation, before the image capture module of the vehicle is controlled to capture a video stream, the method further includes: searching for a Bluetooth device with a preset identifier via a Bluetooth module of the vehicle; in response to searching for the Bluetooth device with the preset identifier, waking up a face recognition module of the vehicle; and the controlling the image capture module of the vehicle to capture a video stream includes: controlling the image capture module to capture a video stream via the woken-up face recognition module.
[0175] In a possible implementation, after the face recognition result is obtained, the method further includes: in response to the face recognition result being face recognition failure, activating a password unlocking module of the vehicle to start a password unlocking process.
[0176] In this implementation, password unlocking is an alternative to face recognition unlocking. Reasons for face recognition failure can include at least one of a live body detection result being a human fake body, face authentication failure, image capture failure (for example, camera failure), and a number of recognitions exceeding a predetermined number. When the person does not pass face recognition, the password unlocking process is started. For example, a password input by a user can be obtained through a touch screen on the B pillar. In one example, after M consecutive incorrect passwords are input, the password unlocking is disabled, for example, M is equal to 5.
[0177] In a possible implementation, the living body detection based on the first image and the first depth map comprises: updating the first depth map based on the first image to obtain a second depth map; and determining a living body detection result based on the first image and the second depth map.
[0178] In this implementation, the depth value of one or more pixels in the first depth map can be updated based on the first image to obtain a second depth map.
[0179] In a possible implementation, the updating of the first depth map based on the first image to obtain a second depth map comprises: updating the depth value of a depth failure pixel in the first depth map based on the first image to obtain the second depth map.
[0180] In the depth map, a depth failure pixel can refer to a pixel whose depth value is invalid, i.e., a pixel whose depth value is inaccurate or obviously inconsistent with the actual situation. The number of depth failure pixels can be one or more. By updating the depth value of at least one depth failure pixel in the depth map, the depth value of the depth failure pixel is more accurate, which helps to improve the accuracy of living body detection.
[0181] In some embodiments, the first depth map is a depth map with missing values, and the second depth map is obtained by repairing the first depth map based on the first image. Optionally, the repairing of the first depth map comprises determining or supplementing the depth value of a pixel with a missing value, but the embodiments of the present disclosure are not limited thereto.
[0182] In the embodiments of the present disclosure, the first depth map can be updated or repaired in various ways. In some embodiments, the first image is directly used for living body detection, for example, the first depth map is directly updated using the first image. In other embodiments, the first image is preprocessed, and living body detection is performed based on the preprocessed first image. For example, the updating of the first depth map based on the first image comprises: obtaining an image of the face from the first image; and updating the first depth map based on the image of the face.
[0183] The image of the face can be obtained from the first image in various manners. As an example, face detection is performed on the first image to obtain position information of the face, such as position information of a bounding box of the face, and the image of the face is obtained from the first image based on the position information of the face. For example, an image of a region where the bounding box of the face is located is obtained from the first image as the image of the face, and for another example, the bounding box of the face is enlarged by a certain multiple and an image of a region where the enlarged bounding box is located is obtained from the first image as the image of the face. As another example, the obtaining the image of the face from the first image includes: obtaining key point information of the face in the first image; and obtaining the image of the face from the first image based on the key point information of the face.
[0184] Optionally, the obtaining the key point information of the face in the first image includes: performing face detection on the first image to obtain a region where the face is located; and performing key point detection on an image of the region where the face is located to obtain the key point information of the face in the first image.
[0185] Optionally, the key point information of the face can include position information of a plurality of key points of the face. For example, the key points of the face can include one or more of an eye key point, a brow key point, a nose key point, a mouth key point, and a face contour key point. The eye key point can include one or more of an eye contour key point, a corner of the eye key point, and a pupil key point.
[0186] In an example, based on the key point information of the face, a contour of the face is determined, and the image of the face is obtained from the first image according to the contour of the face. Compared with the position information of the face obtained by face detection, the position of the face obtained by the key point information is more accurate, thereby facilitating improvement of the accuracy of subsequent live body detection.
[0187] Optionally, based on the key points of the face in the first image, a contour of the face in the first image can be determined, and an image of a region where the contour of the face in the first image is located or an image of a region obtained after the region is enlarged by a certain multiple is determined as the image of the face. For example, an elliptical region in the first image determined based on the key points of the face can be determined as the image of the face, or a minimum circumscribed rectangular region of the elliptical region in the first image determined based on the key points of the face can be determined as the image of the face, but the embodiments of the present disclosure are not limited thereto.
[0188] In this way, by obtaining the image of the face from the first image and performing live body detection based on the image of the face, interference of background information in the first image on the live body detection can be reduced.
[0189] In the embodiments of the present disclosure, the obtained original depth map can be updated. Alternatively, in some embodiments, the updating the first depth map based on the first image to obtain a second depth map comprises: obtaining a depth map of a face from the first depth map; and updating the depth map of the face based on the first image to obtain the second depth map.
[0190] As an example, the position information of the face in the first image is obtained, and the depth map of the face is obtained from the first depth map based on the position information of the face. Optionally, the first depth map and the first image can be registered or aligned in advance, but the embodiments of the present disclosure are not limited thereto.
[0191] In this way, by obtaining the depth map of the face from the first depth map and updating the depth map of the face based on the first image to obtain the second depth map, the interference of the background information in the first depth map on the live detection can be reduced.
[0192] In some embodiments, after obtaining the first image and the first depth map corresponding to the first image, the first image and the first depth map are aligned according to the parameters of the image sensor and the parameters of the depth sensor.
[0193] As an example, the first depth map can be converted so that the converted first depth map is aligned with the first image. For example, a first conversion matrix can be determined according to the parameters of the depth sensor and the parameters of the image sensor, and the first depth map is converted according to the first conversion matrix. Accordingly, at least a part of the converted first depth map can be updated based on at least a part of the first image to obtain the second depth map. For example, the converted first depth map is updated based on the first image to obtain the second depth map. For another example, the depth map of the face obtained from the first depth map is updated based on the image of the face cropped from the first image to obtain the second depth map, and so on.
[0194] As another example, the first image can be converted so that the converted first image is aligned with the first depth map. For example, a second conversion matrix can be determined according to the parameters of the depth sensor and the parameters of the image sensor, and the first image is converted according to the second conversion matrix. Accordingly, at least a part of the first depth map can be updated based on at least a part of the converted first image to obtain the second depth map.
[0195] Optionally, the parameters of the depth sensor can include intrinsic parameters and / or extrinsic parameters of the depth sensor, and the parameters of the image sensor can include intrinsic parameters and / or extrinsic parameters of the image sensor. By aligning the first depth map and the first image, the corresponding parts in the first depth map and the first image can be located at the same positions in the two images.
[0196] In the above example, the first image is an original image (for example, an RGB or infrared image), and in other embodiments, the first image can also refer to an image of a face cropped from the original image, and similarly, the first depth map can also refer to a depth map of a face cropped from the original depth map, which is not limited in the embodiments of the present disclosure.
[0197] Figure 7 A schematic diagram showing one example of a living body detection method according to an embodiment of the present disclosure is shown. In Figure 7 In the example shown, the first image is an RGB image, the RGB image and the first depth map are aligned and corrected, and the processed image is input into a face key point model for processing to obtain an RGB face map (an image of a face) and a depth face map (a depth map of a face), and the depth face map is updated or repaired based on the RGB face map. In this way, the amount of subsequent data processing can be reduced, and the efficiency and accuracy of living body detection can be improved.
[0198] In the embodiments of the present disclosure, the living body detection result of the face can be that the face is a living body or the face is a fake body.
[0199] In some embodiments, the determining a living body detection result based on the first image and the second depth map includes: inputting the first image and the second depth map into a living body detection neural network for processing to obtain the living body detection result. Alternatively, the first image and the second depth map are processed by other living body detection algorithms to obtain the living body detection result.
[0200] In some embodiments, the determining a living body detection result based on the first image and the second depth map includes: performing feature extraction processing on the first image to obtain first feature information; performing feature extraction processing on the second depth map to obtain second feature information; and determining a living body detection result based on the first feature information and the second feature information.
[0201] Optionally, the feature extraction processing can be implemented by a neural network or other machine learning algorithm, and the type of the extracted feature information can be obtained by learning from samples, which is not limited in the embodiments of the present disclosure.
[0202] In some specific scenarios (such as an outdoor strong light scenario), the acquired depth map (for example, a depth map acquired by a depth sensor) can have a partially invalid area. In addition, under normal lighting, factors such as reflection of glasses, black hair, or black glasses frames can also randomly cause local invalidation of the depth map. Certain special paper can cause a printed face photo to have a large-area or local invalidation effect. In addition, blocking the active light source of the depth sensor can also cause partial invalidation of the depth map, while the imaging of the dummy on the image sensor is normal. Therefore, in some cases of partial or complete invalidation of the depth map, using the depth map to distinguish between a living body and a dummy can cause errors. Therefore, in the embodiments of the present disclosure, the first depth map is repaired or updated, and the repaired or updated depth map is used for living body detection, which is beneficial to improving the accuracy of living body detection.
[0203] Figure 8 FIG. 6 shows an example of determining a living body detection result of a face in a first image based on the first image and a second depth map in a living body detection method according to an embodiment of the present disclosure.
[0204] In this example, the first image and the second depth map are input into a living body detection neural network for living body detection processing to obtain a living body detection result.
[0205] As shown in FIG. 6, the living body detection neural network includes two branches, i.e., a first subnetwork and a second subnetwork. The first subnetwork is configured to perform feature extraction processing on the first image to obtain first feature information, and the second subnetwork is configured to perform feature extraction processing on the second depth map to obtain second feature information. Figure 8 In an optional example, the first subnetwork can include a convolution layer, a down-sampling layer, and a fully connected layer.
[0206] For example, the first subnetwork can include a first-level convolution layer, a first-level down-sampling layer, and a first-level fully connected layer. The first-level convolution layer can include one or more convolution layers, the first-level down-sampling layer can include one or more down-sampling layers, and the first-level fully connected layer can include one or more fully connected layers.
[0207] For another example, the first subnetwork can include multiple levels of convolution layers, multiple levels of down-sampling layers, and a first-level fully connected layer. Each level of convolution layer can include one or more convolution layers, each level of down-sampling layer can include one or more down-sampling layers, and the first-level fully connected layer can include one or more fully connected layers. The i th convolution layer is cascaded with the i th down-sampling layer, the i th down-sampling layer is cascaded with the i+1 th convolution layer, and the n th down-sampling layer is cascaded with the fully connected layer, where i and n are positive integers, 1≤i≤n, and n represents the number of levels of the convolution layers and the down-sampling layers in the depth prediction neural network.
[0208]
[0209] Alternatively, the first sub-network can include a convolutional layer, a down-sampling layer, a normalization layer, and a fully connected layer.
[0210] For example, the first sub-network can include a first convolutional layer, a first normalization layer, a first down-sampling layer, and a first fully connected layer. The first convolutional layer can include one or more convolutional layers, the first down-sampling layer can include one or more down-sampling layers, and the first fully connected layer can include one or more fully connected layers.
[0211] For another example, the first sub-network can include a plurality of convolutional layers, a plurality of normalization layers, and a plurality of down-sampling layers, and a first fully connected layer. Each convolutional layer can include one or more convolutional layers, each down-sampling layer can include one or more down-sampling layers, and the first fully connected layer can include one or more fully connected layers. The i-th convolutional layer is cascaded with the i-th normalization layer, the i-th normalization layer is cascaded with the i-th down-sampling layer, the i-th down-sampling layer is cascaded with the (i+1)-th convolutional layer, and the n-th down-sampling layer is cascaded with the fully connected layer, where i and n are positive integers, 1≤i≤n, and n represents the number of convolutional layers, the number of down-sampling layers, and the number of normalization layers in the first sub-network.
[0212] As an example, the first image is subjected to convolution processing to obtain a first convolutional result, the first convolutional result is subjected to down-sampling processing to obtain a first down-sampled result, and the first down-sampled result is used to obtain first feature information.
[0213] For example, the first image can be subjected to convolution processing and down-sampling processing by a first convolutional layer and a first down-sampling layer. The first convolutional layer can include one or more convolutional layers, and the first down-sampling layer can include one or more down-sampling layers.
[0214] For another example, the first image can be subjected to convolution processing and down-sampling processing by a plurality of convolutional layers and a plurality of down-sampling layers. Each convolutional layer can include one or more convolutional layers, and each down-sampling layer can include one or more down-sampling layers.
[0215] For example, the first convolutional result can be subjected to down-sampling processing to obtain a first down-sampled result, which can include normalizing the first convolutional result to obtain a first normalized result, and subjecting the first normalized result to down-sampling processing to obtain the first down-sampled result.
[0216] For example, the first down-sampled result can be input into a fully connected layer, and the first down-sampled result can be subjected to fusion processing by the fully connected layer to obtain first feature information.
[0217] Optionally, the second sub-network and the first sub-network have the same network structure but different parameters. Alternatively, the second sub-network has a different network structure from the first sub-network, which is not limited in the embodiments of the present disclosure.
[0218] As shown in Figure 8 The living body detection neural network further includes a third sub-network configured to process the first feature information obtained by the first sub-network and the second feature information obtained by the second sub-network to obtain a living body detection result of the face in the first image. Optionally, the third sub-network can include a full connection layer and an output layer. For example, the output layer adopts a softmax function, and if the output of the output layer is 1, it indicates that the face is a living body, and if the output of the output layer is 0, it indicates that the face is a fake body, but the specific implementation of the third sub-network is not limited in the embodiments of the present disclosure.
[0219] As an example, the living body detection result is determined based on the first feature information and the second feature information, including: performing fusion processing on the first feature information and the second feature information to obtain third feature information; and determining the living body detection result based on the third feature information.
[0220] For example, the first feature information and the second feature information are fused by the full connection layer to obtain the third feature information.
[0221] In some embodiments, the living body detection result is determined based on the third feature information, including: obtaining a probability that the face is a living body based on the third feature information; and determining the living body detection result according to the probability that the face is a living body.
[0222] For example, if the probability that the face is a living body is greater than a second threshold value, it is determined that the living body detection result of the face is that the face is a living body. For another example, if the probability that the face is a living body is less than or equal to the second threshold value, it is determined that the living body detection result of the face is that the face is a fake body.
[0223] In other embodiments, a probability that the face is a fake body is obtained based on the third feature information, and the living body detection result of the face is determined according to the probability that the face is a fake body. For example, if the probability that the face is a fake body is greater than a third threshold value, it is determined that the living body detection result of the face is that the face is a fake body. For another example, if the probability that the face is a fake body is less than or equal to the third threshold value, it is determined that the living body detection result of the face is that the face is a living body.
[0224] In one example, the third feature information can be input into a Softmax layer, and the probability that the face is a living body or a fake body is obtained through the Softmax layer. For example, the output of the Softmax layer includes two neurons, one of which represents the probability that the face is a living body, and the other represents the probability that the face is a fake body, but the embodiments of the present disclosure are not limited thereto.
[0225] In the embodiments of the present disclosure, by acquiring a first image and a first depth map corresponding to the first image, updating the first depth map based on the first image to obtain a second depth map, and determining a live body detection result of a face in the first image based on the first image and the second depth map, the depth map can be perfected, and thus the accuracy of live body detection can be improved.
[0226] In a possible implementation, the updating the first depth map based on the first image to obtain a second depth map includes: determining depth prediction values and associated information of a plurality of pixels in the first image based on the first image, where the associated information of the plurality of pixels indicates a degree of association between the plurality of pixels; and updating the first depth map based on the depth prediction values and the associated information of the plurality of pixels to obtain a second depth map.
[0227] Specifically, the depth prediction values of the plurality of pixels in the first image are determined based on the first image, and the first depth map is repaired and perfected based on the depth prediction values of the plurality of pixels.
[0228] Specifically, the depth prediction values of the plurality of pixels in the first image are obtained by processing the first image. For example, the first image is input into a depth prediction neural network for processing to obtain the depth prediction values of the plurality of pixels, for example, a depth prediction map corresponding to the first image, but the embodiments of the present disclosure are not limited thereto.
[0229] In some embodiments, the determining the depth prediction values of the plurality of pixels in the first image based on the first image includes: determining the depth prediction values of the plurality of pixels in the first image based on the first image and the first depth map.
[0230] As an example, the determining the depth prediction values of the plurality of pixels in the first image based on the first image and the first depth map includes: inputting the first image and the first depth map into a depth prediction neural network for processing to obtain the depth prediction values of the plurality of pixels in the first image. Alternatively, the first image and the first depth map are processed by other manners to obtain the depth prediction values of the plurality of pixels, and the embodiments of the present disclosure are not limited thereto.
[0231] Figure 9 A schematic diagram of a depth prediction neural network in the vehicle door control method provided by the embodiments of the present disclosure is shown. As shown in Figure 9 The first image and the first depth map can be input into the depth prediction neural network for processing to obtain an initial depth estimation map. Based on the initial depth estimation map, the depth prediction values of the plurality of pixels in the first image can be determined. For example, the pixel values of the initial depth estimation map are the depth prediction values of the corresponding pixels in the first image.
[0232] The deep prediction neural network can be implemented by various network structures. In one example, the deep prediction neural network comprises an encoding part and a decoding part. Optionally, the encoding part can comprise a convolution layer and a down-sampling layer, and the decoding part can comprise an inverse convolution layer and / or an up-sampling layer. In addition, the encoding part and / or the decoding part can further comprise a normalization layer, and the specific implementation of the encoding part and the decoding part is not limited in the embodiments of the present disclosure. In the encoding part, the resolution of the feature map gradually decreases and the number of feature maps gradually increases with the increase of the number of network layers, so that rich semantic features and image spatial features can be obtained; in the decoding part, the resolution of the feature map gradually increases, and the resolution of the feature map finally output by the decoding part is the same as that of the first depth map.
[0233] In some embodiments, the determining the depth prediction values of the plurality of pixels in the first image based on the first image and the first depth map comprises: performing fusion processing on the first image and the first depth map to obtain a fusion result; and determining the depth prediction values of the plurality of pixels in the first image based on the fusion result.
[0234] In one example, the first image and the first depth map can be concatenated to obtain the fusion result.
[0235] In one example, the fusion result is subjected to convolution processing to obtain a second convolution result; the second convolution result is subjected to down-sampling processing to obtain a first encoding result; and the depth prediction values of the plurality of pixels in the first image are determined based on the first encoding result.
[0236] For example, the fusion result can be subjected to convolution processing by a convolution layer to obtain the second convolution result.
[0237] For example, the second convolution result is subjected to normalization processing to obtain a second normalization result; and the second normalization result is subjected to down-sampling processing to obtain the first encoding result. Here, the second convolution result can be subjected to normalization processing by a normalization layer to obtain the second normalization result; and the second normalization result can be subjected to down-sampling processing by a down-sampling layer to obtain the first encoding result. Alternatively, the second convolution result can be subjected to down-sampling processing by a down-sampling layer to obtain the first encoding result.
[0238] For example, the first encoding result is subjected to inverse convolution processing to obtain a first inverse convolution result; and the first inverse convolution result is subjected to normalization processing to obtain the depth prediction values. Here, the first encoding result can be subjected to inverse convolution processing by an inverse convolution layer to obtain the first inverse convolution result; and the first inverse convolution result can be subjected to normalization processing by a normalization layer to obtain the depth prediction values. Alternatively, the first encoding result can be subjected to inverse convolution processing by an inverse convolution layer to obtain the depth prediction values.
[0239] For example, the first encoding result is upsampled to obtain a first upsampled result; the first upsampled result is then normalized to obtain a depth prediction value. Here, an upsampling layer can be used to upsample the first encoding result to obtain the first upsampled result; a normalization layer can then be used to normalize the first upsampled result to obtain the depth prediction value. Alternatively, an upsampling layer can be used to upsample the first encoding result to obtain the depth prediction value.
[0240] Furthermore, by processing the first image, association information of multiple pixels in the first image is obtained. This association information can include the correlation degree between each pixel in the first image and its surrounding pixels. The surrounding pixels can include at least one adjacent pixel, or multiple pixels spaced no more than a certain value from the pixel. For example, such as... Figure 12 As shown, the surrounding pixels of pixel 5 include its neighboring pixels 1, 2, 3, 4, 6, 7, 8, and 9. Correspondingly, the association information of multiple pixels in the first image includes the correlation degree between pixels 1, 2, 3, 4, 6, 7, 8, and 9 and pixel 5. As an example, the correlation degree between the first pixel and the second pixel can be measured using the correlation between the first pixel and the second pixel. In this embodiment, correlation techniques can be used to determine the correlation between pixels, which will not be elaborated further here.
[0241] In this embodiment of the disclosure, the association information of multiple pixels can be determined in various ways. In some embodiments, determining the association information of multiple pixels in the first image based on the first image includes: inputting the first image into an association detection neural network for processing to obtain the association information of multiple pixels in the first image. For example, obtaining the association feature map corresponding to the first image. Alternatively, the association information of multiple pixels can also be obtained by other algorithms, which are not limited in this embodiment of the disclosure.
[0242] Figure 10 This diagram illustrates a correlation detection neural network in the door control method provided in an embodiment of this disclosure. For example... Figure 10 As shown, the first image is input into the correlation detection neural network for processing, resulting in multiple correlation feature maps. Based on these multiple correlation feature maps, the correlation information of multiple pixels in the first image can be determined. For example, the surrounding pixels of a pixel refer to pixels with a distance of 0 from that pixel; that is, the surrounding pixels of a pixel refer to pixels adjacent to that pixel. The correlation detection neural network can then output 8 correlation feature maps. For example, in the first correlation feature map, pixel P... i,j The pixel value = pixel P in the first image i-1,j-1 With pixel Pi,j the correlation degree between pixel P i,j represents the pixel in the i-th row and the j-th column; in the second correlation feature map, the pixel value of pixel P i,j represents the correlation degree between pixel P i-1,j and pixel P i,j ; in the third correlation feature map, the pixel value of pixel P i,j represents the correlation degree between pixel P i-1,j+1 and pixel P i,j ; in the fourth correlation feature map, the pixel value of pixel P i,j represents the correlation degree between pixel P i,j-1 and pixel P i,j ; in the fifth correlation feature map, the pixel value of pixel P i,j represents the correlation degree between pixel P i,j+1 and pixel P i,j ; in the sixth correlation feature map, the pixel value of pixel P i,j represents the correlation degree between pixel P i+1,j-1 and pixel P i,j ; in the seventh correlation feature map, the pixel value of pixel P i,j represents the correlation degree between pixel P i+1,j and pixel P i,j ; in the eighth correlation feature map, the pixel value of pixel P i,j represents the correlation degree between pixel P i+1,j+1 and pixel P i,j .
[0243] The correlation degree detection neural network can be implemented through various network structures. As an example, the correlation degree detection neural network can include an encoding part and a decoding part. The encoding part can include a convolutional layer and a down-sampling layer, and the decoding part can include an inverse convolutional layer and / or an up-sampling layer. The encoding part can also include a normalization layer, and the decoding part can also include a normalization layer. In the encoding part, the resolution of the feature map gradually decreases, and the number of feature maps gradually increases, thereby obtaining rich semantic features and image spatial features; in the decoding part, the resolution of the feature map gradually increases, and the resolution of the feature map finally output by the decoding part is the same as that of the first image. In the embodiments of the present disclosure, the correlation information can be an image, or other data forms such as a matrix.
[0244] As an example, inputting the first image into the correlation detection neural network for processing to obtain the correlation information of the plurality of pixels in the first image can include: performing convolution processing on the first image to obtain a third convolution result; performing down-sampling processing based on the third convolution result to obtain a second encoding result; and obtaining the correlation information of the plurality of pixels in the first image based on the second encoding result.
[0245] In one example, the first image can be convoluted by a convolution layer to obtain the third convolution result.
[0246] In one example, the down-sampling processing based on the third convolution result to obtain the second encoding result can include: performing normalization processing on the third convolution result to obtain a third normalization result; and performing down-sampling processing on the third normalization result to obtain the second encoding result. In this example, the third convolution result can be normalized by a normalization layer to obtain the third normalization result, and the third normalization result can be down-sampled by a down-sampling layer to obtain the second encoding result. Alternatively, the third convolution result can be down-sampled by a down-sampling layer to obtain the second encoding result.
[0247] In one example, determining the correlation information based on the second encoding result can include: performing deconvolution processing on the second encoding result to obtain a second deconvolution result; and performing normalization processing on the second deconvolution result to obtain the correlation information. In this example, the second encoding result can be deconvoluted by a deconvolution layer to obtain the second deconvolution result, and the second deconvolution result can be normalized by a normalization layer to obtain the correlation information. Alternatively, the second encoding result can be deconvoluted by a deconvolution layer to obtain the correlation information.
[0248] In one example, determining the correlation information based on the second encoding result can include: performing up-sampling processing on the second encoding result to obtain a second up-sampling result; and performing normalization processing on the second up-sampling result to obtain the correlation information. In this example, the second encoding result can be up-sampled by an up-sampling layer to obtain the second up-sampling result, and the second up-sampling result can be normalized by a normalization layer to obtain the correlation information. Alternatively, the second encoding result can be up-sampled by an up-sampling layer to obtain the correlation information.
[0249] The current TOF, structured light and other 3D sensors are easily affected by sunlight outdoors, resulting in large areas of missing holes in the depth map, thereby affecting the performance of the 3D living body detection algorithm. The 3D living body detection algorithm based on depth map self-completion proposed in the embodiments of the present disclosure improves the performance of the 3D living body detection algorithm by repairing the depth map detected by the 3D sensor.
[0250] In some embodiments, after obtaining the depth prediction values and the associated information of the plurality of pixels, the first depth map is updated based on the depth prediction values and the associated information of the plurality of pixels to obtain a second depth map. Figure 11 An exemplary schematic diagram of depth map updating in the vehicle door control method provided by the embodiments of the present disclosure is shown. Figure 11 In the example shown, the first depth map is a depth map with missing values, and the obtained depth prediction values and associated information of the plurality of pixels are an initial depth estimation map and an associated feature map, respectively. At this time, the depth map with missing values, the initial depth estimation map, and the associated feature map are input into a depth map updating module (for example, a depth updating neural network) for processing to obtain a final depth map, i.e., a second depth map.
[0251] In a possible implementation, the updating of the first depth map based on the depth prediction values and the associated information of the plurality of pixels to obtain a second depth map includes: determining a depth failure pixel in the first depth map; obtaining a depth prediction value of the depth failure pixel and depth prediction values of a plurality of surrounding pixels of the depth failure pixel from the depth prediction values of the plurality of pixels; obtaining an association degree between the depth failure pixel and the plurality of surrounding pixels of the depth failure pixel from the associated information of the plurality of pixels; and determining an updated depth value of the depth failure pixel based on the depth prediction value of the depth failure pixel, the depth prediction values of the plurality of surrounding pixels of the depth failure pixel, and the association degree between the depth failure pixel and the surrounding pixels of the depth failure pixel.
[0252] In the embodiments of the present disclosure, the depth failure pixel in the depth map can be determined in various ways. As an example, a pixel with a depth value equal to 0 in the first depth map is determined as a depth failure pixel, or a pixel without a depth value in the first depth map is determined as a depth failure pixel.
[0253] In this example, for the part with a value in the first depth map with missing values (i.e., the depth value is not 0), it is considered that the depth value is correct and reliable, and this part is not updated, and the original depth value is retained. The depth value of the pixel with a depth value of 0 in the first depth map is updated.
[0254] As another example, the depth sensor can set the depth value of the depth failure pixel to one or more preset values or a preset range. In an example, a pixel with a depth value equal to a preset value or belonging to a preset range in the first depth map can be determined as a depth failure pixel.
[0255] The embodiments of the present disclosure can also determine the depth failure pixel in the first depth map based on other statistical methods, which are not limited in the embodiments of the present disclosure.
[0256] In the implementation, a depth value of a pixel in the first image that is same as the depth failure pixel position can be determined as the depth prediction value of the depth failure pixel, and similarly, a depth value of a pixel in the first image that is same as a surrounding pixel position of the depth failure pixel can be determined as the depth prediction value of the surrounding pixel of the depth failure pixel.
[0257] As an example, a distance between the surrounding pixel of the depth failure pixel and the depth failure pixel is less than or equal to a first threshold value.
[0258] Figure 12 A schematic diagram of the surrounding pixels in the vehicle door control method provided by the embodiments of the present disclosure is shown. For example, the first threshold value is 0, and only the neighbor pixels are used as the surrounding pixels. For example, the neighbor pixels of the pixel 5 include the pixel 1, the pixel 2, the pixel 3, the pixel 4, the pixel 6, the pixel 7, the pixel 8, and the pixel 9, and only the pixel 1, the pixel 2, the pixel 3, the pixel 4, the pixel 6, the pixel 7, the pixel 8, and the pixel 9 are used as the surrounding pixels of the pixel 5.
[0259] Figure 13 Another schematic diagram of the surrounding pixels in the vehicle door control method provided by the embodiments of the present disclosure is shown. For example, the first threshold value is 1, and in addition to the neighbor pixels, the neighbor pixels of the neighbor pixels are also used as the surrounding pixels. That is, in addition to the pixel 1, the pixel 2, the pixel 3, the pixel 4, the pixel 6, the pixel 7, the pixel 8, and the pixel 9, the pixels 10 to 25 are also used as the surrounding pixels of the pixel 5.
[0260] As an example, the determining the updated depth value of the depth failure pixel based on the depth prediction value of the depth failure pixel, the depth prediction values of the surrounding pixels of the depth failure pixel, and the correlation degrees between the depth failure pixel and the surrounding pixels of the depth failure pixel includes: determining a depth correlation value of the depth failure pixel based on the depth prediction values of the surrounding pixels of the depth failure pixel and the correlation degrees between the depth failure pixel and the surrounding pixels of the depth failure pixel; and determining the updated depth value of the depth failure pixel based on the depth prediction value of the depth failure pixel and the depth correlation value.
[0261] As another example, an effective depth value of a surrounding pixel of the depth failure pixel is determined based on a depth prediction value of the surrounding pixel and a correlation between the depth failure pixel and the surrounding pixel; an updated depth value of the depth failure pixel is determined based on the effective depth values of the surrounding pixels of the depth failure pixel and the depth prediction value of the depth failure pixel. For example, a product of the depth prediction value of a surrounding pixel of the depth failure pixel and a correlation corresponding to the surrounding pixel can be determined as the effective depth value of the surrounding pixel for the depth failure pixel, where the correlation corresponding to the surrounding pixel refers to the correlation between the surrounding pixel and the depth failure pixel. For example, a product of a sum of the effective depth values of the surrounding pixels of the depth failure pixel and a first preset coefficient can be determined as a first product; a product of the depth prediction value of the depth failure pixel and a second preset coefficient can be determined as a second product; and a sum of the first product and the second product can be determined as the updated depth value of the depth failure pixel. In some embodiments, a sum of the first preset coefficient and the second preset coefficient is 1.
[0262] In one example, the determining the depth correlation value of the depth failure pixel based on the depth prediction values of the surrounding pixels of the depth failure pixel and the correlations between the depth failure pixel and the surrounding pixels of the depth failure pixel comprises: performing weighted sum processing on the depth prediction values of the surrounding pixels of the depth failure pixel by taking the correlation between the depth failure pixel and each surrounding pixel as a weight of the each surrounding pixel, to obtain the depth correlation value of the depth failure pixel. For example, pixel 5 is the depth failure pixel, and the depth correlation value of the depth failure pixel 5 is the updated depth value of the depth failure pixel 5 wherein, w i represents the correlation between pixel i and pixel 5, F i represents the depth prediction value of pixel i.
[0263] In another example, a product of the correlation between each surrounding pixel of the depth failure pixel and the depth failure pixel and the depth prediction value of the each surrounding pixel is determined; and a maximum value of the products is determined as the depth correlation value of the depth failure pixel.
[0264] In one example, a sum of the depth prediction value of the depth failure pixel and the depth correlation value is determined as the updated depth value of the depth failure pixel.
[0265] In another example, the product of the predicted depth value of the depth-failed pixel and a third preset coefficient is determined to obtain a third product; the product of the depth correlation value and a fourth preset coefficient is determined to obtain a fourth product; the sum of the third product and the fourth product is determined as the updated depth value of the depth-failed pixel. In some embodiments, the sum of the third preset coefficient and the fourth preset coefficient is 1.
[0266] In some embodiments, the depth value of a non-depth-failure pixel in the second depth map is equal to the depth value of the non-depth-failure pixel in the first depth map.
[0267] In other embodiments, the depth values of non-depth-failed pixels can also be updated to obtain a more accurate second depth map, thereby further improving the accuracy of liveness detection.
[0268] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further.
[0269] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0270] In addition, this disclosure also provides a door control device, electronic device, computer-readable storage medium, and program, all of which can be used to implement any of the door control methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding descriptions in the method section and will not be repeated here.
[0271] Figure 14 A block diagram of a door control device according to an embodiment of the present disclosure is shown. Figure 14 As shown, the vehicle door control device includes: a first control module 21, used to control an image acquisition module installed in the vehicle to acquire a video stream; a face recognition module 22, used to perform face recognition based on at least one image in the video stream to obtain a face recognition result; a first determination module 23, used to determine control information corresponding to at least one door of the vehicle based on the face recognition result; a first acquisition module 24, used to acquire the status information of the door if the control information includes controlling any door of the vehicle to open; a second control module 25, used to control the door to unlock and open if the status information of the door is unlocked; and / or, control the door to open if the status information of the door is unlocked and not opened.
[0272] In a possible implementation, the apparatus further includes a second determination module configured to determine door opening intention information based on the video stream; and the first determination module 23 is configured to determine the control information corresponding to the at least one door of the vehicle based on the face recognition result and the door opening intention information.
[0273] In a possible implementation, the second determination module is configured to determine an intersection over union of images of adjacent frames in the video stream; and determine the door opening intention information according to the intersection over union of the images of the adjacent frames.
[0274] In a possible implementation, the second determination module is configured to determine the intersection over union of the images of the adjacent frames in the video stream as the intersection over union of the bounding boxes of the human bodies in the images of the adjacent frames.
[0275] In a possible implementation, the second determination module is configured to cache intersection over unions of images of N groups of adjacent frames that are newly collected, where N is an integer greater than 1; determine an average value of the cached intersection over unions; and determine that the door opening intention information is intentional door opening if the average value is greater than a first preset value for a duration reaching a first preset time length.
[0276] In a possible implementation, the second determination module is configured to determine areas of human body regions in a plurality of images that are newly collected in the video stream; and determine the door opening intention information according to the areas of the human body regions in the plurality of images that are newly collected.
[0277] In a possible implementation, the second determination module is configured to determine that the door opening intention information is intentional door opening if the areas of the human body regions in the plurality of images that are newly collected are all greater than a first preset area.
[0278] In a possible implementation, the second determination module is configured to determine that the door opening intention information is intentional door opening if the areas of the human body regions in the plurality of images that are newly collected gradually increase.
[0279] In a possible implementation, the first determination module 23 is configured to determine that the control information includes controlling the at least one door of the vehicle to be opened if the face recognition result is that face recognition is successful and the door opening intention information is intentional door opening.
[0280] In a possible implementation, the apparatus further includes a third determination module configured to perform object detection on at least one image in the video stream to determine object carrying information of a person; and the first determination module 23 is configured to determine the control information corresponding to the at least one door of the vehicle based on the face recognition result and the object carrying information of the person.
[0281] In a possible implementation, the first determining module 23 is configured to: if the face recognition result is face recognition success, and the object carrying information of the person is that the person carries an object, determine that the control information comprises control at least one door of the vehicle to open.
[0282] In a possible implementation, the first determining module 23 is configured to: if the face recognition result is face recognition success, and the object carrying information of the person is that the person carries an object of a preset category, determine that the control information comprises control the trunk door of the vehicle to open.
[0283] In a possible implementation, the apparatus further includes a third determining module configured to perform object detection on at least one image in the video stream to determine object carrying information of the person; and the first determining module 23 is configured to determine the control information corresponding to at least one door of the vehicle based on the face recognition result, the door opening intention information, and the object carrying information of the person.
[0284] In a possible implementation, the first determining module 23 is configured to: if the face recognition result is face recognition success, the door opening intention information is intentional door opening, and the object carrying information of the person is that the person carries an object, determine that the control information comprises control at least one door of the vehicle to open.
[0285] In a possible implementation, the first determining module 23 is configured to: if the face recognition result is face recognition success, the door opening intention information is intentional door opening, and the object carrying information of the person is that the person carries an object of a preset category, determine that the control information comprises control the trunk door of the vehicle to open.
[0286] In a possible implementation, the third determining module is configured to: perform object detection on at least one image in the video stream to obtain an object detection result; and determine the object carrying information of the person based on the object detection result.
[0287] In a possible implementation, the third determining module is configured to: detect a bounding box of a human body in at least one image in the video stream; and perform object detection on a region corresponding to the bounding box to obtain an object detection result.
[0288] In a possible implementation, the third determining module is configured to: if the object detection result is that an object is detected, obtain a distance between the object and a hand of the person; and determine the object carrying information of the person based on the distance.
[0289] In a possible implementation, the third determining module is configured to: if the object detection result is that an object is detected, acquire a size of the object; and determine the object carrying information of the person based on the distance and the size.
[0290] In a possible implementation, the third determining module is configured to: if the distance is less than or equal to a preset distance, and the size is greater than or equal to a preset size, determine that the object carrying information of the person is that the person carries an object.
[0291] In a possible implementation, the third determining module is configured to: if the object detection result is that an object is detected, acquire a size of the object; and determine the object carrying information of the person based on the size.
[0292] In a possible implementation, the first control module 21 is configured to: control an image acquisition module arranged on a trunk door of the vehicle to acquire a video stream.
[0293] In a possible implementation, the apparatus further includes a third control module configured to: control the trunk door to open, in a case where it is determined that the person leaves a cabin interior of the vehicle according to a video stream acquired by an image acquisition module arranged in the cabin interior of the vehicle, or in a case where it is detected that the door opening intention information of the person is an intention to get off the vehicle.
[0294] In a possible implementation, the first determining module 23 is configured to: if the face recognition result is that face recognition is successful and the person is not a driver, the door opening intention information is an intention to open the door, and the object carrying information of the person is that an object is carried, determine that the control information includes control of at least one non-driver seat door of the vehicle to be opened.
[0295] In a possible implementation, the apparatus further includes a fourth control module configured to: control the vehicle door to be closed, or control the vehicle door to be closed and locked, in a case where an automatic door closing condition is met.
[0296] In a possible implementation, the automatic door closing condition includes one or more of the following: the door opening intention information for controlling the vehicle door to be opened is an intention to get on the vehicle, and it is determined that a person with the intention to get on the vehicle has been seated according to a video stream acquired by an image acquisition module arranged in a cabin interior of the vehicle; the door opening intention information for controlling the vehicle door to be opened is an intention to get off the vehicle, and it is determined that a person with the intention to get off the vehicle has left the cabin interior according to a video stream acquired by an image acquisition module arranged in the cabin interior of the vehicle; and a duration for which the vehicle door is opened reaches a second preset duration.
[0297] In a possible implementation, the face recognition includes face authentication; and the face recognition module 22 is configured to perform face authentication based on the first image in the video stream and pre-registered face features.
[0298] In a possible implementation, the face recognition further includes living body detection; and the face recognition module 22 is configured to acquire, by using a depth sensor in the image acquisition module, a first depth map corresponding to the first image in the video stream, and perform living body detection based on the first image and the first depth map.
[0299] In a possible implementation, the face recognition further includes permission authentication; and the face recognition module 22 is configured to acquire, based on the first image in the video stream, door opening permission information of the person, and perform permission authentication based on the door opening permission information of the person.
[0300] In a possible implementation, the door opening permission information of the person includes one or more of the following: information of a door for which the person has door opening permission, a time for which the person has door opening permission, and a number of times of door opening permission corresponding to the person.
[0301] In a possible implementation, the information of the door for which the person has door opening permission is all doors or a trunk door.
[0302] In a possible implementation, the apparatus further includes one or both of the following: a first registration module configured to perform user registration based on a face image acquired by the image acquisition module; and a second registration module configured to perform remote registration based on a face image acquired or uploaded by a first terminal, and send registration information to the vehicle, wherein the first terminal is a terminal corresponding to a vehicle owner, and the registration information includes the acquired or uploaded face image.
[0303] In a possible implementation, the face image uploaded by the first terminal includes a face image sent by a second terminal to the first terminal, and the second terminal is a terminal corresponding to a temporary user; and the registration information further includes door opening permission information corresponding to the uploaded face image.
[0304] In a possible implementation, the first control module 21 is configured to control the image acquisition module arranged outside a cabin of the vehicle to acquire a video stream outside the vehicle.
[0305] In a possible implementation, the first control module 21 is configured to control the image acquisition module arranged inside the cabin of the vehicle to acquire a video stream inside the vehicle.
[0306] In a possible implementation, the first control module 21 is configured to: in a case where the vehicle is at a speed of 0 and there is a person in the vehicle, control an image acquisition module arranged in the interior of the vehicle to acquire a video stream of the vehicle.
[0307] In a possible implementation, the device further includes a second acquisition module configured to acquire information about a seat adjustment of a passenger of the vehicle; and a generation or update module configured to generate or update the seat preference information corresponding to the passenger according to the information about the seat adjustment of the passenger.
[0308] In a possible implementation, the generation or update module is configured to generate or update the seat preference information corresponding to the passenger according to position information of a seat on which the passenger is seated and the information about the seat adjustment of the passenger.
[0309] In a possible implementation, the device further includes a third acquisition module configured to acquire seat preference information corresponding to a passenger of the vehicle based on the face recognition result; and a seat adjustment module configured to adjust a seat on which the passenger is seated according to the seat preference information corresponding to the passenger.
[0310] In a possible implementation, the seat adjustment module is configured to determine position information of a seat on which the passenger is seated; and adjust the seat on which the passenger is seated according to the position information of the seat on which the passenger is seated and the seat preference information corresponding to the passenger.
[0311] In a possible implementation, the device further includes a search module configured to search, by a Bluetooth module arranged in the vehicle, for a Bluetooth device of a preset identifier; an establishment module configured to, in response to searching for the Bluetooth device of the preset identifier, establish a Bluetooth pairing connection between the Bluetooth module and the Bluetooth device of the preset identifier; and a wake-up module configured to, in response to the Bluetooth pairing connection being successful, wake up a face recognition module arranged in the vehicle; and the first control module 21 is configured to control the image acquisition module to acquire the video stream by the face recognition module that is woken up.
[0312] In a possible implementation, the search module is configured to search, by a Bluetooth module arranged in the vehicle, for a Bluetooth device of a preset identifier in a case where the vehicle is in an off state or in an off state and a door lock state.
[0313] In a possible implementation, the device further includes a seventh control module configured to, if a face image is not acquired within a preset time, control the face recognition module to enter a dormant state.
[0314] In a possible implementation, the apparatus further includes a fifth control module configured to control the face recognition module to enter a dormant state if face recognition is not passed within a preset time.
[0315] In a possible implementation, the apparatus further includes a sixth control module configured to control the face recognition module to enter a dormant state if the driving speed of the vehicle is not 0.
[0316] In a possible implementation, the face recognition module 22 is configured to update the first depth map based on the first image to obtain a second depth map, and determine a live body detection result based on the first image and the second depth map.
[0317] In a possible implementation, the face recognition module 22 is configured to update a depth value of a depth failure pixel in the first depth map based on the first image to obtain the second depth map.
[0318] In a possible implementation, the face recognition module 22 is configured to determine depth prediction values and associated information of a plurality of pixels in the first image based on the first image, where the associated information of the plurality of pixels indicates a correlation degree between the plurality of pixels, and update the first depth map based on the depth prediction values and the associated information of the plurality of pixels to obtain a second depth map.
[0319] In a possible implementation, the face recognition module 22 is configured to determine a depth failure pixel in the first depth map, obtain a depth prediction value of the depth failure pixel and depth prediction values of a plurality of surrounding pixels of the depth failure pixel from the depth prediction values of the plurality of pixels, obtain a correlation degree between the depth failure pixel and the plurality of surrounding pixels of the depth failure pixel from the associated information of the plurality of pixels, and determine an updated depth value of the depth failure pixel based on the depth prediction value of the depth failure pixel, the depth prediction values of the plurality of surrounding pixels of the depth failure pixel, and the correlation degree between the depth failure pixel and the surrounding pixels of the depth failure pixel.
[0320] In a possible implementation, the face recognition module 22 is configured to determine a depth correlation value of the depth failure pixel based on the depth prediction values of the surrounding pixels of the depth failure pixel and the correlation degree between the depth failure pixel and the plurality of surrounding pixels of the depth failure pixel, and determine the updated depth value of the depth failure pixel based on the depth prediction value of the depth failure pixel and the depth correlation value.
[0321] In a possible implementation, the face recognition module 22 is configured to: take the correlation degree between the depth failure pixel and each surrounding pixel as a weight of the each surrounding pixel, and perform weighted sum processing on depth prediction values of the surrounding pixels of the depth failure pixel to obtain a depth correlation value of the depth failure pixel.
[0322] In a possible implementation, the face recognition module 22 is configured to: determine, based on the first image and the first depth map, the depth prediction values of the pixels in the first image.
[0323] In a possible implementation, the face recognition module 22 is configured to: input the first image and the first depth map into a depth prediction neural network for processing to obtain the depth prediction values of the pixels in the first image.
[0324] In a possible implementation, the face recognition module 22 is configured to: perform fusion processing on the first image and the first depth map to obtain a fusion result; and determine, based on the fusion result, the depth prediction values of the pixels in the first image.
[0325] In a possible implementation, the face recognition module 22 is configured to: input the first image into a correlation degree detection neural network for processing to obtain the correlation information of the pixels in the first image.
[0326] In a possible implementation, the face recognition module 22 is configured to: obtain an image of a face from the first image; and update the first depth map based on the image of the face.
[0327] In a possible implementation, the face recognition module 22 is configured to: obtain key point information of a face in the first image; and obtain the image of the face from the first image based on the key point information of the face.
[0328] In a possible implementation, the face recognition module 22 is configured to: perform face detection on the first image to obtain a face region; and perform key point detection on an image of the face region to obtain the key point information of the face in the first image.
[0329] In a possible implementation, the face recognition module 22 is configured to: obtain a depth map of a face from the first depth map; and update the depth map of the face based on the first image to obtain the second depth map.
[0330] In a possible implementation, the face recognition module 22 is configured to: input the first image and the second depth map into a living body detection neural network for processing to obtain a living body detection result.
[0331] In one possible implementation, the face recognition module 22 is used to: perform feature extraction processing on the first image to obtain first feature information; perform feature extraction processing on the second depth map to obtain second feature information; and determine the liveness detection result based on the first feature information and the second feature information.
[0332] In one possible implementation, the face recognition module 22 is used to: fuse the first feature information and the second feature information to obtain third feature information; and determine the liveness detection result based on the third feature information.
[0333] In one possible implementation, the face recognition module 22 is used to: obtain the probability that the face is a live object based on the third feature information; and determine the liveness detection result based on the probability that the face is a live object.
[0334] In one possible implementation, the device further includes: a password unlocking module, configured to activate the password unlocking module installed in the vehicle to initiate the password unlocking process in response to the face recognition result being a face recognition failure.
[0335] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0336] Figure 15 A block diagram of a door control system provided in an embodiment of this disclosure is shown. (As follows) Figure 15 As shown, the door control system includes: a memory 41, an object detection module 42, a face recognition module 43, and an image acquisition module 44; the face recognition module 43 is connected to the memory 41, the object detection module 42, and the image acquisition module 44 respectively, and the object detection module 42 is connected to the image acquisition module 44; the face recognition module 43 is also provided with a communication interface for connecting to the door domain controller, and the face recognition module sends control information for unlocking and opening the door to the door domain controller through the communication interface.
[0337] In one possible implementation, the door control system further includes a Bluetooth module 45 connected to the face recognition module 43; the Bluetooth module 45 includes waking up the microprocessor 451 of the face recognition module 43 and the Bluetooth sensor 452 connected to the microprocessor 451 when it successfully pairs and connects with a Bluetooth device with a preset identifier or when the Bluetooth device with the preset identifier is found.
[0338] In a possible implementation, the memory 41 can include at least one of a Flash and a DDR3 (Double Date Rate 3) memory.
[0339] In a possible implementation, the face recognition module 43 can be implemented by using an SoC (System on Chip).
[0340] In a possible implementation, the face recognition module 43 is connected to a door domain controller through a CAN (Controller Area Network) bus.
[0341] In a possible implementation, the image acquisition module 44 includes an image sensor and a depth sensor.
[0342] In a possible implementation, the depth sensor includes at least one of a binocular infrared sensor and a TOF (Time of Flight) sensor.
[0343] In a possible implementation, the depth sensor includes a binocular infrared sensor, and two infrared cameras of the binocular infrared sensor are arranged on two sides of a camera of the image sensor. For example, in the example shown in FIG. 1, the image sensor is an RGB sensor, the camera of the image sensor is an RGB camera, the depth sensor is a binocular infrared sensor, the depth sensor includes two IR (infrared) cameras, and the two infrared cameras of the binocular infrared sensor are arranged on two sides of the RGB camera of the image sensor. Figure 4a
[0344] In a possible implementation, the image acquisition module 44 further includes at least one fill light, the at least one fill light is arranged between an infrared camera of the binocular infrared sensor and a camera of the image sensor, and the at least one fill light includes at least one of a fill light for the image sensor and a fill light for the depth sensor. For example, if the image sensor is an RGB sensor, the fill light for the image sensor can be a white light; if the image sensor is an infrared sensor, the fill light for the image sensor can be an infrared light; and if the depth sensor is a binocular infrared sensor, the fill light for the depth sensor can be an infrared light. In the example shown in FIG. 1, an infrared light is arranged between the infrared camera of the binocular infrared sensor and the camera of the image sensor. For example, the infrared light can use infrared rays of 940 nm. Figure 4a
[0345] In one example, the fill light can be in a constant-on mode. In this example, when the camera of the image acquisition module is in a working state, the fill light is in an on state.
[0346] In another example, the light compensation lamp can be turned on when the light is insufficient. For example, the ambient light intensity can be obtained by the ambient light sensor, and it is determined that the light is insufficient when the ambient light intensity is lower than the light intensity threshold, and the light compensation lamp is turned on.
[0347] In a possible implementation, the image acquisition module 44 further comprises a laser, which is arranged between the camera of the depth sensor and the camera of the image sensor. For example, in the example shown in the figure, the image sensor is an RGB sensor, the camera of the image sensor is an RGB camera, the depth sensor is a TOF sensor, and the laser is arranged between the camera of the TOF sensor and the camera of the RGB sensor. For example, the laser can be a VCSEL, and the TOF sensor can acquire a depth map based on the laser emitted by the VCSEL. Figure 4b
[0348] In one example, the depth sensor is connected with the face recognition module 43 through an LVDS (Low-Voltage Differential Signaling) interface.
[0349] In a possible implementation, the vehicle-mounted face unlocking system further comprises a password unlocking module 46 for unlocking the vehicle door, wherein the password unlocking module 46 is connected with the face recognition module 43.
[0350] In a possible implementation, the password unlocking module 46 comprises one or both of a touch screen and a keyboard.
[0351] In one example, the touch screen is connected with the face recognition module 43 through an FPD-Link (Flat Panel Display Link).
[0352] In a possible implementation, the vehicle-mounted face unlocking system further comprises a battery module 47, wherein the battery module 47 is connected with the face recognition module 43. In one example, the battery module 47 is further connected with the microprocessor 451.
[0353] In a possible implementation, the memory 41, the face recognition module 43, the Bluetooth module 45 and the battery module 47 can be built on an ECU (Electronic Control Unit).
[0354] Figure 16 A schematic diagram of a vehicle door control system according to an embodiment of the present disclosure is shown. In the example shown in the figure, the vehicle door control system comprises a face recognition module 43, a microprocessor 451, a memory 41, a Bluetooth module 45, a light compensation lamp 42, a password unlocking module 46, a battery module 47, and a vehicle door 48. Figure 16 In the shown example, the face recognition module is implemented by the SoC 101, the memory includes a flash memory 102 and a DDR3 memory 103, the Bluetooth module includes a Bluetooth sensor 104 and a microcontroller unit (MCU) 105, the SoC 101, the flash memory 102, the DDR3 memory 103, the Bluetooth sensor 104, the MCU 105 and a power management module 106 are built on the ECU 100, the image acquisition module includes a depth sensor (3D camera) 200, the depth sensor 200 is connected to the SoC 101 through an LVDS interface, the password unlocking module includes a touch screen 300, the touch screen 300 is connected to the SoC 101 through an FPD-Link, and the SoC 101 is connected to a door area controller 400 through a CAN bus.
[0355] Figure 17 A schematic diagram of a vehicle is shown. As shown in the figure, Figure 17 The vehicle includes a door control system 51, and the door control system 51 is connected to a door area controller 52 of the vehicle.
[0356] In a possible implementation, the image acquisition module is arranged on an outdoor part of the vehicle.
[0357] In a possible implementation, the image acquisition module is arranged at at least one of the following positions: a B-pillar of the vehicle, at least one door of the vehicle, and at least one rearview mirror of the vehicle.
[0358] In a possible implementation, the image acquisition module is arranged on an indoor part of the vehicle.
[0359] In a possible implementation, the face recognition module is arranged in the vehicle, and the face recognition module is connected to the door area controller through a CAN bus.
[0360] The embodiment of the present disclosure further provides a computer readable storage medium, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the above method. The computer readable storage medium can be a non-volatile computer readable storage medium or a volatile computer readable storage medium.
[0361] The embodiment of the present disclosure further provides a computer program product, which includes computer readable code, and when the computer readable code is run on a device, a processor in the device executes instructions for implementing the door control method provided in any of the above embodiments.
[0362] The embodiment of the present disclosure further provides another computer program product for storing computer readable instructions, which, when executed, cause a computer to perform the operations of the vehicle door control method provided by any of the above embodiments.
[0363] The embodiment of the present disclosure further provides an electronic device, comprising: one or more processors; a memory for storing executable instructions; wherein the one or more processors are configured to invoke the executable instructions stored in the memory to perform the above method.
[0364] The electronic device can be provided as a terminal, a server or other forms of devices.
[0365] Figure 18 A block diagram of an electronic device 800 is shown. The electronic device 800 can be a terminal such as a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like.
[0366] Reference Figure 18 The electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0367] The processing component 802 usually controls overall operations of the electronic device 800, such as operations associated with displaying, making phone calls, data communications, camera operations and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or part of steps of the above methods. Further, the processing component 802 can include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0368] The memory 804 is configured to store various types of data to support operations of the electronic device 800. Examples of these data include instructions for any application or method operating on the electronic device 800, contact data, phonebook data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0369] The power component 806 provides power to the various components of the electronic device 800. The power component 806 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0370] The multimedia component 808 includes a screen providing an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0371] The audio component 810 is configured to output and / or input an audio signal. For example, the audio component 810 includes a microphone (MIC) configured to receive an external audio signal when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting an audio signal.
[0372] The I / O interface 812 provides an interface between the processing component 802 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0373] The sensor component 814 includes one or more sensors for providing status assessments for various aspects of the electronic device 800. For example, the sensor component 814 can detect an open / closed position of the electronic device 800, relative positioning of components of the electronic device 800, such as a display and a keypad of the electronic device 800, a change in position of the electronic device 800 or a component of the electronic device 800, presence or absence of user contact with the electronic device 800, orientation or acceleration / deceleration / g-force and temperature changes of the electronic device 800. The sensor component 814 can include an optical sensor that is configured to detect ambient light, a proximity sensor configured to detect the presence of nearby objects without any physical touch, or a CMOS or CCD image sensor for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0374] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as Wi-Fi, 2G, 3G, 4G / LTE, 5G, or a combination thereof. In an example embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) techniques, infrared data association (IrDA) techniques, ultra-wideband (UWB) techniques, Bluetooth (BT) techniques, and other techniques.
[0375] In an example embodiment, the electronic device 800 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements to perform the above-described methods.
[0376] In an example embodiment, a non-transitory computer-readable storage medium, such as the memory 804 including computer program instructions, is also provided, which can be executed by the processor 820 of the electronic device 800 to complete the above-described methods.
[0377] Figure 19 A block diagram of an electronic device 1900 is shown, which is provided by an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server. Referring to FIG. 19, the electronic device 1900 includes one or more processors 1910, memory 1920 and a bus 1930. Figure 19The electronic device 1900 includes a processing component 1922, which is further composed of one or more processors, and a memory resource represented by the memory 1932 for storing instructions, such as an application program, executable by the processing component 1922. The application program stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above method.
[0378] The electronic device 1900 can further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows Mac OS or the like.
[0379] In an exemplary embodiment, a non-transitory computer readable storage medium, such as the memory 1932 including computer program instructions, is also provided, which can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0380] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0381] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a
[0382] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0383] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, for example, through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0384] The computer readable program instructions can also be loaded onto a computing / processing device, other programmable data processing apparatus, or other device to cause a series of operations to be performed on the computing / processing device, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computing / processing device, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0385] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0386] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0387] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0388] The computer program product can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0389] Having described above several embodiments of the disclosure, any modifications and variations that fall within the scope of the described embodiments are also intended to be within the scope of the disclosure. As will be apparent to those skilled in the art, some modifications and variations to the embodiments described above can be practiced while staying within the scope and spirit of the described embodiments. The foregoing description of the described embodiments has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the described embodiments to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. It is intended that the disclosed embodiments be limited only by the claims.
Claims
1. A vehicle door control method characterized by, The method comprises: controlling an image acquisition module arranged on a vehicle to acquire a video stream; performing face recognition based on at least one image in the video stream to obtain a face recognition result; determining an intersection-over-union of images of adjacent frames in the video stream, and determining door opening intention information according to the intersection-over-union of the images of the adjacent frames; and / or, determining an area of a human body region in a plurality of newly-acquired images in the video stream, and determining the door opening intention information according to the area of the human body region in the plurality of newly-acquired images; determining control information corresponding to at least one door of the vehicle based on the face recognition result and the door opening intention information; if the control information comprises control information for opening any door of the vehicle, obtaining state information of the door; if the state information of the door is not unlocked, controlling the door to be unlocked and opened; and / or, if the state information of the door is unlocked but not opened, controlling the door to be opened; wherein the determining of the door opening intention information according to the intersection-over-union of the images of the adjacent frames comprises: caching intersection-overs of N groups of adjacent frames of newly-acquired images, wherein N is an integer greater than 1; determining an average value of the cached intersection-overs; if the average value is greater than a first preset value for a duration reaching a first preset time length, determining that the door opening intention information is intentional door opening.
2. The method of claim 1, wherein, The determining of the intersection-over-union of the images of the adjacent frames in the video stream comprises: determining an intersection-over-union of a bounding box of a human body in images of adjacent frames in the video stream as the intersection-over-union of the images of the adjacent frames.
3. The method of claim 1, wherein, The determining of the door opening intention information according to the area of the human body region in the plurality of newly-acquired images comprises: if the area of the human body region in the plurality of newly-acquired images is all greater than a first preset area, determining that the door opening intention information is intentional door opening.
4. The method of claim 1, wherein, The determining of the door opening intention information according to the area of the human body region in the plurality of newly-acquired images comprises: if the area of the human body region in the plurality of newly-acquired images gradually increases, determining that the door opening intention information is intentional door opening.
5. The method according to any one of claims 1 to 4, characterized in that, The determining of the control information corresponding to at least one door of the vehicle based on the face recognition result and the door opening intention information comprises: if the face recognition result is face recognition success and the door opening intention information is intentional door opening, determining that the control information comprises control information for opening at least one door of the vehicle.
6. A vehicle door control device characterized by comprising: The method comprises: a first control module for controlling an image acquisition module arranged on a vehicle to acquire a video stream; a face recognition module for performing face recognition based on at least one image in the video stream to obtain a face recognition result; a second determination module for determining an intersection-over-union of images of adjacent frames in the video stream, and determining door opening intention information according to the intersection-over-union of the images of the adjacent frames, and / or, determining an area of a human body region in a plurality of newly-acquired images in the video stream, and determining the door opening intention information according to the area of the human body region in the plurality of newly-acquired images; a first determination module for determining control information corresponding to at least one door of the vehicle based on the face recognition result and the door opening intention information; The first obtaining module is configured to obtain state information of the door if the control information comprises control of any door of the vehicle to be opened. The second control module is configured to control the door to be unlocked and opened if the state information of the door is that the door is not unlocked. And / or, control the door to be opened if the state information of the door is that the door is unlocked but not opened. The second determining module is further configured to: cache an intersection over union of N groups of adjacent frames of images collected most recently, wherein N is an integer greater than 1; determine an average value of the cached intersection over union; determine that the door opening intention information is intentional door opening if the average value is greater than a first preset value for a duration reaching a first preset time length.
7. A vehicle door control system characterized by comprising: The vehicle comprises the door control system of claim 7, and the door control system is connected with a door domain controller of the vehicle. The vehicle comprises: a memory, an object detection module, a face recognition module, and an image acquisition module; the face recognition module is connected with the memory, the object detection module, and the image acquisition module, and the object detection module is connected with the image acquisition module; the face recognition module is further provided with a communication interface for connection with a door domain controller, and the face recognition module sends control information for unlocking and opening the door to the door domain controller through the communication interface; the image acquisition module is configured to acquire a video stream; the face recognition module is configured to: perform face recognition based on at least one image in the video stream to obtain a face recognition result; determine an intersection over union of adjacent frames of images in the video stream, and determine door opening intention information according to the intersection over union of the adjacent frames of images; determine an area of a human body region in a plurality of frames of images collected most recently, and determine the door opening intention information according to the area of the human body region in the plurality of frames of images collected most recently; and determine control information corresponding to at least one door of the vehicle based on the face recognition result and the door opening intention information. The face recognition module is further configured to: cache an intersection over union of N groups of adjacent frames of images collected most recently, wherein N is an integer greater than 1; determine an average value of the cached intersection over union; 8. A vehicle characterized by comprising: determine that the door opening intention information is intentional door opening if the average value is greater than a first preset value for a duration reaching a first preset time length.
9. An electronic device, comprising: The vehicle comprises the door control system of claim 7, and the door control system is connected with a door domain controller of the vehicle. The vehicle comprises: a processor; a memory for storing processor-executable instructions; 10. A computer-readable storage medium having stored thereon computer program instructions, wherein, wherein the processor is configured to execute the method of any one of claims 1 to 5. The computer program instructions, when executed by the processor, implement the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Vehicle electric tail door opening method based on target detection and motion recognition
CN109882019A
Car door unlocking methods and device, system, car, electronic equipment and memory media
CN110335389A