Automobile safe starting method, system and equipment based on face recognition and medium

By acquiring facial images to generate text information, extracting features, and predicting age values, the safety hazards of traditional car starting methods are solved, realizing safe car starting based on facial recognition, and improving the safety and anti-theft capabilities of vehicle use.

CN121799338APending Publication Date: 2026-04-07DONGFENG MOTOR GRP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing car starting methods pose safety risks. Keys or electronic keys are easily lost or stolen, and traditional facial recognition technology cannot authorize relatives and friends to use the car and does not take into account driving age restrictions, leading to safety hazards.

Method used

By acquiring the facial image of the target object, generating the first text information, extracting image and text features, predicting the age value, and performing facial correction and authorized facial matching to start the vehicle when the driver meets the driving age, the face recognition is optimized using a multimodal CLIP model and knowledge distillation technology to increase the safety of vehicle use.

Benefits of technology

It improves the safety and anti-theft features of car starting by predicting age values ​​through image and text features and performing facial correction matching, thereby increasing the safety and anti-theft features during vehicle use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121799338A_ABST
    Figure CN121799338A_ABST
Patent Text Reader

Abstract

The invention provides an automobile safe starting method, system and device based on face recognition and a medium, and belongs to the technical field of intelligent automobiles. The method comprises the steps that a face image of a target object is acquired, and first text information is generated based on the face image of the target object; extracting image features in the face image and text features in the first text information, aligning the image features and the text features, and predicting an age value of the target object; and when the age value of the target object accords with the driving age, performing face correction on the face image of the target object, matching the corrected face image with an authorized face image, and starting the vehicle when the matching is successful. Through the technical scheme in the embodiment of the invention, whether the target object conforms to the driving age or not can be accurately predicted in combination with the image features and the text features, whether the target object is an authorized object or not is judged, and the safety and the anti-theft performance in the vehicle using process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent automobiles, and in particular to a vehicle safety starting method and system based on face recognition, a device, and a medium. BACKGROUND

[0002] With the large-scale development of the automobile industry, the number of vehicles has increased significantly, and automobile safety has been increasingly concerned.

[0003] Currently, the starting mode of a vehicle mainly relies on a key or an electronic key, but the starting mode of the vehicle mainly using the key or the electronic key has certain safety hazards. For example, the key can be stolen or lost, resulting in the vehicle being stolen. The traditional face recognition technology only acts on the vehicle owner, and the vehicle still needs to be started by the key when a non-vehicle owner starts the vehicle.

[0004] Therefore, it is urgent to upgrade the existing face recognition system to improve the safety of the vehicle and authorize the relatives and friends meeting certain conditions to use it. SUMMARY

[0005] The present application aims to solve at least one of the technical problems in the prior art, and provides a vehicle safety starting method and system based on face recognition, a device, and a storage medium.

[0006] In a first aspect, an embodiment of the present application provides a vehicle safety starting method based on face recognition, comprising:

[0007] obtaining a face image of a target object, and generating first text information based on the face image of the target object;

[0008] extracting image features in the face image and text features in the first text information, aligning the image features and the text features, and predicting an age value of the target object;

[0009] when the age value of the target object meets a driving age, performing face correction on the face image of the target object, matching the corrected face image with an authorized face image, and starting the vehicle when the matching is successful.

[0010] In some embodiments, the first text information is generated based on the face image of the target object, comprising:

[0011] determining all first-type facial key points in the face image;

[0012] recording each unobstructed first-type facial key point and coordinate position in the face image;

[0013] extracting image features in the face image and text features in the first text information, aligning the image features and the text features, and predicting an age value of the target object;generate an age-related semantic hint based on the information of each un-occluded first-type facial key point and the information of each occluded first-type facial key point;

[0014] generate the first text information based on the age-related semantic hint and the coordinates of each un-occluded first-type facial key point.

[0015] In some embodiments, the process of predicting the age value of the target object comprises:

[0016] constructing a multi-modal CLIP model comprising a key point regression branch and an age regression branch;

[0017] aligning the image feature and the text feature, and obtaining second-type facial key points and coordinates through the key point regression branch;

[0018] predicting the age value of the target object through the age regression branch after obtaining the second-type facial key points and coordinates.

[0019] In some embodiments, after obtaining the second-type facial key points and coordinates, the process further comprises:

[0020] generating second text information based on the second-type facial key points and coordinates;

[0021] calculating the similarity between the first text information and the second text information;

[0022] obtaining the confidence of each second-type facial key point and coordinate according to the similarity between the first text information and the second text information.

[0023] In some embodiments, the process further comprises: when resources are limited, lightening the multi-modal CLIP model through knowledge distillation.

[0024] In some embodiments, before performing face correction on the face image of the target object, the process further comprises:

[0025] generating a second face image based on the second-type facial key points and coordinates;

[0026] determining whether the second face image is a valid face image according to the confidence of each second-type facial key point and coordinate;

[0027] if the second face image is determined to be a valid face image, covering the face image of the target object with the second face image and performing face correction on the second face image;

[0028] If it is determined that the second face image is not a valid face image, face correction of the face image of the target object is terminated, a next frame of the face image of the target object is acquired, the first text information is regenerated for face recognition.

[0029] In some embodiments, before face correction of the face image of the target object, the method further includes:

[0030] extracting a plurality of preselected key points from the second type of face key points;

[0031] projecting the plurality of preselected key points into a three-dimensional space;

[0032] calculating a head deflection posture angle of the target object based on a positional relationship of the plurality of preselected key points in the three-dimensional space;

[0033] If the head deflection posture angle of the target object meets a pre-designated range, face correction of the face image of the target object is performed.

[0034] If the head deflection posture angle of the target object does not meet the pre-designated range, face correction of the face image of the target object is terminated, a next frame of the face image of the target object is acquired, and the first text information is regenerated for face recognition.

[0035] In a second aspect, an embodiment of the present application provides an automobile safety starting system based on face recognition, which includes:

[0036] an image acquisition unit configured to acquire a face image of a target object and generate first text information based on the face image of the target object;

[0037] a feature extraction unit configured to extract image features in the face image and text features in the first text information, align the image features and the text features, and predict an age value of the target object;

[0038] a face correction unit configured to perform face correction of the face image of the target object when the age value of the target object meets a driving age, match the corrected face image with an authorized face image, and start a vehicle when the matching is successful.

[0039] In a third aspect, an embodiment of the present application provides an electronic device, which includes:

[0040] at least one processor; and a memory connected to the at least one processor in communication;

[0041] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the steps of the method of any of the embodiments of the application.

[0042] In a fourth aspect, an embodiment of the application provides a computer readable storage medium storing computer instructions for enabling a processor to perform the steps of the method of any of the embodiments of the application when the computer instructions are executed by the processor.

[0043] Compared with the prior art, the application has the following beneficial effects:

[0044] The face recognition-based automobile safety starting method provided by the application first acquires a face image of a target object, and generates first text information based on the face image of the target object; then extracts image features in the face image and text features in the first text information, aligns the image features and the text features, and predicts an age value of the target object; finally, when the age value of the target object meets a driving age, the face image of the target object is corrected, the corrected face image is matched with an authorized face image, and the vehicle is started when the matching is successful. Through the technical solution of the application, the age value of the target object can be predicted through image and text features, and the face image is matched with an authorized user through correction, thereby increasing the safety and anti-theft property of the vehicle in use. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only preferred embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0046] Figure 1 A flowchart of the face recognition-based automobile safety starting method provided by the embodiment of the application is shown in the figure.

[0047] Figure 2 A flowchart of the method for generating first text information provided by the embodiment of the application is shown in the figure.

[0048] Figure 3 A flowchart of another face recognition-based automobile safety starting method provided by the embodiment of the application is shown in the figure.

[0049] Figure 4 A structure block diagram of a face recognition-based automobile safety starting system provided by the embodiment of the application is shown in the figure.

[0050] Figure 5 A structure block diagram of an electronic device provided by the embodiment of the application is shown in the figure. Detailed Implementation

[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] To enable those skilled in the art to better understand the technical solutions of the present invention, exemplary embodiments of the present invention are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0053] Where there is no conflict, the various embodiments of the present invention and the features thereof may be combined with each other.

[0054] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0055] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Terms such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0056] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having the meaning consistent with their meaning in the context of the relevant art and the invention, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.

[0057] In the technical solution of this invention, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information all comply with relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution follows relevant national laws and regulations (e.g., the "Information Security Technology - Personal Information Security Specification"). For example: appropriate measures are taken for personal information access control; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; and explicit identity targeting is eliminated when using personal information to avoid precisely locating a specific individual.

[0058] Before introducing the technical solution of this invention, it should be noted that the following face verification schemes exist in related technologies:

[0059] 1) Option 1: Capture the facial features of people entering the vehicle, calculate the similarity between the person's facial features and those of the vehicle owner, and then determine whether the person entering the vehicle is the owner. However, this option does not consider facial recognition for authorized users, which would prevent relatives or family members from starting the vehicle. It also does not consider age restrictions for authorized drivers, potentially resulting in unqualified individuals appearing on the authorized driver list, leading to improper vehicle start-up and potential driving safety hazards.

[0060] 2) Option Two: Obtain facial images and facial landmarks of the occupants inside the vehicle, align them, and perform detection to obtain multiple gender and age detection values, thereby determining the occupants' gender and age. However, this option is susceptible to the effects of lighting, occlusion, and pose, causing landmark extraction and alignment to fail, affecting detection accuracy. Furthermore, errors in landmark detection in this option propagate to the alignment stage, leading to cropping region shifts (e.g., misalignment of the eyes or mouth), thus reducing the performance of the detection model. For example, wearing glasses can cause deviations in the localization of the nose bridge landmark, resulting in misalignment of the aligned facial region and excessively inaccurate age estimation results.

[0061] To address at least one of the technical problems existing in the aforementioned related technologies, the present invention provides a car safe starting method based on face recognition. Figure 1 This is a flowchart illustrating a car safety start method based on face recognition, provided by an embodiment of the present invention. This method is particularly suitable for situations where a car owner authorizes other drivers to use the car. The method can be executed by a car safety start system based on face recognition, which can be implemented in software and / or hardware and can be configured in an electronic device.

[0062] like Figure 1 As shown, the method specifically includes:

[0063] S1, acquire a face image of a target object, and generate first text information based on the face image of the target object.

[0064] In the above process, the target object is a driver who is currently performing face recognition, the face image can be acquired by a camera located at an A-pillar, and the size of the face image is standardized to 224x224. The first text information refers to a natural language description recording all key point information in the face image.

[0065] Figure 2 A flowchart of a method for generating first text information according to an embodiment of the present application is shown in FIG. 1. The method includes the following steps. Figure 2

[0066] S1011, determine all first-type facial key points in the face image.

[0067] For example, the eye, mouth, nose, and other parts of the face are all first-type facial key points. It should be noted that there can be a situation of occlusion in the facial key points in the face image, for example, the target object wears sunglasses or a mask, and in this case, the facial key points about the eye or about the mouth and nose in the face image are occluded.

[0068] It should be noted that the first-type facial key points serve as prior knowledge, and their role is to guide the model to pay attention to the area where the second-type facial key points are located in the subsequent process. The second-type facial key points are one of the results output by the model and are the predicted values that need to be optimized.

[0069] The number of first-type facial key points cannot be set too high, which will cause the complexity of the recognition process to exceed the limit, nor can it be set too low, which will cause the accuracy of the recognition process to be substandard.

[0070] In a specific application scenario, 19 first-type facial key points that are representative in the face image are selected.

[0071] S1012, record each unoccluded first-type facial key point and coordinate position in the face image.

[0072] For example, in a certain face image, the unoccluded first-type facial key points include the eye, nose, and the like, the coordinates of the center of the left eye are recorded as (45, 78), the coordinates of the center of the right eye are recorded as (162, 80), the coordinates of the tip of the nose are recorded as (103, 145), and so on.

[0073] It should be noted that the coordinate system and the coordinate origin in the face image can be randomly set by humans.

[0074] ​Furthermore, for the first type of facial key points that are occluded in the face image, only the first type of facial key points need to be recorded, without recording the coordinate position of the first type of facial key points in the face image.

[0075] S1013, combine the information of each unobstructed first-class facial key point and the information of each obstructed first-class facial key point to generate age-related semantic prompts.

[0076] For example, an age-related semantic prompt could be generated such as "obvious wrinkles around the eyes, deep nasolabial folds, consistent with characteristics of a 50-year-old woman." Alternatively, the model could be guided to focus on the visible area in cases of occlusion, generating an age-related semantic prompt such as "20-year-old male wearing sunglasses, key points of both eyes are occluded."

[0077] S1014, combine each unobstructed first-type facial key point and its coordinate position with age-related semantic prompts to generate first text information.

[0078] For example, the first text information could be such as "obvious wrinkles around the eyes, deep nasolabial folds, consistent with the characteristics of a 50-year-old woman, facial key points: left eye center located at (45,78), right eye center located at (162,80), nose tip located at (103,145)...". Or, the first text information could be such as "20-year-old male wearing sunglasses, left and right eye key points are obscured, facial key points: nose tip located at (103,140)...".

[0079] S2, extract image features from the face image and text features from the first text information, align the image features and text features, and predict the age value of the target object.

[0080] Understandably, the first text information contains age-related semantic cues (such as "obvious wrinkles around the eyes, deep nasolabial folds, consistent with characteristics of a 50-year-old woman"), along with the coordinates of unobstructed first-type facial key points. The coordinates of these unobstructed first-type facial key points are extracted from the first text information and used as text features to highlight age-sensitive areas such as the eye area and corners of the mouth. After multiplying the text features with the image features, the age regression branch focuses on details in age-sensitive areas and related regions. Comparative learning of age between text and image allows the model to learn the alignment of key regions with the text; for example, focusing on the eye area combined with wrinkle descriptions increases the probability of predicting a higher age for the target object.

[0081] In some embodiments, the process of predicting the age of a target object includes:

[0082] S2011, construct a multimodal CLIP model, which includes a keypoint regression branch and an age regression branch.

[0083] S2012, after aligning image features and text features, obtains the second type of facial key points and their coordinate positions through the key point regression branch.

[0084] S2013, after obtaining the second type of facial key points and coordinate positions, predicts the age value of the target object through the age regression branch.

[0085] By fusing image and text features and inputting them into a multimodal CLIP model, joint features are generated. Simultaneously, textual cues guide the model to ignore missing keypoints and focus on age features in unoccluded areas, providing a more flexible and reliable solution for age detection in complex in-vehicle environments.

[0086] For example, in the multimodal CLIP model, ViT extracts image features, Transformer processes the first text information to generate text features, and contrastive learning loss is used to align the image features and text features.

[0087] Furthermore, after obtaining the second type of facial key points and their coordinates, it also includes:

[0088] S20131, Based on the second type of facial key points and coordinate positions, reverse the generation of second text information.

[0089] S20132, Calculate the similarity between the first text information and the second text information.

[0090] S20133, Calculate the confidence level of each second-class facial key point and coordinate position based on the similarity between the first text information and the second text information.

[0091] Furthermore, the loss function in the multimodal CLIP model consists of the loss from image feature and text feature alignment, the loss from the keypoint regression branch, and the loss from the age regression branch. The regularization term in the multimodal CLIP model is composed of the similarity between the first and second text information.

[0092] It is understandable that the smaller the loss function value and the larger the regularization term value, the more accurate the prediction results of the multimodal CLIP model.

[0093] In some embodiments, when resources are limited, the multimodal CLIP model is lightweighted through knowledge distillation.

[0094] In a specific application scenario, image features are extracted using ViT in the CLIP dual-editor. A Transformer processes the first text information to generate text features. The coordinates of unoccluded first-type facial keypoints are then parsed from the text features to generate spatial attention, which is multiplied with the image features to highlight the region containing second-type facial keypoints. Contrastive learning loss is used to align image-text features, and keypoint regression and age regression branches are added after the CLIP image encoder ViT. Lightweight convolutional layers predict keypoint coordinates, and fully connected layers predict age values. The loss function consists of contrastive loss, age regression loss, and keypoint regression loss. Training is divided into two phases: a pre-training phase using contrastive learning loss to align image-text features, and a fine-tuning phase fixing the CLIP backbone to train age estimation and keypoint detection tasks. Text-guided keypoint correction is implemented, generating second-type text information based on the predicted second-type facial keypoints and their coordinates. The similarity between the first and second text information is calculated as a regularization term.

[0095] Since this function needs to be implemented on resource-constrained edge devices, knowledge distillation can be used to transfer the trained multimodal CLIP model as a teacher model to a student model using a lightweight model. The ViT in CLIP is replaced with a lightweight convolutional neural network, simplifying the number of Transformer layers and attention heads in the text encoder. The lightweight multimodal model is then trained and outputs the second type of facial key points and their coordinates, confidence scores, and age prediction results for 19 driver face images, including eyes, mouth, nose tip, and nose corner points.

[0096] It should be noted that knowledge distillation is a model compression technique. Its core idea is to transfer the knowledge learned by a powerful but complex teacher model to a simpler and more efficient student model.

[0097] The teacher model refers to the well-trained multimodal CLIP model (with a large number of parameters and high accuracy), while the student model refers to the lightweight model (replacing ViT with a lightweight CNN and simplifying the text encoder Transformer).

[0098] By using knowledge distillation, student models can be made as close as possible to the performance of teacher models while remaining lightweight.

[0099] S3, when the age of the target object meets the driving age, performs facial correction on the target object's facial image, matches the corrected facial image with the authorized facial image, and starts the vehicle when the match is successful.

[0100] In some embodiments, before performing face correction on the facial image of the target object, the method further includes:

[0101] S3011, Generate a second face image based on the second type of facial key points and coordinate positions.

[0102] S3012, based on the confidence level of each second type of facial key point and its coordinate position, determine whether the second face image is a valid face image.

[0103] S3013, if the second face image is determined to be a valid face image, then the second face image is used to cover the face image of the target object, and face correction is performed on the second face image.

[0104] S3014, if it is determined that the second face image is not a valid face image, then the face correction of the target object's face image is terminated, and the next frame of the target object's face image is obtained, and the first text information is regenerated for face recognition.

[0105] In some embodiments, before performing face correction on the facial image of the target object, the method further includes:

[0106] S3021, Extract multiple pre-selected key points from the second type of facial key points.

[0107] S3022 projects multiple pre-selected key points into three-dimensional space.

[0108] S3023, calculates the head deflection angle of the target object based on the positional relationship of multiple pre-selected key points in three-dimensional space.

[0109] S3024, If the head tilt angle of the target object meets the pre-calibrated range, then perform face correction on the face image of the target object.

[0110] S3025, if the head deflection angle of the target object does not conform to the pre-calibrated range, then the face correction of the target object's face image is terminated, and the next frame of the target object's face image is obtained, and the first text information is regenerated for face recognition.

[0111] It should be noted that the process of determining the head tilt angle can occur after the confidence level determination process.

[0112] Figure 3 A flowchart illustrating another car safe start method based on face recognition provided in an embodiment of the present invention is shown below. Figure 3As shown, in a specific application scenario, the image from a camera located on the A-pillar is captured and standardized to 224×224 pixels. The preprocessed image is then input into a distilled lightweight model. The model's inference results are decoded to obtain the keypoint coordinates, confidence level, and age prediction of the driver in the image. The age prediction result determines whether to proceed to the next step. If the current driver is under 18 years old, the vehicle cannot be started; if the current driver is over 18 years old, the subsequent processing logic is executed.

[0113] Based on the second type of facial key points, their coordinates, and confidence levels, the bounding box coordinates of the complete second face image are obtained. The confidence level of the key points is used to determine whether the second face image in this frame is a valid face image. If the confidence level of the second type of facial key points is large or lower than the preset value, the image frame is skipped and the next frame is processed directly.

[0114] Nine pre-defined key points from the second type of facial key points in the face image are selected. The rotation matrix from the camera to the 3D coordinate system is calculated, and the rotation matrix is ​​decomposed to obtain the head yaw (heading angle, the horizontal left-right rotation angle of the head) and pitch (pitch angle, the vertical up-down rotation angle of the head). The driver's head posture is judged based on the yaw and pitch angles to determine if they meet the requirements. If the yaw and pitch angles meet the pre-calibrated range (v1≤yaw≤v2 and v3≤pitch≤v4), the driver's head posture in the current frame is considered to meet the requirements. Otherwise, the image frame is skipped, and the next frame is processed directly. If the driver's head posture does not meet the requirements, face correction is performed on the face image based on the second type of facial key points and the bounding box coordinates of the complete face information, resulting in a corrected face image, which is then feature-encoded.

[0115] In some embodiments, vehicle owners can log in to their accounts and register their faces via a mobile app, and authorize their relatives and friends to log in. The facial information of the vehicle owner and other registered authorized vehicle users is stored in the cloud. The vehicle system and the cloud communicate to update the facial information of authorized vehicle users in real time, and then save this information to the vehicle's infotainment system. When the predicted driver's age is greater than 18, the obtained feature code is compared with the facial information of the authorized vehicle user in the vehicle's infotainment system to determine if the current driver is an authorized vehicle user. If so, the vehicle is activated; otherwise, activation is not performed.

[0116] By adding multiple authorized user accounts for vehicles and a driver age detection model, and using the result of the driver age detection model as the activation condition for the facial recognition module, the system will perform age detection on the driver each time they get in the vehicle. If the driver meets the usage requirements, the system will then perform facial recognition on the authorized user of the vehicle, thereby increasing the security and anti-theft measures during vehicle use.

[0117] The technical solution in this embodiment focuses facial key points on the unobstructed visible area using first text information, and then re-obstructs new second-type facial key points within that visible area. On one hand, the age of the target object is predicted based on these second-type facial key points; on the other hand, second text information is generated in reverse based on these second-type facial key points. By calculating the similarity between the text information, the accuracy of the second-type facial key points can be verified. Furthermore, a complete second face image is generated based on these second-type facial key points. Then, based on the confidence level of each second-type facial key point, the completed face image is validated for effectiveness and head tilt angle.

[0118] In addition, this solution proposes a multimodal CLIP model, and can perform lightweight processing of the multimodal CLIP model when resources are limited, so as to predict the age of the target object through the lightweight model.

[0119] Furthermore, this solution can improve driving safety by adding driver age detection and vehicle authorization user facial recognition technology to the current car startup process through software algorithm optimization and post-processing strategies, without incurring additional costs.

[0120] Based on the same inventive concept, this invention also provides a car safety start system based on facial recognition. Figure 4 A structural block diagram of a car safety start system based on face recognition provided in an embodiment of the present invention is shown below. Figure 4 As shown, the system specifically includes:

[0121] The image acquisition unit 100 is used to acquire a face image of a target object and generate first text information based on the face image of the target object;

[0122] The feature extraction unit 200 is used to extract image features from the face image and text features from the first text information, align the image features and text features, and predict the age value of the target object;

[0123] The face correction unit 300 is used to correct the face image of the target object when the target object's age value meets the driving age, match the corrected face image with the authorized face image, and start the vehicle when the match is successful.

[0124] Based on the same inventive concept, embodiments of the present invention also provide an electronic device.Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Figure 5 As shown, an embodiment of the present invention provides an electronic device including: one or more processors 101, a memory 102, and one or more I / O interfaces 103. The memory 102 stores one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement any of the face recognition-based car safety start methods described in the above embodiments; the one or more I / O interfaces 103 are connected between the processor and the memory, configured to enable information interaction between the processor and the memory.

[0125] The processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 102 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read / write interface) 103 is connected between the processor 101 and the memory 102, and can realize information interaction between the processor 101 and the memory 102, including but not limited to a data bus (BUS).

[0126] In some embodiments, the processor 101, memory 102, and I / O interface 103 are interconnected via bus 104, and thus connected to other components of the computing device.

[0127] In some embodiments, the one or more processors 101 include a field-programmable gate array.

[0128] This invention also provides a computer-readable medium. The computer-readable medium stores a computer program, which, when executed by a processor, implements the steps of any of the face recognition-based car secure start methods described in the above embodiments. The computer-readable storage medium can be volatile or non-volatile.

[0129] This invention also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the above-described car safety start method based on face recognition.

[0130] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).

[0131] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0132] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0133] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.

[0134] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0135] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0136] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0137] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0139] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.

Claims

1. A method for safe car starting based on facial recognition, characterized in that, include: Acquire a facial image of the target object, and generate first text information based on the facial image of the target object; Extract image features from the face image and text features from the first text information, align the image features and the text features, and predict the age value of the target object; When the age of the target object meets the driving age requirement, the facial image of the target object is corrected, the corrected facial image is matched with the authorized facial image, and the vehicle is started when the match is successful.

2. The method according to claim 1, characterized in that, Generate first text information based on the facial image of the target object, including: Identify all first-class facial key points in the face image; Record the first type of facial key points and their coordinate positions that are not occluded in the face image; By combining the information of each unobstructed first-class facial key point and the information of each obstructed first-class facial key point, age-related semantic prompts are generated. The first text information is generated by combining the first type of facial key points and coordinate positions that are not obscured with the age-related semantic prompts.

3. The method according to claim 1, characterized in that, The process of predicting the age of the target object includes: Construct a multimodal CLIP model, which includes a keypoint regression branch and an age regression branch; After aligning the image features and the text features, the second type of facial key points and their coordinate positions are obtained through the key point regression branch; After obtaining the second type of facial key points and their coordinate positions, the age value of the target object is predicted through the age regression branch.

4. The method according to claim 3, characterized in that, After obtaining the second type of facial key points and their coordinates, the process also includes: Based on the second type of facial key points and coordinate positions, the second text information is generated in reverse; Calculate the similarity between the first text information and the second text information; Based on the similarity between the first text information and the second text information, the confidence level of each second type of facial key point and its coordinate position is obtained.

5. The method according to claim 3, characterized in that, Also includes: When resources are limited, the multimodal CLIP model is lightweighted through knowledge distillation.

6. The method according to claim 3, characterized in that, Before performing facial correction on the target object's facial image, the following steps are also included: A second face image is generated based on the second type of facial key points and their coordinate positions; Based on the confidence level of each second type of facial key point and its coordinate position, determine whether the second face image is a valid face image; If the second face image is determined to be a valid face image, then the second face image is overlaid on the face image of the target object, and face correction is performed on the second face image; If the second face image is determined to be an invalid face image, then the face correction of the target object's face image is terminated, and the next frame of the target object's face image is obtained, and the first text information is regenerated for face recognition.

7. The method according to claim 3, characterized in that, Before performing facial correction on the target object's facial image, the following steps are also included: Extract multiple pre-selected key points from the second type of facial key points; Project the pre-selected key points into three-dimensional space; Based on the positional relationship of the pre-selected multiple key points in the three-dimensional space, the head deflection posture angle of the target object is calculated; If the head deflection angle of the target object meets the pre-calibrated range, then face correction is performed on the face image of the target object; If the head deflection angle of the target object does not conform to the pre-calibrated range, then the face correction of the target object's face image is terminated, and the next frame of the target object's face image is obtained, and the first text information is regenerated for face recognition.

8. A car safety start system based on facial recognition, characterized in that, The system is configured to implement the method according to any one of claims 1-7, the system comprising: An image acquisition unit is used to acquire a facial image of a target object and generate first text information based on the facial image of the target object; A feature extraction unit is used to extract image features from the face image and text features from the first text information, align the image features and the text features, and predict the age value of the target object. The face correction unit is used to correct the face image of the target object when the target object's age value meets the driving age requirement, match the corrected face image with the authorized face image, and start the vehicle when the match is successful.

9. An electronic device, characterized in that, The electronic device includes: At least one processor, and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to perform the steps of the method according to any one of claims 1-7.