Information binding method and device
By binding the passenger's human body, face and hand information in the passenger monitoring system, the inconvenience problem of the need to set up detection solutions for different monitoring locations in the prior art is solved, and more efficient monitoring and the effect of saving development resources is achieved.
Patent Information
- Application Number
- CN202311616671.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
When monitoring passengers in the cabin, existing passenger monitoring technologies usually only focus on specific information, which leads to developers needing to set up corresponding detection solutions for different monitoring locations, which is very inconvenient.
Provide an information binding method, by obtaining target images, detecting passengers' human body information, face information and hand information, and binding these information in combination with preset matching conditions to obtain binding results.
All passengers' detection boxes can be obtained through one inspection, avoiding the need to set corresponding detection solutions for different scenarios, and saving development costs and energy.
Smart Images

Figure CN120071306A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an information binding method and apparatus. Background Art
[0002] With the development of intelligent vehicles, more and more intelligent vehicles monitor passengers in the vehicle cabin based on an Occupancy Monitoring System (OMS) inside the vehicle, so as to further improve the safety performance of the vehicle by monitoring the perception data of the passengers in the cabin.
[0003] Currently, in the field of passenger monitoring, passengers in the vehicle cabin are usually identified and monitored through visual perception. However, in the prior art, when monitoring passengers in the vehicle cabin, most attention is paid only to specific information. For example, in a gesture interaction scenario, only the key points of the passengers' hands are monitored; in a face recognition scenario, only the key points of the passengers' faces are monitored. Therefore, when designing the solution in the prior art, only a corresponding solution is designed for the monitored part, resulting in the need for developers to set corresponding detection solutions for different parts, which is very inconvenient. Summary of the Invention
[0004] In view of this, embodiments of the present application provide an information binding method and apparatus, which are used to save development costs and the energy of developers.
[0005] In a first aspect, embodiments of the present application provide an information binding method, including:
[0006] Obtain a target image; the target image includes multiple passengers in the vehicle cabin;
[0007] Detect the target image to obtain the body information, face information, and hand information of the multiple passengers;
[0008] Obtain the identifier of the body information corresponding to the multiple passengers;
[0009] Combine preset matching conditions to obtain the face information and hand information that match the body information, and bind the matching body information, face information, and hand information to obtain a binding result; the preset matching conditions are used to obtain the face information and hand information that match the body information;
[0010] Perform identification processing on the face information and hand information in the binding result according to the identifier of the body information, so as to obtain a target binding result.
[0011] As an optional implementation manner of the embodiment of the present application, the human body information includes a human body detection frame and a sequence of human body key points, the face information includes a face detection frame and a sequence of face key points, and the human hand information includes a human hand detection frame and a sequence of human hand key points;
[0012] The step of obtaining the face information and the human hand information that match the human body information according to the preset matching conditions and binding the matched human body information, face information, and human hand information to obtain a binding result includes:
[0013] Obtain the face detection frame, the sequence of face key points, the human hand detection frame, and the sequence of human hand key points that match the human body detection frame and the sequence of human body key points according to the preset matching conditions, and bind the matched human body detection frame, the sequence of human body key points, the face detection frame, the sequence of face key points, the human hand detection frame, and the sequence of human hand key points to obtain a binding result.
[0014] As an optional implementation manner of the embodiment of the present application, the step of obtaining the identifiers of the human body information corresponding to the multiple passengers respectively includes:
[0015] Based on the time sequence tracking algorithm, match the human body detection frames corresponding to the multiple passengers with multiple historical human body detection frames;
[0016] For each of the multiple passengers, obtain the historical human body detection frame corresponding to the human body detection frame; the historical human body detection frame is the human body detection frame corresponding to the passenger in the historical image frame; the historical human body detection frame carries a corresponding identifier;
[0017] Based on the identifiers carried by the historical human body detection frames corresponding to the human body detection frames of the multiple passengers respectively, obtain the identifiers of the human body information corresponding to the multiple passengers in the target image. As an optional implementation manner of the embodiment of the present application, the preset matching conditions include a first preset distance and a second preset distance; the step of obtaining the face information and the human hand information that match the human body information according to the preset matching conditions and binding the matched human body information, face information, and human hand information to obtain a binding result includes:
[0018] For each passenger, based on the position information of the multiple sequences of face key points in the target image, obtain a target sequence of face key points that meets the first preset distance, and based on the position information of the multiple sequences of human hand key points in the target image, obtain a target sequence of human hand key points that meets the second preset distance;
[0019] Based on the position information of the target human face key point sequence and the target human hand key point sequence, obtain a target human face detection frame and a target human hand detection frame corresponding to the passenger;
[0020] Bind the target human face key point sequence, the target human hand key point sequence, the target human face detection frame, the target human hand detection frame, and the target human hand detection frame with the human body detection frame and the human body key point sequence to obtain a binding result.
[0021] As an optional implementation manner of an embodiment of the present application, the hand information in the passenger information includes: left hand information and right hand information. The left hand information includes a left hand detection frame corresponding to the passenger and a left hand key point sequence; the right hand information includes a right hand detection frame corresponding to the passenger and a right hand key point sequence;
[0022] After obtaining the target binding result, the method further includes:
[0023] According to the identifier of the human body information corresponding to the passenger, obtain the hand information corresponding to the passenger;
[0024] According to the left hand information and the right hand information in the hand information, determine the gestures corresponding to the left hand and the right hand of the passenger respectively, so as to obtain corresponding instructions according to the gestures of the passenger.
[0025] As an optional implementation manner of an embodiment of the present application, after obtaining the target binding result, the method further includes:
[0026] When performing passenger occupancy detection on the vehicle cabin, according to the target binding result, based on the identifier corresponding to the passenger's human body
[0027] information, obtain the human body detection frames corresponding to each passenger;
[0028] According to the detection frames corresponding to each passenger, determine the occupancy corresponding to each passenger;
[0029] When performing face orientation detection in the vehicle cabin, according to the target binding result, based on the identifier corresponding to the passenger's human body information, obtain the human face detection frames and the human face key point sequences corresponding to each passenger;
[0030] Analyze the human face detection frames and the human face key point sequences corresponding to each passenger to determine the face orientations corresponding to each passenger.
[0031] In a second aspect, an embodiment of the present application provides an information binding device, including:
[0032] An image acquisition unit, configured to acquire a target image; the target image includes multiple passengers in the vehicle cabin;
[0033] A detection unit for detecting the target image to obtain the body information, face information, and hand information of the multiple passengers;
[0034] An identification acquisition unit for acquiring the identifications of the body information corresponding to the multiple passengers;
[0035] A binding unit for obtaining the face information and hand information that match the body information in combination with preset matching conditions, and binding the matching body information, face information, and hand information to obtain a binding result; the preset matching conditions are used to obtain the face information and hand information that match the body information;
[0036] A processing unit for performing identification processing on the face information and hand information in the binding result according to the identification of the body information to obtain a target binding result.
[0037] As an optional implementation manner of an embodiment of the present application, the body information includes a body detection frame and a body key point sequence, the face information includes a face detection frame and a face key point sequence, and the hand information includes a hand detection frame and a hand key point sequence; the binding unit is specifically configured to obtain the face detection frame, the face key point sequence, the hand detection frame, and the hand key point sequence that match the body detection frame and the body key point sequence in combination with preset matching conditions, and bind the matching body detection frame, body key point sequence, face detection frame, face key point sequence, hand detection frame, and hand key point sequence to obtain a binding result.
[0038] As an alternative implementation manner of an embodiment of the present application, the identification acquisition unit is specifically configured to match the human detection frames corresponding to the multiple passengers with multiple historical human detection frames based on a timing tracking algorithm; for each of the multiple passengers, obtain the historical human detection frame corresponding to the human detection frame; the historical human detection frame is the human detection frame corresponding to a passenger in a historical image frame; the historical human detection frame carries a corresponding identification; based on the identifications carried by the historical human detection frames respectively corresponding to the human detection frames of the multiple passengers, obtain the identifications of the human information corresponding to the multiple passengers in the target image. As an alternative implementation manner of an embodiment of the present application, the preset matching condition includes a first preset distance and a second preset distance; the binding unit is specifically configured to, for each passenger, obtain a target face key point sequence that meets the first preset distance based on the position information of multiple face key point sequences in the target image, and obtain a target hand key point sequence that meets the second preset distance based on the position information of multiple hand key point sequences in the target image; based on the position information of the target face key point sequence and the target hand key point sequence, obtain a target face detection frame and a target hand detection frame corresponding to the passenger; bind the target face key point sequence, the target hand key point sequence, the target face detection frame, the target hand detection frame, and the target hand detection frame with the human detection frame and the human key point sequence to obtain a binding result.
[0039] As an alternative implementation manner of an embodiment of the present application, the hand information in the passenger information includes: left hand information and right hand information, the left hand information includes a left hand detection frame corresponding to the passenger and a left hand key point sequence; the right hand information includes a right hand detection frame corresponding to the passenger and a right hand key point sequence; after obtaining the target binding result, the processing unit is further configured to obtain the hand information corresponding to the passenger according to the identification of the human information corresponding to the passenger; determine the gestures corresponding to the left hand and the right hand of the passenger according to the left hand information and the right hand information in the hand information, so as to obtain corresponding instructions according to the gestures of the passenger.
[0040] As an alternative implementation manner of an embodiment of the present application, when performing passenger occupancy detection on a vehicle cabin, according to the target binding result, based on the identification corresponding to the passenger, obtain the human detection frame corresponding to each passenger; according to the detection frames corresponding to each passenger, determine the occupancy corresponding to each passenger; when performing face orientation detection in the vehicle cabin, according to the target binding result, based on the identification corresponding to the passenger, obtain the face detection frame and the face key point sequence corresponding to each passenger; analyze the face detection frame and the face key point sequence corresponding to each passenger to determine the face orientation corresponding to each passenger.
[0041] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, where the memory is used to store a computer program; and the processor is used to cause the electronic device to implement the information binding method described in any one of the above embodiments when executing the computer program.
[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computing device, the computing device is caused to implement the information binding method described in any one of the above embodiments.
[0043] In a fifth aspect, an embodiment of the present application provides a vehicle, including: the information binding device described in the second aspect or the electronic device described in the third aspect.
[0044] The information binding method provided by the embodiment of the present application is specifically as follows: obtain a target image; the target image includes a plurality of passengers in the vehicle cabin; detect the target image to obtain the body information, face information, and hand information of the plurality of passengers; obtain the identifier of the body information corresponding to the plurality of passengers; combine preset matching conditions to obtain the face information and hand information that match the body information, and bind the matching body information, face information, and hand information to obtain a binding result; the preset matching conditions are used to obtain the face information and hand information that match the body information; perform identification processing on the face information and hand information in the binding result according to the identifier of the body information to obtain a target binding result. By binding the detection frames corresponding to each passenger, the embodiment of the present application can obtain all the detection frames corresponding to the passenger through one detection, avoiding the need to set corresponding detection schemes for different scenarios as in the prior art. Therefore, by binding the body information, face information, and hand information that match the passenger, the embodiment of the present application obtains the target binding result, and then directly obtains the required passenger information according to the target binding result, without the need to set corresponding detection schemes, thereby saving development costs and the energy of developers. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 One of the step flowcharts of the information binding method provided by the embodiments of the present application;
[0048] Figure 2 A schematic diagram of the information binding method provided by the embodiments of the present application;
[0049] Figure 3 Another step flowchart of the information binding method provided by the embodiments of the present application;
[0050] Figure 4 A framework diagram of the target detection model of the information binding method provided by the embodiments of the present application;
[0051] Figure 5 A structural schematic diagram of the information binding device provided by the embodiments of the present application;
[0052] Figure 6 A hardware structural schematic diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0053] In order to be able to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0054] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.
[0055] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner. In addition, in the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more.
[0056] It should be noted that in this article, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.
[0057] An embodiment of the present application provides an information binding method, and the execution subject of the information binding method is: an information binding device, which can specifically be a device such as a computer or an image processing server, or can also be a partial functional module in a device such as a computer or an image processing server. Refer to Figure 1 As shown, the information binding method includes the following steps S101 - S105:
[0058] S101. Obtain a target image.
[0059] Among them, the target image includes multiple passengers in the vehicle cabin.
[0060] In some embodiments, an image in the vehicle cabin can be collected by an image collection device arranged in the vehicle cabin as the target image; specifically, the image collection device can be a camera, and the interior of the vehicle cabin can be video-collected by the camera, and the target image can be obtained by intercepting a video frame of the video.
[0061] In the embodiment of the present application, the target image is obtained by the image collection device at the vehicle end, and then the target image can be transmitted to the information binding device in the cloud for target detection first, and then information binding is performed. Setting the information binding device in the cloud can avoid setting the information binding device at the vehicle end, resulting in an excessive load at the vehicle end and thus affecting the information binding efficiency.
[0062] S102. Detect the target image to obtain the human body information, face information, and hand information of the multiple passengers.
[0063] In some embodiments, when detecting the target image, the main detection is of multiple passengers in the target image to monitor the body, face, and hand movements of the passengers in the target image. Specifically, for the passengers in the target image, corresponding human body information, face information, and hand information are obtained. The human body information includes a human body detection box and a sequence of human body key points. The face information includes a face detection box and a sequence of face key points. The hand information includes a hand detection box and a sequence of hand key points to describe the positions and action postures of the passengers. It should be noted that when obtaining each detection box in the embodiments of the present application, information such as the size and coordinates of the detection box can also be obtained.
[0064] Optionally, for the multiple passengers in the target image, a target detection model based on the YOLO (an algorithm for object detection using a convolutional neural network) algorithm can be used to identify the parts of each passenger that appear in the target image, specifically including: three parts: the human body, the face, and the hand, and the corresponding parts are marked with detection boxes to obtain the detection box corresponding to the human body of the passenger, the detection box corresponding to the face, and the detection box corresponding to the hand. It should be noted that the multiple detection boxes in the present application include but are not limited to the detection box corresponding to the human body, the detection box corresponding to the face, and the detection box corresponding to the hand described above, and may also include detection boxes for other human body parts. The present application does not make any limitation in this regard.
[0065] Optionally, for the multiple passengers in the target image, key point detection can be performed on the target image to obtain key point sequences for three parts: the human body, the face, and the hand. Among them, to obtain the human body key point sequence, a Human Keypoints Detection (HKD) algorithm can be used, which is also known as human pose estimation. It is a relatively basic task in computer vision and a prerequisite task for human action recognition, behavior analysis, human-computer interaction, etc. At the same time, after completing the key point detection, key point tracking will also be performed, which is also known as human pose tracking. To obtain the face key point sequence, it can be obtained based on the Multi-task Cascaded Convolutional Networks (MTCNN). This neural network can perform tasks such as face detection, key point localization, and pose estimation simultaneously. To obtain the hand key point sequence, a hand pose estimation model can be used to obtain the key point sequence of the passenger's hand.
[0066] In some embodiments, the human detection frame is a detection frame that can include the entire body trunk of the passenger. Exemplarily, assuming that the entire body part of the passenger is presented in the target image, then the detection frame is the smallest rectangular frame that can enclose the head, two arms, an entire trunk, legs, and feet of the passenger's body; at the same time, the range inside the human detection frame can include: a sequence of key points of the human body, such as 21 key points corresponding to parts such as the head, chin, neck, shoulders, elbows, hands, hips, knees, and feet; assuming that only a partial body part of the human body can be presented in the target image, then the human detection frame can only be the smallest rectangular frame that includes a partial trunk of the human body, and a sequence of key points of the partial human body.
[0067] In some embodiments, the human detection frame can include the presentation range of the face of the passenger in the target image. Exemplarily, assuming that the complete face of the passenger can be presented in the target image, then the face detection frame is the smallest rectangular frame that can enclose the eyes, ears, nose, and chin of the human face. Assuming that only partial information of the human face can be presented in the target image, the face detection frame is the smallest rectangular frame that can enclose partial organs of the human face. Specifically, the face detection frame further includes face key points, and the face key points can be 68 main key points of the face, including key points of each facial organ such as the left and right eyes, nose, chin, and mouth.
[0068] In some embodiments, the human hand detection frame is a frame that can include the presentation range of the hand of the passenger in the target image. Exemplarily, assuming that the entire part of the human hand can be presented in the target image, the human hand detection frame can include the smallest rectangular frame of each fingertip, the connection points of each phalanx, etc. of the human hand. Assuming that only partial information of the human hand can be presented in the target image, the human hand detection frame is the smallest rectangular frame that can enclose partial fingertips and partial phalanx connection points of the human hand. Specifically, the human hand detection frame further includes human hand key points, and the human hand key points can be 21 main bone nodes of the hand, including fingertips, connection points of each phalanx, etc.
[0069] Exemplarily, referring to Figure 2 As shown, it is a schematic diagram of the detection frame provided by the embodiment of the present application. Through target detection of the image, a human detection frame, a face detection frame, and a human hand detection frame corresponding to each passenger in the image will be obtained. When the image includes a little girl as shown in Figure 2 As shown, 3 detection frames can be obtained through detection, including: a human detection frame 21, a face detection frame 22, and a human hand detection frame 23.
[0070] S103. Obtain the identifiers of the human body information corresponding to the multiple passengers.
[0071] In some embodiments, the identifier of the human body information corresponding to the passenger can be obtained through the identifier corresponding to the human detection frame corresponding to the passenger. Among them, the method of obtaining the identifier corresponding to the human detection frame corresponding to the passenger can be to obtain the identifiers of the human body information corresponding to each passenger based on historical image frames. If passenger A has never appeared in the historical frames, an identifier that has not appeared in the current frame can be assigned to passenger A.
[0072] S104. Obtain the face information and the hand information that match the human body information in combination with preset matching conditions, and bind the matching human body information, face information, and hand information to obtain a binding result.
[0073] Among them, the preset matching conditions are used to obtain the face information and the hand information that match the human body information.
[0074] In some embodiments, since the human body information includes a human detection frame and a sequence of human body key points, the face information includes a face detection frame and a sequence of face key points, and the hand information includes a hand detection frame and a sequence of hand key points; then in the above step S104, the specific implementation method of obtaining the face information and the hand information that match the human body information in combination with preset matching conditions and binding the matching human body information, face information, and hand information to obtain a binding result can be:
[0075] Obtain the face detection frame, the sequence of face key points, the hand detection frame, and the sequence of hand key points that match the human detection frame and the sequence of human body key points in combination with preset matching conditions, and bind the matching human detection frame, sequence of human body key points, face detection frame, sequence of face key points, hand detection frame, and sequence of hand key points to obtain a binding result.
[0076] Therefore, the binding result will include: the human detection frame, sequence of human body key points, face detection frame, sequence of face key points, hand detection frame, and sequence of hand key points corresponding to each passenger.
[0077] S105. Perform identification processing on the face information and the hand information in the binding result according to the identifier of the human body information to obtain a target binding result.
[0078] In some embodiments, after obtaining the identifiers of the human body information corresponding to the multiple passengers through the above step 103, in step S105, it is necessary to perform identification processing on the face information and the hand information in the binding result according to the identifier of the human body information to obtain a target binding result, so as to obtain the corresponding face information and hand information according to the identifier of the human body information.
[0079] Exemplarily, the identification 1 corresponding to the human body information of passenger 1 is combined, and the human body detection frame and human body key points corresponding to passenger 1 are bound. The specific method can be: displaying its identification at a preset position corresponding to the human body detection frame. The identification is displayed at the upper left corner of the human body detection frame corresponding to passenger 1 as: ID: 1; at the same time, the identification 1 of passenger 1 can also be added to the name of each key point coordinate in the human body key point sequence. For example, human body 1.1(x1, y1, z1) represents the coordinate of the first key point in the human body key point sequence of passenger A, and human body 1.2(x2, y2, z2) represents the coordinate of the second key point in the human body key point sequence of passenger A. The same applies to other key points.
[0080] Exemplarily, the identification 1 corresponding to passenger 1 is combined, and the identification is displayed at the upper left corner of the human body detection frame corresponding to passenger 1 as: ID: 1. Then, the identification corresponding to the upper left corner of the face detection frame can be: ID: 1.1; at the same time, the identification 1 corresponding to passenger A and the face identification 1 of passenger A can also be added to the name of each key point coordinate in the face key point sequence. For example, face 1.1.1: (x1, y1, z1) represents the coordinate of the first key point in the face key point sequence of passenger A, and human body 1.1.2(x2, y2, z2) represents the coordinate of the second key point in the face key point sequence of passenger A. The same applies to other key points.
[0081] Exemplarily, the identification 1 corresponding to passenger 1 is combined, and the identification is displayed at the upper left corner of the human body detection frame corresponding to passenger 1 as: ID: 1. Then, in order to distinguish the hand detection frame from the face detection frame, the identification corresponding to the upper left corner of one of the hand detection frames can be: ID: 1.2; at the same time, the identification 1 corresponding to passenger A and the hand identification 2 of passenger A can also be added to the name of each key point coordinate in the face key point sequence. For example, hand 1.2.1(x1, y1, z1) represents the coordinate of the first key point in the hand key point sequence of passenger A, and hand 1.2.2(x2, y2, z2) represents the coordinate of the second key point in the hand key point sequence of passenger A. The same applies to other key points.
[0082] In one embodiment, if there are two passengers in the target image, namely passenger A and passenger B, and after detection, a total of 8 detection frames of the human bodies, faces, and human hands of passenger A and passenger B are obtained. The above step S104 is to bind the human body detection frame, face detection frame, and two human hand detection frames corresponding to passenger A. When the identifier of passenger A is 1, the identifier of the corresponding human body detection frame is 1, the identifier of the corresponding face detection frame is 1.1, and the corresponding human hand detection frames are 1.2 and 1.3 respectively; bind the human body detection frame, face detection frame, and two human hand detection frames corresponding to passenger B. When the identifier of passenger B is 2, the identifier of the corresponding human body detection frame is 2, the identifier of the corresponding face detection frame is 2.1, and the corresponding human hand detection frames are 2.2 and 2.3 respectively; to clarify the detection frames corresponding to each passenger for differentiation. This facilitates the subsequent recognition of the limbs, expressions, and hand movements of passenger A and passenger B, improving efficiency.
[0083] In some embodiments, based on the target binding result and in combination with the identifier of the passenger corresponding human body information, all the detection frames bound to the passenger and the key point sequences in the detection frames can be obtained. Therefore, by obtaining all the detection frames corresponding to each passenger and the key point sequences in the detection frames, whether it is for scenarios of gesture interaction, expression recognition, or limb movement recognition, etc., different scenarios can be obtained through the target binding result obtained in this application, thereby improving the efficiency of target image recognition and facilitating the development of other downstream tasks.
[0084] The information binding method provided by the embodiments of this application is specifically as follows: Obtain a target image; the target image includes multiple passengers in the vehicle cabin; detect the target image to obtain the body information, face information, and hand information of the multiple passengers; obtain the identifiers of the body information corresponding to the multiple passengers; combine preset matching conditions to obtain the face information and hand information that match the body information, and bind the matching body information, face information, and hand information to obtain a binding result; the preset matching conditions are used to obtain the face information and hand information that match the body information; perform identification processing on the face information and hand information in the binding result according to the identifier of the body information to obtain a target binding result. By binding the detection frames corresponding to each passenger, the embodiments of this application can obtain all the detection frames corresponding to the passenger through one detection, avoiding the need to set corresponding detection schemes for different scenarios as in the prior art. Therefore, by binding the body information, face information, and hand information that match the passenger, the embodiments of this application obtain the target binding result, and then directly obtain the required passenger information according to the target binding result, without the need to set corresponding detection schemes, thus saving development costs and the energy of developers.
[0085] As an extension and refinement of the above embodiment, refer to Figure 3 As shown, it is an information binding method provided by the embodiments of this application. The information binding method includes the following steps:
[0086] S301. Obtain a target image.
[0087] Among them, the target image includes multiple passengers in the vehicle cabin.
[0088] S302. Detect the target image to obtain the body information, face information, and hand information of the multiple passengers.
[0089] Optionally, the implementation method of detecting the target image to obtain the body information, face information, and hand information of the multiple passengers can be: input the target image into a target detection model to obtain a set of detection frames and a key point sequence of multiple passengers in the target image.
[0090] Among them, the target detection model can be obtained by training a machine model based on sample data. The sample data includes: multiple sample images, as well as sample detection frames and sample key point sequences corresponding to the multiple sample images.
[0091] In some embodiments, the passengers in the target image can be detected by a trained object detection model to determine a corresponding plurality of detection boxes. Exemplarily, the detection boxes can be human body detection boxes, face detection boxes, and human hand detection boxes of the passengers in the target image.
[0092] In some embodiments, the human body detection boxes, face detection boxes, and human hand detection boxes of the passengers in the target image can be generated based on the object detection model. Specifically, referring to Figure 4 as shown, the object detection model 400 in step S302 above can include the following modules:
[0093] A human body detection module 41, configured to obtain the human body detection boxes corresponding to the plurality of passengers in the target image and the key point sequence corresponding to the human body.
[0094] A face detection module 42, configured to obtain the face detection boxes corresponding to the plurality of passengers in the target image and the key point sequence corresponding to the face.
[0095] A human hand detection module 43, configured to obtain the human hand detection boxes corresponding to the plurality of passengers in the target image and the key point sequence corresponding to the human hand.
[0096] Specifically, the object detection model in the embodiments of the present application can be obtained based on an object detection algorithm. For example: Fast Region Convolutional Neural Networks (Fast R-CNN), Faster Region Convolutional Neural Networks (Faster R-CNN), or Region based Fully Convolutional Network (R-FCN) are used to detect the human body, face, and human hand of the passengers in the target image to obtain the human body detection box of the human body, the face detection box of the face, and the human hand detection box of the human hand; among them, when using Faster R-CNN for object detection, it can be achieved through the following four basic steps, such as: candidate region generation, feature extraction, classification, and position refinement; in this way, it is possible to quickly and conveniently give human body detection boxes, face detection boxes, and human hand detection boxes with higher accuracy; at the same time, the use of object detection algorithms can reduce the computational amount.
[0097] It should be noted that in the embodiments of the present application, the object detection model can be a complete large model, or can be composed of three models that respectively detect human body information, face information, and human hand information.
[0098] S303. Based on the temporal tracking algorithm, match the human detection boxes corresponding to the multiple passengers with multiple historical human detection boxes.
[0099] In some embodiments, matching the detection boxes in the target image with the detection boxes in the previous historical image frame based on the temporal tracking algorithm may utilize the Kalman filter algorithm, the Hungarian algorithm, and cascade matching to track the passengers in the previous historical image frame, and then obtain the identifiers corresponding to the human information of each passenger in the previous frame for assignment to the corresponding passengers in the target image.
[0100] Among them, the main function of the Kalman filter algorithm is to predict the motion variables at the next moment using a series of current motion variables, but the first detection result is used to initialize the motion variables of the Kalman filter. The role of the Hungarian algorithm is to solve the assignment problem, that is, to assign a group of detection boxes and the boxes predicted by the Kalman filter to make the boxes predicted by the Kalman filter find the detection box that best matches itself to achieve the tracking effect.
[0101] Specifically, the main idea of the embodiments of this application is to combine the two tasks of object detection and object tracking. First, use an object detection algorithm (such as Faster R-CNN, etc.) to detect the position and bounding box of the target object in each frame. Then, extract the feature representation of the target through a deep learning model (such as CNN) and match each target with the targets that have been tracked in the previous frame. Factors such as the feature similarity and motion consistency of the target are considered during the matching process to determine the identity and trajectory of the target. One of the key contributions of DeepSORT is the use of a powerful appearance feature descriptor that can accurately distinguish the similarity between different targets.
[0102] S304. For each of the multiple passengers, obtain the historical human detection box corresponding to the human detection box.
[0103] Among them, the historical human detection box is the human detection box corresponding to the passenger in the historical image frame; the historical human detection box carries the corresponding identifier.
[0104] In some embodiments, after the detection box in the target image of the current frame successfully matches any detection box in the set of first detection boxes in the previous frame, the identifier of any detection box in the set of first detection boxes in the previous frame can be obtained and then used as the identifier of the detection box in the target image of the current frame.
[0105] Optionally, the target tracking algorithm matches multiple detection boxes in the target image with the detection boxes in the first detection box set to obtain the matching results of each detection box in the target image with the multiple detection boxes in the first detection box set, which is used to track the passengers in the target image of the current frame and can also be implemented by a multi-object tracking algorithm. Specifically, Multiple Object Tracking (MOT) aims to associate target objects across video frames to obtain the entire motion trajectory.
[0106] S305. Obtain the identifiers of the human body information corresponding to multiple passengers in the target image based on the identifiers carried by the historical human body detection boxes respectively corresponding to the human body detection boxes of the multiple passengers. In some embodiments, after obtaining the historical human body detection boxes corresponding to the human body detection boxes of each passenger in the target image in the above step S304, the identifiers of the human body information corresponding to each passenger in the target image can be obtained according to the identifiers corresponding to the corresponding historical human body detection boxes.
[0107] S306. Combine the preset matching conditions to obtain the face detection box, the face key point sequence, the hand detection box, and the hand key point sequence that match the human body detection box and the human body key point sequence, and bind the matching human body detection box, the human body key point sequence, the face detection box, the face key point sequence, the hand detection box, and the hand key point sequence to obtain the binding result.
[0108] In some embodiments, the specific implementation method of combining the preset matching conditions to obtain the face information and the hand information that match the human body information and binding the matching human body information, face information, and hand information to obtain the binding result may include the following steps 1 to 3:
[0109] Step 1. For each passenger, based on the position information of multiple face key point sequences in the target image, obtain the target face key point sequence that meets the first preset distance, and based on the position information of multiple hand key point sequences in the target image, obtain the target hand key point sequence that meets the second preset distance.
[0110] Specifically, the first preset distance may be a threshold distance set for the chin key point in the human body key point sequence to the chin key point in the face key point sequence; the second preset distance may be a threshold distance set for the wrist key point in the human body key point sequence to the wrist key point in the hand key point sequence.
[0111] Exemplarily, the distance L between the key point A at the chin in the human key point sequence and the key point B at the chin in the human face key point sequence can be calculated. By comparing the size of the distance L with a preset threshold, if the distance L is less than the corresponding preset threshold, it indicates that the distance between the key point A at the chin in the human key points and the key point B at the chin in the human face key points is very close, and thus it is of the same passenger. That is, the target human face key point sequence corresponding to the human information can be obtained in this way.
[0112] The left and right wrist key points C and D of the human key points, and the key point E at the wrist in the human hand key point sequence can be obtained, and the distances M and N between CE and DE are calculated. Similarly, M and N also need to be compared with the preset threshold. If the M and N are less than the corresponding preset threshold, it indicates that the distances between the left and right wrist key points C and D in the human key points and the key point E at the wrist in the human hand key point sequence are very close, and thus it is of the same passenger. That is, the target human hand key point sequence corresponding to the human information can be obtained in this way.
[0113] Step 2: Based on the position information of the target human face key point sequence and the target human hand key point sequence, obtain the target human face detection frame and the target human hand detection frame corresponding to the passenger.
[0114] In some embodiments, after obtaining the target human face key point sequence and the target human hand key point sequence corresponding to the passenger through the above Step 1, the corresponding target human face detection frame and target human hand detection frame can be further determined according to the position information of the target human face key point sequence and the target human hand key point sequence.
[0115] Step 3: Bind the target human face key point sequence, the target human hand key point sequence, the target human face detection frame, the target human hand detection frame, and the target human hand detection frame with the human detection frame and the human key point sequence to obtain a binding result.
[0116] In some embodiments, after obtaining the target human face key point sequence, the target human hand key point sequence, the target human face detection frame, the target human hand detection frame, and the target human hand detection frame corresponding to each passenger, as well as the human detection frame and the human key point sequence, binding them can obtain the binding result.
[0117] S307: Perform identification processing on the face information and the hand information in the binding result according to the identifier of the human information to obtain a target binding result.
[0118] As an extension and refinement of the above embodiments, after obtaining the target binding result, it is also necessary to perform temporal fusion on the detection boxes in the binding result. Specifically, in traditional perception algorithms, temporal fusion is the key to improving the accuracy and continuity of perception algorithms. It can make up for the limitations of single-frame perception, increase the receptive field, improve the problem of frame-to-frame jumps and target occlusions in object detection, more accurately judge the target movement speed, and also play an important role in target prediction and tracking.
[0119] Specifically, the temporal fusion may include some filtering and smoothing operations, such as smoothing of detection boxes, smoothing of key points, etc.; finally, structured information can be obtained; the resultant information can be represented in JSON format. It should be noted that JSON is a data format according to the syntax of JavaScript objects. Although it is based on JavaScript syntax, it is independent of JavaScript, which is why many program environments can read (interpret) and generate JSON.
[0120] By binding the human detection box, face detection box, and hand detection box corresponding to the passenger in the embodiment of the present application, each passenger can be described, such as where the human body is, where the face is, where the left and right hands are, etc. Then, based on the above information, it can be conveniently used for downstream tasks and extended: for example, in the passenger occupancy task, the key point information of the human body can be used to determine which seat the passenger is sitting on; for child detection, the human body picture can be cropped out for classification and recognition of children or adults; for gesture recognition, the temporal information of the left and right hands of the person can be obtained for static or dynamic gesture recognition, and at the same time, it can also be conveniently known which person is making the gesture, whether it is the left hand or the right hand, etc., so as to save the development cost and the energy of developers.
[0121] Optionally, in the embodiment of the present application, the hand information in the passenger information includes: left hand information and right hand information. The left hand information includes the left hand detection box corresponding to the passenger and the left hand key point sequence; the right hand information includes the right hand detection box corresponding to the passenger and the right hand key point sequence.
[0122] Therefore, after obtaining the target binding result, the target binding method further includes: obtaining the hand information corresponding to the passenger according to the identifier of the human body information corresponding to the passenger; determining the gestures corresponding to the left and right hands of the passenger according to the left hand information and the right hand information in the hand information, so as to obtain the corresponding instructions according to the gestures of the passenger.
[0123] Specifically, gesture recognition can be performed on the gestures of passengers from the target binding result to obtain the timing detection frames and key point sequences of the left and right hands of the passengers, and gesture recognition is performed. At the same time, it is also possible to clearly obtain which passenger is making the gesture and whether it is the left hand or the right hand based on the identification.
[0124] In the embodiment of the present application, the target binding result obtained based on the information binding method can be applied in: when performing passenger occupancy detection on the vehicle cabin, according to the target binding result, based on the identification corresponding to the passenger, obtain the human detection frames corresponding to each passenger; then, according to the detection frames corresponding to each passenger, determine the occupancy corresponding to each passenger; and when performing face orientation detection in the vehicle cabin, according to the target binding result, based on the identification corresponding to the passenger, obtain the face detection frames and face key point sequences corresponding to each passenger; then analyze the face detection frames and face key point sequences corresponding to each passenger to determine the face orientation corresponding to each passenger.
[0125] Based on the same inventive concept, as an implementation of the above method, the embodiment of the present application also provides an information binding device. This embodiment corresponds to the foregoing method embodiment. For the convenience of reading, the details in the foregoing method embodiment will not be described one by one in this embodiment, but it should be clear that the information binding device in this embodiment can correspondingly implement all the contents in the foregoing method embodiment.
[0126] The embodiment of the present application provides an information binding device, Figure 5 For the structural schematic diagram of this information binding device, as Figure 5 shown, the information binding device 500 includes:
[0127] An image acquisition unit 501, configured to acquire a target image; the target image includes a plurality of passengers in the cabin;
[0128] A detection unit 502, configured to detect the target image to obtain the human body information, face information, and human hand information of the plurality of passengers;
[0129] An identification acquisition unit 503, configured to acquire the identifications of the human body information corresponding to the plurality of passengers;
[0130] A binding unit 504, configured to combine preset matching conditions to obtain the face information and human hand information that match the human body information, and bind the matching human body information, face information, and human hand information to obtain a binding result; the preset matching conditions are used to obtain the face information and human hand information that match the human body information;
[0131] A processing unit 505, configured to perform identification processing on the face information and the hand information in the binding result according to the identification of the human body information, so as to obtain a target binding result.
[0132] As an optional implementation manner of an embodiment of the present application, the human body information includes a human body detection frame and a sequence of human body key points, the face information includes a face detection frame and a sequence of face key points, and the hand information includes a hand detection frame and a sequence of hand key points; the binding unit is specifically configured to obtain the face detection frame, the sequence of face key points, the hand detection frame, and the sequence of hand key points that match the human body detection frame and the sequence of human body key points in combination with a preset matching condition, and bind the matching human body detection frame, the sequence of human body key points, the face detection frame, the sequence of face key points, the hand detection frame, and the sequence of hand key points to obtain a binding result.
[0133] As an optional implementation manner of an embodiment of the present application, the identification acquisition unit is specifically configured to match the human body detection frames corresponding to the multiple passengers with multiple historical human body detection frames based on a time series tracking algorithm; for each of the multiple passengers, obtain the historical human body detection frame corresponding to the human body detection frame; the historical human body detection frame is the human body detection frame corresponding to a passenger in a historical image frame; the historical human body detection frame carries a corresponding identification; based on the identifications carried by the historical human body detection frames respectively corresponding to the human body detection frames of the multiple passengers, obtain the identifications of the human body information corresponding to the multiple passengers in the target image. As an optional implementation manner of an embodiment of the present application, the preset matching condition includes a first preset distance and a second preset distance; the binding unit is specifically configured to, for each passenger, obtain a target face key point sequence that satisfies the first preset distance based on the position information of multiple face key point sequences in the target image, and obtain a target hand key point sequence that satisfies the second preset distance based on the position information of multiple hand key point sequences in the target image; based on the position information of the target face key point sequence and the target hand key point sequence, obtain a target face detection frame and a target hand detection frame corresponding to the passenger; bind the target face key point sequence, the target hand key point sequence, the target face detection frame, the target hand detection frame, and the target hand detection frame with the human body detection frame and the sequence of human body key points to obtain a binding result.
[0134] As an optional implementation manner of an embodiment of the present application, the hand information in the passenger information includes: left - hand information and right - hand information. The left - hand information includes a left - hand detection frame corresponding to the passenger and a left - hand key - point sequence; the right - hand information includes a right - hand detection frame corresponding to the passenger and a right - hand key - point sequence. After obtaining the target binding result, the processing unit is further configured to obtain the hand information corresponding to the passenger according to the identifier of the human body information corresponding to the passenger; determine the gestures corresponding to the left and right hands of the passenger according to the left - hand information and the right - hand information in the hand information, so as to obtain corresponding instructions according to the gestures of the passenger.
[0135] As an optional implementation manner of an embodiment of the present application, when performing passenger occupancy detection on a vehicle cabin, according to the target binding result, based on the identifier corresponding to the passenger, obtain the human - body detection frame corresponding to each passenger; determine the occupancy corresponding to each passenger according to the detection frame corresponding to each passenger. When performing face - orientation detection in the vehicle cabin, according to the target binding result, based on the identifier corresponding to the passenger, obtain the face - detection frame and the face - key - point sequence corresponding to each passenger; analyze the face - detection frame and the face - key - point sequence corresponding to each passenger to determine the face orientation corresponding to each passenger.
[0136] Based on the same inventive concept, an embodiment of the present disclosure also provides an electronic device. Figure 6 The structural schematic diagram of the electronic device provided by the embodiment of the present disclosure is as Figure 6 shown. The electronic device provided in this embodiment includes: a memory 601 and a processor 602. The memory 601 is used to store a computer program; the processor 602 is configured to execute the information - binding method provided in the above - mentioned embodiment when executing the computer program.
[0137] Based on the same inventive concept, an embodiment of the present application also provides a computer - readable storage medium. A computer program is stored on the computer - readable storage medium. When the computer program is executed by a processor, the computing device is enabled to implement the information - binding method provided in the above - mentioned embodiment.
[0138] Based on the same inventive concept, an embodiment of the present application also provides a vehicle, and the vehicle includes the information - binding device provided in the above - mentioned embodiment or the electronic device provided in the above - mentioned embodiment.
[0139] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer - program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer - program product implemented on one or more computer - usable storage media containing computer - usable program code.
[0140] The processor can be a Central Processing Unit (CPU), or it can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0141] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0142] Computer-readable media includes permanent and non-permanent, removable and non-removable storage media. The storage media can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules or other data. Examples of the computer's storage media include, but are not limited to, Phase Change Memory (PRAM), Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), other types of Random Access Memory (RAM), Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash memory or other memory technologies, Compact Disc Read-Only Memory (CD-ROM), Digital Versatile Disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0143] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An information binding method, characterized in that, it includes: Obtain a target image; multiple passengers in the vehicle cabin are included in the target image; Detect the target image to obtain the body information, face information, and hand information of the multiple passengers; Obtain the identifiers of the body information corresponding to the multiple passengers; Combine preset matching conditions to obtain the face information and hand information that match the body information, and bind the matching body information, face information, and hand information to obtain a binding result; The preset matching conditions are used to obtain the face information and hand information that match the body information; Perform identification processing on the face information and hand information in the binding result according to the identifier of the body information to obtain a target binding result.
2. The method according to claim 1, characterized in that, The body information includes a body detection frame and a sequence of body key points, the face information includes a face detection frame and a sequence of face key points, and the hand information includes a hand detection frame and a sequence of hand key points; The step of combining preset matching conditions to obtain the face information and hand information that match the body information, and binding the matching body information, face information, and hand information to obtain a binding result includes: Combine preset matching conditions to obtain the face detection frame, the sequence of face key points, the hand detection frame, and the sequence of hand key points that match the body detection frame and the sequence of body key points, and bind the matching body detection frame, the sequence of body key points, the face detection frame, the sequence of face key points, the hand detection frame, and the sequence of hand key points to obtain a binding result.
3. The method according to claim 1, characterized in that, The step of obtaining the identifiers of the body information corresponding to the multiple passengers respectively includes: Based on a temporal tracking algorithm, match the body detection frames corresponding to the multiple passengers with multiple historical body detection frames; For each of the multiple passengers, obtain the historical body detection frame corresponding to the body detection frame; the historical body detection frame is the body detection frame corresponding to the passenger in the historical image frame; the historical body detection frame carries a corresponding identifier; Based on the identifiers carried by the historical body detection frames corresponding to the body detection frames of the multiple passengers respectively, obtain the identifiers of the body information corresponding to the multiple passengers in the target image.
4. The method according to claim 2, characterized in that, The preset matching conditions include a first preset distance and a second preset distance; the step of combining preset matching conditions to obtain the face detection frame, the sequence of face key points, the hand detection frame, and the sequence of hand key points that match the body detection frame and the sequence of body key points, and binding the matching body detection frame, the sequence of body key points, the face detection frame, the sequence of face key points, the hand detection frame, and the sequence of hand key points to obtain a binding result includes: For each passenger, based on the position information of the multiple face key point sequences in the target image, obtain a target face key point sequence that meets the first preset distance, and based on the position information of the multiple hand key point sequences in the target image, obtain a target hand key point sequence that meets the second preset distance; Based on the position information of the target face key point sequence and the target hand key point sequence, obtain a target face detection frame and a target hand detection frame corresponding to the passenger; Bind the target face key point sequence, the target hand key point sequence, the target face detection frame, the target hand detection frame, and the target hand detection frame with the human body detection frame and the human body key point sequence to obtain a binding result.
5. The method according to claim 1, wherein, the hand information in the passenger information includes: left hand information and right hand information, the left hand information includes a left hand detection frame corresponding to the passenger and a left hand key point sequence; the right hand information includes a right hand detection frame corresponding to the passenger and a right hand key point sequence; After obtaining the target binding result, the method further includes: According to the identifier of the human body information corresponding to the passenger, obtain the hand information corresponding to the passenger; According to the left hand information and the right hand information in the hand information, determine the gestures corresponding to the left hand and the right hand of the passenger, so as to obtain corresponding instructions according to the gestures of the passenger.
6. The method according to claim 1, wherein, After obtaining the target binding result, the method further includes: When performing passenger occupancy detection on the vehicle cabin, according to the target binding result, based on the identifier corresponding to the passenger's human body information, obtain the human body detection frames corresponding to each passenger; According to the detection frames corresponding to each passenger, determine the occupancy corresponding to each passenger; When performing face orientation detection in the vehicle cabin, according to the target binding result, based on the identifier corresponding to the passenger's human body information, obtain the face detection frames and face key point sequences corresponding to each passenger; Analyze the face detection frames and face key point sequences corresponding to each passenger to determine the face orientation corresponding to each passenger.
7. An information binding device, wherein, comprising: an image acquisition unit for acquiring a target image; the target image includes multiple passengers in the vehicle cabin; a detection unit for detecting the target image to obtain the human body information, face information, and hand information of the multiple passengers; an identifier acquisition unit for acquiring the identifiers of the human body information corresponding to the multiple passengers; a binding unit for combining preset matching conditions to obtain the face information and hand information that match the human body information, and binding the matching human body information, face information, and hand information to obtain a binding result; the preset matching conditions are used to obtain the face information and hand information that match the human body information; a processing unit for performing identifier processing on the face information and hand information in the binding result according to the identifier of the human body information to obtain a target binding result.
8. An electronic device, characterized in that, comprising: a memory and a processor, the memory being used for storing a computer program; the processor being used for, when executing the computer program, enabling the electronic device to implement the information binding method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a computing device, enabling the computing device to implement the information binding method according to any one of claims 1-6.
10. A vehicle, characterized in that, comprising: the information binding device according to claim 7 or the electronic device according to claim 8.