Cabin image detection method and device, cabin, vehicle and storage medium

By determining the target area in the cockpit image detection and performing preset processing, combining the fusion screening of the keyframe frame extraction algorithm and the human posture estimation algorithm, the problem of low image detection accuracy in the vehicle-mounted cockpit environment is solved, and higher image detection and human motion detection accuracy are achieved.

CN120148013APending Publication Date: 2025-06-13GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510272360.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the on-board cockpit environment, multiple target objects may exist in the same image, resulting in a large number of interference items in the image detection task, affecting the identification of key areas of the image, thereby reducing the accuracy of cockpit image detection.

Method used

By acquiring detection tasks and perceived information in the cockpit, the target area in the cockpit is determined, and the cockpit image is preset to extract the detection area information corresponding to the target area. Then, the detection area information of at least two frames of images is performed to identify the keyframes, and the keyframe frame extraction algorithm and the human posture estimation algorithm are combined to determine the image keyframes.

Benefits of technology

By focusing on the detection area corresponding to the target area in the cockpit image, the environmental interference is significantly reduced, the accuracy of image detection is improved, and the accuracy of human motion detection is improved. At the same time, the fusion screening process makes the selection of image keyframes more intuitive and interpretable, improving the accuracy and reliability of human posture detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148013A_ABST
    Figure CN120148013A_ABST
Patent Text Reader

Abstract

The invention relates to a cabin image detection method and device, a cabin, a vehicle and a storage medium. The cockpit image detection method comprises the following steps: acquiring a detection task and sensing information in a cockpit; determining a target area in the cabin according to the detection task and the sensing information in the cabin; acquiring a cabin image; performing preset processing on the cabin image according to the target area in the cabin to obtain detection area information corresponding to the target area in the cabin in the cabin image; and carrying out key frame identification processing on the detection area information of the at least two frames of cabin images, and determining an image key frame according to an identification processing result. According to the scheme provided by the invention, the accuracy of cabin image detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and in particular, to a cockpit image detection method, device, cockpit, vehicle, and storage medium. Background Art

[0002] With the continuous development of autonomous driving technology, vehicle cockpits have become more intelligent. For example, the control of functions such as Bluetooth, navigation, voice playback, and voice interaction can be set in the cockpit.

[0003] In the cockpit, it is necessary to detect the motion state of target objects in the cockpit, such as the human body. In order to detect the motion state of the target object, it is necessary to perform image detection on the images in the cockpit, and further perform human motion state detection based on the image detection results. In the cockpit image detection methods of related technologies, generally, key frames are extracted from the images using a key frame extraction algorithm and directly output as image key frames.

[0004] However, in the in-vehicle cockpit environment, there may be multiple target objects in the same image, such as multiple people moving simultaneously. Too many complex factors result in a large number of interference items in the image detection task; these interference items affect the recognition of the key regions of the image, thereby reducing the accuracy of cockpit image detection. Summary of the Invention

[0005] To solve or partially solve the problems existing in the related technologies, the present application provides a cockpit image detection method, device, cockpit, vehicle, and storage medium, which can improve the accuracy of cockpit image detection.

[0006] The first aspect of the present application provides a cockpit image detection method, including: Obtaining a detection task and perception information in the cockpit; Determining a target area in the cockpit according to the detection task and the perception information in the cockpit; Obtaining a cockpit image; Performing a preset process on the cockpit image according to the target area in the cockpit to obtain detection area information corresponding to the target area in the cockpit in the cockpit image; Performing key frame recognition processing on the detection area information of at least two frames of the cockpit image, and determining an image key frame according to the recognition processing result.

[0007] In an embodiment, the performing a preset process on the cockpit image according to the target area in the cockpit to obtain detection area information corresponding to the target area in the cockpit in the cockpit image includes: Performing mask processing on the cockpit image in an attention mask manner according to the target area in the cockpit to obtain detection area information corresponding to the target area in the cockpit in the cockpit image.

[0008] In one embodiment, masking the cockpit image according to the target area in the cockpit by using an attention mask method to obtain detection area information corresponding to the target area in the cockpit in the cockpit image, includes: Masking the cockpit image according to the target area in the cockpit by using an attention mask method, and extracting detection area coding information token corresponding to the target area in the cockpit in the cockpit image.

[0009] In one embodiment, performing key frame recognition processing on the detection area information of at least two frames of the cockpit image, and determining an image key frame according to the recognition processing result, includes: Performing recognition processing on the detection area information of at least two frames of the cockpit image by using a key frame extraction algorithm, and determining a first image key frame according to the recognition processing result; Performing recognition processing on the detection area information of at least two frames of the cockpit image by using a human pose estimation algorithm, and determining a second image key frame according to the recognition processing result; Performing fusion screening on the first image key frame according to the second image key frame, and using the first image key frame obtained after screening as the determined image key frame.

[0010] In one embodiment, performing fusion screening on the first image key frame according to the second image key frame, and using the first image key frame obtained after screening as the determined image key frame, includes: Determining the corresponding second image key frame according to the first image key frame; When a change in human pose key points is detected in the second image key frame, screening the first image key frame as the determined image key frame.

[0011] In one embodiment, performing key frame recognition processing on the detection area information of at least two frames of the cockpit image, and determining an image key frame according to the recognition processing result, includes: Performing recognition processing on the detection area information of at least two frames of the cockpit image by using a key frame extraction algorithm, and determining an image key frame according to the recognition processing result; or, Performing recognition processing on the detection area information of at least two frames of the cockpit image by using a human pose estimation algorithm, and determining an image key frame according to the recognition processing result.

[0012] A second aspect of the present application provides a cockpit image detection device, including: A task and perception information acquisition module, configured to acquire a detection task and perception information in the cockpit; The cockpit target area determination module is used to determine the target area in the cockpit according to the detection task and the in-cockpit perception information; The image acquisition module is used to acquire cockpit images; The image detection area information determination module is used to perform preset processing on the cockpit image according to the target area in the cockpit to obtain the detection area information corresponding to the target area in the cockpit in the cockpit image; The key frame determination module is used to perform key frame recognition processing on the detection area information of at least two frames of the cockpit images, and determine the image key frames according to the recognition processing results.

[0013] In one embodiment, the image detection area information determination module performs masking processing on the cockpit image in an attention mask manner according to the target area in the cockpit to obtain the detection area information corresponding to the target area in the cockpit in the cockpit image.

[0014] In one embodiment, the key frame determination module includes: The first algorithm processing sub-module is used to perform recognition processing on the detection area information of at least two frames of the cockpit images using a key frame extraction algorithm, and determine the first image key frames according to the recognition processing results; The second algorithm processing sub-module is used to perform recognition processing on the detection area information of at least two frames of the cockpit images using a human pose estimation algorithm, and determine the second image key frames according to the recognition processing results; The fusion screening sub-module is used to perform fusion screening on the first image key frames according to the second image key frames, and use the first image key frames obtained after screening as the determined image key frames.

[0015] The third aspect of the present application provides a cockpit, including the cockpit image detection device as described above.

[0016] The fourth aspect of the present application provides a vehicle, including: A processor; and A memory, on which executable code is stored, and when the executable code is executed by the processor, the processor is caused to execute the method as described above.

[0017] The fifth aspect of the present application provides a computer-readable storage medium, on which executable code is stored, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the method as described above.

[0018] The technical solution provided by the present application may include the following beneficial effects: The technical solution of the present application can determine the target area in the cockpit according to the detection task and the perception information in the cockpit. In this way, after obtaining the cockpit image, the cockpit image can be pre-processed according to the target area in the cockpit to obtain the detection area information corresponding to the target area in the cockpit image. Then, key frame recognition processing is performed on the detection area information of at least two frames of the cockpit images, and the image key frames are determined according to the recognition processing results. Through the above processing, for the obtained cockpit image, the present application only needs to focus on the detection area corresponding to the target area in the cockpit image, and other irrelevant areas in the cockpit image can be ignored, thereby significantly reducing the influence of environmental interference (such as light changes, non-target area interference, etc.), improving the accuracy of image detection, and further improving the accuracy of human motion detection.

[0019] Further, the technical solution of the present application can combine and fuse the results of the recognition processing using the key frame extraction algorithm and the results of the recognition processing using the human pose estimation algorithm, making full use of the respective advantages of the two algorithms, making the selection process of the image key frames more intuitive and interpretable, further improving the accuracy and reliability of human pose detection in a complex cockpit environment, and also supporting custom focus points based on scene requirements to meet the application requirements of diverse scenarios.

[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] By describing the exemplary embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more apparent. Among them, in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.

[0022] Figure 1 is a schematic flowchart of the cockpit image detection method shown in the present application; Figure 2 is another schematic flowchart of the cockpit image detection method shown in the present application; Figure 3 is a schematic application diagram of the cockpit image detection method shown in the present application; Figure 4 is a schematic diagram of image partitioning in the cockpit image detection method shown in the present application; Figure 5 is a schematic structural diagram of the cockpit image detection device shown in the present application; Figure 6 is another schematic structural diagram of the cockpit image detection device shown in the present application; Figure 7It is a schematic structural diagram of the vehicle shown in the present application. Detailed implementation manners

[0023] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0024] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0025] It should be understood that although the terms "first", "second", "third", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.

[0026] In the related art, there may be multiple target objects in the same image, such as multiple people moving simultaneously. Too many complex factors result in a large number of interference items in the image detection task, affecting the recognition of the key area of the image, thereby reducing the accuracy of cockpit image detection.

[0027] In view of the above problems, the present application provides a cockpit image detection method, which can improve the accuracy of cockpit image detection.

[0028] The technical solutions of the present application will be described in detail below with reference to the accompanying drawings.

[0029] See Figure 1 , a cockpit image detection method shown in the present application includes: S110, obtaining a detection task and perception information in the cockpit.

[0030] Among them, a detection task input by the user can be obtained and the perception information of each perception point in the cockpit can be obtained.

[0031] S120. Determine the target area inside the cockpit according to the detection task and the perception information inside the cockpit.

[0032] Among them, according to the detection task and the perception information inside the cockpit, the target area inside the cockpit corresponding to the detection task and the perception information inside the cockpit can be determined in the pre-divided cockpit areas. For example, the cockpit positions can be pre-divided into: the driver's position, the armrest box position, the co-pilot's position, the left rear position, the middle rear position, the right rear position, etc.

[0033] For example, the detection task is to detect all passenger positions (excluding the driver), but since the perception information of the perception points indicates that only the co-pilot's position and the left rear position are occupied, only these two partition areas need to be used as the target areas inside the cockpit. Subsequently, the action changes corresponding to these two partition areas in the cockpit image can be detected, and there is no need to detect the action changes in other areas.

[0034] S130. Obtain the cockpit image.

[0035] It should be noted that there is no sequential relationship between step S130 and step S110.

[0036] S140. Perform preset processing on the cockpit image according to the target area inside the cockpit to obtain the detection area information corresponding to the target area inside the cockpit in the cockpit image.

[0037] In one embodiment, the cockpit image is masked according to the target area inside the cockpit using the attention mask method to obtain the detection area information corresponding to the target area inside the cockpit in the cockpit image. For example, the detection area coding information token corresponding to the target area inside the cockpit in the cockpit image is extracted.

[0038] In this application, the attention range is restricted through the Attention Mask mechanism. For example, the cockpit image is masked using the attention mask method, so that only these two partition areas in the cockpit image can be focused on as the detection areas that need to be detected, and other partition areas in the cockpit image can be ignored, thereby extracting the detection area coding information token corresponding to these two partition areas in the cockpit image.

[0039] S150. Perform key frame recognition processing on the detection area information of at least two frames of cockpit images, and determine the image key frames according to the recognition processing results.

[0040] In one embodiment, the detection area information of at least two frames of cockpit images can be recognized using the key frame extraction algorithm, and the image key frames are determined according to the recognition processing results; or, the detection area information of at least two frames of cockpit images is recognized using the human pose estimation algorithm, and the image key frames are determined according to the recognition processing results.

[0041] As can be seen from this example, the cockpit image detection method of the present application can determine the target area in the cockpit according to the detection task and the perception information in the cockpit. In this way, after obtaining the cockpit image, the cockpit image can be pre-processed according to the target area in the cockpit to obtain the detection area information corresponding to the target area in the cockpit in the cockpit image. Then, key frame recognition processing is performed on the detection area information of at least two frames of cockpit images, and image key frames are determined according to the recognition processing results. Through the above processing, for the obtained cockpit image, the present application only needs to focus on the detection area corresponding to the target area in the cockpit in the cockpit image, and other irrelevant areas in the cockpit image can be ignored, thereby significantly reducing the influence of environmental interference (such as light changes, non-target area interference, etc.), improving the accuracy of image detection, and further improving the accuracy of human motion detection.

[0042] See Figure 2 , another schematic flowchart of the cockpit image detection method shown in the present application, and at the same time, see Figure 3 The application schematic diagram of the cockpit image detection method shown.

[0043] As Figure 2 shown, the method includes: S210, obtain the detection task and the perception information in the cockpit.

[0044] In this step S210, the detection task input by the user can be obtained. For example, the detection task can be "detect whether the passenger is sleeping" or "detect whether the hand is out of the window", etc.

[0045] In this step S210, the perception information in the cockpit can be obtained. Various sensors can be set in the cockpit, and the positions where the sensors are installed can be used as perception points. Therefore, the perception information of each perception point in the cockpit can be obtained. For example, the perception information of the perception point is "a person is sitting in the driver's seat, a person is sitting in the co-pilot seat, a person is sitting in the left rear seat" or "a person is sitting in the driver's seat, a person is sitting in the co-pilot seat, a person is sitting in the middle rear seat", etc.

[0046] S220, according to the detection task and the perception information in the cockpit, determine the target area in the cockpit corresponding to the detection task and the perception information in the cockpit in the pre-divided cockpit area.

[0047] See Figure 4 , the present application can pre-divide the cockpit area according to the cockpit position. The cockpit position can be divided into: the driver's seat position, the armrest box position, the co-pilot seat position, the left rear seat position, the middle rear seat position, the right rear seat position, etc. For example, Figure 4Among them, 1 is the driver's seat position, 2 is the armrest box position, 3 is the co-driver's seat position, 4 is the left rear seat position, 5 is the middle rear seat position, and 6 is the right rear seat position. It should be noted that the area division here is only for illustration and is not limited to this. Other division settings can also be made according to needs.

[0048] In this step S220, according to the detection task and the perception information in the cockpit, the in-cockpit target area corresponding to the detection task and the perception information in the cockpit can be determined in the pre-divided cockpit area.

[0049] For example, according to the detection task of "detecting whether a passenger is sleeping", and the perception information of the perception point is that "the driver's seat is occupied, the co-driver's seat is occupied, and the left rear seat is occupied", in the pre-divided cockpit area, the corresponding in-cockpit target areas can be determined as the co-driver's seat position and the left rear seat position, without considering the other 4 areas such as the driver's seat position, the armrest box position, the middle rear seat position, and the right rear seat position.

[0050] That is to say, the detection task is to detect all passenger positions (excluding the driver's seat), but since the perception information of the perception point indicates that only the co-driver's seat position and the left rear seat are occupied, only the action changes in these two divided areas need to be detected, and there is no need to detect the action changes in other areas.

[0051] For example, according to the detection task of "detecting whether a hand is stretched out of the window", and the perception information of the perception point is that "the driver's seat is occupied, the co-driver's seat is occupied, and the middle rear seat is occupied", in the pre-divided cockpit area, the corresponding in-cockpit target areas can be determined as the driver's seat position and the co-driver's seat position, without considering the other 4 areas such as the armrest box position, the middle rear seat position, the left rear seat position, and the right rear seat position.

[0052] That is to say, the detection task is to detect all positions where a hand can be stretched out of the window and is occupied. The perception information of the perception point indicates that there are people in the driver's seat position, the co-driver's seat position, and the middle rear seat position, but the hand cannot be stretched out of the window in the middle rear seat position. Therefore, only the action changes in the driver's seat position and the co-driver's seat position need to be detected, and there is no need to detect the action changes in other areas.

[0053] Therefore, the present application can screen out the corresponding divided areas according to different detection task requirements, and combine the perception information of the perception points in the vehicle cockpit, and use a large model to predict the divided areas that need to be concerned for different requirements as the in-cockpit target areas, reducing the unnecessary interference generated by other areas.

[0054] S230, obtain the cockpit image.

[0055] In step S230, consecutive frame cockpit images captured by a sensor can be obtained. For example, cockpit images captured by an in-cockpit camera or a binocular camera can be obtained. Consecutive frame cockpit images intercepted from a video stream can also be obtained in this step.

[0056] It should be noted that there is no sequential relationship between step S230 and step S210.

[0057] S240. According to the target area in the cockpit, the cockpit image is masked using the attention mask method, and the detection area coding information token corresponding to the target area in the cockpit is extracted from the cockpit image.

[0058] According to the above steps, the partition areas that need attention can be obtained, that is, the target areas in the cockpit can be obtained. For example, the target areas in the cockpit are the driver's position and the co-driver's position (corresponding to Figure 4 zones 1 and 3 in the figure). For the obtained cockpit image, only the areas related to the driver's position and the co-driver's position in the cockpit image need to be concerned, and the areas in the cockpit image that are not related to the driver's position and the co-driver's position can be ignored. Then, in this step S240, the present application restricts the attention range through the Attention Mask mechanism. For example, the cockpit image is masked using the attention mask method, so that only these two partition areas in the image need to be concerned as the detection areas to be detected for the cockpit image, and other partition areas in the cockpit image can be ignored, thereby extracting the detection area coding information token corresponding to these two partition areas in the cockpit image.

[0059] In the large model, information finally exists in the form of tokens. A token is the basic unit for the model to process and understand input information. The token of the present application can refer to the encoded information converted from the picture information, and this step S240 can extract the tokens contained in the pixel blocks corresponding to the target area of the cockpit image.

[0060] The present application can mask the positions in the cockpit image that do not need to be concerned through the Attention Mask, and set the attention weights corresponding to the positions that do not need to be concerned to a very small value (such as negative infinity), so as to ignore them when calculating the attention distribution. For example, by applying a mask on the weight matrix of the Attention calculation to control the model's attention to different elements in the sequence, if an element should be ignored during the Attention calculation, the corresponding weight will be set to a very small negative number. In this way, after being processed by the softmax function (normalized exponential function), the weights of these positions will be close to 0, achieving the ignoring effect.

[0061] This application can perform masking processing on consecutive frames (e.g., at least two frames) of cockpit images for a detection task using an Attention Mask, thereby calculating tokens for multiple consecutive frames of processed cockpit images. After obtaining the tokens of the consecutive frames of cockpit images, subsequent steps S250 (key frame extraction algorithm processing) and S260 (human pose estimation algorithm processing) can be performed respectively to obtain the required key frames.

[0062] S250, use the key frame extraction algorithm to perform recognition processing on the token of the encoded information of the detection area of at least two frames of cockpit images, and determine the first image key frame according to the recognition processing result.

[0063] The process of this application using the key frame extraction algorithm to perform recognition processing on the tokens extracted from the cockpit images can be as follows but is not limited to this: This application can calculate the change of frames in a consecutive frame image through a corner detection algorithm. Assume that the input consecutive frame cockpit images are , where represents the t-th image.

[0064] Image features can express the main information of the objects in the image. Image features can include: corners, edges, textures, etc. This application takes the image feature as a corner for recognition and analysis as an example but is not limited to this. A corner is a pixel area corresponding to the maximum value of the gray level gradient. The image gray level is step-like, gradually progressing from white to gray. The position of the point corresponding to the maximum value of the gradient obtained by calculating the gradient of the gray level can be qualitatively defined as its corner feature. A corner is the intersection of two edges in an image and usually has two main features: high response and multi-directionality. Corner feature detection algorithms can include different detection algorithms, such as the Harris corner detection algorithm, the Shi-Tomasi corner detection algorithm, the CSS (Curvature-Scale-Space) corner detection algorithm, etc.

[0065] The Harris corner detection algorithm mainly moves a local window on the image to determine whether the gray level changes significantly. If the gray level values (on the gradient map) within the window all change significantly, then there are corners in the area where this window is located. That is to say, the Harris corner detection algorithm is a local area-based detection method. It detects corners by calculating the gray level change in the area around each pixel in the image. This algorithm uses eigenvalues to determine whether a pixel point is a corner. When the eigenvalue is large, it indicates that there are corners around this point. The Harris corner detection algorithm has good rotation invariance and illumination invariance, so it is widely used in image registration and target tracking.

[0066] The embodiments of the present application can use the Harris corner detection algorithm in related technologies for calculation, and there is no limitation thereto.

[0067] For each The detection area coding information token in can use the corner detection algorithm (Harris corner detection) to extract the corner set :

[0068] Among them, represents the number of corners detected in the t-th image, i is greater than or equal to 1, represents the corner set of the t-th image, represents the x-axis coordinate of the i-th corner, represents the y-axis coordinate of the i-th corner.

[0069] Then, compare the corner changes of two adjacent cockpit images

[0070]

[0071] Among them, represents the number of corners detected in the t-th image, represents the number of corners detected in the (t - 1)-th image, represents the corner set of the t-th image, represents the corner set of the (t - 1)-th image.

[0072] Among them, the absolute value of the intersection of the corner set and the corner set can be taken, and divided by the larger value from and to obtain the corner change of two adjacent cockpit images .

[0073] If is smaller, it means that the change between these two cockpit images is greater.

[0074] The present application can set a corner change threshold . If (the corner change of two adjacent cockpit images is less than the corner change threshold), then determine that these two cockpit images are candidate key frames, that is, as the first image key frame.

[0075] Among them, the threshold can have a value range between 0.1 and 0.4, and the present application can take 0.25 but is not limited thereto.

[0076] Therefore, in step S250 of the present application, the detection area information of at least two cockpit images is identified and processed using a key frame extraction algorithm. According to the corner point change between two adjacent cockpit images being less than the corner point change threshold, the two adjacent cockpit images are determined as the first image key frames.

[0077] S260, the detection area encoding information token of at least two cockpit images is identified and processed using a human pose estimation algorithm. According to the recognition processing result, the second image key frames are determined.

[0078] The HPE (Human Pose Estimation) algorithm can be used to identify and classify the human body and joints, capture a set of coordinates for each joint, and generate key points (keypoints) describing the human pose. Human pose estimation models include skeleton-based models, contour-based models, volume-based models, etc.

[0079] Among them, the skeleton-based model is also called the kinematic model. This model includes a set of key points (joints), such as ankles, knees, shoulders, elbows, wrists, and limb directions, and is mainly used for 3D and 2D pose estimation. The contour-based model is also called the planar model and is used for two-dimensional pose estimation. It consists of the contours and approximate widths of the body, torso, and limbs. The volume-based model is also called the volume model and is used for 3D pose estimation. It consists of multiple popular 3D human models and poses represented by human geometric meshes and shapes, and is usually used for 3D human pose estimation based on deep learning.

[0080] In the related art, a deep learning-based human pose estimation algorithm can be adopted, which may include the OpenPose model algorithm, the AlphaPose model algorithm, the PoseNet model algorithm, etc.

[0081] The present application can adopt the OpenPose model algorithm of the related art for recognition and processing, but is not limited thereto. The OpenPose model algorithm is a bottom-up method. The neural network first detects body parts or key points in the image and then assembles them into a person. The OpenPose model mainly detects human key points such as hands, faces, and feet of multiple people in real-time scenarios. The principle of the OpenPose model is to use CNN (Convolutional Neural Networks) to perform deep learning processing on the image, so as to estimate human pose key points in real-time, including body and hand postures. OpenPose adopts a unique dual-branch network structure. One branch focuses on human pose detection, and the other branch is responsible for hand pose detection. Each branch contains multiple stages, and each stage extracts and refines image features through convolutional layers and pooling layers. Among them, In this application, first, the set of human pose key points of the detection area encoding information token in the cockpit image of the t-th frame can be obtained according to the OpenPose model:

[0082] where represents the number of human pose key points detected in the t-th image, i is greater than or equal to 1, represents the x-axis coordinate of the i-th human pose key point in the t-th image, represents the y-axis coordinate of the i-th human pose key point in the t-th image.

[0083] Then, the displacement change of the human pose key points between two adjacent frames of cockpit images is calculated respectively, and the joint angle change of the human pose key points between two adjacent frames of cockpit images is calculated.

[0084] Among them, calculating the displacement change of the human pose key points between two adjacent frames of cockpit images can be calculated in the following way:

[0085] Among them, calculating the joint angle change of the human pose key points between two adjacent frames of cockpit images can be calculated in the following way: This application can currently consider calculating the angles of the elbow joint / wrist joint / knee joint of the target object such as the human pose. The joint angle can be defined by three points. For example, the elbow joint is defined by three key points: shoulder / elbow / arm (if these three points cannot be found, this set of joints is not calculated).

[0086] Among them, assuming that the three joint points are defined as , then the joint angle has the following calculation formula:

[0087] The calculation formula for calculating the joint angle change of the human pose key points between two adjacent frames of cockpit images is as follows:

[0088] where n is the number of joints counted in the t-th image or the (t - 1)-th image, is the joint angle of the human pose key points in the t-th cockpit image, is the joint angle of the human pose key points in the (t - 1)-th cockpit image.

[0089] This application can respectively set the key point displacement change threshold and the key point joint angle change threshold .

[0090] If the displacement of the human body pose key points in two adjacent cockpit images changes , or the joint angles of the human body pose key points in two adjacent cockpit images change , it is considered that the human body pose key points have changed, and these two cockpit images are determined as candidate key frames, that is, as the second image key frames.

[0091] It should be noted that for the threshold of the joint angle change of the key points , the general value can be between 10° and 15°; for the threshold of the displacement change of the key points , the value is generally related to the image size and can be between 5 and 20 pixel points.

[0092] Therefore, in step S260 of the present application, the detection area information of at least two cockpit images is processed by using a human body pose estimation algorithm. According to the displacement change of the human body pose key points in two adjacent cockpit images being greater than the key point displacement change threshold, or the joint angle change of the human body pose key points in two adjacent cockpit images being greater than the key point joint angle change threshold, two adjacent cockpit images are determined as the second image key frames.

[0093] It should be noted that step S260 and step S250 have no sequential relationship and can be executed separately.

[0094] S270, fuse and screen the first image key frames according to the second image key frames, and use the screened first image key frames as the determined image key frames.

[0095] The present application combines and considers the results of step S260 and step S250 for secondary screening.

[0096] Suppose N groups of candidate key frames (the first image key frames) are obtained through the key frame extraction algorithm in step S250; for these N groups of candidate key frames, further refer to the results (the second image key frames) obtained by processing in step S260 using the human body pose estimation algorithm for fusion screening to obtain the finally determined image key frames, so as to exclude the influence of environmental factors, improve the accuracy of determining the image key frames, and thus improve the image detection accuracy.

[0097] Among them, the processing process of the fusion screening in the present application may include: Determine the corresponding second image key frames according to the first image key frames; When a change in the human body pose key points is detected in the second image key frames, screen the first image key frames as the determined image key frames.

[0098] That is to say, for a set of candidate key frames (the first image key frames), if it is found that the detected human body pose key points have changed according to the corresponding second image key frames of this set of candidate key frames, then this set of candidate key frames is selected into the final set of image key frames , that is, this set of candidate key frames (the first image key frames) is screened as the determined image key frames; For a set of candidate key frames (the first image key frames), if it is found that the detected human body pose key points have not changed according to the corresponding second image key frames of this set of candidate key frames, it is considered that the change of this set of candidate key frames is caused by environmental factors and can be discarded. At this time, this set of candidate key frames (the first image key frames) is not screened as the determined image key frames; After all N groups of candidate key frames are fused and screened, the finally determined set of image key frames is output , and the image key frames in the set are all the image key frames after fusion and screening.

[0099] Through the above processing, the present application can find more accurate image key frames to improve the accuracy of image detection. After obtaining the final image key frames, the VLM (Vision Language Model) can be used in combination with other relevant algorithms to perform the detection of the motion state of the target object, so as to judge the motion state of the human body in the cockpit. Among them, the detection of the motion state of the target object according to the determined image key frames can be implemented by using relevant technical solutions, and the present application does not limit this.

[0100] It should be noted that the algorithm fusion in step S270 of the present application is to improve the accuracy of the screening effect. According to needs, the present application may also not perform the algorithm fusion process. For example, after determining the detection area coding information token, the key frame extraction algorithm can be directly used in step S250 to identify the detection area coding information token of at least two cockpit images, and the image key frames are determined according to the recognition result; or, the human body pose estimation algorithm can be directly used in step S260 to identify the detection area coding information token of at least two cockpit images, and the image key frames are determined according to the recognition result.

[0101] From this example, it can be seen that the present application proposes an image detection method based on position-based Attention Mask, and combines a key frame extraction algorithm and a human pose estimation algorithm for fusion screening, ultimately improving the accuracy of cockpit image detection. Using the method of the present application, the interference of environmental complexity can be reduced, and the accuracy of image detection can be improved. By designing a position-based Attention Mask mechanism, the present application restricts the information in non-concerned areas of the image, significantly reducing the influence of environmental interferences (such as light changes, non-target area interferences, etc.), thereby improving the accuracy of image detection, and further improving the accuracy of human motion detection.

[0102] Using the method of the present application, the image recognition and detection results can be further optimized. In the related art, when only the key frame extraction algorithm is used, the interpretability of the key frame extraction algorithm is poor. Since the key frame extraction algorithm usually relies on complex feature extraction and clustering algorithms, the overall process is difficult to interpret and difficult to meet the user's customized requirements for detection needs in some scenarios. In the related art, when only the human pose estimation algorithm is used, the accuracy of the human pose estimation algorithm is low. In the actual application scenario in the cockpit, the deep learning-based human pose estimation algorithm shows low accuracy in complex environments (such as occlusion, space limitation), and cannot meet the high-precision detection requirements. The technical solution of the present application combines the key frame extraction algorithm and the deep learning estimation algorithm based on human pose, making the selection process of image key frames more intuitive and interpretable, and also improving the accuracy and reliability of human pose detection in complex cockpit environments. At the same time, it supports customizing the focus points based on scenario requirements, meeting the application requirements of diverse scenarios.

[0103] Corresponding to the foregoing method embodiments for implementing application functions, the present application also provides a cockpit image detection device, a vehicle, and corresponding embodiments.

[0104] Figure 5 It is a schematic structural diagram of the cockpit image detection device shown in the present application.

[0105] See Figure 5 , the cockpit image detection device 50 provided by the present application includes: a task and perception information acquisition module 51, a cockpit target area determination module 52, an image acquisition module 53, an image detection area information determination module 54, and a key frame determination module 55.

[0106] The task and perception information acquisition module 51 is used to acquire detection tasks and perception information in the cockpit. Among them, the task and perception information acquisition module 51 can acquire the detection tasks input by the user and acquire the perception information of each perception point in the cockpit.

[0107] The cockpit target area determination module 52 is configured to determine the cockpit target area according to the detection task and the in-cockpit perception information. Among them, the cockpit target area determination module 52 can determine the in-cockpit target area corresponding to the detection task and the in-cockpit perception information in the pre-partitioned cockpit area. For example, the cockpit positions can be pre-divided into: the driver's position, the armrest box position, the co-driver's position, the left rear position, the middle rear position, the right rear position, etc. For example, the detection task is to detect all passenger positions (excluding the driver), but since the perception information of the perception points indicates that only the co-driver's position and the left rear position are occupied, only the action changes in these two partition areas in the cockpit image need to be detected, and there is no need to detect the action changes in other areas.

[0108] The image acquisition module 53 is configured to acquire the cockpit image.

[0109] The image detection area information determination module 54 is configured to perform a preset process on the cockpit image according to the in-cockpit target area to obtain the detection area information corresponding to the in-cockpit target area in the cockpit image. The image detection area information determination module 54 can perform a masking process on the cockpit image in an attention mask manner according to the in-cockpit target area to obtain the detection area information corresponding to the in-cockpit target area in the cockpit image. For example, the detection area coding information token corresponding to the in-cockpit target area in the cockpit image is extracted.

[0110] The key frame determination module 55 is configured to perform key frame recognition processing on the detection area information of at least two frames of cockpit images, and determine the image key frame according to the recognition processing result. The key frame determination module 55 can perform recognition processing on the detection area information of at least two frames of cockpit images using a key frame extraction algorithm, and determine the image key frame according to the recognition processing result; or, perform recognition processing on the detection area information of at least two frames of cockpit images using a human pose estimation algorithm, and determine the image key frame according to the recognition processing result.

[0111] For the acquired cockpit image, the device provided in this application only needs to focus on the detection area corresponding to the in-cockpit target area in the cockpit image, and other irrelevant areas in the cockpit image can be ignored, so that the influence of environmental interference (such as light changes, non-target area interference, etc.) can be significantly reduced, thereby improving the accuracy of image detection, and further improving the accuracy of human motion detection.

[0112] Figure 6 It is another structural schematic diagram of the cockpit image detection device shown in this application.

[0113] See Figure 6, the cockpit image detection device 50 provided by this application includes: a task and perception information acquisition module 51, a cockpit target area determination module 52, an image acquisition module 53, an image detection area information determination module 54, and a key frame determination module 55.

[0114] In one embodiment, the image detection area information determination module 54 performs masking processing on the cockpit image in an attention mask manner according to the target area in the cockpit to obtain the detection area information corresponding to the target area in the cockpit in the cockpit image.

[0115] In one embodiment, the key frame determination module 55 includes: a first algorithm processing sub-module 551, a second algorithm processing sub-module 552, and a fusion screening sub-module 553.

[0116] The first algorithm processing sub-module 551 is configured to perform recognition processing on the detection area information of at least two frames of cockpit images using a key frame extraction algorithm, and determine the first image key frame according to the recognition processing result; The second algorithm processing sub-module 552 is configured to perform recognition processing on the detection area information of at least two frames of cockpit images using a human pose estimation algorithm, and determine the second image key frame according to the recognition processing result; The fusion screening sub-module 553 is configured to perform fusion screening on the first image key frame according to the second image key frame, and use the first image key frame obtained after screening as the determined image key frame for performing target object motion state detection according to the determined image key frame. The fusion screening sub-module 553 determines the corresponding second image key frame according to the first image key frame; when it is detected that the human pose key points have changed in the second image key frame, the first image key frame is screened as the determined image key frame; when it is detected that the human pose key points have not changed in the second image key frame, the first image key frame is discarded, that is, the first image key frame is not used as the determined image key frame.

[0117] It should be noted that the key frame determination module 55 can also directly perform recognition processing on the detection area information of at least two frames of cockpit images using a key frame extraction algorithm, and determine the image key frame according to the recognition processing result; or, directly perform recognition processing on the detection area information of at least two frames of cockpit images using a human pose estimation algorithm, and determine the image key frame according to the recognition processing result.

[0118] The device provided by this application can significantly reduce the influence of environmental interferences (such as light changes, non-target area interferences, etc.), thereby improving the accuracy of image detection, and further improving the accuracy of human motion detection. Additionally, the results of recognition processing using the key frame extraction algorithm and the results of recognition processing using the human pose estimation algorithm can be combined and screened, making full use of the respective advantages of the two algorithms, making the process of selecting image key frames more intuitive and interpretable, and further improving the accuracy and reliability of human pose detection in complex cockpit environments. At the same time, it can also support custom focus points based on scene requirements, meeting the application requirements of diverse scenarios.

[0119] This application also provides a cockpit, which may include the cockpit image detection device as described above Figure 5 or Figure 6 shown.

[0120] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0121] Figure 7 is a schematic structural diagram of a vehicle shown in an embodiment of this application.

[0122] Refer to Figure 7 , the vehicle 1000 includes a memory 1010 and a processor 1020.

[0123] The processor 1020 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0124] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage device can be a readable and writable storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 1010 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 1010 can include removable storage devices that are readable and / or writable, such as compact discs (CDs), read-only digital versatile discs (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray discs, ultra-density discs, flash memory cards (such as SD cards, min SD cards, Micro-SD cards, etc.), magnetic floppy disks, etc. Computer-readable storage media do not include carrier waves and instantaneous electronic signals transmitted wirelessly or wired.

[0125] Executable code is stored on the memory 1010, and when the executable code is processed by the processor 1020, it can cause the processor 1020 to execute some or all of the methods described above.

[0126] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.

[0127] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When the executable code (or the computer program or computer instruction code) is executed by a processor of an electronic device (or a server, etc.), it causes the processor to execute some or all of the steps of the above method according to the present application.

[0128] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A cockpit image detection method, characterized in that: include: Obtain detection tasks and in-cabin perception information; Determining a target area in the cabin according to the detection task and the in-cabin perception information; Get cockpit image; Performing preset processing on the cockpit image according to the target area in the cockpit to obtain detection area information corresponding to the target area in the cockpit in the cockpit image; A key frame recognition process is performed on the detection area information of at least two frames of the cockpit image, and an image key frame is determined according to the recognition process result.

2. The method according to claim 1, characterized in that The performing preset processing on the cockpit image according to the target area in the cockpit to obtain detection area information corresponding to the target area in the cockpit in the cockpit image includes: The cockpit image is masked by using an attention mask method according to the target area in the cockpit, so as to obtain detection area information corresponding to the target area in the cockpit in the cockpit image.

3. The method according to claim 2, characterized in that The step of performing mask processing on the cockpit image by using an attention mask method according to the target area in the cockpit to obtain detection area information corresponding to the target area in the cockpit in the cockpit image includes: The cockpit image is masked by using an attention mask method according to the target area in the cockpit, and a detection area coding information token corresponding to the target area in the cockpit is extracted from the cockpit image.

4. The method according to any one of claims 1 to 3, characterized in that: The step of performing key frame recognition processing on the detection area information of at least two frames of the cockpit images and determining the image key frame according to the recognition processing result includes: Using a key frame extraction algorithm to perform recognition processing on the detection area information of at least two frames of the cockpit image, and determining a first image key frame according to the recognition processing result; Using a human posture estimation algorithm to perform recognition processing on the detection area information of at least two frames of the cockpit image, and determining a second image key frame according to the recognition processing result; The first image key frame is fused and screened according to the second image key frame, and the first image key frame obtained after screening is used as the determined image key frame.

5. The method according to claim 4, characterized in that The step of fusing and screening the first image key frame according to the second image key frame and taking the first image key frame obtained after screening as the determined image key frame includes: Determine the corresponding second image key frame according to the first image key frame; When the second image key frame detects that the key point of the human body posture changes, the first image key frame is selected as the determined image key frame.

6. The method according to any one of claims 1 to 3, characterized in that: The step of performing key frame recognition processing on the detection area information of at least two frames of the cockpit images and determining the image key frame according to the recognition processing result includes: Using a key frame extraction algorithm to perform recognition processing on the detection area information of at least two frames of the cockpit image, and determining the image key frame according to the recognition processing result; or, The detection area information of at least two frames of the cockpit image is recognized and processed using a human posture estimation algorithm, and an image key frame is determined according to the recognition processing result.

7. A cockpit image detection device, characterized in that: include: Task and perception information acquisition module, used to obtain detection tasks and perception information in the cockpit; A cockpit target area determination module, used to determine the target area in the cockpit according to the detection task and the cockpit perception information; An image acquisition module, used for acquiring cockpit images; An image detection area information determination module, configured to perform preset processing on the cockpit image according to the target area in the cockpit, and obtain detection area information corresponding to the target area in the cockpit in the cockpit image; The key frame determination module is used to perform key frame recognition processing on the detection area information of at least two frames of the cockpit image, and determine the image key frame according to the recognition processing result.

8. The device according to claim 7, characterized in that: The image detection area information determination module performs mask processing on the cockpit image in an attention mask manner according to the target area in the cockpit, and obtains detection area information corresponding to the target area in the cockpit in the cockpit image.

9. The device according to claim 7 or 8, characterized in that The key frame determination module comprises: A first algorithm processing submodule is used to perform recognition processing on the detection area information of at least two frames of the cockpit image using a key frame extraction algorithm, and determine a first image key frame according to the recognition processing result; A second algorithm processing submodule is used to perform recognition processing on the detection area information of at least two frames of the cockpit image using a human posture estimation algorithm, and determine a second image key frame according to the recognition processing result; The fusion and screening submodule is used to perform fusion and screening on the first image key frame according to the second image key frame, and use the first image key frame obtained after screening as the determined image key frame.

10. A cockpit, characterized in that: Comprising a cockpit image detection device as described in any one of claims 7 to 9.

11. A vehicle, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 6.

12. A computer-readable storage medium having executable codes stored thereon, which, when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1 to 6.