Behavior detection method and device, vehicle, electronic equipment and storage medium

By combining target object and region information in the image for feature fusion processing, the problem of insufficient accuracy in behavior detection in vehicle driving scenarios is solved, and more efficient behavior detection is achieved.

CN116863448BActive Publication Date: 2026-05-08BEIJING CO WHEELS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CO WHEELS TECH CO LTD
Filing Date
2022-03-21
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In vehicle driving scenarios, existing behavior detection methods have poor accuracy, resulting in poor detection of non-standard behaviors.

Method used

Behavior detection is performed by combining target object information and image region information in the image to be detected, identifying the features of target objects and image regions in the image to be detected, and determining human behavior information by combining feature fusion processing.

Benefits of technology

It improves the accuracy and efficiency of behavior detection, and enhances the robustness and applicability of behavior detection methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116863448B_ABST
    Figure CN116863448B_ABST
Patent Text Reader

Abstract

The present disclosure provides a behavior detection method and device, a vehicle, an electronic device and a storage medium. The method comprises: obtaining a to-be-detected image, identifying target object information and image region information in the to-be-detected image, and determining human body behavior information according to the target object information and the image region information. According to the present disclosure, the to-be-detected image is detected by combining the target object information and the image region information in the to-be-detected image, which can effectively improve the accuracy and detection efficiency of behavior detection and recognition, thereby effectively improving the robustness and applicability of the behavior detection method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a behavior detection method, device, vehicle, electronic device, and storage medium. Background Technology

[0002] In vehicle driving scenarios, regulating driver behavior is an important aspect of addressing vehicle safety hazards. Therefore, there is an urgent need to propose a behavior detection method to detect non-standard driver behaviors (such as smoking, drinking water, and making or receiving phone calls) in vehicle driving scenarios, thereby regulating driver behavior accordingly.

[0003] In related technologies, the accuracy of behavior detection is poor, and the behavior detection effect is not good. Summary of the Invention

[0004] This disclosure aims to at least partially address one of the technical problems in the related art.

[0005] Therefore, the purpose of this disclosure is to propose a behavior detection method, device, vehicle, electronic device, and storage medium. Since it combines the target object information and image region information in the image to be detected to perform behavior detection on the image to be detected, it can effectively improve the accuracy and efficiency of behavior detection and recognition, thereby effectively improving the robustness and applicability of the behavior detection method.

[0006] The behavior detection method proposed in the first aspect of this disclosure includes: acquiring an image to be detected; identifying target object information and image region information in the image to be detected; and determining human behavior information based on the target object information and image region information.

[0007] The behavior detection method proposed in the first aspect of this disclosure acquires an image to be detected, identifies target object information and image region information in the image, and determines human behavior information based on the target object information and image region information. Therefore, since behavior detection is performed on the image by combining the target object information and image region information, the accuracy and efficiency of behavior detection and recognition can be effectively improved, thereby significantly enhancing the robustness and applicability of the behavior detection method.

[0008] The behavior detection device proposed in the second aspect of this disclosure includes: an acquisition module for acquiring an image to be detected; an identification module for identifying target object information and image region information in the image to be detected; and a determination module for determining human behavior information based on the target object information and image region information.

[0009] The behavior detection apparatus proposed in the second aspect of this disclosure acquires an image to be detected, identifies target object information and image region information in the image, and determines human behavior information based on the target object information and image region information. Therefore, since behavior detection is performed on the image by combining the target object information and image region information, the accuracy and efficiency of behavior detection and recognition can be effectively improved, thereby significantly enhancing the robustness and applicability of the behavior detection method.

[0010] According to a third aspect of this disclosure, a vehicle is provided, including: a behavior detection device as described in the embodiments of the second aspect of this disclosure.

[0011] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the behavior detection method of the first aspect of this disclosure.

[0012] According to a fifth aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, the computer instructions being used to cause a computer to execute the behavior detection method of the first aspect of this disclosure.

[0013] According to a sixth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the behavior detection method of the first aspect of this disclosure.

[0014] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0015] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which:

[0016] Figure 1 This is a schematic flowchart of a behavior detection method proposed in an embodiment of this disclosure;

[0017] Figure 2 This is a schematic flowchart of a behavior detection method proposed in another embodiment of this disclosure;

[0018] Figure 3 This is a schematic flowchart of a behavior detection method proposed in another embodiment of this disclosure;

[0019] Figure 4 This is a schematic diagram of the structure of a behavior detection device according to an embodiment of this disclosure;

[0020] Figure 5 This is a schematic diagram of the structure of a behavior detection device according to another embodiment of this disclosure;

[0021] Figure 6 This is a schematic diagram of the structure of a vehicle according to an embodiment of the present disclosure;

[0022] Figure 7 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0023] Embodiments of this disclosure are described in detail below, with examples of embodiments illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are used only to explain this disclosure, and should not be construed as limiting this disclosure. Rather, embodiments of this disclosure include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.

[0024] Figure 1 This is a schematic flowchart of a behavior detection method proposed in one embodiment of this disclosure.

[0025] It should be noted that the execution subject of the behavior detection method in this embodiment is a behavior detection device, which can be implemented by software and / or hardware. The device can be configured in an electronic device, which may include, but is not limited to, a terminal, a server, etc.

[0026] like Figure 1 As shown, the behavior detection method includes:

[0027] S101: Acquire the image to be detected.

[0028] It should be noted that the images to be detected in the embodiments of this disclosure are obtained after authorization by the relevant users, and the process complies with the provisions of relevant laws and regulations and does not violate public order and good morals.

[0029] The image for which behavior detection is to be performed can be referred to as the image to be detected. The image to be detected can be a face image captured from a face, or it can be an image captured from a part of the human body (e.g., the hand, without limitation).

[0030] In this embodiment of the disclosure, the number of images to be tested can be one or more. The images to be tested can be, for example, a single image or an image corresponding to a video frame in a video. The images to be tested can also be two-dimensional images or three-dimensional images, and there are no limitations on this.

[0031] In some embodiments, the image to be detected can be obtained by using a camera device with video recording function, such as a mobile phone or camera, to capture a face image. Alternatively, it can be obtained by parsing a video stream to obtain multiple video frame images, selecting a video frame image containing a face from the multiple video frame images, and using this video frame image as the image to be detected. Then, subsequent behavior detection methods can be performed based on the aforementioned obtained image to be detected. There are no restrictions on this.

[0032] S102: Identify target object information and image region information in the image to be detected.

[0033] In this embodiment of the disclosure, when performing behavior detection on the aforementioned obtained image to be detected, the image to be detected can be divided into multiple image regions. Taking the image to be detected as a face image as an example, the image to be detected can be divided into face regions (e.g., eye regions, mouth regions) and object regions (e.g., glasses regions, straw regions when drinking water). There are no restrictions on this.

[0034] The information used to describe the face region in the image to be detected can be called image region information. This image region information can specifically be, for example, image region pixel information, image region feature information, etc., without limitation.

[0035] The information used to describe the target object in the image to be detected can be called target object information. Specifically, the target object information can be, for example, the position information of the target object relative to the image to be detected, the feature information of the target object, etc., without limitation.

[0036] In some embodiments, identifying target object information and image region information in the image to be detected can be achieved by using a feature parsing method (e.g., Local Binary Patterns (LBP) algorithm, which is not limited to this method) to perform feature recognition on the aforementioned acquired image to be detected, so as to identify multiple feature information from the image to be detected. Then, the multiple feature information obtained can be parsed and processed, and the information used to describe the object among the multiple feature information can be used as target object information, and the information used to describe the face among the multiple feature information can be used as image region information, which is not limited to this method.

[0037] Alternatively, any other possible method can be used to identify target object information and image region information in the image to be detected, such as model recognition, convolutional neural network recognition, etc., without any restrictions.

[0038] S103: Determine human behavior information based on target object information and image region information.

[0039] Information used to describe human behavior can be called human behavior information. This human behavior information can specifically include, for example, information on the category of human behavior, information on the evaluation of human behavior, etc., without any limitation.

[0040] In this embodiment of the disclosure, after acquiring the image to be detected and identifying the target object information and image region information in the image to be detected, human behavior information can be determined by combining the target object information and image region information.

[0041] In some embodiments, determining human behavior information based on target object information and image region information can be achieved by first determining the target object information and image region information, then performing a fusion process on the target object information and image region information to obtain a corresponding fusion result, and then determining the human behavior information based on the fusion result.

[0042] For example, the target object information (target object features) and image region information (image region features) obtained above can be subjected to feature fusion processing to obtain corresponding fused features. Then, human behavior information can be determined by combining the fused features, without any restrictions.

[0043] In other embodiments, target object information and image region information can be used as dual references to determine human behavior information. For example, the behavior that may occur in the image region and the object information corresponding to the behavior can be determined first based on the image region information. Then, it can be determined whether the aforementioned target object information matches the behavior in the image region and the object information corresponding to the behavior. When the target object information matches the behavior in the image region and the object information corresponding to the behavior, the behavior corresponding to the object information is taken as human behavior information, without any limitation.

[0044] For example, assuming the image region information identified from the image to be detected is the mouth image region information, it can be determined that the behavior that may occur in the mouth image region could be smoking, drinking water, etc., and the object information corresponding to smoking behavior is determined to be cigarette, and the object information corresponding to drinking behavior is water cup. Then, the previously identified target object information can be compared with cigarette or water cup, and when the target object information is determined to be cigarette, it is determined that smoking behavior exists, and when the target object information is determined to be water cup, it is determined that drinking behavior exists, without any restrictions.

[0045] In this embodiment, by acquiring the image to be detected and identifying the target object information and image region information in the image, and determining human behavior information based on the target object information and image region information, the method effectively improves the accuracy and efficiency of behavior detection and recognition by combining the target object information and image region information in the image to be detected. This significantly enhances the robustness and applicability of the behavior detection method.

[0046] Figure 2 This is a flowchart illustrating a behavior detection method proposed in another embodiment of this disclosure.

[0047] like Figure 2 As shown, the behavior detection method includes:

[0048] S201: Acquire an initial image, wherein the initial image is captured by the vehicle's camera device and the initial image corresponds to the driving area annotation information.

[0049] In the initial stage of the behavior detection method, the image obtained for behavior detection can be referred to as the initial image.

[0050] In this embodiment of the disclosure, the initial image may be an image captured by a camera device of the vehicle, which may be pre-installed in the vehicle or integrated with the vehicle.

[0051] In this embodiment of the disclosure, when the vehicle's camera device acquires the initial image, in order to effectively reduce the amount of data in subsequent image processing, corresponding annotation information can be pre-set for the initial image. For example, a corresponding driving area can be pre-set for the initial image. Then, after the initial image is captured by the vehicle's camera device, subsequent behavior detection methods can be performed on the image within the driving area. This can effectively reduce the amount of data in subsequent image processing, effectively save computing resources, and effectively help improve behavior detection efficiency.

[0052] The information used to describe the driving area pre-set for the initial image can be called driving area standard information. This driving area labeling information can specifically be, for example, the location standard information of the driving area, without limitation.

[0053] In other words, a specific application scenario of the behavior detection method described in this disclosure embodiment can be, for example, in a vehicle driving scenario, using the vehicle's camera device to acquire an initial image, and then combining it with preset vehicle driving area annotation information to determine the image to be detected from the initial image, and performing behavior detection on the image to be detected to determine whether the driver has violated driving regulations (e.g., smoking, making or receiving phone calls, without limitation).

[0054] It should be noted that the explanation of the embodiments of this disclosure can be based on the above-described vehicle driving application scenario. In addition, the embodiments of this disclosure can also be applied to any other possible behavior detection application scenario, and there are no limitations on this.

[0055] S202: Input the initial image into the face detection model, and the face detection model performs face detection on the initial image to obtain the face detection result.

[0056] Among them, the face detection model can be used to perform face recognition detection on the initial image and output the corresponding detection results, which can be called face detection results.

[0057] The face detection model can be an artificial intelligence model, such as a neural network model or a machine learning model, or any other possible model capable of performing face detection tasks. There are no restrictions on this.

[0058] In other words, after obtaining the initial image, the embodiments of this disclosure can input the initial image into the face detection model, which will then perform face detection on the initial image and output the corresponding face detection results. Subsequently, the image to be detected can be obtained based on the face detection results. For details, please refer to the following embodiments.

[0059] S203: If the face detection result indicates that a face exists in the initial image, then the image to be detected is determined from the initial image based on the driving area annotation information.

[0060] It is understandable that the initial image obtained during the execution of the behavior detection method may not always contain valid image information. Taking the behavior detection method as an example in the above-mentioned vehicle driving scenario, when the vehicle's camera device acquires the initial image, there may be cases where there are blanks in the data, such as when the driver is not in the car or when the driver is leaning over. In such cases, the acquired initial image cannot represent the driver's behavior information. If the initial image is processed at this time, there may be a significant waste of resources. Therefore, after acquiring the initial image, face detection can be performed on the initial image to ensure that the subsequent images to be detected can contain valid face information.

[0061] In this embodiment of the disclosure, after inputting an initial image into a face detection model and performing face detection on the initial image to obtain the face detection result, subsequent steps can be triggered according to the indication of the face detection result.

[0062] In this embodiment of the disclosure, if the face detection result indicates that there is a face in the initial image, it means that the initial image can support subsequent behavior detection methods. At this time, the initial image can be segmented according to the driving area annotation information to determine the image to be detected from the initial image. There are no restrictions on this.

[0063] In this embodiment, an initial image is acquired, and then a face detection model is used to perform face detection on the initial image. When the face detection result indicates that a face exists in the initial image, the image to be detected is determined from the initial image by combining the driving area annotation information set for the initial image. Since face detection is performed on the initial image by combining the face detection model, it can be ensured that the image to be detected has valid face information. Therefore, when performing behavior detection on the image to be detected, the effectiveness of the behavior detection operation can be effectively guaranteed, effectively saving the resource waste caused by invalid operations. In addition, since the image to be detected is determined from the initial image by combining the driving area annotation information set for the initial image, the amount of data for subsequent image processing can be effectively reduced while effectively ensuring the integrity of the initial image information, thereby effectively saving computing resources and effectively helping to improve the efficiency of behavior detection.

[0064] S204: Identify target object information and image region information in the image to be detected.

[0065] S205: Determine human behavior information based on target object information and image region information.

[0066] The descriptions of S204-S205 can be found in the above embodiments, and will not be repeated here.

[0067] In this embodiment, an initial image is acquired, and then a face detection model is used to perform face detection on the initial image. When the face detection result indicates that a face exists in the initial image, the image to be detected is determined from the initial image by combining the driving area annotation information set for the initial image. Since face detection is performed on the initial image by combining the face detection model, it can be ensured that the image to be detected has valid face information. Therefore, when performing behavior detection on the image to be detected, the effectiveness of the behavior detection operation can be effectively guaranteed, effectively saving the resource waste caused by invalid operations. In addition, since the image to be detected is determined from the initial image by combining the driving area annotation information set for the initial image, the amount of data for subsequent image processing can be effectively reduced while effectively ensuring the integrity of the initial image information, thereby effectively saving computing resources and effectively assisting in improving the efficiency of behavior detection. Then, the target object information and image region information in the image to be detected are identified, and human behavior information is determined based on the target object information and image region information. This can improve the accuracy and efficiency of behavior detection and recognition in vehicle driving scenarios, thereby effectively meeting the application requirements of behavior detection and recognition in vehicle driving scenarios and effectively improving the robustness and applicability of behavior detection methods.

[0068] Figure 3 This is a flowchart illustrating a behavior detection method proposed in another embodiment of this disclosure.

[0069] like Figure 3 As shown, the behavior detection method includes:

[0070] S301: Acquire the image to be detected.

[0071] For a detailed description of S301, please refer to the above embodiments, which will not be repeated here.

[0072] S302: Identify the target object region and part image region in the image to be detected.

[0073] The image region corresponding to the target object in the image to be detected can be called the target object region. For example, the target object region can be the image region corresponding to the cigarette when smoking, or the image region corresponding to the phone when answering a call. Similarly, the image region corresponding to the face in the image to be detected can be called the part image region. For example, the part image region can be the mouth region or the eye region, without limitation.

[0074] In some embodiments, identifying the target object region and part image region in the image to be detected can be achieved by performing image segmentation processing on the image to be detected based on the target object information and image region information after determining the target object information and image region information in the image to be detected, so as to obtain the target object region and part image region. Alternatively, any other possible method can be used to identify the target object region and part image region in the image to be detected, such as detection box recognition method, image recognition method, etc., without limitation.

[0075] Optionally, in some embodiments, identifying the target object region in the image to be detected can be achieved by inputting the image to be detected into a pre-trained object region detection model, which then performs object region detection on the image to be detected to obtain the target object region. Since the object region detection model is used to perform object region detection on the image to be detected, the interference of other subjective factors on object region detection can be effectively reduced during the object region detection process, thereby effectively improving both the efficiency and accuracy of object region detection.

[0076] Among them, the object region detection model can be used to detect object regions in the image to be detected. The object region detection model can be an artificial intelligence model, such as a neural network model or a machine learning model, or any other possible model that can perform the object region detection task. There are no restrictions on this.

[0077] In other words, in this embodiment of the present disclosure, identifying the target object region in the image to be detected can be achieved by inputting the image to be detected into an object region detection model, which then performs object region detection on the image to be detected and outputs the target object region.

[0078] Optionally, in some embodiments, identifying the part image region in the image to be detected may involve inputting the image to be detected into a pre-trained keypoint detection model, which performs keypoint detection on the image to be detected to obtain multiple keypoint information. Then, based on the multiple keypoint information, the image to be detected is cropped to obtain the part image region. Since keypoint detection is performed on the image to be detected using a keypoint detection model, the robustness and applicability of keypoint detection can be effectively improved, as can the accuracy of keypoint detection. Therefore, when cropping the image to be detected based on the keypoint information obtained from keypoint detection, the cropped part image region can have better accuracy, effectively improving the recognition effect of the part image region.

[0079] Specifically, these can be key points that can be used to characterize parts of the human body, such as the center point of the eye. Correspondingly, the information used to describe these key points can be called key point information. This key point information can be specifically, for example, the position coordinates of the key points relative to the image to be detected, without any limitation.

[0080] The pre-trained model used to identify key points in the image to be detected can be called a key point detection model. This key point detection model can be an artificial intelligence model, such as a neural network model or a machine learning model, or any other possible model capable of performing key point detection tasks. There are no restrictions on this.

[0081] In other words, after acquiring the image to be detected, the embodiments of this disclosure can input the image to be detected into a pre-trained key point detection model, which will then perform key point detection on the image to be detected and output the corresponding key point information. Afterward, the image to be detected can be cropped by combining the key point information output by the key point detection model to obtain the part image region.

[0082] S303: Determine the target object information based on the target object region.

[0083] In this embodiment of the present disclosure, after determining the target object region from the image to be detected, the target object information corresponding to the target object region can be determined based on the target object region.

[0084] In some embodiments, determining target object information based on the target object region can be achieved by parsing the target object region after it has been determined, in order to determine the target object information corresponding to the target object region.

[0085] For example, parsing the target object region can be done by performing feature parsing on the target object region after determining the target object region, and using the aforementioned feature parsing results as the target object information; or, parsing the target object region can be done by performing model parsing on the target object region after determining the target object region, and using the aforementioned model parsing results as the target object information. There is no limitation on this.

[0086] Optionally, in some embodiments, determining the target object information based on the target object region can also involve obtaining the region category information of the target object region and using the region category information as the target object information. Since the region category information of the target object region is used as the target object information, sufficient reference information can be provided for the execution process of the subsequent behavior detection method. This enables the complex target object regions to be organized into target object regions with region category information, thereby effectively simplifying the operation logic of subsequent behavior detection based on the region category information, and thus effectively improving the efficiency of behavior detection.

[0087] In this embodiment of the disclosure, the target object region can be divided into multiple region categories according to the type of target object corresponding to the target object region. For example, the region category corresponding to smoke, the region category corresponding to telephone, and the information used to describe the region category can be called region category information. The region category information can be specifically, for example, the position information of the region corresponding to smoke relative to the image to be detected, without limitation.

[0088] In this embodiment of the disclosure, the target object information is determined based on the target object region. Alternatively, after determining the target object region, the region category information corresponding to each target object region can be determined, and the region category information can be used as the target object information. That is, the target object region can be input into a pre-trained region category determination model, the region category determination model can classify the target object region, and output the region category information corresponding to the target object region, and the region category information can be used as the target object information. There are no restrictions on this.

[0089] S304: Determine image region information based on the image region of the affected area.

[0090] In this embodiment of the present disclosure, after determining the part image region from the image to be detected, image region information corresponding to the part image region can be determined based on the part image region.

[0091] In some embodiments, determining image region information based on the part image region can be achieved by parsing the part image region after it has been determined, in order to determine the image region information corresponding to the part image region.

[0092] For example, parsing a part of an image region can be performed by first determining the part of the image region, then performing feature parsing on the part of the image region, and using the result of the feature parsing as the image region information; or, parsing a part of an image region can be performed by first determining the part of the image region, then performing model parsing on the part of the image region, and using the result of the model parsing as the image region information. There are no restrictions on which approach is being taken.

[0093] Optionally, in some embodiments, determining image region information based on the part image region can be achieved by determining the part category based on the part image region and using the part category as the image region information. Since the part category corresponding to the part image region is used as the image region information, sufficient reference information can be provided for the subsequent execution process of the behavior detection method. This enables the complex part image regions to be organized into part image regions with corresponding part categories, thereby effectively simplifying the operation logic of subsequent behavior detection based on the part category and thus effectively improving the efficiency of behavior detection.

[0094] In this embodiment of the disclosure, the images can be divided into multiple body part categories. For example, the images of the left eye and the right eye can be classified together as the eye category, and there is no limitation on this.

[0095] In other words, in this embodiment of the present disclosure, after determining the part image region, the part image corresponding to the part image region may be classified to determine the part category corresponding to the part image region, and the part category may be used as the image region information, without limitation.

[0096] S305: Input the target object information and image region information into the behavior information classification model.

[0097] The information classification model can support behavioral information classification processing of target object information and image region information. Specifically, the behavioral information classification model can be, for example, a secondary classification model, or it can be configured as any other possible model capable of performing behavioral information classification operations, without any restrictions.

[0098] S306: Based on the behavior information classification model, determine whether the matching relationship between the target object information and the image region information and the candidate association relationship meet the set conditions.

[0099] Among them, there can be a corresponding relationship between the target object information and the image region information, and this relationship can be called the matching relationship.

[0100] Among them, the pre-set constraints on the matching relationship between target object information and image region information, and the candidate association relationship, can be called the setting conditions. These setting conditions can be adaptively configured in combination with the behavior detection requirements in actual business scenarios, and there are no restrictions on them.

[0101] In this embodiment of the disclosure, based on the behavior information classification model, determining whether the matching relationship between the target object information and the image region information and the candidate association relationship meet the set conditions can be achieved by inputting the target object information and the image region information into the behavior information classification model to obtain the matching relationship between the target object information and the image region information output by the behavior information classification model. Then, the matching relationship and the candidate association relationship can be compared to determine whether the matching relationship between the target object information and the image region information and the candidate association relationship meet the set conditions.

[0102] For example, it can be to determine the similarity between the relationship to be matched and the candidate relationship, compare the similarity with a pre-set similarity threshold, and determine that the relationship between the target object information and the image region information meets the set conditions when the similarity is greater than the similarity threshold, without any restrictions.

[0103] In this embodiment of the present disclosure, when the set conditions are met between the matching relationship and the candidate association relationship between the target object information and the image region information, the candidate behavior information can be used as human behavior information. Since the set conditions are met by combining the behavior information classification model, the interference caused by other subjective factors can be effectively reduced, and the timing of the satisfaction of the set conditions can be accurately determined, thereby effectively improving the detection and recognition effect of human behavior information.

[0104] S307: When the candidate association relationship between the candidate object information and the candidate region information corresponding to the candidate behavior information meets the set conditions, the candidate behavior information is regarded as human behavior information.

[0105] The candidate behavior information may include a variety of behavior information. That is, the behavior detection method described in this embodiment can combine target object information and image region information to determine the candidate behavior information corresponding to the image to be detected from the candidate behavior information, and detect the human behavior information obtained therefrom.

[0106] Among them, candidate behavior information can have corresponding object information, which can be called candidate object information, and candidate behavior information can have corresponding region information, which can be called candidate region information.

[0107] Among them, there can be a corresponding relationship between candidate object information and candidate region information. This relationship can be called a candidate relationship. The candidate relationship can be specifically, for example, a semantic relationship, a feature relationship, etc., without any restrictions.

[0108] In other words, in this embodiment of the present disclosure, before the behavior detection method starts to execute, it is possible to pre-configure multiple candidate behavior information and parse and process the multiple candidate behavior information respectively to determine the candidate object information and candidate region information corresponding to the multiple candidate behavior information. Then, based on the candidate object information and candidate region information, the human behavior information corresponding to the image to be detected can be determined. For details, please refer to the following embodiments.

[0109] In this embodiment of the disclosure, when the set conditions are met between the matching relationship between the target object information and the image region information and the candidate association relationship, the candidate behavior information is used as human behavior information. This can effectively narrow the search range of human behavior information, thereby enabling the rapid determination of human behavior information from candidate behavior information and thus effectively improving the determination efficiency of human behavior information.

[0110] In this embodiment, by acquiring the image to be detected and identifying the target object region and part image region in the image, and then determining the target object information based on the target object region, sufficient reference information can be provided for the subsequent execution process of the behavior detection method. This process organizes the complex target object regions into target object regions with region category information, thereby effectively simplifying the subsequent behavior detection operation logic based on the region category information, and thus effectively improving the behavior detection efficiency. Furthermore, by determining the part image region, sufficient reference information can be provided for the subsequent execution process of the behavior detection method, and the complex part image regions can be organized into part image regions with corresponding part categories, thereby effectively simplifying the subsequent behavior detection operation logic based on the part category. The operational logic of behavior detection effectively improves behavior detection efficiency. By combining a behavior information classification model, it determines whether the matching relationship between the target object information and the image region information meets the set conditions. This effectively reduces the interference caused by other subjective factors and accurately determines when the set conditions are met, thus effectively improving the detection and recognition effect of human behavior information. Since the candidate behavior information is used as human behavior information when the matching relationship between the target object information and the image region information meets the set conditions, the search range for human behavior information can be effectively narrowed. This enables the rapid determination of human behavior information from candidate behavior information, thereby effectively improving the determination efficiency of human behavior information.

[0111] Figure 4 This is a schematic diagram of the structure of a behavior detection device according to an embodiment of this disclosure.

[0112] like Figure 4 As shown, the behavior detection device 40 includes:

[0113] The acquisition module 401 is used to acquire the image to be detected;

[0114] Recognition module 402 is used to identify target object information and image region information in the image to be detected; and

[0115] The determination module 403 is used to determine human behavior information based on the target object information and image region information.

[0116] In some embodiments of this disclosure, such as Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of a behavior detection device according to another embodiment of the present disclosure, wherein the recognition module 402 includes:

[0117] The recognition submodule 4021 is used to identify the target object region and part image region in the image to be detected;

[0118] The first determining submodule 4022 is used to determine the target object information based on the target object region;

[0119] The second determining submodule 4023 is used to determine image region information based on the image region of the part.

[0120] In some embodiments of this disclosure, the first determining submodule 4022 is further configured to:

[0121] Obtain the region category information of the target object region;

[0122] Use region category information as target object information.

[0123] In some embodiments of this disclosure, the second determining submodule 4023 is further configured to:

[0124] Determine the body part category based on the image region;

[0125] Use body part category as image region information.

[0126] In some embodiments of this disclosure, the determining module 403 is further configured to:

[0127] When the candidate association relationship between the candidate object information and the candidate region information corresponding to the candidate behavior information meets the set conditions, the candidate behavior information is regarded as human behavior information.

[0128] Among them, target association relationship is the association relationship between target object information and image region information.

[0129] In some embodiments of this disclosure, the determining module 403 is further configured to:

[0130] The target object information and image region information are input into the behavior information classification model;

[0131] Based on the behavioral information classification model, the matching relationship between target object information and image region information and the candidate association relationship are determined to meet the set conditions.

[0132] In some embodiments of this disclosure, the identification submodule 4021 is further configured to:

[0133] The image to be detected is input into a pre-trained object region detection model, which then performs object region detection on the image to obtain the target object region.

[0134] In some embodiments of this disclosure, the identification submodule 4021 is further configured to:

[0135] The image to be detected is input into a pre-trained keypoint detection model, which then performs keypoint detection on the image to obtain multiple keypoint information.

[0136] The image to be detected is cropped based on information from multiple key points to obtain the local image region.

[0137] In some embodiments of this disclosure, the acquisition module 401 is further configured to:

[0138] Acquire an initial image, which is captured by the vehicle's camera device and corresponds to the driving area annotation information;

[0139] The initial image is input into the face detection model, which then performs face detection on the initial image to obtain the face detection result.

[0140] If the face detection result indicates that a face exists in the initial image, then the image to be detected is determined from the initial image based on the driving area annotation information.

[0141] With the above Figures 1 to 3 Corresponding to the behavior detection method provided in the embodiments, this disclosure also provides a behavior detection device. Because the behavior detection device provided in the embodiments of this disclosure is similar to the one described above… Figures 1 to 3 The behavior detection method provided in the embodiments corresponds to the behavior detection method provided in the embodiments of this disclosure, and therefore the implementation of the behavior detection method is also applicable to the behavior detection device provided in the embodiments of this disclosure, and will not be described in detail in the embodiments of this disclosure.

[0142] In this embodiment, by acquiring the image to be detected and identifying the target object information and image region information in the image, and determining human behavior information based on the target object information and image region information, the method effectively improves the accuracy and efficiency of behavior detection and recognition by combining the target object information and image region information in the image to be detected. This significantly enhances the robustness and applicability of the behavior detection method.

[0143] Figure 6 This is a schematic diagram of the structure of a vehicle according to an embodiment of the present disclosure.

[0144] like Figure 6 As shown, the vehicle 60 includes: the behavior detection device 40 in the above embodiment.

[0145] In this embodiment, by acquiring the image to be detected and identifying the target object information and image region information in the image, and determining human behavior information based on the target object information and image region information, the method effectively improves the accuracy and efficiency of behavior detection and recognition by combining the target object information and image region information in the image to be detected. This significantly enhances the robustness and applicability of the behavior detection method.

[0146] To implement the above embodiments, this disclosure also proposes a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the behavior detection method proposed in the foregoing embodiments of this disclosure.

[0147] To implement the above embodiments, this disclosure also proposes an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the behavior detection method proposed in the foregoing embodiments of this disclosure.

[0148] To implement the above embodiments, this disclosure also proposes a computer program product that, when the instruction processor in the computer program product is executed, performs the behavior detection method as proposed in the foregoing embodiments of this disclosure.

[0149] Figure 7 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. Figure 7 The electronic device 12 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0150] like Figure 7As shown, electronic device 12 is represented in the form of a general-purpose computing device. Components of electronic device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processing unit 16). Bus 18 represents one or more of several bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus structures.

[0151] For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnection (PCI) bus.

[0152] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0153] Memory 28 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 30 and / or cache memory 32. Electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 7 Not shown; usually referred to as a "hard drive".

[0154] although Figure 7Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disc drive for reading and writing to a removable non-volatile optical disc (e.g., a compact disc read-only memory (CD-ROM), a digital video disc read-only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.

[0155] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of this disclosure.

[0156] Electronic device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with electronic device 12, and / or with any device that enables electronic device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, electronic device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of electronic device 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0157] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the behavior detection method mentioned in the foregoing embodiments.

[0158] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0159] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

[0160] It should be noted that in the description of this disclosure, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this disclosure, unless otherwise stated, "a plurality of" means two or more.

[0161] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.

[0162] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0163] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0164] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0165] The storage media mentioned above can be read-only memory, disk, or optical disk, etc.

[0166] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0167] Although embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A behavior detection method, characterized in that, include: Acquire the image to be detected; Identify target object information and image region information in the image to be detected; as well as Based on the target object information and the image region information, human behavior information is determined; The method further includes: The target object information and the image region information are input into the behavior information classification model; Based on the behavioral information classification model, it is determined that the unmatched relationship and the candidate association relationship between the target object information and the image region information meet the set conditions. Wherein, the relationship to be matched is the association between the target object information and the image region information; The step of determining human behavior information based on the target object information and the image region information includes: When the candidate association relationship between the candidate object information and the candidate region information corresponding to the candidate behavior information meets the set conditions, the candidate behavior information is used as the human behavior information. The target association relationship is the association relationship between the target object information and the image region information.

2. The method as described in claim 1, characterized in that, The identification of target object information and image region information in the image to be detected includes: Identify the target object region and part image region in the image to be detected; The target object information is determined based on the target object region; Based on the image region of the described part, the image region information is determined.

3. The method as described in claim 2, characterized in that, The step of determining the target object information based on the target object region includes: Obtain the region category information of the target object region; The region category information is used as the target object information.

4. The method as described in claim 2, characterized in that, The step of determining the image region information based on the image region of the affected area includes: Based on the image region of the described part, determine the part category; The category of the body part is used as the image region information.

5. The method as described in claim 2, characterized in that, The process of identifying the target object region in the image to be detected includes: The image to be detected is input into a pre-trained object region detection model, which then performs object region detection on the image to obtain the target object region.

6. The method as described in claim 2, characterized in that, The identification of the part image region in the image to be detected includes: The image to be detected is input into a pre-trained keypoint detection model, which performs keypoint detection on the image to obtain multiple keypoint information. The image to be detected is cropped based on the information of the multiple key points to obtain the image region of the specified location.

7. The method as described in claim 1, characterized in that, The acquisition of the image to be detected includes: Acquire an initial image, wherein the initial image is captured by the vehicle's camera device and the initial image corresponds to driving area annotation information; The initial image is input into the face detection model, which performs face detection on the initial image to obtain the face detection result. If the face detection result indicates that a face exists in the initial image, then the image to be detected is determined from the initial image based on the driving area annotation information.

8. A behavior detection device, characterized in that, The device includes: The acquisition module is used to acquire the image to be detected; The recognition module is used to identify target object information and image region information in the image to be detected; and The determination module is used to determine human behavior information based on the target object information and the image region information; The determining module is further configured to input the target object information and the image region information into the behavior information classification model; Based on the behavioral information classification model, it is determined that the unmatched relationship and the candidate association relationship between the target object information and the image region information meet the set conditions. Wherein, the relationship to be matched is the association between the target object information and the image region information; The determining module is further configured to: When the candidate association relationship between the candidate object information and the candidate region information corresponding to the candidate behavior information meets the set conditions, the candidate behavior information is used as the human behavior information. The target association relationship is the association relationship between the target object information and the image region information.

9. The apparatus as claimed in claim 8, characterized in that, The identification module includes: The recognition submodule is used to identify the target object region and part image region in the image to be detected; The first determining submodule is used to determine the target object information based on the target object region; The second determining submodule is used to determine the image region information based on the image region of the said part.

10. The apparatus as claimed in claim 9, characterized in that, The first determining submodule is further configured to: Obtain the region category information of the target object region; The region category information is used as the target object information.

11. The apparatus as claimed in claim 9, characterized in that, The second determining submodule is further configured to: Based on the image region of the described part, determine the part category; The category of the body part is used as the image region information.

12. The apparatus as claimed in claim 9, characterized in that, The identification submodule is also used for: The image to be detected is input into a pre-trained object region detection model, which then performs object region detection on the image to obtain the target object region.

13. The apparatus as claimed in claim 9, characterized in that, The identification submodule is also used for: The image to be detected is input into a pre-trained keypoint detection model, which performs keypoint detection on the image to obtain multiple keypoint information. The image to be detected is cropped based on the information of the multiple key points to obtain the image region of the specified location.

14. The apparatus as claimed in claim 8, characterized in that, The acquisition module is also used for: Acquire an initial image, wherein the initial image is captured by the vehicle's camera device and the initial image corresponds to driving area annotation information; The initial image is input into the face detection model, which performs face detection on the initial image to obtain the face detection result. If the face detection result indicates that a face exists in the initial image, then the image to be detected is determined from the initial image based on the driving area annotation information.

15. A vehicle, characterized in that, include: The behavior detection device as described in any one of claims 8-14 above.

16. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the behavior detection method according to any one of claims 1-7.

17. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, in, The computer instructions are used to cause the computer to execute the behavior detection method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Driving behavior detection method and device and readable storage medium

    CN112183356A

  • Abnormal driving behavior identification and early warning method and electronic equipment

    CN112613441A