A behavior recognition method and an electronic device

By adjusting the focal length and image acquisition equipment according to the target distance, the problem of inaccurate smoking behavior recognition was solved, achieving high accuracy and fast behavior recognition, and reducing security risks.

CN116844223BActive Publication Date: 2026-05-26HISENSE GRP HLDG CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HISENSE GRP HLDG CO LTD
Filing Date
2023-03-31
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies for identifying smoking behavior are inaccurate, leading to safety hazards, especially in smoke-free areas where it is difficult to quickly and accurately identify smokers, resulting in false positives and false negatives.

Method used

By acquiring the image to be identified and its disparity map from the acquisition device, the target detection model is used to determine the area of ​​the target to be detected. The target distance is calculated based on the disparity map and the parameters of the acquisition device. The focal length of the acquisition device is adjusted to obtain a clear target image, and then the target is identified in combination with the behavior recognition model.

Benefits of technology

It improves the accuracy of identifying smoking behavior, reduces safety hazards, lightens the workload of staff, and promptly stops potential dangers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844223B_ABST
    Figure CN116844223B_ABST
Patent Text Reader

Abstract

This application provides a behavior recognition method and electronic device to address the problem of inaccurate recognition of smoking behavior in existing technologies, which poses a safety hazard. In this application embodiment, after acquiring an image to be recognized, the electronic device determines the target distance between the target in the image and the acquisition device, detects targets related to the target behavior, and adjusts the focal length of the acquisition device according to the target distance. This results in a clearer image of the target, with its enlarged area within the image. Behavior recognition is then performed based on this target image, accurately identifying the presence of the corresponding behavior and reducing safety risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of small target detection technology, and in particular to a behavior recognition method and electronic device. Background Technology

[0002] Smoking in certain no-smoking areas can lead to the combustion or explosion of flammable and explosive materials, causing major social accidents. These no-smoking areas include subway stations, gas stations, and airports. For example, the open flame from smoking in a gas station coming into contact with the gasoline could potentially cause an explosion. To reduce the risks associated with smoking, related technologies have proposed installing smoke detectors in no-smoking areas, such as numerous smoke detectors in subway stations. Smoke from smoking in a subway station could trigger these detectors. Once the detected smoke concentration reaches the detector's alarm threshold, it will send an alarm signal. However, by this time, a major accident may have already occurred. Even if no major accident occurs, subway station staff cannot quickly pinpoint the cause of the alarm, wasting considerable time investigating the problem. This increases the workload of staff and disrupts citizens' travel.

[0003] To detect smokers, related technologies propose inputting captured images into a recognition model to detect the presence of smokers. However, images captured by the acquisition device typically contain many people, and for pedestrians at a distance, the area where the pedestrian is located in the image is not clear enough, leading to inaccurate recognition results. Furthermore, cigarette butts and smoke have even smaller pixels, making it difficult to improve the accuracy of the recognition model. In addition, due to the small size of cigarette butts, the recognition model may mistakenly identify drinking straws, lollipops, etc., as cigarettes. To further improve detection accuracy, related technologies propose using cascaded matching to identify the presence of smokers. Specifically, pedestrians who are likely to smoke are first screened, and then smoking detection is performed on the screened individuals. While this reduces false positives to some extent, it still cannot fundamentally solve the problem of the small number of pixels and difficulty in detecting cigarettes. Because the recognition of smoking behavior in these technologies is inaccurate, there are potential safety hazards. Summary of the Invention

[0004] This application provides a behavior recognition method and electronic device to solve the problem that the recognition of smoking behavior in the prior art is inaccurate, which leads to safety hazards.

[0005] In a first aspect, embodiments of this application provide a behavior recognition method, the method comprising:

[0006] Acquire the image to be identified captured by the acquisition device and the disparity map corresponding to the image to be identified;

[0007] Based on the target detection model, the region containing the target in the image to be identified is obtained, wherein the target is a behavior-related target; the target distance between the target and the acquisition device is determined according to the disparity at the target region corresponding to the region in the disparity map and the parameters of the acquisition device.

[0008] Based on the pre-saved correspondence between distance and focal length, the target focal length corresponding to the target distance is determined; the target focal length is sent to the acquisition device so that the acquisition device can acquire a target image with the focal length of the target focal length; if the target image acquired by the acquisition device is received, it is determined whether there is a corresponding recognized behavior in the target image based on the behavior recognition model.

[0009] Secondly, embodiments of this application also provide an electronic device, which includes at least a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the steps of the behavior recognition method as described in any of the preceding claims.

[0010] In this embodiment, the electronic device acquires the image to be identified and the corresponding disparity map captured by the acquisition device. Based on a pre-trained target detection model, it acquires the region containing the target in the image to be identified. According to the distance to the target region corresponding to the target region in the disparity map, it determines the target distance between the target and the acquisition device. Based on a pre-saved correspondence between distance and focal length, it determines the target focal length corresponding to the target distance and sends the target focal length to the acquisition device so that the acquisition device acquires the target image at the target focal length. If the target image acquired by the acquisition device is received, it determines whether there is a corresponding behavior in the target image based on the behavior recognition model. In this embodiment, after acquiring the image to be identified, the electronic device determines the target distance between the target and the acquisition device, detects targets related to target behavior, and adjusts the focal length of the acquisition device according to the target distance. This results in a clearer target image with an enlarged area in the image. Behavior recognition is then performed based on the target image, accurately identifying whether there is a corresponding behavior and reducing security risks. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A schematic diagram of a behavior recognition process provided in an embodiment of this application;

[0013] Figure 2a A schematic diagram of an image to be identified provided in an embodiment of this application;

[0014] Figure 2b A schematic diagram of a disparity map provided in an embodiment of this application;

[0015] Figure 3 This is a schematic diagram of the relevant geometric relationships in a binocular acquisition device provided in an embodiment of this application;

[0016] Figure 4 A schematic diagram illustrating the relationship between distance and parallax is provided for an embodiment of this application;

[0017] Figure 5a A schematic diagram of a region containing a detection target provided in an embodiment of this application;

[0018] Figure 5b A schematic diagram of a target region in a disparity map provided in an embodiment of this application;

[0019] Figure 6 A schematic diagram of a camera coordinate system provided in an embodiment of this application;

[0020] Figure 7a This application provides a schematic diagram of a region in an image to be identified that contains a target for detection.

[0021] Figure 7b A schematic diagram of a target image provided in an embodiment of this application;

[0022] Figures 8a-8d These are schematic diagrams of images acquired by the acquisition devices at different focal lengths according to the embodiments of this application;

[0023] Figure 9 A schematic diagram illustrating the relationship between focal length and the FV of the corresponding acquired image, provided for an embodiment of this application;

[0024] Figure 10a A schematic diagram of an image to be identified provided in an embodiment of this application;

[0025] Figure 10b Another schematic image provided for an embodiment of this application;

[0026] Figure 11 A detailed process diagram of behavior recognition provided for an embodiment of this application;

[0027] Figure 12 This is a schematic diagram of the structure of a behavior recognition device provided in an embodiment of this application;

[0028] Figure 13This is a schematic diagram of an electronic device structure provided in an embodiment of this application. Detailed Implementation

[0029] The present application will now be described in further detail with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present application.

[0030] To improve the accuracy of identifying smoking behavior, embodiments of this application provide a behavior recognition method and an electronic device.

[0031] This behavior recognition method is primarily used for smoking detection in smoke-free scenarios such as rail transit and smart cities. The method includes: an electronic device acquiring an image to be recognized and a corresponding disparity map from a data acquisition device; based on a target detection model, acquiring regions in the image containing the target to be recognized, where the target is a behavior-related target; determining the target distance between the target and the data acquisition device based on the disparity at the target region in the disparity map and the parameters of the data acquisition device; determining the target focal length corresponding to the target distance based on a pre-saved relationship between distance and focal length; sending the target focal length to the data acquisition device so that the data acquisition device can acquire the target image at the target focal length; and if the target image acquired by the data acquisition device is received, determining whether the target image contains a corresponding behavior based on the behavior recognition model, thereby improving the accuracy of behavior recognition.

[0032] Figure 1 A schematic diagram of a behavior recognition process provided in this application embodiment includes the following steps:

[0033] S101: Acquire the image to be identified acquired by the acquisition device and the disparity map corresponding to the image to be identified.

[0034] The behavior recognition method provided in this application is applied to an electronic device, such as a PC or server, which is a smart device.

[0035] In order to perform behavior recognition, electronic devices can acquire images to be recognized from acquisition devices. This can be done by the electronic device controlling the connected acquisition device to acquire the image to be recognized after receiving an alarm signal from a smoke detector, and then sending the acquired image to the electronic device. Alternatively, the acquisition device can acquire the image to be recognized in real time and send it to the electronic device, thus enabling the electronic device to obtain the image to be recognized.

[0036] To perform behavior recognition, the electronic device also acquires a disparity map corresponding to the image to be recognized. Specifically, the acquisition device can be a binocular acquisition device. The electronic device can acquire the image to be recognized and another image acquired by the binocular acquisition device, and jointly determine the disparity map based on the image to be recognized and the other image. The image to be recognized and the other image were acquired by the left (left and right) acquisition device and the right (left and right) acquisition device of the binocular acquisition device, respectively. Which image was acquired by the left (left and right) acquisition device is not limited here. How to determine the disparity map based on the two images acquired by the binocular acquisition device is existing technology and will not be elaborated here. The image to be recognized and the disparity map have the same size.

[0037] Figure 2a This is a schematic diagram of an image to be identified, provided as an embodiment of this application.

[0038] Depend on Figure 2a It can be seen that the distance between the target to be detected in the image and the acquisition device is relatively far, and the model alone cannot accurately identify whether there is a person smoking, or which person is smoking.

[0039] Figure 2b This is a schematic diagram of a disparity map provided in an embodiment of this application.

[0040] in, Figure 2b for Figure 2a The disparity map corresponding to the image to be identified is shown.

[0041] S102: Based on the target detection model, obtain the region containing the target in the image to be identified, wherein the target is a behavior-related target; determine the target distance between the target and the acquisition device according to the disparity at the target region corresponding to the region in the disparity map and the parameters of the acquisition device;

[0042] To improve the accuracy of behavior recognition, the electronic device locally stores a pre-trained target detection model. After acquiring an image to be recognized, the electronic device can input the image into this pre-trained target detection model and obtain its output. This output can be an image showing the region of the detected target within the image, or it can be the location information of the detected target's region within the image. The detected target is a behavior-related target, which can be a human body part. Specifically, the behavior-related human body part can be the human body, the head, the hand, or a head or hand exhibiting behavior. The head or hand exhibiting behavior is identified by the target detection model as a possible head or hand exhibiting behavior. In essence, the electronic device acquires the region in the image containing the detected target. The electronic device then uses the target detection model, which can be a YOLO model, to detect the target of interest.

[0043] After acquiring the region containing the target in the image to be recognized, the electronic device can acquire the target region corresponding to that region in the disparity map. The disparity map corresponding to the image to be recognized is the same size as the image to be recognized, and each pixel in the image to be recognized has a corresponding pixel in the disparity map. Therefore, it can be acquired that any region in the image to be recognized has a corresponding target region in the disparity map. After acquiring the target region, the electronic device can acquire the disparity at the target region in the disparity map. The target region may contain multiple disparities, and the electronic device can determine the median, mean, or mode of these multiple disparities. Based on the acquired disparities and the parameters of the acquisition device, the target distance between the target and the acquisition device is determined.

[0044] Specifically, electronic devices can determine the target distance between the detection target and the acquisition device using the following formula:

[0045]

[0046] Where Z is the target distance between the target being detected and the acquisition device, B is the baseline of the binocular acquisition device, f is the focal length of the binocular acquisition device (the two cameras of the binocular acquisition device have the same focal length), and d is the parallax.

[0047] Specifically, the formula is determined as follows: The principle of binocular stereo vision is similar to how human eyes observe objects, calculating the distance between the object and the acquisition device based on the object's parallax. Assume a perfectly calibrated binocular acquisition device with identical parameters for both cameras. After calibration, the images captured by the two cameras are aligned horizontally, meaning the left camera (the left and right cameras in the binocular acquisition device) captures the image reference image at any point P. referenceThe corresponding point P on the target image acquired by the right camera (left and right in a binocular acquisition device) target In the pixel coordinate system, pixels with the same row coordinates are collinear. In the world coordinate system, pixels with the same row coordinates on the two imager planes are collinear. Figure 3 This is a schematic diagram of the relevant geometric relationships in a binocular acquisition device provided in an embodiment of this application. Figure 3 Let P be a point in space. r Let P be the projection point of point P in space onto the left (left and right) imager in the binocular acquisition device. t Let O be the projection point of point P in space onto the right (left and right) imaging instrument in a binocular acquisition device. r The projection center of the left (left and right) camera in a binocular acquisition device, O t The projection center of the right camera.

[0048] With projection point P r and P t The column coordinates are X r and X t Both cameras have a focal length of f, and P is perpendicular to the line O. r O t Taking the distance Z as an example, it is clear that triangle PP... r P t and triangle PO r O t They are similar triangles. According to the principle of similar triangles, it is easy to conclude that: X r -X t For the projection point P on the target image r Point P on the reference image t The pixel difference between the two cameras, i.e., the parallax of point P. The distance O between the projection centers of the two cameras. r O t The baseline of the stereo acquisition device is defined by the symbol B. The distance from point P to the stereo acquisition device is defined as the distance from point P to the line O. r O t The distance is denoted by Z. Specifically, it can be determined using the following formula:

[0049]

[0050] The relationship between distance and parallax can be determined using this formula.

[0051] Figure 4 This is a schematic diagram illustrating the relationship between distance and parallax, provided as an embodiment of this application.

[0052] Figure 4This diagram illustrates the relationship between the distance Z from the object to the acquisition device and the parallax d, given a fixed focal length f of the acquisition device and a fixed baseline B of the binocular acquisition device. Figure 4 It can be seen that the distance Z is inversely proportional to the parallax d. Figure 9 Middle left ( Figure 4 The left and right sides shown respectively include spatial points P1, P2, and P3. (From...) Figure 4 It can be seen that the distance between spatial points P1, P2 and P3 and the acquisition device gradually decreases, and the corresponding parallax gradually increases.

[0053] Figure 5a This is a schematic diagram of a region containing a detection target, provided as an embodiment of this application.

[0054] in, Figure 5a It can be the output of an object detection model. Figure 5a The area highlighted in the box represents a detected target. Figure 5a The target of the detection is the human head.

[0055] Figure 5b This is a schematic diagram of a target region in a disparity map provided in an embodiment of this application.

[0056] Figure 5b The area highlighted in the box is the target region corresponding to the region containing the target to be detected in the image to be identified.

[0057] S103: Based on the pre-saved correspondence between distance and focal length, determine the target focal length corresponding to the target distance; send the target focal length to the acquisition device so that the acquisition device can acquire the target image with the focal length of the target focal length; if the target image acquired by the acquisition device is received, determine whether there is a corresponding recognized behavior in the target image based on the behavior recognition model.

[0058] Since the distance between the smoker and the acquisition device is relatively far, the area of ​​the target to be detected in the acquired image is small. If the recognition is performed directly by the behavior recognition model, the recognition may be inaccurate. In this embodiment, the focal length of the acquisition device can be adjusted so that the acquisition device can acquire a target image with a larger area of ​​the target to be detected. The recognition of whether the smoking behavior exists can be performed based on the target image, thereby improving the accuracy of behavior recognition.

[0059] Specifically, to improve the accuracy of behavior recognition, the electronic device pre-stores the correspondence between distance and focal length. After adjusting the focal length of the acquisition device to the focal length corresponding to a certain distance, the acquisition device can accurately acquire an image containing objects that are separated from the acquisition device by that distance. Specifically, the area occupied by the object separated from the acquisition device by that distance in the image is relatively large. Therefore, the electronic device can determine the target focal length corresponding to the target distance based on the pre-stored correspondence between distance and focal length. To obtain a target image where the detected target occupies a large area in the image, the electronic device sends the target focal length to the acquisition device after determining it. Upon receiving the target focal length from the electronic device, the acquisition device adjusts its focal length to the target focal length, acquires the target image, and sends the acquired target image back to the electronic device. It should be noted that this target image is the image where the detected target occupies a large area in the image. To achieve focal length adjustment, the acquisition device described in this embodiment is a long focal length acquisition device.

[0060] To improve the accuracy of behavior recognition, the electronic device locally stores a pre-trained behavior recognition model. After receiving a target image, the electronic device inputs the target headshot into the behavior recognition model and obtains the model's output, which indicates whether the target image contains a corresponding behavior. If the target image contains the corresponding behavior, the electronic device can send the output of the target detection model described above to a preset device. This preset device can be a device used by staff, who can then determine which user is smoking based on the target detection model's output, thereby stopping the smoking in time and preventing dangerous accidents. In this embodiment, the detection of smoking is improved by magnifying a small target, which can reduce the workload of staff and eliminate safety hazards.

[0061] In this embodiment, a small target is magnified using optical zoom before being fed into a behavior recognition model for detection. Specifically, a larger target related to the small target is detected first (the target described in this embodiment), then the small target (the cigarette described in this embodiment) is focused using optical zoom, and finally the magnified small target is input into the behavior recognition model to obtain a high-confidence detection result.

[0062] It should be noted that if the target detection model identifies a region in the image to be identified that contains multiple detection targets, the subsequent electronic device will sequentially acquire the corresponding target image for each region, and for each acquired target image, determine whether there is a corresponding recognition behavior, that is, whether there is a smoking behavior.

[0063] In this embodiment, after acquiring the image to be identified, the electronic device determines the target distance between the target to be detected and the acquisition device in the image to be identified, detects the target related to the target behavior, and adjusts the focal length of the acquisition device according to the target distance, thereby acquiring a target image with a clearer target and an enlarged area in the image. Based on the target image, behavior recognition is performed, thereby accurately identifying whether there is a corresponding behavior and reducing security risks.

[0064] To improve the accuracy of behavior recognition, based on the above embodiments, in this embodiment, after acquiring the region containing the target in the image to be recognized and before sending the target focal length to the acquisition device, the method further includes:

[0065] Based on the location information of the region in the image to be identified, the target position information of the detected target in the camera coordinate system corresponding to the acquisition device is determined;

[0066] Based on the target location information, determine the target yaw angle and target pitch angle of the acquisition device when the detected target at the target location information is located at the acquisition center of the acquisition device;

[0067] The step of sending the target focal length to the acquisition device includes:

[0068] The target focal length, the target yaw angle, and the target pitch angle are sent to the acquisition device so that when the optical axis of the acquisition device is rotated at the target yaw angle and the target pitch angle respectively, the target image at the target focal length is acquired.

[0069] Since the area of ​​the target to be detected may not be at the acquisition center of the acquisition device, if only the focal length of the acquisition device is adjusted, the target image acquired by the acquisition device may not contain the target. In other words, the target may be outside the acquisition area of ​​the acquisition device. In order to accurately acquire the target image and thus improve the accuracy of behavior recognition, the electronic device can determine the target position information of the target in the camera coordinate system corresponding to the acquisition device, and then adjust the angle of the acquisition device according to the target position information so that the target is located at the acquisition center of the acquisition device.

[0070] Electronic devices can determine the target's position in the camera coordinate system corresponding to the acquisition device based on the target's location information within the image to be recognized. Specifically, how to determine the corresponding position information in the camera coordinate system given the known location information in the image is existing technology and will not be elaborated upon here. The x-coordinate and y-coordinate of the target's location information in the image to be recognized can be the average of the x-coordinate and y-coordinate of the location information of each pixel within the target's location area.

[0071] After determining the target position information of the target in the camera coordinate system corresponding to the acquisition device, the electronic device determines the target yaw angle and target pitch angle of the acquisition device when the target is located at the acquisition center of the acquisition device. In order to include the target in the target image, the electronic device sends the target yaw angle and target pitch angle to the acquisition device together with the target focal length. After receiving the target yaw angle and target pitch angle, the acquisition device rotates its own optical axis to the target yaw angle in the direction of the yaw angle and to the target pitch angle in the direction of the pitch angle, and adjusts the focal length to the target focal length to acquire the corresponding target image.

[0072] It's important to note that stereo vision typically employs four coordinate systems: pixel coordinates, image plane coordinates, camera coordinates, and world coordinates. The pixel coordinate system has its origin at the top-left vertex of the image (top, bottom, left, and right as shown in the image), and its unit is pixels. The coordinates of pixel p(u,v) in the image are (u,v), indicating that the pixel's position in the image is column u and row v. The image plane coordinate system has its origin at the intersection of the optical axis and the image plane, and its unit is meters or millimeters. The x-axis is in the same direction as the pixel row, and the y-axis is in the same direction as the pixel column.

[0073] The formula for the matrix transformation from the imaging plane coordinate system to the pixel coordinate system is:

[0074]

[0075] Where u0 and v0 are the x and y coordinates of a pixel in the pixel coordinate system, respectively, and s x s y Let be the number of pixels per unit size on the x-axis and y-axis, respectively. If the size of a unit pixel on the x-axis is dX, then s x = 1 / dX, x i Let y be the x-coordinate of the pixel in the imaging plane coordinate system. i Let z be the ordinate of the pixel in the imaging plane coordinate system. i Let be the vertical coordinate of the pixel in the imaging plane coordinate system, and u and v be the column number and row number of the image, respectively.

[0076] The camera coordinate system has its origin at the camera's optical center and its Z-axis as the optical axis. The formula for the matrix transformation from the camera coordinate system to the imaging plane coordinate system is:

[0077]

[0078] Where s is a proportionality coefficient related to the parameters of the acquisition device, and x c y c and z cThese represent the x, y, and y coordinates of the pixel in the camera coordinate system, where f is the focal length and x is the focal length. i and y i These are the horizontal and vertical coordinates of the pixel in the imaging plane coordinate system, respectively.

[0079] As described above, the formula for the matrix transformation from the pixel coordinate system to the camera coordinate system is:

[0080]

[0081] Where u0 and v0 are the x and y coordinates of a pixel in the pixel coordinate system, respectively, and s x s y Let be the number of pixels per unit size on the x-axis and y-axis, respectively. If the size of a unit pixel on the x-axis is dX, then s x = 1 / dX, x c y c and z c , , and , respectively, represent the x, y, and y coordinates of the pixel in the camera coordinate system, and u and v represent the column and row numbers of the image, respectively.

[0082] Specifically, given a known positional information in the pixel coordinate system, how to determine the positional information in the camera coordinate system is an existing technology, which will not be elaborated here.

[0083] Because long focal length acquisition devices have shallow depth of field and are difficult to focus, when the focal length differs greatly from the theoretical focal length, the image of the object is very blurry, making it difficult to detect behavior. Furthermore, the target is prone to shifting out of the image after being magnified. Therefore, the core of focusing a long focal length acquisition device is to find the focal length and determine the adjustment angle. This application embodiment fundamentally solves the problem of small targets having few pixels and being difficult to detect by adjusting the angle and focal length of the acquisition device.

[0084] To accurately determine the target position information of the detected target in the camera coordinate system, based on the above embodiments, in this embodiment, determining the target position information of the detected target in the camera coordinate system corresponding to the acquisition device based on the position information of the region in the image to be identified includes:

[0085] Based on the position information of the preset points in the region in the image to be identified, the target position information of the detected target in the camera coordinate system corresponding to the acquisition device is determined.

[0086] To determine the target position information of the detection target in the camera coordinate system corresponding to the acquisition device, after acquiring the area of ​​the target in the image to be recognized, the electronic device can determine the position information of the preset point of the area in the image to be recognized. The preset point can be the center point, the leftmost and topmost point (top, bottom, left, right) of the image to be recognized, or other points. After determining the position information of the preset point of the area in the image to be recognized, the electronic device can determine the target position information of the corresponding target in the camera coordinate system corresponding to the acquisition device based on the position information of the preset point of the area in the image to be recognized.

[0087] To accurately determine the target yaw angle and target pitch angle, based on the above embodiments, in this embodiment, when determining that the target position information is located at the acquisition center of the acquisition device, the target yaw angle and target pitch angle at the acquisition device include:

[0088]

[0089]

[0090] Where θ is the target yaw angle, y is the ordinate of the target position information center point, x is the abscissa of the target position information center point, Φ is the target pitch angle, and z is the ordinate of the target position information center point.

[0091] When determining the target yaw angle and target pitch angle, electronic equipment can use the following formula:

[0092]

[0093]

[0094] Where θ is the target yaw angle, y is the ordinate of the target position information center point, x is the abscissa of the target position information center point, Φ is the target pitch angle, and z is the ordinate of the target position information center point.

[0095] Figure 6 This is a schematic diagram of a camera coordinate system provided in an embodiment of this application.

[0096] Figure 6 The coordinate system shown is the camera coordinate system. Figure 6 Point N shown is the point at the target location. Figure 6 In this context, θ represents the target yaw angle during rotation, and Φ represents the target pitch angle during rotation.

[0097] In this embodiment, the electronic device determines the angle of rotation of the optical axis of the acquisition device when it is aligned with the target by detecting the target position information of the target in the camera coordinate system, so that the optical axis of the camera is aligned with the target, thereby avoiding the problem of the target moving out of the picture due to magnification.

[0098] Figure 7a This is a schematic diagram of a region containing a detection target in an image to be identified, provided in an embodiment of this application.

[0099] in, Figure 7a To identify the image Figure 2a A magnified image of the region containing the detected target, by Figure 7a It can be seen that simply magnifying the region containing the target to be detected results in a low-resolution image.

[0100] Figure 7b This is a schematic diagram of a target image provided in an embodiment of this application.

[0101] in, Figure 7b This is a schematic diagram of the target image acquired after the acquisition device adjusts its focus and angle. Figure 7b It can be seen that the target image contains the detected target and has high clarity. Figure 7a and Figure 7b As can be seen, the method provided in this application embodiment can obtain images with high clarity containing the detection target, thereby improving the accuracy of behavior recognition.

[0102] To improve the accuracy of behavior recognition, based on the above embodiments, in this embodiment, after receiving the target image acquired by the acquisition device, and before determining whether a corresponding behavior exists in the target image based on the behavior recognition model, the method further includes:

[0103] Repeat the following steps:

[0104] The target focal length is increased by a first preset focal length to obtain a first candidate focal length, and the first candidate focal length is sent to the acquisition device so that the acquisition device acquires a first candidate image with the focal length of the first candidate focal length; the first candidate image is received, and if the first focus quality (Face Value, FV) of the first candidate image is greater than the second FV of the target image, the target image is replaced with the first candidate image, and the target focal length is replaced with the first candidate focal length;

[0105] If the first FV is less than the second FV, then the target focal length is reduced by a second preset focal length to obtain a second candidate focal length, wherein the second preset focal length is less than the first preset focal length. The second candidate focal length is sent to the acquisition device so that the acquisition device acquires a second candidate image with the focal length of the second candidate focal length. The second candidate image is received. If the third FV of the second candidate image is greater than the second FV, then the target image is replaced by the second candidate image, and the target focal length is replaced by the second candidate focal length.

[0106] This continues until both the first and third FV determined for the target image are smaller than the second FV.

[0107] Since the accuracy of the determined target distance is not high, the target focal length may also be inaccurate. Furthermore, the image quality of the target image acquired by the acquisition device based on the target focal length may not be high. In order to obtain a more accurate focal length and then acquire a target image with higher image quality based on the more accurate focal length, in this embodiment of the application, the electronic device can fine-tune the target focal length corresponding to the target distance and acquire the image based on the fine-tuned focal length, thereby improving the image quality of the acquired image and thus improving the accuracy of behavior recognition.

[0108] Specifically, the electronic device may repeat the following steps until the first FV of the first candidate image and the third FV of the second candidate image determined for the target image are both less than the second FV of the target image:

[0109] The electronic device increases the target focal length by a first preset focal length, obtains the increased first candidate focal length, and sends the first candidate focal length to the acquisition device. Upon receiving the first candidate focal length from the electronic device, the acquisition device adjusts its own focal length to the first candidate focal length and acquires a first candidate image based on the first candidate focal length. The first candidate image is then sent to the electronic device. After acquiring the first candidate image, the electronic device determines whether the first FV of the first candidate image is greater than the second FV of the target image. If the first FV of the first candidate image is greater than the second FV of the target image, it indicates that the first candidate image is of better quality. Therefore, the first candidate image is used to replace the target image, and the first candidate focal length is used to replace the target focal length. It should be noted that the better the image focus, the larger the image's FV.

[0110] If the first focal length (FV) of the first candidate image is not greater than the second FV of the target image, it indicates that the quality of the first candidate image is inferior to that of the target image. Therefore, the target focal length is reduced by a second preset focal length to obtain the reduced second candidate focal length. Specifically, the second candidate focal length is less than the first preset focal length, and can be half of the first preset focal length. After obtaining the second candidate focal length, it is sent to the acquisition device. Upon receiving the second candidate focal length from the electronic device, the acquisition device adjusts its own focal length to the second candidate focal length and acquires a second candidate image based on the second candidate focal length. The second candidate image is then sent to the electronic device. After acquiring the second candidate image, the electronic device determines whether the third FV of the second candidate image is greater than the second FV of the target image. If the third FV of the second candidate image is greater than the second FV of the target image, it indicates that the quality of the second candidate image is better. Therefore, the second candidate image is used to replace the target image, and the second candidate focal length is used to replace the target focal length.

[0111] The electronic device repeats the above steps until the first FV and the third FV determined for the target image are both less than the second FV.

[0112] In this embodiment of the application, the method by which the electronic device determines the focal length can be called a hill-climbing algorithm, that is, the electronic device refines the focus and precisely adjusts the focal length of the acquisition device through the hill-climbing algorithm.

[0113] Figures 8a-8d These are schematic diagrams of images acquired by the acquisition devices at different focal lengths according to the embodiments of this application.

[0114] Among them, the data acquisition equipment collects Figures 8a-8d When the focal lengths are similar, Figures 8a-8d It can be seen that images with similar focal lengths all contain the target being detected, but their sharpness differs.

[0115] In this embodiment of the application, the electronic device can determine the corresponding FV based on the image gradient of the image. The specific method of determination is prior art and will not be described in detail here. Figure 8a The corresponding FV is 56.77. Figure 8b The corresponding FV is 15.15. Figure 8c The corresponding FV is 7.05. Figure 8d The corresponding FV is 6.41.

[0116] Figure 9 This is a schematic diagram illustrating the relationship between focal length and the FV of the corresponding acquired image, provided as an embodiment of this application.

[0117] Figure 9 The horizontal axis represents the focal length. Figure 9 The vertical axis represents the field of view (FV) of the acquired image, and is derived from... Figure 9It can be seen that as the focal length increases, the FV of the corresponding acquired image first increases and then decreases.

[0118] To accurately obtain the disparity map corresponding to the image to be identified, based on the above embodiments, in this embodiment, obtaining the disparity map corresponding to the image to be identified includes:

[0119] Acquire another image captured by the acquisition device; wherein, the acquisition device is a binocular acquisition device;

[0120] A binocular stereo matching algorithm is used to process the image to be identified and the other image to obtain a disparity map.

[0121] To obtain the disparity map corresponding to the image to be identified, the acquisition device described in this application embodiment is a binocular acquisition device. After acquiring the image to be identified and another image, the binocular acquisition device can send the image to be identified and the other image to an electronic device, which can then acquire the other image. After acquiring the image to be identified and the other image, a binocular stereo matching algorithm is used to process the image to be identified and the other image to obtain the disparity map. The binocular stereo matching algorithm can be BM, SGB, Var, etc., and can employ dense algorithms, sparse algorithms, or even match only the center points of the target detection boxes of the two cameras. Specifically, how to obtain the disparity map after acquiring the images from the binocular cameras is existing technology and will not be elaborated here.

[0122] When determining the parallax at the target area, the electronic device can also acquire the regions containing the target in the two images acquired by the binocular acquisition device, and determine the position information of preset points in the two acquired regions. These preset points can be the center point or the leftmost and topmost points (top, bottom, left, and right of the image to be recognized). The corresponding parallax is determined based on the acquired position information, which can be specifically determined using the following formula:

[0123] d=C xl -C xr

[0124] Where d represents the corresponding disparity, and C xl and C xr These are the x-coordinates of preset points in the two regions on the image. Specifically, how to determine the parallax of a certain region using images captured by two cameras is existing technology and will not be elaborated upon here.

[0125] Figure 10a This is a schematic diagram of an image to be identified provided in an embodiment of this application. Figure 10b Another schematic image provided for an embodiment of this application.

[0126] in, Figure 10aImages captured by the left (left and right) camera of a binocular acquisition device. Figure 10b The images captured by the right (left and right) camera of the binocular acquisition device are... Figure 10a and Figure 10b However, the images captured by the two cameras of the binocular acquisition device are similar in content.

[0127] To accurately determine the distance between the detection target and the acquisition device, based on the above embodiments, in this embodiment, determining the target distance between the detection target and the acquisition device according to the distance at the target area corresponding to the area in the disparity map includes:

[0128] The distance at a preset point in the target region corresponding to the region in the disparity map is determined as the target distance between the detection target and the acquisition device.

[0129] After acquiring the target area in the disparity map, the electronic device can obtain the distance to a preset point in the target area. This preset point can be the center point, the leftmost and topmost point (the top, bottom, left, and right of the image to be recognized), or other points. After obtaining the distance to the preset point in the target area, the electronic device can determine this distance as the target distance between the detection target and the acquisition device.

[0130] Figure 11 A detailed process diagram of behavior recognition provided in this application embodiment is shown, which includes the following steps:

[0131] Depend on Figure 11 As can be seen, the electronic device first acquires the image to be recognized and another image acquired by the binocular acquisition device. It then processes the image to be recognized and the other image using a binocular stereo matching algorithm to obtain a disparity map. The image to be recognized is then input into a target detection model to obtain the region containing the target object output by the target detection model. The target region corresponding to this region in the disparity map is determined. Based on the disparity at the target region in the disparity map and the parameters of the acquisition device, the target distance between the target object and the acquisition device is determined. Finally, the corresponding target focal length is determined based on the target distance, and the target image acquired by the acquisition device at the target focal length is obtained.

[0132] After acquiring the target image, the electronic device uses a ladder-climbing algorithm to obtain a target image with better focus quality, and then inputs the target image into the behavior recognition model to obtain whether the behavior recognition model outputs whether the corresponding behavior exists.

[0133] In order to accurately obtain the target detection model, based on the above embodiments, the target detection model in this application embodiment is determined in the following way:

[0134] Obtain any first sample image from the first sample set, and the first location information of each detected target contained in the first sample image;

[0135] The first sample image is input into the original target detection model to obtain the second location information of each detected target contained in the first sample image;

[0136] The original target detection model is trained based on the first location information and the second location information.

[0137] In order to train the target detection model, this application embodiment stores a first sample set for training. The first sample images in the first sample set include images of different users. For example, the first sample images in the first sample set include users wearing different colored clothes, of different ages, and of different genders. The first sample set also stores the first location information of each detected target contained in the first sample image for each first sample image.

[0138] After obtaining any first sample image in the first sample set and the first location information of each detected target contained in the first sample image, the first sample image is input into the original target detection model, and the original target detection model outputs the second location information of each detected target contained in the first sample image.

[0139] After the original target detection model determines the second location information of each detected target contained in the first sample image, the original target detection model is trained based on the first location information of each detected target contained in the first sample image and the second location information output by the original target detection model.

[0140] The original object detection model is trained using the above method. When a preset condition is met, a trained object detection model is obtained. This preset condition may be that the number of first sample images in the first sample set whose second position information matches the first position information obtained after training the original object detection model is greater than a set number; or it may be that the number of iterations for training the original object detection model reaches a set maximum number of iterations, etc. Specifically, this application embodiment does not impose any limitations on this.

[0141] In order to obtain the trained behavior recognition model, based on the above embodiments, the behavior recognition model in this application embodiment is determined in the following way:

[0142] Obtain any second sample image from the second sample set, and the annotation result indicating whether there is a corresponding recognized behavior in the second sample image;

[0143] The second sample image is input into the original behavior recognition model to obtain the output result of whether there is a corresponding behavior to be recognized in the second sample image;

[0144] The original behavior recognition model is trained based on the annotation results and the output results.

[0145] To train the behavior recognition model, this application embodiment stores a second sample set for training. The second sample images in the second sample set include images of different users. For example, the second sample images in the second sample set include users wearing different colored clothes, of different ages, different genders, smoking or not smoking, etc. The second sample set also stores a labeling result for each second sample image, indicating whether there is a corresponding recognized behavior in the second sample image. For example, 00 can be used to indicate the presence of a corresponding recognized behavior, and 01 can be used to indicate the absence of a corresponding recognized behavior.

[0146] After obtaining any second sample image in the second sample set and the annotation result of whether there is a corresponding recognized behavior in the second sample image, the second sample image is input into the original behavior recognition model, and the original behavior recognition model outputs the output result of whether there is a corresponding recognized behavior in the second sample image.

[0147] After the original behavior recognition model determines whether there is a corresponding behavior in the second sample image, the original behavior recognition model is trained based on the output result of whether there is a corresponding behavior in the second sample image and the annotation result of whether there is a corresponding behavior in the second sample image.

[0148] The original behavior recognition model is trained using the above method. When preset conditions are met, a trained behavior recognition model is obtained. These preset conditions may include: the number of second sample images in the second sample set whose output results match the labeled results after training the original behavior recognition model is greater than a set number; or the number of iterations for training the original behavior recognition model reaching a set maximum number of iterations, etc. Specifically, this application embodiment does not impose limitations on these conditions.

[0149] Figure 12 This application provides a schematic diagram of a behavior recognition device, which includes:

[0150] The acquisition and determination module 1201 is used to acquire the image to be identified and the disparity map corresponding to the image to be identified acquired by the acquisition device; based on the target detection model, acquire the region in the image to be identified that contains the target to be detected, wherein the target to be detected is a behavior-related target; and determine the target distance between the target to be detected and the acquisition device according to the disparity at the target region corresponding to the region in the disparity map and the parameters of the acquisition device.

[0151] The processing module 1202 is used to determine the target focal length corresponding to the target distance based on the pre-saved correspondence between distance and focal length; send the target focal length to the acquisition device so that the acquisition device can acquire the target image with the focal length of the target focal length; if the target image acquired by the acquisition device is received, determine whether there is a corresponding recognized behavior in the target image based on the behavior recognition model.

[0152] Furthermore, the processing module 1202 is also used to determine the target position information of the detected target in the camera coordinate system corresponding to the acquisition device based on the position information of the region in the image to be identified; and to determine the target yaw angle and target pitch angle of the acquisition device when the detected target at the target position information is located at the acquisition center of the acquisition device based on the target position information.

[0153] The processing module 1202 is specifically used to send the target focal length, the target yaw angle, and the target pitch angle to the acquisition device, so that when the optical axis of the acquisition device is rotated at the target yaw angle and the target pitch angle respectively, the target image at the target focal length is acquired.

[0154] Furthermore, the processing module 1202 is specifically used to determine the target position information of the detected target in the camera coordinate system corresponding to the acquisition device based on the position information of the preset points of the region in the image to be identified.

[0155] Further, the processing module 1202 is also configured to repeatedly execute the following steps: increasing the target focal length by a first preset focal length to obtain a first candidate focal length, sending the first candidate focal length to the acquisition device so that the acquisition device acquires a first candidate image with the focal length of the first candidate focal length; receiving the first candidate image, if the first focus quality (FV) of the first candidate image is greater than the second FV of the target image, then replacing the target image with the first candidate image and replacing the target focal length with the first candidate focal length; if the first FV is less than the second FV, then decreasing the target focal length by a second preset focal length to obtain a second candidate focal length, wherein the second preset focal length is less than the first preset focal length, sending the second candidate focal length to the acquisition device so that the acquisition device acquires a second candidate image with the focal length of the second candidate focal length; receiving the second candidate image, if the third FV of the second candidate image is greater than the second FV, then replacing the target image with the second candidate image and replacing the target focal length with the second candidate focal length; until both the first FV and the third FV determined for the target image are less than the second FV.

[0156] Furthermore, the acquisition and determination module 1201 is specifically used to acquire another image acquired by the acquisition device; wherein, the acquisition device is a binocular acquisition device; and a binocular stereo matching algorithm is used to process the image to be identified and the other image to obtain a disparity map.

[0157] Furthermore, the acquisition and determination module 1201 is specifically used to determine the distance at a preset point in the target area corresponding to the region in the disparity map as the target distance between the detection target and the acquisition device.

[0158] Furthermore, the processing module 1202 is also used to acquire any first sample image in the first sample set, and the first position information of each detected target contained in the first sample image; input the first sample image into the original target detection model to acquire the second position information of each detected target contained in the first sample image; and train the original target detection model according to the first position information and the second position information.

[0159] Furthermore, the processing module 1202 is also used to acquire any second sample image in the second sample set, and the annotation result of whether there is a corresponding recognized behavior in the second sample image; input the second sample image into the original behavior recognition model, and acquire the output result of whether there is a corresponding recognized behavior in the second sample image; and train the original behavior recognition model according to the annotation result and the output result.

[0160] Figure 13This application provides a schematic diagram of an electronic device structure based on an embodiment of the present application. In addition to the above embodiments, this application also provides an electronic device, such as... Figure 13 As shown, it includes: processor 1301, communication interface 1302, memory 1303 and communication bus 1304, wherein processor 1301, communication interface 1302 and memory 1303 communicate with each other through communication bus 1304;

[0161] The memory 1303 stores a computer program. When the program is executed by the processor 1301, the processor 1301 performs the following steps:

[0162] Acquire the image to be identified captured by the acquisition device and the disparity map corresponding to the image to be identified;

[0163] Based on the target detection model, the region containing the target in the image to be identified is obtained, wherein the target is a behavior-related target; the target distance between the target and the acquisition device is determined according to the disparity at the target region corresponding to the region in the disparity map and the parameters of the acquisition device.

[0164] Based on the pre-saved correspondence between distance and focal length, the target focal length corresponding to the target distance is determined; the target focal length is sent to the acquisition device so that the acquisition device can acquire a target image with the focal length of the target focal length; if the target image acquired by the acquisition device is received, it is determined whether there is a corresponding recognized behavior in the target image based on the behavior recognition model.

[0165] Furthermore, the processor 1301 is also configured to determine the target position information of the detected target in the camera coordinate system corresponding to the acquisition device based on the position information of the region in the image to be identified;

[0166] Based on the target location information, determine the target yaw angle and target pitch angle of the acquisition device when the detected target at the target location information is located at the acquisition center of the acquisition device;

[0167] The processor 1301 is specifically used to send the target focal length, the target yaw angle, and the target pitch angle to the acquisition device, so that when the optical axis of the acquisition device is rotated at the target yaw angle and the target pitch angle respectively, the target image at the target focal length is acquired.

[0168] Furthermore, the processor 1301 is specifically used to determine the target position information of the detected target in the camera coordinate system corresponding to the acquisition device based on the position information of the preset points of the region in the image to be identified.

[0169] Furthermore, the processor 1301 is also configured to repeatedly execute the following steps:

[0170] Increase the target focal length by a first preset focal length to obtain a first candidate focal length, and send the first candidate focal length to the acquisition device so that the acquisition device acquires a first candidate image with the focal length of the first candidate focal length; receive the first candidate image, and if the first focus quality (FV) of the first candidate image is greater than the second FV of the target image, then replace the target image with the first candidate image and replace the target focal length with the first candidate focal length.

[0171] If the first FV is less than the second FV, then the target focal length is reduced by a second preset focal length to obtain a second candidate focal length, wherein the second preset focal length is less than the first preset focal length. The second candidate focal length is sent to the acquisition device so that the acquisition device acquires a second candidate image with the focal length of the second candidate focal length. The second candidate image is received. If the third FV of the second candidate image is greater than the second FV, then the target image is replaced by the second candidate image, and the target focal length is replaced by the second candidate focal length.

[0172] This continues until both the first and third FV determined for the target image are smaller than the second FV.

[0173] Furthermore, the processor 1301 is specifically used to acquire another image acquired by the acquisition device; wherein, the acquisition device is a binocular acquisition device;

[0174] A binocular stereo matching algorithm is used to process the image to be identified and the other image to obtain a disparity map.

[0175] Furthermore, the processor 1301 is specifically used to determine the distance at a preset point in the target region corresponding to the region in the disparity map as the target distance between the detection target and the acquisition device.

[0176] Furthermore, the processor 1301 is also configured to acquire any first sample image in the first sample set, and the first location information of each detected target contained in the first sample image;

[0177] The first sample image is input into the original target detection model to obtain the second location information of each detected target contained in the first sample image;

[0178] The original target detection model is trained based on the first location information and the second location information.

[0179] Furthermore, the processor 1301 is also used to acquire any second sample image in the second sample set, and whether there is a labeling result for the corresponding recognized behavior in the second sample image;

[0180] The second sample image is input into the original behavior recognition model to obtain the output result of whether there is a corresponding behavior to be recognized in the second sample image;

[0181] The original behavior recognition model is trained based on the annotation results and the output results.

[0182] The communication bus mentioned in the above server can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0183] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0184] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0185] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0186] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program executable by an electronic device. When the program is run on the electronic device, the electronic device performs the following steps:

[0187] The memory stores a computer program that, when executed by the processor, causes the processor to perform the following steps:

[0188] Acquire the image to be identified captured by the acquisition device and the disparity map corresponding to the image to be identified;

[0189] Based on the target detection model, the region containing the target in the image to be identified is obtained, wherein the target is a behavior-related target; the target distance between the target and the acquisition device is determined according to the disparity at the target region corresponding to the region in the disparity map and the parameters of the acquisition device.

[0190] Based on the pre-saved correspondence between distance and focal length, the target focal length corresponding to the target distance is determined; the target focal length is sent to the acquisition device so that the acquisition device can acquire a target image with the focal length of the target focal length; if the target image acquired by the acquisition device is received, it is determined whether there is a corresponding recognized behavior in the target image based on the behavior recognition model.

[0191] In one possible implementation, after acquiring the region containing the target in the image to be identified and before sending the target focal length to the acquisition device, the method further includes:

[0192] Based on the location information of the region in the image to be identified, the target position information of the detected target in the camera coordinate system corresponding to the acquisition device is determined;

[0193] Based on the target location information, determine the target yaw angle and target pitch angle of the acquisition device when the detected target at the target location information is located at the acquisition center of the acquisition device;

[0194] The step of sending the target focal length to the acquisition device includes:

[0195] The target focal length, the target yaw angle, and the target pitch angle are sent to the acquisition device so that when the optical axis of the acquisition device is rotated at the target yaw angle and the target pitch angle respectively, the target image at the target focal length is acquired.

[0196] In one possible implementation, determining the target position information of the detected target in the camera coordinate system corresponding to the acquisition device based on the position information of the region in the image to be identified includes:

[0197] Based on the position information of the preset points in the region in the image to be identified, the target position information of the detected target in the camera coordinate system corresponding to the acquisition device is determined.

[0198] In one possible implementation, when determining that the target location information is located at the acquisition center of the acquisition device based on the target location information, the target yaw angle and target pitch angle of the acquisition device include:

[0199]

[0200]

[0201] Where θ is the target yaw angle, y is the ordinate of the target position information center point, x is the abscissa of the target position information center point, Φ is the target pitch angle, and z is the ordinate of the target position information center point.

[0202] In one possible implementation, after receiving the target image acquired by the acquisition device, and before determining whether a corresponding recognized behavior exists in the target image based on the behavior recognition model, the method further includes:

[0203] Repeat the following steps:

[0204] Increase the target focal length by a first preset focal length to obtain a first candidate focal length, and send the first candidate focal length to the acquisition device so that the acquisition device acquires a first candidate image with the focal length of the first candidate focal length; receive the first candidate image, and if the first focus quality (FV) of the first candidate image is greater than the second FV of the target image, then replace the target image with the first candidate image and replace the target focal length with the first candidate focal length.

[0205] If the first FV is less than the second FV, then the target focal length is reduced by a second preset focal length to obtain a second candidate focal length, wherein the second preset focal length is less than the first preset focal length. The second candidate focal length is sent to the acquisition device so that the acquisition device acquires a second candidate image with the focal length of the second candidate focal length. The second candidate image is received. If the third FV of the second candidate image is greater than the second FV, then the target image is replaced by the second candidate image, and the target focal length is replaced by the second candidate focal length.

[0206] This continues until both the first and third FV determined for the target image are smaller than the second FV.

[0207] In one possible implementation, obtaining the disparity map corresponding to the image to be identified includes:

[0208] Acquire another image captured by the acquisition device; wherein, the acquisition device is a binocular acquisition device;

[0209] A binocular stereo matching algorithm is used to process the image to be identified and the other image to obtain a disparity map.

[0210] In one possible implementation, determining the target distance between the detection target and the acquisition device based on the distance at the target region corresponding to the region in the disparity map includes:

[0211] The distance at a preset point in the target region corresponding to the region in the disparity map is determined as the target distance between the detection target and the acquisition device.

[0212] In one possible implementation, the target detection model is determined in the following manner:

[0213] Obtain any first sample image from the first sample set, and the first location information of each detected target contained in the first sample image;

[0214] The first sample image is input into the original target detection model to obtain the second location information of each detected target contained in the first sample image;

[0215] The original target detection model is trained based on the first location information and the second location information.

[0216] In one possible implementation, the behavior recognition model is determined in the following way:

[0217] Obtain any second sample image from the second sample set, and the annotation result indicating whether there is a corresponding recognized behavior in the second sample image;

[0218] The second sample image is input into the original behavior recognition model to obtain the output result of whether there is a corresponding behavior to be recognized in the second sample image;

[0219] The original behavior recognition model is trained based on the annotation results and the output results.

[0220] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0221] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0222] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0223] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes. Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, this application also intends to include such modifications and variations if they fall within the scope of the claims of this application and their equivalents.

Claims

1. A behavior recognition method, characterized by, The method includes: Acquire the image to be identified captured by the acquisition device and the disparity map corresponding to the image to be identified; Based on the target detection model, the region containing the target in the image to be identified is obtained, wherein the target is a behavior-related target; the target distance between the target and the acquisition device is determined according to the disparity at the target region corresponding to the region in the disparity map and the parameters of the acquisition device. Based on the pre-saved correspondence between distance and focal length, the target focal length corresponding to the target distance is determined; the target focal length is sent to the acquisition device so that the acquisition device can acquire the target image with the focal length of the target focal length; if the target image acquired by the acquisition device is received, it is determined whether there is a corresponding recognized behavior in the target image based on the behavior recognition model; The method further includes, after acquiring the region containing the target in the image to be identified and before sending the target focal length to the acquisition device: Based on the location information of the region in the image to be identified, the target position information of the detected target in the camera coordinate system corresponding to the acquisition device is determined; Based on the target location information, determine the target yaw angle and target pitch angle of the acquisition device when the detected target at the target location information is located at the acquisition center of the acquisition device; The step of sending the target focal length to the acquisition device includes: The target focal length, the target yaw angle, and the target pitch angle are sent to the acquisition device so that when the optical axis of the acquisition device is rotated at the target yaw angle and the target pitch angle respectively, the target image at the target focal length is acquired.

2. The method of claim 1, wherein, The step of determining the target position information of the detected target in the camera coordinate system corresponding to the acquisition device based on the position information of the region in the image to be identified includes: Based on the position information of the preset points in the region in the image to be identified, the target position information of the detected target in the camera coordinate system corresponding to the acquisition device is determined.

3. The method according to claim 1 or 2, characterized in that, When determining that the target location information is located at the acquisition center of the acquisition device based on the target location information, the target yaw angle and target pitch angle of the acquisition device include: Where θ is the target yaw angle, y is the ordinate of the target position information center point, x is the abscissa of the target position information center point, Φ is the target pitch angle, and z is the ordinate of the target position information center point.

4. The method of claim 1, wherein, After receiving the target image acquired by the acquisition device, and before determining whether a corresponding recognized behavior exists in the target image based on the behavior recognition model, the method further includes: Repeat the following steps: Increase the target focal length by a first preset focal length to obtain a first candidate focal length, and send the first candidate focal length to the acquisition device so that the acquisition device acquires a first candidate image with the focal length of the first candidate focal length; receive the first candidate image, and if the first focus quality (FV) of the first candidate image is greater than the second FV of the target image, then replace the target image with the first candidate image and replace the target focal length with the first candidate focal length. If the first FV is less than the second FV, then the target focal length is reduced by a second preset focal length to obtain a second candidate focal length, wherein the second preset focal length is less than the first preset focal length. The second candidate focal length is sent to the acquisition device so that the acquisition device acquires a second candidate image with the focal length of the second candidate focal length. The second candidate image is received. If the third FV of the second candidate image is greater than the second FV, then the target image is replaced by the second candidate image, and the target focal length is replaced by the second candidate focal length. This continues until both the first and third FV determined for the target image are smaller than the second FV.

5. The method of claim 1, wherein, Obtaining the disparity map corresponding to the image to be identified includes: Acquire another image captured by the acquisition device; wherein, the acquisition device is a binocular acquisition device; A binocular stereo matching algorithm is used to process the image to be identified and the other image to obtain a disparity map.

6. The method of claim 1, wherein, Determining the target distance between the detection target and the acquisition device based on the distance at the target region corresponding to the region in the disparity map includes: The distance at a preset point in the target region corresponding to the region in the disparity map is determined as the target distance between the detection target and the acquisition device.

7. The method of claim 1, wherein, The target detection model is determined in the following way: Obtain any first sample image from the first sample set, and the first location information of each detected target contained in the first sample image; The first sample image is input into the original target detection model to obtain the second location information of each detected target contained in the first sample image; The original target detection model is trained based on the first location information and the second location information.

8. The method of claim 1, wherein, The behavior recognition model is determined in the following way: Obtain any second sample image from the second sample set, and the annotation result indicating whether there is a corresponding recognized behavior in the second sample image; The second sample image is input into the original behavior recognition model to obtain the output result of whether there is a corresponding behavior to be recognized in the second sample image; The original behavior recognition model is trained based on the annotation results and the output results.

9. An electronic device, comprising: The electronic device includes at least a processor and a memory, the processor being used to execute a computer program stored in the memory to implement the steps of the behavior recognition method as described in any one of claims 1-8.