Target standing detection method, system, device and storage medium
By training and identifying models under different lighting, combined with camera parameters and video frame processing, the problem of low target stand detection accuracy under the influence of ambient light is solved, and higher detection accuracy and accuracy are achieved.
Patent Information
- Application Number
- CN202210242497.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-03-11
AI Technical Summary
When the prior art achieves target stand detection, it is greatly affected by changes in ambient light, resulting in low accuracy of stand detection.
By obtaining the video frame to be detected, the camera parameters are determined, and the area to be detected is enlarged in a magnification proportion to generate the image to be detected. Then call the trained recognition model for humanoid detection to determine the height of the target person. When the height is greater than the preset height threshold, the target person is determined to be in the standing state. The recognition model can effectively reduce the impact of ambient light on detection by training under different lights.
It effectively reduces the impact of ambient light on target stand detection, improves the accuracy of target detection, and reduces the impact of other activity interference through smooth filtering, improving the accuracy and accuracy of target stand detection.
Smart Images

Figure CN114708529B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a method, system, device and storage medium for detecting a standing target. Background Art
[0002] With the continuous development of communication technology, concepts such as online education and remote office have been further promoted. For example, in an education recording and broadcasting system or a conference recording and broadcasting system, classroom or conference recordings can be saved for users to learn. In order to achieve the purpose of analyzing the classroom performance of teachers and students, these classroom image resources or conference image resources can be used to track speakers, that is, to track standing targets in the video.
[0003] In the related art, the optical flow method is generally used to realize the standing detection of a person. However, when realizing the standing detection, the related art is greatly affected by the change of the ambient light, resulting in low accuracy of the standing detection. Summary of the invention
[0004] The present application aims to solve one of the technical problems in the related art at least to a certain extent. To this end, the present application proposes a target standing detection method, system, device and storage medium.
[0005] In a first aspect, an embodiment of the present application provides a method for detecting a standing target, comprising: obtaining a video frame to be detected; determining the camera parameters corresponding to the video frame to be detected; enlarging the area to be detected in the video frame to be detected at a first magnification ratio to generate an image to be detected; calling a trained recognition model to perform human figure detection on the image to be detected, and determining a first human figure area in the image to be detected; determining the height of a target person corresponding to the first human figure area according to the first human figure area, the first magnification ratio and the camera parameters; when the height of the target person is greater than a preset height threshold, determining that the target person is in a standing state; wherein, the training steps of the recognition model are as follows: obtaining a training image; wherein, the training image includes images of a person standing or sitting in a fixed scene taken by cameras at different positions under different lighting conditions; inputting the training image into a deep learning model for training, and obtaining the recognition model after the training is completed; wherein, the recognition model is used to identify the first human figure area in the image to be detected.
[0006] Optionally, the area to be detected includes a fixed reference object, and the fixed reference object includes at least one of a podium and a projection screen.
[0007] Optionally, the camera parameters include a shooting height and a shooting angle of the camera and setting parameters of the camera, and the setting parameters include a lens focal length and image sensor parameters.
[0008] Optionally, determining the height of a target person corresponding to the first human-shaped area according to the first human-shaped area, the first magnification ratio and the camera parameters includes: determining a second human-shaped area of the target person in the video frame to be detected according to the first human-shaped area and the first magnification ratio; determining the height of the target person according to the second human-shaped area and the camera parameters.
[0009] Optionally, the method also includes: determining a video frame set including multiple consecutive video frames to be detected; determining human features corresponding to all the video frames to be detected in the video frame set; wherein the human features include the second human-shaped area and the height of the target person; and performing smoothing filtering on the human features to determine the standing status of the target person in the video frame set.
[0010] Optionally, the smoothing filtering of the character features to determine the standing status of the target character in the video frame set includes: extracting one of the video frames to be detected from the video frame set every fixed number of frames; determining the character features corresponding to the extracted video frame to be detected; weighting the character features according to the order of the extracted video frames to be detected; smoothing filtering the character features to which the weights have been assigned; and determining the standing status of the target character in the video frame set.
[0011] Optionally, the method further includes: when the height of the target person is less than or equal to a preset height threshold, determining that the target person is in a non-standing state.
[0012] In the second aspect, an embodiment of the present application provides a target standing detection system, including: a first module, used to obtain a video frame to be detected; a second module, used to determine the camera parameters corresponding to the video frame to be detected; a third module, used to enlarge the area to be detected in the video frame to be detected at a first magnification ratio, and generate an image to be detected; a fourth module, used to call a trained recognition model to perform human figure detection on the image to be detected, and determine the first human figure area in the image to be detected; a fifth module, used to determine the height of a target person corresponding to the first human figure area according to the first human figure area, the first magnification ratio and the camera parameters; a sixth module, used to determine that the target person is in a standing state when the height of the target person is greater than a preset height threshold.
[0013] In a third aspect, an embodiment of the present application provides a target standing detection device, comprising: at least one processor; at least one memory for storing at least one program; when the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned target standing detection method.
[0014] In a fourth aspect, an embodiment of the present application provides a computer storage medium, in which a program executable by a processor is stored. When the program executable by the processor is executed by the processor, it is used to implement the above-mentioned target standing detection method.
[0015] The beneficial effects of the embodiments of the present application are as follows: first, a video frame to be detected is obtained; the camera parameters corresponding to the video frame to be detected are determined; the area to be detected in the video frame to be detected is enlarged at a first magnification ratio to generate an image to be detected; then the trained recognition model is called to perform human shape detection on the image to be detected, and the first human shape area in the image to be detected is determined; according to the first human shape area, the first magnification ratio and the camera parameters, the height of the target person corresponding to the first human shape area is determined; when the height of the target person is greater than the preset height threshold, the target person is determined to be in a standing state. Among them, the recognition model used in the embodiments of the present application is trained using images of people standing or sitting in a fixed scene taken by cameras at different positions under different lighting. Therefore, the embodiments of the present application can effectively reduce the influence of ambient light on target standing detection and improve the accuracy of target detection by training the recognition model with a variety of training images under different lighting. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0017] Figure 1 is a flowchart of the steps of the target standing detection method provided in an embodiment of the present application;
[0018] Figure 2 A flowchart of the training steps of the recognition model provided in the embodiment of the present application;
[0019] Figure 3 A schematic diagram of a target standing detection system provided in an embodiment of the present application;
[0020] Figure 4 A schematic diagram of a target standing detection device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0022] It should be noted that, although the functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0023] With the continuous development of communication technology, concepts such as online education and remote office have been further promoted. For example, in an education recording and broadcasting system or a conference recording and broadcasting system, classroom or conference recordings can be saved for users to learn. In order to achieve the purpose of analyzing the classroom performance of teachers and students, these classroom image resources or conference image resources can be used to track speakers, that is, to track standing targets in the video. In related technologies, human standing detection is generally achieved through optical flow method. However, when implementing standing detection, the related technology is greatly affected by changes in ambient light, resulting in low accuracy of standing detection.
[0024] Based on this, the present application provides a method, system, device and storage medium for detecting a standing target, the method comprising: first, obtaining a video frame to be detected; determining the camera parameters corresponding to the video frame to be detected; amplifying the area to be detected in the video frame to be detected at a first magnification ratio to generate an image to be detected; then calling a trained recognition model to perform human figure detection on the image to be detected, and determining the first human figure area in the image to be detected; according to the first human figure area, the first magnification ratio and the camera parameters, determining the height of the target person corresponding to the first human figure area; when the height of the target person is greater than a preset height threshold, determining that the target person is in a standing state. Among them, the recognition model used in the embodiment of the present application is trained using images of people standing or sitting in a fixed scene taken by cameras at different positions under different lighting. Therefore, the embodiment of the present application trains the recognition model by using a variety of training images under different lighting, which can effectively reduce the influence of ambient light on target standing detection and improve the accuracy of target detection.
[0025] The embodiments of the present application are further described below in conjunction with the accompanying drawings.
[0026] refer to Figure 1 , Figure 1 is a flowchart of the steps of the target standing detection method provided in an embodiment of the present application, the method includes but is not limited to steps S100-S150:
[0027] S100, obtaining a video frame to be detected;
[0028] Specifically, first, multiple cameras are set up in a fixed place where video shooting is required. For example, when it is necessary to record a video of a teacher teaching in a classroom, and it is necessary to perform target standing detection on the teacher standing near the podium in the video, it is necessary to obtain the video to be detected captured by the camera, and obtain the video frame to be detected in which the teacher is captured in the video to be detected.
[0029] It is understandable that, since a plurality of cameras are provided at the shooting location, a plurality of videos to be detected that are shot from different angles of the teacher can be obtained, and therefore cameras at different locations correspond to a plurality of groups of video frames to be detected.
[0030] S110, determining the camera parameters corresponding to the video frame to be detected;
[0031] Specifically, as mentioned in the above steps, cameras at different positions correspond to multiple groups of video frames to be detected. During the camera debugging process, the imaging principle of the camera can actually be combined to determine the correspondence between the height and size of objects in the real environment and the height and size of the image in the camera.
[0032] That is to say, in the process of installing and debugging the camera, the setting parameters of the camera can be obtained first, and the setting parameters include the focal length of the lens when the camera is shooting, and then the shooting height and shooting angle of the camera can be determined according to the actual installation position of the camera. Use the camera to shoot a fixed reference object in a fixed scene. Since the height and size of the fixed reference object can be known in advance through measurement, the user can determine the corresponding relationship between the height and size of the object in the real environment and the image sensor parameters of the camera (that is, the height and size of the image captured by the image sensor) within the shooting range of the camera. Therefore, in the same way, when the target person enters the camera, since the relationship between the height of the real object and the image height of the object in the camera has been determined in advance, the height of the target person can be inferred through the image in the camera.
[0033] Therefore, in order to determine the height of the target person in the video frame to be detected, it is necessary to determine the camera parameters corresponding to the video frame to be detected. In the embodiment of the present application, the camera parameters include the shooting height, shooting angle and setting parameters of the camera.
[0034] S120, enlarging the area to be detected in the video frame to be detected at a first enlargement ratio to generate an image to be detected;
[0035] Specifically, due to the limitations of the shooting position of the camera, the hardware shooting conditions of the camera, etc., when the target person stands far away from the camera, the image of the target person in the camera may be too small, making it difficult to determine the exact height of the target person. In order to improve the accuracy of target recognition, it is proposed in the embodiment of the present application to first enlarge the area to be detected in the video frame to be detected by a first magnification ratio. In the embodiment of the present application, the area to be detected refers to the area including the target person in the video frame to be detected. In the embodiment of the present application, the position where the target person stands in a fixed scene is usually in a fixed area. For example, security personnel generally need to stand in a sentry box. For example, when a teacher maintains a standing state in a classroom, he generally needs to stand near the podium. Therefore, objects such as sentry boxes, podiums, and projection screens can be used as fixed reference objects for preliminarily determining the standing detection range of the target person.
[0036] Furthermore, since the positions of multiple cameras in the present application are fixed in advance, the scenes in the same group of video frames to be detected are basically fixed. Therefore, the sizes of the areas to be detected in the same group of video frames to be detected are the same. After determining the first magnification ratio required for target standing detection, all areas to be detected corresponding to all video frames to be detected in the same group can be determined. These areas to be detected are magnified at the first magnification ratio to generate images to be detected, and these images to be detected are used for more accurate target standing detection.
[0037] It is understandable that when the camera is set far away from the target person, after the detection area is enlarged according to the preset first magnification ratio, an image to be detected containing a larger target person can be obtained, thereby improving the accuracy and recognition rate of the recognition model for small object detection in subsequent steps.
[0038] S130, calling the trained recognition model to perform human shape detection on the image to be detected, and determining a first human shape region in the image to be detected;
[0039] Specifically, after the image to be detected is generated by step S120, it is necessary to detect the image to be detected containing the target person to determine whether the target person exists in the image to be detected and whether the target person is standing in the image to be detected. In the embodiment of the present application, a trained recognition model is used to perform human figure detection on the image to be detected to determine the human figure outline in the image to be detected, and the human figure outline is called the first human figure area.
[0040] Reference Figure 2 , Figure 2 The flowchart of the training steps of the recognition model provided in the embodiment of the present application includes but is not limited to steps S200-S210:
[0041] S200, obtaining a training image;
[0042] Specifically, in order for the recognition model to be able to detect the target person in the image, a large number of training images are needed to train the recognition model. In the related art, the recognition of people in the image is often affected by the lighting, resulting in the recognition model's recognition effect on the person in the image being blurred and the accuracy being low. In an embodiment of the present application, images of people standing or sitting in a fixed scene taken by cameras at different positions under different lighting conditions are used as training images. For example, in a fixed classroom, three time periods are selected: nine in the morning, twelve at noon, and three in the afternoon. In situations such as switching lights and curtains, multiple images of people standing near the podium and sitting near the podium are taken, and these images are used as training images to train the recognition model.
[0043] S210, inputting the training image into a deep learning model for training, and obtaining a recognition model after the training is completed;
[0044] Specifically, YOLO V3 is a target detection model based on a neural network, which has the advantages of fast speed, low background false detection rate and high versatility. Therefore, the YOLO V3 model is used as the basis in this application, and the YOLO V3 model is trained by a large number of training images obtained in step S200, and finally a recognition model capable of identifying the target person is obtained. It is understandable that since the training images include a large number of images of people standing or sitting under different lighting conditions, the recognition model obtained by training can effectively resist the influence of changing light on the detection results, and obtain a more accurate target person recognition result.
[0045] It should be noted that the content to be identified in the present application is the target person in the image. Therefore, the recognition model is used to identify the first human-shaped area in the image to be detected. The first human-shaped area represents the target person in the video frame to be detected.
[0046] Through the above steps S200-S210, the embodiment of the present application provides specific training steps for the recognition model, and the model can be used to recognize the first human-shaped area in the image to be detected. Through the above content, step S130 has been explained, and step S140 will be explained below.
[0047] S140, determining the height of the target person corresponding to the first human-shaped area according to the first human-shaped area, the first magnification ratio and the camera parameters;
[0048] Specifically, after the recognition model recognizes the first human-shaped area in the image to be detected, the first human-shaped area can be restored to the second human-shaped area in the video frame to be detected according to the first magnification ratio. It can be understood that the first magnification ratio at this time refers to the magnification multiple of the first human-shaped area following the area to be detected, and restoring the first human-shaped area to the second human-shaped area is to reduce the first human-shaped area by a corresponding multiple, determine the second human-shaped area in the video frame to be detected, that is, determine the height and size of the second human-shaped area in the video frame to be detected.
[0049] According to the above step S110, the corresponding relationship between the height and size of the object in the real environment and the height and size of the image in the camera can be determined in combination with the imaging principle of the camera. Therefore, the height of the target person can be determined according to the second human-shaped area and the camera parameters.
[0050] In the target standing detection method proposed in the embodiment of the present application, the main purpose is to detect the target person who maintains a standing posture for a considerable period of time. But in fact, in addition to the target person, there will be other moving people in the fixed scene. For example, when a teacher is teaching, there will be students leaving their seats to answer questions, or students standing up from their seats to answer questions. Although these students have appeared in a standing posture for a short period of time, it is obvious that these students are not the target people to be captured by the detection method of the embodiment of the present application. Therefore, the embodiment of the present application proposes to perform smoothing filtering on the second human-shaped area to reduce the impact of people who occasionally stand or walk on the target standing detection.
[0051] In some embodiments, the specific process of smoothing the second human-shaped area is as follows: first, determine a video frame set containing multiple continuous video frames to be detected, for example, obtain all continuous video frames to be detected corresponding to a video with a total duration of 10 minutes as a video frame set; then, determine the character features corresponding to all video frames to be detected in the video frame set. In the embodiment of the present application, the character features include the height of the second human-shaped area and the target person. The character features can be obtained by the above steps S100 and S140. Then, extract a video frame to be detected from the video frame set every fixed number of frames, for example, extract the video frames to be detected in a multiple of 10 in order, then extract the 1st, 10th, 20th... After extracting the required video frames to be detected, determine the character features corresponding to the extracted video frames to be detected, and according to the order of the extracted video frames to be detected, weight the character features, assign a heavier weight to the character features in the last extracted video frame to be detected, and then smooth the character features after completing the weight allocation; determine the standing situation of the target person in the video frame set.
[0052] It is understandable that the second human-shaped area in the character feature can represent the position of the character in the video frame to be detected. If the second human-shaped area of the character in the extracted video frame to be detected moves significantly, or has moved out of the pre-defined target character standing range, the character is considered to be an active interfering character and is filtered out using a smoothing filter. Similarly, the height of the target character in the character feature can represent the posture of the character in the video frame to be detected. If the height of the character changes significantly in a relatively close video frame to be detected, it means that the character stands and sits down in a short period of time. Therefore, the character is also determined to be an interfering character and is filtered out using a smoothing filter.
[0053] Therefore, the embodiment of the present application effectively improves the accuracy of target person recognition through smoothing filtering.
[0054] Step S140 has been explained through the above content, and step S150 will be explained below.
[0055] S150: When the height of the target person is greater than a preset height threshold, it is determined that the target person is in a standing state;
[0056] Specifically, in the embodiment of the present application, the height of the target person is calculated to determine whether the target person is in a standing state. For example, the height threshold is set to 1.65m. When the height of the target person is less than or equal to 1.65m, it is determined that the target person is in a non-standing state, which may be a sitting or squatting position; and when the height of the target person is greater than 1.65m, it is determined that the target person is in a standing state, thereby completing the standing detection of the target person.
[0057] Through steps S100-S150, the embodiment of the present application provides a method for detecting a standing target. First, a video frame to be detected is obtained; the camera parameters corresponding to the video frame to be detected are determined; the area to be detected in the video frame to be detected is enlarged at a first magnification ratio to generate an image to be detected; then the trained recognition model is called to perform human shape detection on the image to be detected, and the first human shape area in the image to be detected is determined; according to the first human shape area, the first magnification ratio and the camera parameters, the height of the target person corresponding to the first human shape area is determined; when the height of the target person is greater than the preset height threshold, the target person is determined to be in a standing state. Among them, the recognition model used in the embodiment of the present application is trained by using images of people standing or sitting in a fixed scene taken by cameras at different positions under different lighting. Therefore, the embodiment of the present application trains the recognition model by using a variety of training images under different lighting, which can effectively reduce the influence of ambient light on target standing detection and improve the accuracy of target detection. In addition, the embodiment of the present application also proposes to reduce the influence of interfering people in other activities by smoothing filtering, so as to further improve the accuracy and precision of target standing detection.
[0058] Reference Figure 3 , Figure 3 The schematic diagram of the target standing detection system provided by the embodiment of the present application, the system 300 includes a first module 310, a second module 320, a third module 330, a fourth module 340, a fifth module 350 and a sixth module 360. The first module is used to obtain a video frame to be detected; the second module is used to determine the camera parameters corresponding to the video frame to be detected; the third module is used to magnify the area to be detected in the video frame to be detected at a first magnification ratio to generate an image to be detected; the fourth module is used to call the trained recognition model to perform human shape detection on the image to be detected and determine the first human shape area in the image to be detected; the fifth module is used to determine the height of the target person corresponding to the first human shape area according to the first human shape area, the first magnification ratio and the camera parameters; the sixth module is used to determine that the target person is in a standing state when the height of the target person is greater than a preset height threshold.
[0059] refer to Figure 4 , Figure 4 A schematic diagram of a target standing detection device provided in an embodiment of the present application, wherein the device 400 includes at least one processor 410 and at least one memory 420 for storing at least one program; Figure 4 A processor and a memory are taken as an example.
[0060] The processor and the memory may be connected via a bus or other means. Figure 4 The example of connecting through bus is taken in the following.
[0061] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0062] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0063] The embodiment of the present application further discloses a computer storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to implement the method proposed in the present application when executed by the processor.
[0064] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0065] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions without violating the spirit of the present application. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A target standing detection method, characterized in that: include: Get the video frame to be detected; Determine the camera parameters corresponding to the video frame to be detected; Amplifying the area to be detected in the video frame to be detected at a first magnification ratio to generate an image to be detected; Calling the trained recognition model to perform human shape detection on the image to be detected, and determining a first human shape area in the image to be detected; Determining a height of a target person corresponding to the first human-shaped area according to the first human-shaped area, the first magnification ratio and the camera parameters; When the height of the target person is greater than a preset height threshold, it is determined that the target person is in a standing state; The training steps of the recognition model are as follows: Get training images; The training images include images of people standing or sitting in a fixed scene, taken by cameras at different positions under different lighting conditions; Inputting the training image into a deep learning model for training, and obtaining the recognition model after the training is completed; Wherein, the recognition model is used to recognize the first human-shaped area in the image to be detected; The step of determining the height of the target person corresponding to the first human-shaped area according to the first human-shaped area, the first magnification ratio and the camera parameter includes: Determine a second human-shaped region of the target person in the to-be-detected video frame according to the first human-shaped region and the first magnification ratio; Determining the height of the target person according to the second human-shaped area and the camera parameters; Determine a video frame set including a plurality of continuous video frames to be detected; Determine the human features corresponding to all the to-be-detected video frames in the video frame set; Wherein, the character features include the second human-shaped area and the height of the target character; Performing smoothing filtering on the character features to determine the standing condition of the target character in the set of video frames; The step of smoothing and filtering the character features to determine the standing situation of the target character in the video frame set includes: Extracting one of the to-be-detected video frames from the video frame set at fixed frame intervals; Determine the human feature corresponding to the extracted video frame to be detected; According to the order of the extracted video frames to be detected, weights are assigned to the character features; Smoothing filtering is performed on the character features after weight allocation; and the standing condition of the target character in the video frame set is determined.
2. The target standing detection method according to claim 1, characterized in that: The area to be detected includes a fixed reference object, and the fixed reference object includes at least one of a podium and a projection screen.
3. The target standing detection method according to claim 2, characterized in that: The camera parameters include the shooting height and shooting angle of the camera and the setting parameters of the camera, and the setting parameters include the focal length of the lens and the image sensor parameters.
4. The target standing detection method according to claim 1, characterized in that: The method further comprises: When the height of the target person is less than or equal to a preset height threshold, it is determined that the target person is in a non-standing state.
5. A target standing detection system, characterized in that: include: The first module is used to obtain the video frame to be detected; The second module is used to determine the camera parameters corresponding to the video frame to be detected; The third module is used to enlarge the area to be detected in the video frame to be detected at a first enlargement ratio to generate an image to be detected; The fourth module is used to call the trained recognition model to perform human shape detection on the image to be detected, and determine the first human shape area in the image to be detected; A fifth module, configured to determine a height of a target person corresponding to the first human-shaped area according to the first human-shaped area, the first magnification ratio and the camera parameters; The sixth module is used to determine that the target person is in a standing state when the height of the target person is greater than a preset height threshold; The method for determining the height of the target person corresponding to the first human-shaped area according to the first human-shaped area, the first magnification ratio and the camera parameter includes: Determine a second human-shaped region of the target person in the to-be-detected video frame according to the first human-shaped region and the first magnification ratio; Determining the height of the target person according to the second human-shaped area and the camera parameters; Determine a video frame set including a plurality of continuous video frames to be detected; Determine the human features corresponding to all the to-be-detected video frames in the video frame set; Wherein, the character features include the second human-shaped area and the height of the target character; Performing smoothing filtering on the character features to determine the standing condition of the target character in the set of video frames; The method for performing smooth filtering on the character features to determine the standing condition of the target character in the video frame set includes: Extracting one of the to-be-detected video frames from the video frame set at fixed frame intervals; Determine the human feature corresponding to the extracted video frame to be detected; According to the order of the extracted video frames to be detected, weights are assigned to the character features; Smoothing filtering is performed on the character features after weight allocation; and the standing condition of the target character in the video frame set is determined.
6. A target standing detection device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the target standing detection method as described in any one of claims 1 to 4.
7. A computer storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement the target standing detection method as described in any one of claims 1 to 4 when executed by the processor.
Citation Information
Patent Citations
Target identification and tracking method and device, equipment and storage medium
CN113705510A