Image processing method, apparatus and device
Patent Information
- Application Number
- CN202210464887.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2042-04-29
AI Technical Summary
[0005]本申请提供一种图像处理方法、装置及设备,用以解决用户的识别码追踪难度较大导致无法获取识别码对应的行为信息的技术问题
[0065] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
Smart Images

Figure CN117036418B_ABST
Abstract
Description
Technical Field
[0001] This application relates to image processing technology, and more particularly to an image processing method, apparatus, and device. Background Technology
[0002] Currently, in order to understand students' classroom behavior, it is necessary to track and identify each student.
[0003] In existing technologies, when tracking and identifying each student, target detection and action recognition technologies are used to identify students' classroom behavior, and the number of actions performed by each student in the whole class is calculated based on the target tracking algorithm.
[0004] However, in existing technologies, when calculating the number of actions each student performs in a lesson based on the target tracking algorithm, students may exhibit disruptive behaviors in the classroom, such as looking down or blocking each other's view. This makes it impossible to accurately identify the student's identification code, and consequently, to obtain information about the student's classroom behavior. Summary of the Invention
[0005] This application provides an image processing method, apparatus, and device to solve the technical problem that the difficulty in tracking user identification codes leads to the inability to obtain behavioral information corresponding to the identification codes.
[0006] In a first aspect, this application provides an image processing method, comprising:
[0007] Multiple frames of images are processed for recognition to obtain a parameter list for each image; wherein the parameter list includes parameter information of the target to be recognized in the image;
[0008] For two adjacent frames in the multi-frame image, the target to be identified in the next frame is determined based on the parameter list of the previous frame and the parameter list of the next frame; wherein, the target to be identified in the next frame is the target to be identified corresponding to the target to be identified in the previous frame.
[0009] An identification code is generated for the target to be identified in the last frame of the multi-frame image, and target information corresponding to the identification code is generated; wherein, the identification code represents the identity of the target to be identified, and the target information represents the number of times the target to be identified performs a preset action.
[0010] Further, for two adjacent frames in the multi-frame image, based on the parameter list of the previous frame and the parameter list of the next frame, the target to be identified in the next frame is determined, including:
[0011] Based on the parameter list of the first frame image in the multi-frame image, the Gaussian distribution corresponding to the target to be identified in the first frame image is determined; wherein, the Gaussian distribution represents the position information of the target to be identified;
[0012] For two adjacent frames in the multi-frame image, the parameter list of the next frame in the two adjacent frames is used to perform probability calculation with the Gaussian distribution corresponding to the target to be identified in the previous frame in the two adjacent frames, to obtain the probability result information corresponding to the target to be identified in the next frame in the two adjacent frames; wherein, the probability result information is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame in the two adjacent frames.
[0013] Based on the probability result information, the target to be identified in the next frame of the two adjacent frames is determined.
[0014] Further, based on the probability result information, determining the target to be identified in the next frame of the two adjacent frames includes:
[0015] If the probability result information is determined to be a likelihood probability value, wherein the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames, and the likelihood probability values between the target and multiple Gaussian distributions in the previous frame of two adjacent frames, then the maximum likelihood probability value corresponding to the target to be identified is determined; wherein the maximum likelihood probability value is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames.
[0016] Based on the maximum likelihood probability value corresponding to the target to be identified, the target to be identified in the next frame of two adjacent frames is determined.
[0017] Furthermore, the parameter information corresponding to the target to be identified includes the face width value; the method also includes:
[0018] Based on the preset first constraint information, the distance between the target to be identified in the next frame of two adjacent frames and the location of the Gaussian distribution corresponding to the target to be identified is determined; wherein, the first constraint information represents the location information corresponding to the target to be identified.
[0019] If it is determined that the distance is greater than the face width of the target to be identified, then the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames is replaced with the Gaussian distribution corresponding to the target to be identified in the previous frame of the two adjacent frames, so as to obtain the updated Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames.
[0020] Based on the Gaussian distribution corresponding to the target to be identified in the next frame of the updated two adjacent frames, the target to be identified in the next frame of the two adjacent frames is determined.
[0021] Further, based on the probability result information, determining the target to be identified in the next frame of the two adjacent frames includes:
[0022] If the probability result information is determined to be no likelihood probability value, wherein the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames, and the likelihood probability value between the target and multiple Gaussian distributions in the previous frame of two adjacent frames, then the target to be identified is determined to be a new target.
[0023] Based on the parameter information corresponding to the newly added target, determine the Gaussian distribution corresponding to the newly added target in the next frame of two adjacent frames;
[0024] The target to be identified in the next frame of the two adjacent frames is determined based on the Gaussian distribution corresponding to the newly added target in the next frame of the two adjacent frames.
[0025] Further, based on the probability result information, determining the target to be identified in the next frame of the two adjacent frames includes:
[0026] If the probability result information is determined to be that no parameter information corresponding to the Gaussian distribution is obtained, then the target to be identified corresponding to the Gaussian distribution is determined to be in a disappeared state;
[0027] A predetermined number of targets to be identified are located adjacent to the location of the Gaussian distribution, along with the parameter information corresponding to the targets to be identified; wherein the Gaussian distribution includes a first face area, and the parameter information corresponding to the targets to be identified includes a second face area;
[0028] Based on the preset second constraint information, a target face area equal to the first face area is determined from a preset number of second face areas; wherein, the second constraint information represents the pixel ratio of the face area of the target to be identified;
[0029] Based on the parameter information corresponding to the target face area, determine the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames;
[0030] The target to be identified in the next frame of the two adjacent frames is determined based on the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames.
[0031] Furthermore, the multi-frame images are processed for recognition to obtain a parameter list for each image, including:
[0032] According to the preset depth target detection network, multiple frames of images are processed for recognition to obtain a parameter list for each image; wherein, the depth target detection network is used to indicate target recognition information, and the parameter list includes parameter information of the target to be identified in the image.
[0033] Furthermore, the parameter information corresponding to the target to be identified includes face information and body information. The face information includes the horizontal coordinate of the face center, the vertical coordinate of the face center, the face width value, the face height value, and the face area. The body information includes the horizontal coordinate of the body center, the vertical coordinate of the body center, the body width value, the body height value, and the body area.
[0034] Secondly, this application provides an image processing apparatus, comprising:
[0035] The recognition unit is used to perform recognition processing on multiple frames of images to obtain a parameter list for each image; wherein the parameter list includes parameter information of the target to be recognized in the image;
[0036] The first determining unit is configured to, for two adjacent frames in the multi-frame image, determine the target to be identified in the next frame of the two adjacent frames based on the parameter list of the previous frame and the parameter list of the next frame of the two adjacent frames; wherein the target to be identified in the next frame of the two adjacent frames is the target to be identified corresponding to the target to be identified in the previous frame of the two adjacent frames.
[0037] The first generation unit is used to generate the identification code corresponding to the target to be identified in the last frame of the multi-frame images;
[0038] The second generation unit is used to generate target information corresponding to the identification code; wherein the identification code represents the identity identifier of the target to be identified, and the target information represents the number of times the target to be identified has performed a preset action.
[0039] Further, the first determining unit includes:
[0040] The first determining module is used to determine the Gaussian distribution corresponding to the target to be identified in the first frame image based on the parameter list of the first frame image in the multi-frame images; wherein the Gaussian distribution represents the position information of the target to be identified;
[0041] The calculation module is used to perform probability calculation on the parameter list of the next frame in the multi-frame images and the Gaussian distribution corresponding to the target to be identified in the previous frame in the multi-frame images, to obtain probability result information corresponding to the target to be identified in the next frame in the multi-frame images; wherein, the probability result information is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame in the multi-frame images.
[0042] The second determining module is used to determine the target to be identified in the next frame of the two adjacent frames based on the probability result information.
[0043] Further, the second determining module includes:
[0044] The first determining submodule is configured to determine the maximum likelihood probability value corresponding to the target to be identified if the probability result information is determined to be a likelihood probability value, wherein the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames, and the likelihood probability values between the target and multiple Gaussian distributions in the previous frame of two adjacent frames; wherein the maximum likelihood probability value is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames.
[0045] The second determination submodule is used to determine the target to be identified in the next frame of two adjacent frames based on the maximum likelihood probability value corresponding to the target to be identified.
[0046] Furthermore, the parameter information corresponding to the target to be identified includes the face width value; the method further includes:
[0047] The second determining unit is used to determine the distance between the target to be identified in the next frame of two adjacent frames and the location of the Gaussian distribution corresponding to the target to be identified, based on the preset first constraint information; wherein, the first constraint information represents the location information corresponding to the target to be identified.
[0048] The update unit is used to replace the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames with the Gaussian distribution corresponding to the target to be identified in the previous frame of the two adjacent frames if it is determined that the distance is greater than the face width value of the target to be identified, so as to obtain the updated Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames.
[0049] The third determining unit is used to determine the target to be identified in the next frame of two adjacent frames based on the Gaussian distribution corresponding to the target to be identified in the next frame of the updated two adjacent frames.
[0050] Further, the second determining module includes:
[0051] The third determining submodule is used to determine that if the probability result information is that no likelihood probability value is obtained, wherein the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames, and the likelihood probability value between the target and multiple Gaussian distributions in the previous frame of two adjacent frames, the target to be identified is a new target.
[0052] The fourth determination submodule is used to determine the Gaussian distribution corresponding to the newly added target in the next frame of two adjacent frames based on the parameter information corresponding to the newly added target.
[0053] The fifth determination submodule is used to determine the target to be identified in the next frame of the two adjacent frames based on the Gaussian distribution corresponding to the newly added target in the next frame of the two adjacent frames.
[0054] Further, the second determining module includes:
[0055] The sixth determining submodule is used to determine that the target to be identified corresponding to the Gaussian distribution is in a disappeared state if the probability result information is determined to be that no parameter information corresponding to the Gaussian distribution is obtained.
[0056] The seventh determination submodule is used to determine a preset number of targets to be identified that are adjacent to the location of the Gaussian distribution, as well as the parameter information corresponding to the targets to be identified; wherein, the Gaussian distribution includes a first face area, and the parameter information corresponding to the targets to be identified includes a second face area;
[0057] The eighth determining submodule is used to determine a target face area that is equal to the first face area from a preset number of second face areas based on preset second constraint information; wherein, the second constraint information represents the pixel ratio of the face area of the target to be identified.
[0058] The ninth determining submodule is used to determine the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames based on the parameter information corresponding to the target face area.
[0059] The tenth determination submodule is used to determine the target to be identified in the next frame of the two adjacent frames based on the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames.
[0060] Furthermore, the identification unit is specifically used for:
[0061] According to the preset depth target detection network, multiple frames of images are processed for recognition to obtain a parameter list for each image; wherein, the depth target detection network is used to indicate target recognition information, and the parameter list includes parameter information of the target to be identified in the image.
[0062] Furthermore, the parameter information corresponding to the target to be identified includes face information and body information. The face information includes the horizontal coordinate of the face center, the vertical coordinate of the face center, the face width value, the face height value, and the face area. The body information includes the horizontal coordinate of the body center, the vertical coordinate of the body center, the body width value, the body height value, and the body area.
[0063] Thirdly, this application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the method described in the first aspect.
[0064] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in the first aspect.
[0065] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0066] This application provides an image processing method, apparatus, and device that performs recognition processing on multiple frames of images to obtain a parameter list for each image; wherein the parameter list includes parameter information of the target to be identified in the image. For two adjacent frames in the multi-frame image set, the target to be identified in the next frame is determined based on the parameter list of the preceding frame and the parameter list of the following frame; wherein the target to be identified in the next frame corresponds to the target to be identified in the preceding frame. An identification code corresponding to the target to be identified in the last frame of the multi-frame image set is generated, and target information corresponding to the identification code is generated; wherein the identification code represents the identity of the target to be identified, and the target information represents the number of times the target to be identified performs a preset action. In this scheme, the image includes multiple targets to be identified, and the multi-frame image set is processed to obtain a parameter list for each image, wherein the parameter list includes parameter information of the target to be identified in the image. Then, based on two adjacent frames in a multi-frame image dataset, the target to be identified in the next frame of those two adjacent frames can be determined to be the same target to be identified as a certain target to be identified in the previous frame of those two adjacent frames. The target to be identified in the next frame of those two adjacent frames is then used as the target to be identified in the previous frame of the next set of adjacent frames. Combined with the parameter list in the next frame of the next set of adjacent frames, the target to be identified in the next frame of the next set of adjacent frames is determined, and so on, until the target to be identified in the last frame of the multi-frame image dataset is determined. Finally, an identification code corresponding to the target to be identified in the last frame of the multi-frame image dataset is generated, along with target information corresponding to the identification code. Based on this target information, the number of times the target to be identified performs a preset action is displayed. Therefore, the same target to be identified in multiple frames can be identified, and an identification code corresponding to the target to be identified can be generated, greatly improving the stability of tracking the target to be identified and solving the technical problem that the difficulty in tracking the user's identification code leads to the inability to obtain the behavioral information corresponding to the identification code. Attached Figure Description
[0067] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0068] Figure 1 A schematic flowchart of an image processing method provided in an embodiment of this application;
[0069] Figure 2 A flowchart illustrating another image processing method provided in an embodiment of this application;
[0070] Figure 3This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application;
[0071] Figure 4 This is a schematic diagram of another image processing apparatus provided in an embodiment of this application;
[0072] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0073] Figure 6 This is a block diagram of an electronic device provided in an embodiment of this application.
[0074] The accompanying drawings have illustrated specific embodiments of this disclosure, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this disclosure to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0075] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure.
[0076] In one example, to understand students' classroom behavior, it is necessary to track and identify each student. Current technology uses object detection and action recognition techniques to identify student behavior and calculates the number of actions each student performs throughout the lesson based on object tracking algorithms. However, in this existing technology, students may exhibit disruptive behaviors during the lesson, such as looking down or obstructing each other's view, making it difficult to accurately identify student identifiers and thus hindering the acquisition of student classroom behavior data.
[0077] This application provides an image processing method, apparatus, and device, which aims to solve the above-mentioned technical problems in the prior art.
[0078] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0079] Figure 1 This is a schematic flowchart of an image processing method provided in an embodiment of this application, as shown below. Figure 1 As shown, it includes:
[0080] 101. Perform recognition processing on multiple frames of images to obtain a parameter list for each image; wherein, the parameter list includes parameter information of the target to be recognized in the image.
[0081] For example, the executing entity of this embodiment can be an electronic device, a terminal device, an image processing device or apparatus, or other apparatus or apparatus capable of executing this embodiment, and there is no limitation thereto. In this embodiment, the executing entity is described as an electronic device.
[0082] First, multiple frames of images need to be acquired, and then these frames are processed for recognition. Images can be captured by camera, retrieved from memory, or received from other devices. Then, the multiple frames are processed to obtain a parameter list for each image. This parameter list includes parameter information of the target to be identified in the image. For example, the recognition process may include processing through a pre-set deep target monitoring network. The technical solution of this application is a multi-target tracking algorithm. Besides the technical solution of this application, multi-target tracking algorithms also include the SORT algorithm and the DeepSORT algorithm.
[0083] For example, based on a preset depth target detection network, the electronic device performs recognition processing on multiple frames of images, detects human faces in each frame, and obtains the bounding box (bbox) coordinates and width and height of each human face. Each frame outputs a parameter list bbox, and each parameter list has 10 dimensional features: x-coordinate of face center, y-coordinate of face center, face bbox width, face bbox height, face bbox area, x-coordinate of human body center, y-coordinate of human body center, human body bbox width, human body bbox height, and human body bbox area.
[0084] 102. For two adjacent frames in a multi-frame image, determine the target to be identified in the next frame in the two adjacent frames based on the parameter list of the previous frame in the two adjacent frames and the parameter list of the next frame in the two adjacent frames; wherein, the target to be identified in the next frame in the two adjacent frames is the target to be identified corresponding to the target to be identified in the previous frame in the two adjacent frames.
[0085] For example, firstly, based on the parameter list X1 of the first frame in the multi-frame image, where the length of parameter list X1 is n, initialize N (N = n + 10, where 10 is a reserved placeholder) multidimensional (10-dimensional) Gaussian distributions G, where Gaussian distribution G represents the mean and variance of each target to be identified. For example, Gaussian distribution G includes: G1, G2, G3, ..., where G1 represents the Gaussian distribution of the first person, G2 represents the Gaussian distribution of the second person, and G3 represents the Gaussian distribution of the third person.
[0086] Then, for two adjacent frames in a multi-frame image, based on a preset formula, the parameter list of the latter frame in the two adjacent frames is used to perform probability calculation with the Gaussian distribution corresponding to the target to be identified in the former frame in the two adjacent frames. This yields the probability result information corresponding to the target to be identified in the latter frame in the two adjacent frames. The probability result information is used to indicate the Gaussian distribution corresponding to the target to be identified in the latter frame in the two adjacent frames. The preset formula is as follows:
[0087]
[0088] Among them, P i,j X represents the likelihood probability between the i-th target to be identified in the latter frame of two adjacent frames and the j-th Gaussian distribution in the former frame of two adjacent frames; a b G represents the parameter information of the b-th target to be identified in the a-th frame image. i This represents the Gaussian distribution of the i-th target to be identified in the previous frame of two adjacent frames.
[0089] Finally, based on the probability results, the target to be identified in the next frame of two adjacent frames is determined. The target to be identified in the next frame of two adjacent frames is the target to be identified corresponding to the target to be identified in the previous frame of two adjacent frames. That is, the target to be identified in the next frame of two adjacent frames is replaced with the Gaussian distribution corresponding to the target to be identified in the previous frame of two adjacent frames.
[0090] 103. Generate the identification code corresponding to the target to be identified in the last frame of the multi-frame image, and generate the target information corresponding to the identification code; wherein, the identification code represents the identity of the target to be identified, and the target information represents the number of times the target to be identified performs a preset action.
[0091] For example, the identification code represents the identity of the target to be identified, and the preset action refers to the action pre-stored by the electronic device, such as raising a hand. The electronic device can generate the identification code corresponding to the target to be identified in the last frame of the multi-frame image based on the pre-stored total number of people and the Gaussian distribution corresponding to the target to be identified in the last frame of the multi-frame image, and generate target information corresponding to each identification code. The target information represents the number of times the target to be identified performs the preset action.
[0092] For example, if the total number of people is 30, the electronic device can generate the identification code corresponding to the target in the last frame of the multi-frame image based on the Gaussian distribution corresponding to the 30 people and the 30 targets to be identified in the last frame of the multi-frame image. For example, the identification code (i.e., ID) of the first target to be identified is 1, the identification code (i.e., ID) of the second target to be identified is 2, and so on, and generate the target information corresponding to each identification code.
[0093] In this embodiment, multiple frames of images are processed for recognition to obtain a parameter list for each image; the parameter list includes parameter information of the target to be identified in the image. For two adjacent frames in the multiple images, the target to be identified in the next frame is determined based on the parameter list of the previous frame and the parameter list of the next frame; the target to be identified in the next frame is the target to be identified corresponding to the target to be identified in the previous frame. An identification code corresponding to the target to be identified in the last frame of the multiple images is generated, and target information corresponding to the identification code is generated; the identification code represents the identity of the target to be identified, and the target information represents the number of times the target to be identified performs a preset action. In this scheme, the image includes multiple targets to be identified, and multiple frames of images are processed for recognition to obtain a parameter list for each image, which includes parameter information of the target to be identified in the image. Then, based on two adjacent frames in a multi-frame image dataset, the target to be identified in the next frame of those two adjacent frames can be determined to be the same target to be identified as a certain target to be identified in the previous frame of those two adjacent frames. The target to be identified in the next frame of those two adjacent frames is then used as the target to be identified in the previous frame of the next set of adjacent frames. Combined with the parameter list in the next frame of the next set of adjacent frames, the target to be identified in the next frame of the next set of adjacent frames is determined, and so on, until the target to be identified in the last frame of the multi-frame image dataset is determined. Finally, an identification code corresponding to the target to be identified in the last frame of the multi-frame image dataset is generated, along with target information corresponding to the identification code. Based on this target information, the number of times the target to be identified performs a preset action is displayed. Therefore, the same target to be identified in multiple frames can be identified, and an identification code corresponding to the target to be identified can be generated, greatly improving the stability of tracking the target to be identified and solving the technical problem that the difficulty in tracking the user's identification code leads to the inability to obtain the behavioral information corresponding to the identification code.
[0094] Figure 2 A flowchart illustrating another image processing method provided in this application embodiment is shown below. Figure 2 As shown, the method includes:
[0095] 201. Based on the preset depth target detection network, perform recognition processing on multiple frames of images to obtain a parameter list for each image; wherein, the depth target detection network is used to indicate target recognition information, and the parameter list includes the parameter information of the target to be identified in the image.
[0096] In one example, the parameter information corresponding to the target to be identified includes face information and body information. The face information includes the horizontal coordinate of the face center, the vertical coordinate of the face center, the face width value, the face height value, and the face area. The body information includes the horizontal coordinate of the body center, the vertical coordinate of the body center, the body width value, the body height value, and the body area.
[0097] For example, this step can be referred to Figure 1 Step 101 in the text will not be repeated here.
[0098] 202. Based on the parameter list of the first frame image in the multi-frame image, determine the Gaussian distribution corresponding to the target to be identified in the first frame image; whereby the Gaussian distribution represents the position information of the target to be identified.
[0099] For example, the electronic device initializes N (N = n + 10, where 10 is a reserved placeholder) multidimensional (10-dimensional) Gaussian distributions G based on the parameter list X1 of the first frame of the multi-frame image. The parameter list X1 has a length of n. The Gaussian distribution G represents the mean and variance of each target to be identified. For example, the Gaussian distribution G includes G1, G2, G3, ..., where G1 represents the Gaussian distribution of the first person, G2 represents the Gaussian distribution of the second person, and G3 represents the Gaussian distribution of the third person.
[0100] 203. For two adjacent frames in a multi-frame image, perform probability calculation on the parameter list of the next frame in the two adjacent frames and the Gaussian distribution corresponding to the target to be identified in the previous frame in the two adjacent frames to obtain the probability result information corresponding to the target to be identified in the next frame in the two adjacent frames; wherein, the probability result information is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame in the two adjacent frames.
[0101] 204. Based on the probability results, determine the target to be identified in the next frame of two adjacent frames.
[0102] Step 204 includes three implementation methods:
[0103] The first implementation of step 204: If the probability result information is determined to be the likelihood probability value, where the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames, and the likelihood probability values between the target and multiple Gaussian distributions in the previous frame of two adjacent frames, then the maximum likelihood probability value corresponding to the target to be identified is determined; where the maximum likelihood probability value is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames; based on the maximum likelihood probability value corresponding to the target to be identified, the target to be identified in the next frame of two adjacent frames is determined.
[0104] In one example, the parameter information corresponding to the target to be identified includes the face width value; the first implementation of step 204 further includes: determining the distance between the target to be identified in the next frame of the two adjacent frames and the location of the Gaussian distribution corresponding to the target to be identified, based on the preset first constraint information; wherein, the first constraint information represents the location information corresponding to the target to be identified; if the distance is determined to be greater than the face width value of the target to be identified, then the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames is replaced with the Gaussian distribution corresponding to the target to be identified in the previous frame of the two adjacent frames, to obtain the updated Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames; based on the updated Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames, the target to be identified in the next frame of the two adjacent frames is determined.
[0105] The second implementation of step 204: If the probability result information is determined to be no likelihood probability value, where the likelihood probability value represents the parameter information of the target to be identified in the next frame of the two adjacent frames, and the likelihood probability values between the target and multiple Gaussian distributions in the previous frame of the two adjacent frames, then the target to be identified is determined to be a new target; based on the parameter information corresponding to the new target, the Gaussian distribution corresponding to the new target in the next frame of the two adjacent frames is determined; based on the Gaussian distribution corresponding to the new target in the next frame of the two adjacent frames, the target to be identified in the next frame of the two adjacent frames is determined.
[0106] The third implementation of step 204: If the probability result information indicates that no parameter information corresponding to the Gaussian distribution has been obtained, then the target to be identified corresponding to the Gaussian distribution is determined to be in a disappeared state; a preset number of targets to be identified are determined to be adjacent to the location of the Gaussian distribution, and the parameter information corresponding to the targets to be identified is determined; wherein, the Gaussian distribution includes the first face area, and the parameter information corresponding to the target to be identified includes the second face area; according to the preset second constraint information, among the preset number of second face areas, the target face area equal to the first face area is determined; wherein, the second constraint information represents the pixel ratio of the face area of the target to be identified; according to the parameter information corresponding to the target face area, the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames is determined; according to the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames, the target to be identified in the next frame of the two adjacent frames is determined.
[0107] For example, for two adjacent frames in a multi-frame image, based on a preset formula, the electronic device performs probability calculations on the parameter list of the subsequent frame and the Gaussian distribution corresponding to the target to be identified in the preceding frame, to obtain probability result information for the target to be identified in the subsequent frame. This probability result information indicates the Gaussian distribution corresponding to the target to be identified in the subsequent frame. Based on the probability result information, the target to be identified in the subsequent frame is determined, where the target to be identified in the subsequent frame is the target to be identified corresponding to the target to be identified in the preceding frame.
[0108] For example, in the first implementation of step 204, if the probability result information is determined to be a likelihood probability value, where the likelihood probability value represents the likelihood probability value between the parameter information of each target to be identified in the later frame of the two adjacent frames and multiple Gaussian distributions in the earlier frame of the two adjacent frames, then the maximum likelihood probability value corresponding to the target to be identified is determined among the multiple likelihood probability values. The maximum likelihood probability value indicates the Gaussian distribution corresponding to the target to be identified in the later frame of the two adjacent frames. Finally, based on the maximum likelihood probability value corresponding to the target to be identified, the target to be identified in the later frame of the two adjacent frames is determined, wherein the target to be identified in the later frame of the two adjacent frames is a target to be identified whose Gaussian distribution is equal to that of the target to be identified in the earlier frame of the two adjacent frames.
[0109] After identifying the target in the next frame of two adjacent images, the Gaussian distribution of the target in the next frame can be further verified. Specifically, the parameter information corresponding to the target includes the face width value. The electronic device determines the distance between the target in the next frame and the location of the corresponding Gaussian distribution of the target, based on preset first constraint information. The first constraint information represents the location information of the target. Then, the distance to the target is compared with the face width value of the target. If the distance to the target is greater than the face width value, it indicates that the displacement of the target in the next frame is too large, and the target is not the same as the target in the previous frame. Therefore, the Gaussian distribution corresponding to the target in the next frame is replaced with the Gaussian distribution corresponding to the target in the previous frame, resulting in an updated Gaussian distribution corresponding to the target in the next frame. Finally, based on the Gaussian distribution corresponding to the target to be identified in the next frame of the updated two adjacent frames, the target to be identified in the next frame of the two adjacent frames is determined.
[0110] For example, if the probability result information is determined to be the likelihood probability value, the likelihood probability value represents the likelihood probability between the parameter information of the first target to be identified in the later frame of two adjacent frames and multiple Gaussian distributions in the earlier frame of two adjacent frames. The likelihood probability value includes: the likelihood probability value P between the parameter information of the first target to be identified and the Gaussian distribution G1 in the earlier frame of two adjacent frames. 1,1 The likelihood probability P between the parameter information of the first target to be identified and the Gaussian distribution G2 in the previous frame of the two adjacent frames. 1,2 The likelihood probability P between the parameter information of the first target to be identified and the Gaussian distribution G3 in the previous frame of the two adjacent frames. 1,3 Then, among the multiple likelihood probability values, the maximum likelihood probability value P corresponding to the target to be identified is determined. 1,3 Then, based on the maximum likelihood probability value P corresponding to the target to be identified... 1,3 It can be determined that the Gaussian distribution of the first target to be identified in the second frame of two adjacent frames is G3.
[0111] For example, in the second implementation of step 204, if the probability result information is determined to be that no likelihood probability value was obtained, where the likelihood probability value represents the parameter information of the target to be identified in the later frame of the two adjacent frames, and the likelihood probability values between these parameters and multiple Gaussian distributions in the earlier frame of the two adjacent frames, then the target to be identified in the later frame of the two adjacent frames is determined to be a new target. Then, based on the parameter information corresponding to the new target, the Gaussian distribution corresponding to the new target in the later frame of the two adjacent frames is calculated. Finally, based on the Gaussian distribution corresponding to the new target in the later frame of the two adjacent frames, the target to be identified in the later frame of the two adjacent frames is determined.
[0112] For example, in the third implementation of step 204, if the probability result information is determined to be that no parameter information corresponding to the Gaussian distribution is obtained, then the target to be identified corresponding to the Gaussian distribution is determined to be in a disappearing state. The disappearing state includes the target to be identified being occluded or looking down. A preset number of targets to be identified adjacent to the location of the Gaussian distribution and the parameter information corresponding to the preset number of targets to be identified are determined. The Gaussian distribution includes a first face area, and the parameter information corresponding to the target to be identified includes a second face area. The preset number of targets to be identified can be the three targets to be identified closest to the location of the Gaussian distribution, and there is no limit to the preset number.
[0113] Then, based on the preset second constraint information, the preset number of second face areas are sequentially compared with the first face area to determine the target face area equal to the first face area. The second constraint information represents the pixel proportion of the face area of the target to be identified. Since classroom seating arrangements are usually relatively fixed, forming a matrix layout with multiple rows, when occlusion occurs, students in the front rows typically obscure students in the back rows. Since students in the front rows have a larger pixel proportion in the video frame, the pixel proportion of the face area of the target to be identified can be pre-stored. Finally, based on the parameter information corresponding to the target face area, the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent images is determined. Then, based on the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent images, the target to be identified in the next frame of two adjacent images is determined.
[0114] 205. Generate the identification code corresponding to the target to be identified in the last frame of the multi-frame image, and generate the target information corresponding to the identification code; wherein, the identification code represents the identity of the target to be identified, and the target information represents the number of times the target to be identified performs a preset action.
[0115] For example, this step can be referred to Figure 1 Step 103 in the text will not be repeated here.
[0116] In this embodiment, multiple frames of images are processed for recognition based on a preset depth target detection network to obtain a parameter list for each image. The depth target detection network indicates target recognition information, and the parameter list includes parameter information of the target to be identified in the image. Based on the parameter list of the first frame in the multiple images, a Gaussian distribution corresponding to the target to be identified in the first frame is determined; the Gaussian distribution represents the position information of the target to be identified. For two adjacent frames in the multiple images, the parameter list of the subsequent frame is compared with the Gaussian distribution corresponding to the target to be identified in the preceding frame to obtain probability result information for the target to be identified in the subsequent frame; the probability result information indicates the Gaussian distribution corresponding to the target to be identified in the subsequent frame. Based on the probability result information, the target to be identified in the subsequent frame is determined. The system generates an identification code corresponding to the target in the last frame of a multi-frame image dataset, and generates target information corresponding to the identification code. The identification code represents the identity of the target, and the target information represents the number of times the target performs a preset action. Therefore, by combining multiple constraint conditions, the same target located in multiple frames can be identified, and its corresponding identification code can be generated. This greatly improves the stability of target tracking and solves the technical problem of the difficulty in obtaining behavioral information corresponding to the identification code due to the difficulty in tracking the user's identification code.
[0117] Figure 3 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application, as shown below. Figure 3 As shown, the device includes:
[0118] The recognition unit 31 is used to perform recognition processing on multiple frames of images to obtain a parameter list for each image; wherein the parameter list includes parameter information of the target to be recognized in the image.
[0119] The first determining unit 32 is used to determine the target to be identified in the next frame of two adjacent frames of images based on the parameter list of the previous frame of the two adjacent frames and the parameter list of the next frame of the two adjacent frames; wherein the target to be identified in the next frame of two adjacent frames is the target to be identified corresponding to the target to be identified in the previous frame of the two adjacent frames.
[0120] The first generation unit 33 is used to generate the identification code corresponding to the target to be identified in the last frame of the multi-frame image.
[0121] The second generation unit 34 is used to generate target information corresponding to the identification code; wherein, the identification code represents the identity of the target to be identified, and the target information represents the number of times the target to be identified has performed a preset action.
[0122] The apparatus in this embodiment can execute the technical solutions in the above method. Its specific implementation process and technical principles are the same, and will not be repeated here.
[0123] Figure 4 This is a schematic diagram of the structure of another image processing apparatus provided in an embodiment of this application. Figure 3 Based on the illustrated embodiments, as Figure 4 As shown, the first determining unit 32 includes:
[0124] The first determining module 321 is used to determine the Gaussian distribution corresponding to the target to be identified in the first frame image based on the parameter list of the first frame image in the multi-frame images; wherein the Gaussian distribution represents the position information of the target to be identified.
[0125] The calculation module 322 is used to perform probability calculation on the parameter list of the next frame in the two adjacent frames of the multi-frame image and the Gaussian distribution corresponding to the target to be identified in the previous frame in the two adjacent frames, to obtain the probability result information corresponding to the target to be identified in the next frame in the two adjacent frames; wherein, the probability result information is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame in the two adjacent frames.
[0126] The second determining module 323 is used to determine the target to be identified in the next frame of two adjacent frames based on the probability result information.
[0127] In one example, the second determining module 323 includes:
[0128] The first determining submodule 3231 is used to determine the maximum likelihood probability value corresponding to the target to be identified if the determined probability result information is a likelihood probability value, wherein the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames, and the likelihood probability values between the target and multiple Gaussian distributions in the previous frame of two adjacent frames; wherein the maximum likelihood probability value is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames.
[0129] The second determining submodule 3232 is used to determine the target to be identified in the next frame of two adjacent frames based on the maximum likelihood probability value corresponding to the target to be identified.
[0130] In one example, the parameter information corresponding to the target to be identified includes the face width value; the device also includes:
[0131] The second determining unit 41 is used to determine the distance between the target to be identified in the next frame of two adjacent frames and the location of the Gaussian distribution corresponding to the target to be identified, based on the preset first constraint information; wherein the first constraint information represents the location information corresponding to the target to be identified.
[0132] The update unit 42 is used to replace the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames with the Gaussian distribution corresponding to the target to be identified in the previous frame of the two adjacent frames if the distance is determined to be greater than the face width value of the target to be identified, so as to obtain the updated Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames.
[0133] The third determining unit 43 is used to determine the target to be identified in the next frame of two adjacent frames based on the Gaussian distribution corresponding to the target to be identified in the next frame of the updated two adjacent frames.
[0134] In one example, the second determining module 323 includes:
[0135] The third determining submodule 3233 is used to determine the target to be identified as a new target if the probability result information is that no likelihood probability value is obtained, wherein the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames and the likelihood probability value between the target and multiple Gaussian distributions in the previous frame of two adjacent frames.
[0136] The fourth determination submodule 3234 is used to determine the Gaussian distribution corresponding to the newly added target in the next frame of two adjacent frames based on the parameter information corresponding to the newly added target.
[0137] The fifth determining submodule 3235 is used to determine the target to be identified in the next frame of two adjacent frames based on the Gaussian distribution corresponding to the newly added target in the next frame of two adjacent frames.
[0138] In one example, the second determining module 323 includes:
[0139] The sixth determination submodule 3236 is used to determine that the target to be identified corresponding to the Gaussian distribution is in a disappeared state if the determination probability result information is that no parameter information corresponding to the Gaussian distribution is obtained.
[0140] The seventh determination submodule 3237 is used to determine a preset number of targets to be identified that are adjacent to the location of the Gaussian distribution, as well as the parameter information corresponding to the targets to be identified; wherein, the Gaussian distribution includes the first face area, and the parameter information corresponding to the targets to be identified includes the second face area.
[0141] The eighth determining submodule 3238 is used to determine a target face area that is equal to the first face area from a preset number of second face areas based on preset second constraint information; wherein, the second constraint information represents the pixel ratio of the face area of the target to be identified.
[0142] The ninth determining submodule 3239 is used to determine the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames based on the parameter information corresponding to the target face area.
[0143] The tenth determination submodule 32310 is used to determine the target to be identified in the next frame of two adjacent frames based on the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames.
[0144] In one example, recognition unit 31 is specifically used for:
[0145] Based on the preset depth target detection network, multiple frames of images are processed for recognition to obtain a parameter list for each image; the depth target detection network is used to indicate target recognition information, and the parameter list includes parameter information of the target to be identified in the image.
[0146] In one example, the parameter information corresponding to the target to be identified includes face information and body information. The face information includes the horizontal coordinate of the face center, the vertical coordinate of the face center, the face width value, the face height value, and the face area. The body information includes the horizontal coordinate of the body center, the vertical coordinate of the body center, the body width value, the body height value, and the body area.
[0147] The apparatus in this embodiment can execute the technical solutions in the above method. Its specific implementation process and technical principles are the same, and will not be repeated here.
[0148] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 5 As shown, the electronic device includes: a memory 51 and a processor 52.
[0149] The memory 51 stores a computer program that can run on the processor 52.
[0150] The processor 52 is configured to perform the methods provided in the embodiments described above.
[0151] The electronic device also includes a receiver 53 and a transmitter 54. The receiver 53 is used to receive instructions and data sent by external devices, and the transmitter 54 is used to send instructions and data to external devices.
[0152] Figure 6This is a block diagram of an electronic device provided in an embodiment of this application. The electronic device may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0153] The device 600 may include one or more of the following components: a processing component 602, a memory 604, a power supply component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.
[0154] Processing component 602 typically controls the overall operation of device 600, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 602 may include one or more modules to facilitate interaction between processing component 602 and other components. For example, processing component 602 may include a multimedia module to facilitate interaction between multimedia component 608 and processing component 602.
[0155] Memory 604 is configured to store various types of data to support the operation of device 600. Examples of such data include instructions for any application or method operating on device 600, contact data, phonebook data, messages, pictures, videos, etc. Memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0156] Power supply component 606 provides power to the various components of device 600. Power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 600.
[0157] Multimedia component 608 includes a screen that provides an output interface between device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 608 includes a front-facing camera and / or a rear-facing camera. When device 600 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0158] Audio component 610 is configured to output and / or input audio signals. For example, audio component 610 includes a microphone (MIC) configured to receive external audio signals when device 600 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 604 or transmitted via communication component 616. In some embodiments, audio component 610 also includes a speaker for outputting audio signals.
[0159] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0160] Sensor assembly 614 includes one or more sensors for providing state assessments of various aspects of device 600. For example, sensor assembly 614 may detect the on / off state of device 600, the relative positioning of components such as the display and keypad of device 600, changes in the position of device 600 or a component of device 600, the presence or absence of user contact with device 600, the orientation or acceleration / deceleration of device 600, and temperature changes of device 600. Sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 614 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 614 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0161] Communication component 616 is configured to facilitate wired or wireless communication between device 600 and other devices. Device 600 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 616 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 616 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0162] In an exemplary embodiment, the apparatus 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0163] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, which can be executed by a processor 620 of the device 600 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0164] This application also provides a non-transitory computer-readable storage medium, which, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform the methods provided in the above embodiments.
[0165] This application also provides a computer program product, which includes: a computer program stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to perform the solution provided in any of the above embodiments.
[0166] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0167] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, include: Multiple frames of images are processed for recognition to obtain a parameter list for each image; wherein the parameter list includes parameter information of the target to be recognized in the image; Based on the parameter list of the first frame image in the multi-frame image, the Gaussian distribution corresponding to the target to be identified in the first frame image is determined; wherein, the Gaussian distribution represents the position information of the target to be identified; For two adjacent frames in the multi-frame image, the parameter list of the next frame in the two adjacent frames is used to perform probability calculation with the Gaussian distribution corresponding to the target to be identified in the previous frame in the two adjacent frames, to obtain the probability result information corresponding to the target to be identified in the next frame in the two adjacent frames; wherein, the probability result information is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame in the two adjacent frames. Based on the probability result information, the target to be identified in the next frame of the two adjacent frames is determined; wherein, the target to be identified in the next frame of the two adjacent frames is the target to be identified corresponding to the target to be identified in the previous frame of the two adjacent frames. Generate an identification code corresponding to the target to be identified in the last frame of the multi-frame images, and generate target information corresponding to the identification code; wherein, the identification code represents the identity of the target to be identified, and the target information represents the number of times the target to be identified performs a preset action; Based on the probability result information, the target to be identified in the next frame of the two adjacent frames is determined, including: If the probability result information is determined to be that no parameter information corresponding to the Gaussian distribution is obtained, then the target to be identified corresponding to the Gaussian distribution is determined to be in a disappeared state; A predetermined number of targets to be identified are located adjacent to the location of the Gaussian distribution, along with the parameter information corresponding to the targets to be identified; wherein the Gaussian distribution includes a first face area, and the parameter information corresponding to the targets to be identified includes a second face area; Based on the preset second constraint information, a target face area equal to the first face area is determined from a preset number of second face areas; wherein, the second constraint information represents the pixel ratio of the face area of the target to be identified; Based on the parameter information corresponding to the target face area, determine the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames; The target to be identified in the next frame of the two adjacent frames is determined based on the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames.
2. The method according to claim 1, characterized in that, Based on the probability result information, determining the target to be identified in the next frame of two adjacent frames further includes: If the probability result information is determined to be a likelihood probability value, wherein the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames, and the likelihood probability values between the target and multiple Gaussian distributions in the previous frame of two adjacent frames, then the maximum likelihood probability value corresponding to the target to be identified is determined; wherein the maximum likelihood probability value is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames. Based on the maximum likelihood probability value corresponding to the target to be identified, the target to be identified in the next frame of two adjacent frames is determined.
3. The method according to claim 2, characterized in that, The parameter information corresponding to the target to be identified includes the face width value; the method further includes: Based on the preset first constraint information, the distance between the target to be identified in the next frame of two adjacent frames and the location of the Gaussian distribution corresponding to the target to be identified is determined; wherein, the first constraint information represents the location information corresponding to the target to be identified. If it is determined that the distance is greater than the face width of the target to be identified, then the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames is replaced with the Gaussian distribution corresponding to the target to be identified in the previous frame of the two adjacent frames, so as to obtain the updated Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames. Based on the Gaussian distribution corresponding to the target to be identified in the next frame of the updated two adjacent frames, the target to be identified in the next frame of the two adjacent frames is determined.
4. The method according to claim 1, characterized in that, Based on the probability result information, determining the target to be identified in the next frame of two adjacent frames further includes: If the probability result information is determined to be no likelihood probability value, wherein the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames, and the likelihood probability value between the target and multiple Gaussian distributions in the previous frame of two adjacent frames, then the target to be identified is determined to be a new target. Based on the parameter information corresponding to the newly added target, determine the Gaussian distribution corresponding to the newly added target in the next frame of two adjacent frames; The target to be identified in the next frame of the two adjacent frames is determined based on the Gaussian distribution corresponding to the newly added target in the next frame of the two adjacent frames.
5. The method according to claim 1, characterized in that, The multi-frame image is processed for recognition to obtain a parameter list for each image, including: According to the preset depth target detection network, multiple frames of images are processed for recognition to obtain a parameter list for each image; wherein, the depth target detection network is used to indicate target recognition information, and the parameter list includes parameter information of the target to be recognized in the image.
6. The method according to any one of claims 1-5, characterized in that, The parameter information corresponding to the target to be identified includes face information and body information. The face information includes the horizontal coordinate of the face center, the vertical coordinate of the face center, the face width value, the face height value, and the face area. The body information includes the horizontal coordinate of the body center, the vertical coordinate of the body center, the body width value, the body height value, and the body area.
7. An image processing apparatus, characterized in that, include: The recognition unit is used to perform recognition processing on multiple frames of images to obtain a parameter list for each image; wherein the parameter list includes parameter information of the target to be recognized in the image; The first determining module is used to determine the Gaussian distribution corresponding to the target to be identified in the first frame image based on the parameter list of the first frame image in the multi-frame images; wherein the Gaussian distribution represents the position information of the target to be identified; The calculation module is used to perform probability calculation on the parameter list of the next frame in the multi-frame images and the Gaussian distribution corresponding to the target to be identified in the previous frame in the multi-frame images, to obtain probability result information corresponding to the target to be identified in the next frame in the multi-frame images; wherein, the probability result information is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame in the multi-frame images. The second determining module is used to determine the target to be identified in the next frame of the two adjacent frames based on the probability result information; wherein the target to be identified in the next frame of the two adjacent frames is the target to be identified corresponding to the target to be identified in the previous frame of the two adjacent frames. The first generation unit is used to generate the identification code corresponding to the target to be identified in the last frame of the multi-frame images; The second generation unit is used to generate target information corresponding to the identification code; wherein the identification code represents the identity identifier of the target to be identified, and the target information represents the number of times the target to be identified has performed a preset action; The second determining module includes: The sixth determining submodule is used to determine that the target to be identified corresponding to the Gaussian distribution is in a disappeared state if the probability result information is determined to be that no parameter information corresponding to the Gaussian distribution is obtained. The seventh determination submodule is used to determine a preset number of targets to be identified that are adjacent to the location of the Gaussian distribution, as well as the parameter information corresponding to the targets to be identified; wherein, the Gaussian distribution includes a first face area, and the parameter information corresponding to the targets to be identified includes a second face area; The eighth determining submodule is used to determine a target face area that is equal to the first face area from a preset number of second face areas based on preset second constraint information; wherein, the second constraint information represents the pixel ratio of the face area of the target to be identified. The ninth determining submodule is used to determine the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames based on the parameter information corresponding to the target face area. The tenth determination submodule is used to determine the target to be identified in the next frame of the two adjacent frames based on the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames.
8. The apparatus according to claim 7, characterized in that, The second determining module includes: The first determining submodule is configured to determine the maximum likelihood probability value corresponding to the target to be identified if the probability result information is determined to be a likelihood probability value, wherein the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames, and the likelihood probability values between the target and multiple Gaussian distributions in the previous frame of two adjacent frames; wherein the maximum likelihood probability value is used to indicate the Gaussian distribution corresponding to the target to be identified in the next frame of two adjacent frames. The second determination submodule is used to determine the target to be identified in the next frame of two adjacent frames based on the maximum likelihood probability value corresponding to the target to be identified.
9. The apparatus according to claim 8, characterized in that, The parameter information corresponding to the target to be identified includes the face width value; the device also includes: The second determining unit is used to determine the distance between the target to be identified in the next frame of two adjacent frames and the location of the Gaussian distribution corresponding to the target to be identified, based on the preset first constraint information; wherein, the first constraint information represents the location information corresponding to the target to be identified. The update unit is used to replace the Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames with the Gaussian distribution corresponding to the target to be identified in the previous frame of the two adjacent frames if it is determined that the distance is greater than the face width value of the target to be identified, so as to obtain the updated Gaussian distribution corresponding to the target to be identified in the next frame of the two adjacent frames. The third determining unit is used to determine the target to be identified in the next frame of two adjacent frames based on the Gaussian distribution corresponding to the target to be identified in the next frame of the updated two adjacent frames.
10. The apparatus according to claim 7, characterized in that, The second determining module includes: The third determining submodule is used to determine that if the probability result information is that no likelihood probability value is obtained, wherein the likelihood probability value represents the parameter information of the target to be identified in the next frame of two adjacent frames, and the likelihood probability value between the target and multiple Gaussian distributions in the previous frame of two adjacent frames, the target to be identified is a new target. The fourth determination submodule is used to determine the Gaussian distribution corresponding to the newly added target in the next frame of two adjacent frames based on the parameter information corresponding to the newly added target. The fifth determination submodule is used to determine the target to be identified in the next frame of the two adjacent frames based on the Gaussian distribution corresponding to the newly added target in the next frame of the two adjacent frames.
11. The apparatus according to claim 7, characterized in that, The identification unit is specifically used for: According to the preset depth target detection network, multiple frames of images are processed for recognition to obtain a parameter list for each image; wherein, the depth target detection network is used to indicate target recognition information, and the parameter list includes parameter information of the target to be recognized in the image.
12. The apparatus according to any one of claims 7-11, characterized in that, The parameter information corresponding to the target to be identified includes face information and body information. The face information includes the horizontal coordinate of the face center, the vertical coordinate of the face center, the face width value, the face height value, and the face area. The body information includes the horizontal coordinate of the body center, the vertical coordinate of the body center, the body width value, the body height value, and the body area.
13. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the method of any one of claims 1-6.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.
15. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-6.
Citation Information
Patent Citations
Monitoring method and device based on face recognition, storage medium and computer equipment
CN110580470A