A hand-raising behavior recognition method and device, and an electronic device

By combining the recognition results of the hand-raising training model and the face training model, the problem of misjudgment by the hand-raising behavior tracker in densely populated scenes was solved, and a high accuracy rate of hand-raising behavior recognition was achieved.

CN111382655BActive Publication Date: 2026-01-13SHENZHEN HONGHE INNOVATION INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201910161167.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-03-04
Publication Date
2026-01-13
Estimated Expiration
2039-03-04

AI Technical Summary

Technical Problem

Existing hand-raising behavior trackers are prone to misjudgment and missed judgment in crowded scenes, resulting in a decrease in the accuracy of hand-raising behavior recognition.

Method used

The image is recognized by using a hand-raising training model and a face training model respectively, generating a target hand-raising prediction result set and a target face prediction result set. The two result sets are then matched by matching conditions to output the hand-raising behavior recognition result.

Benefits of technology

It improves the accuracy of hand-raising behavior recognition in crowded scenarios, and can accurately identify the hand-raising behavior of multiple people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111382655B_ABST
    Figure CN111382655B_ABST
Patent Text Reader

Abstract

The application discloses a hand-raising behavior recognition method and device and electronic equipment, comprising: inputting a to-be-recognized image; recognizing the to-be-recognized image by using a hand-raising training model to obtain a target hand-raising prediction result set; a hand-raising behavior tracker respectively labels at least one hand-raising human object according to the target hand-raising prediction result set to generate labeling information of each human object; recognizing the to-be-recognized image by using a face training model to obtain a target face prediction result set; a face tracker respectively labels at least one face object according to the target face prediction result set to generate labeling information of each face object; matching the target hand-raising prediction result set and the target face prediction result set according to a matching condition; and outputting a hand-raising behavior recognition result for the target hand-raising prediction result and the target face prediction result satisfying the matching condition. The application can improve the accuracy of hand-raising behavior recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent identification, in particular to a hand-raising behavior identification method and device and electronic equipment. BACKGROUND

[0002] At present, a hand-raising behavior tracker can be created according to a hand-raising action by using artificial intelligence technology to identify the hand-raising behavior, and the hand-raising behavior tracker is used to track hand actions to realize rapid positioning and identification of the hand-raising behavior. However, the above hand-raising behavior tracker is only applicable to identification of individual hand-raising behavior, and in a crowded scene with many people raising their hands, the above hand-raising behavior tracker will misjudge or miss the judgment, thereby reducing the identification accuracy of the hand-raising behavior. SUMMARY

[0003] Therefore, the present application aims to provide a hand-raising behavior identification method and device and electronic equipment to improve the identification accuracy of the hand-raising behavior.

[0004] To achieve the above purpose, the present application provides a hand-raising behavior identification method, which comprises:

[0005] inputting a to-be-identified image;

[0006] identifying the to-be-identified image by using a hand-raising training model to obtain a target hand-raising prediction result set comprising at least one group of target hand-raising prediction results;

[0007] inputting the target hand-raising prediction result set into a hand-raising behavior tracker, and the hand-raising behavior tracker respectively labels at least one hand-raising human object according to the target hand-raising prediction result set to generate labeling information of each human object;

[0008] identifying the to-be-identified image by using a face training model to obtain a target face prediction result set comprising at least one group of target face prediction results;

[0009] inputting the target face prediction result set into a face tracker, and the face tracker respectively labels at least one face object according to the target face prediction result set to generate labeling information of each face object;

[0010] matching the target hand-raising prediction result set and the target face prediction result set according to a matching condition;

[0011] outputting a hand-raising behavior identification result for the target hand-raising prediction result and the corresponding target face prediction result that meet the matching condition.

[0012] Optionally, the target hand-raising prediction result set comprises at least one target hand-raising prediction result of a human body object, and the target hand-raising prediction result comprises reference point coordinates, height, width, and recognition confidence of the human body object.

[0013] Optionally, the annotation information of each human body object comprises a human body frame of each human body object and a human body tracking identifier of each human body object, and the human body tracking identifier corresponds to the human body object in one-to-one correspondence.

[0014] Optionally, the target face prediction result set comprises at least one target face prediction result of a face object, and the target face prediction result comprises reference point coordinates, height, width, and recognition confidence of the face object.

[0015] Optionally, the annotation information of each face comprises a face frame of each face object and a face tracking identifier of each face object, and the face tracking identifier corresponds to the face object in one-to-one correspondence.

[0016] Optionally, the matching condition is that the face object is within the range of the human body object.

[0017] Optionally, the method further comprises: for the target hand-raising prediction result and the corresponding target face prediction result that satisfy the matching condition, recording the annotation information of the human body object and the annotation information of the face object corresponding thereto, and for a plurality of images to be recognized that are continuously input within a predetermined time, judging whether the plurality of images to be recognized satisfy the matching condition according to the recorded annotation information of the human body object and the annotation information of the face object, and if yes, inputting a hand-raising behavior recognition result.

[0018] The embodiment of the application further provides a hand-raising behavior recognition device, comprising:

[0019] a hand-raising training module configured to recognize the image to be recognized to obtain a target hand-raising prediction result set comprising at least one target hand-raising prediction result;

[0020] a hand-raising behavior tracking module configured to respectively annotate at least one human body object that raises hands according to the input target hand-raising prediction result set to generate annotation information of each human body object;

[0021] a face training module configured to recognize the image to be recognized to obtain a target face prediction result set comprising at least one target face prediction result;

[0022] a face tracking module configured to respectively annotate at least one face object according to the input target face prediction result set to generate annotation information of each face object;

[0023] The matching module is configured to match the target hand-raising prediction result set and the target face prediction result set according to a matching condition.

[0024] The output module is configured to output a hand-raising behavior recognition result for the target hand-raising prediction result and the corresponding target face prediction result that satisfy the matching condition.

[0025] Optionally, the target hand-raising prediction result set includes at least one target hand-raising prediction result of a human object, and the target hand-raising prediction result includes a reference point coordinate, a height, a width, and a recognition confidence of the human object.

[0026] Optionally, the annotation information of each human object includes a human body frame of each human object and a human body tracking identifier of each human object, and the human body tracking identifier is in one-to-one correspondence with the human object.

[0027] Optionally, the target face prediction result set includes at least one target face prediction result of a face object, and the target face prediction result includes a reference point coordinate, a height, a width, and a recognition confidence of the face object.

[0028] Optionally, the annotation information of each face object includes a face frame of each face object and a face tracking identifier of each face object, and the face tracking identifier is in one-to-one correspondence with the face object.

[0029] Optionally, the matching condition is that the face object is within the range of the human object.

[0030] Optionally, the device further includes:

[0031] The recording module is configured to record the annotation information of the human object corresponding to the target hand-raising prediction result and the annotation information of the face object corresponding to the target face prediction result that satisfy the matching condition.

[0032] The output module is configured to output a hand-raising behavior recognition result for the target hand-raising prediction result and the corresponding target face prediction result that satisfy the matching condition.

[0033] The embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the hand-raising behavior recognition method when executing the program.

[0034] From the above, it can be seen that the hand-raising behavior recognition method and device and electronic equipment provided by the application utilize a hand-raising training model and a face training model to respectively recognize a to-be-recognized image, respectively obtain a target hand-raising prediction result set and a target face prediction result set, input the target hand-raising prediction result set and the target face prediction result set into a hand-raising behavior tracker and a face tracker, respectively track and label at least one body object and at least one face object by using the two trackers, match at least one body object according to the target hand-raising prediction result set and the target face prediction result set according to a matching condition, and output a hand-raising behavior recognition result of at least one hand-raising body object for the target hand-raising prediction result and the corresponding target face prediction result that meet the matching condition. The application can realize the recognition of the hand-raising behavior of at least one hand-raising body object, and has a high recognition accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, brief introductions will be given to the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort.

[0036] Figure 1 The method flowchart of the embodiment of the present application is shown in the figure.

[0037] Figure 2 The figure of the body frame and the face frame of the embodiment of the present application is shown in the figure.

[0038] Figure 3 The device structure schematic diagram of the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0039] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0040] It should be noted that all the expressions of "first" and "second" in the embodiments of the present application are used to distinguish two same name but different entities or different parameters. It can be seen that "first" and "second" are only used for the convenience of description, and should not be understood as a limitation of the embodiments of the present application. The subsequent embodiments will not be described one by one.

[0041] The embodiment of the present application provides a hand-raising behavior recognition method, which can realize the recognition of individual hand-raising behavior, and can also realize the recognition of the hand-raising behavior of each hand-raising person in a crowded scene with many hand-raising persons. The hand-raising behavior recognition method comprises:

[0042] inputting a to-be-recognized image;

[0043] recognize the to-be-recognized image by using the hand-raising training model to obtain a target hand-raising prediction result set including at least one group of target hand-raising prediction results;

[0044] input the target hand-raising prediction result set into the hand-raising behavior tracker, and the hand-raising behavior tracker respectively labels at least one hand-raising human object according to the target hand-raising prediction result set to generate labeling information of each human object;

[0045] recognize the to-be-recognized image by using the face training model to obtain a target face prediction result set including at least one group of target face prediction results;

[0046] input the target face prediction result set into the face tracker, and the face tracker respectively labels at least one face object according to the target face prediction result set to generate labeling information of each face object;

[0047] match the target hand-raising prediction result set and the target face prediction result set according to a matching condition, and output a hand-raising behavior recognition result for the target hand-raising prediction result and the corresponding target face prediction result that meet the matching condition.

[0048] The hand-raising behavior recognition method provided in the embodiments of the present application respectively recognizes a to-be-recognized image by using a hand-raising training model and a face training model, and the two models respectively output a target hand-raising prediction result set and a target face prediction result set, wherein the target hand-raising prediction result set includes a hand-raising prediction result of at least one hand-raising human object, and the target face prediction result set includes a face prediction result of at least one face object; the target hand-raising prediction result set and the target face prediction result set are respectively input into a hand-raising behavior tracker and a face tracker, and the two trackers are respectively used to track and label at least one human object and at least one face object; the target hand-raising prediction result set and the target face prediction result set are matched according to a matching condition for at least one human object, and a hand-raising behavior recognition result of at least one hand-raising human object is output for the target hand-raising prediction result and the corresponding target face prediction result that meet the matching condition. The present application can realize the recognition of the hand-raising behavior of at least one hand-raising human object, and has a high recognition accuracy.

[0049] Figure 1 The method flowchart of the embodiments of the present application is shown in the figure. As shown in the figure, the hand-raising behavior recognition method of the embodiments of the present application includes:

[0050] S10: input a to-be-recognized image;

[0051] The to-be-recognized image is each frame image in a video stream. In the embodiment of the present application, an image acquisition device is used to acquire a video stream in its shooting range, each frame image is extracted from the video stream, and each frame image is preprocessed to obtain an image suitable for model recognition processing, and the preprocessed image is taken as the to-be-recognized image. The image acquisition device can be installed in a classroom, a conference room, a lecture hall or the like. For the image acquisition device installed in the classroom, the method of the present application can be used to recognize a plurality of students raising hands in a class, and the subsequent number of students raising hands can be used to evaluate the classroom activity level and the teaching level.

[0052] S11: recognizing the to-be-recognized image by using the hand-raising training model to obtain a target hand-raising prediction result set;

[0053] Based on the deep learning algorithm model, a plurality of hand-raising persons in a crowded place are taken as training samples to train the deep learning algorithm model to generate a hand-raising training model. The number of hand-raising persons in the training samples can be configured according to specific application scenarios.

[0054] The to-be-recognized image is recognized by using the hand-raising training model to obtain a hand-raising prediction result set. The hand-raising prediction result set includes at least one hand-raising prediction result of a human object, and each group of hand-raising prediction results includes the reference point coordinates, height, width, recognition confidence and the like of the human object. According to a preset hand-raising behavior threshold, when the recognition confidence is greater than or equal to the hand-raising behavior threshold, a group of hand-raising prediction results corresponding to the recognition confidence is taken as a target hand-raising prediction result, and the target hand-raising prediction result set is composed of at least one target hand-raising prediction result.

[0055] S12: inputting the target hand-raising prediction result set into a hand-raising behavior tracker, and respectively labeling a corresponding human body frame and a human body tracking identifier for at least one hand-raising human object by using the hand-raising behavior tracker;

[0056] The target hand-raising prediction result set is inputted into the hand-raising behavior tracker, the hand-raising behavior tracker determines the number of hand-raising human objects and the positions of the human objects according to the number of groups of target hand-raising prediction results in the target hand-raising prediction result set, and respectively labels each human object, including labeling a human body frame of each human object and labeling a human body tracking identifier of each human object, the human body tracking identifier corresponding to the human object in one-to-one correspondence.

[0057] The tracking and labeling of the human object by the hand-raising behavior tracker set a moving area, and the movement of the human object within the moving area is also determined as the same human object. The moving area is, for example, within a range of 20 pixels.

[0058] S13: recognizing the to-be-recognized image by using the face training model to obtain a target face prediction result set;

[0059] The deep learning algorithm model is trained based on the face of the multiple hands-up persons in the personnel-intensive place as the training sample, and a face training model is generated.

[0060] The face training model is used to recognize the to-be-recognized image, and a face prediction result set is obtained.

[0061] S14: The target face prediction result set is input into a face tracker, and the face tracker is used to respectively label a corresponding face frame and a face tracking identifier for each face object;

[0062] The target face prediction result set is input into the face tracker, the face tracker determines the number of face objects and the positions of the face objects according to the number of groups of target face prediction results in the target face prediction result set, and labels each face object, including labeling a face frame of each face object and labeling a face tracking identifier of each face object, the face tracking identifier corresponding to the face object.

[0063] S15: According to the target hand-up prediction result set and the target face prediction result set, matching is performed according to a matching condition, and for the target hand-up prediction result and the corresponding target face prediction result that match successfully, step S16 is performed.

[0064] In the embodiment of the application, the target hand-up prediction result set includes at least one group of target hand-up prediction results, and the target face prediction result set includes at least one group of target face prediction results, one group of target hand-up prediction results is matched with each group of target face prediction results, and when a group of target hand-up prediction results and a group of face prediction results satisfy the matching condition, the human body tracking identifier corresponding to the group of target hand-up prediction results and the face tracking identifier corresponding to the group of face prediction results are recorded.

[0065] According to the above process, each group of target hand-up prediction results in the target hand-up prediction result set is sequentially matched with each group of target face prediction results in the target face prediction result set, and for the matching groups that satisfy the matching condition, the corresponding human body tracking identifier and face tracking identifier are recorded.

[0066] S16: The human body tracking identifier and the face tracking identifier respectively corresponding to the target hand-up prediction result and the corresponding target face prediction result that match successfully are recorded.

[0067] S17: performing the recognition and matching process on the continuously inputted to-be-recognized images, and when the hand-raising behavior condition is met, performing step S18, and if the hand-raising recognition condition is not met, continuing the matching;

[0068] The hand-raising recognition condition is set, and when the hand-raising behavior is recognized within a continuous time, it is determined that the hand-raising behavior condition is met, and the recognition result of the hand-raising behavior condition is outputted.

[0069] Within a continuous time (for example, 5 seconds), N to-be-recognized images (for example, N images within the continuous time) are continuously collected, and the N to-be-recognized images collected continuously are recognized and matched according to the above steps S10-S16. In the continuously to-be-recognized images, the target hand-raising prediction result and the target face prediction result corresponding to each to-be-recognized image that meets the matching condition are determined as the recognized hand-raising personnel.

[0070] For the two to-be-recognized images continuously, the hand-raising prediction result corresponding to the body tracking identifier and the face tracking identifier meeting the matching condition in the last to-be-recognized image are recorded, and it is judged whether the hand-raising prediction result corresponding to the body tracking identifier in the current to-be-recognized image is the target hand-raising prediction result, whether the face prediction result corresponding to the face tracking identifier is the target face prediction result, and whether the target hand-raising prediction result and the target face prediction result corresponding to the body tracking identifier and the face tracking identifier meet the matching condition.

[0071] S18: outputting the hand-raising recognition result.

[0072] The outputted hand-raising recognition result includes the information of at least one hand-raising body object, including the position, the image in the body frame, etc.

[0073] In some embodiments, the recognized hand-raising body objects can be sorted according to the time, and the information of the hand-raising body objects ranked in the front is outputted. For example, in the answering scene in the classroom teaching process, the teacher can be assisted to quickly identify and select the students who raise their hands first.

[0074] The recognition and matching process of the embodiments of the present application are exemplarily described below in combination with a specific embodiment. Figure 2 The schematic diagram of the body frame and the face frame of the embodiments of the present application is shown in the figure. The inputted to-be-recognized image is inputted into the hand-raising training model, and the hand-raising training result of at least one body object is outputted. Each group of hand-raising prediction results includes the reference point A coordinate A (x, y) of the body object, the height H A , and the width W A, the recognition confidence; taking a group of hand raising prediction results with the recognition confidence greater than or equal to the hand raising behavior threshold as target hand raising prediction results, and taking at least one group of target hand raising prediction results as a target hand raising prediction result set. The target hand raising prediction result set is input into a hand raising behavior tracker, and the hand raising behavior tracker determines the number of human body objects raising hands and the positions of the human body objects according to the number of groups of target hand raising prediction results, labels a human body frame 20 for each human body object, and labels a human body tracking identifier 21.

[0075] The image to be recognized is input into the face training model, and face training results of at least one face object are output, each group of face training results includes the reference point B coordinate B(x1, y1), height H B , width W B , and recognition confidence of the face object. A The recognition confidence greater than or equal to the face threshold is taken as a target face prediction result, and at least one group of target face prediction results is taken as a target face prediction result set. The target face prediction result set is input into a face tracker, and the face tracker determines the number of face objects and the positions of the face objects according to the number of groups of target face prediction results, labels a face frame 22 for each face object, and labels a face tracking identifier 23. It should be noted that the human body tracking identifier is an identifier of each human body object identified, the human body object and the human body tracking identifier correspond to each other, the face tracking identifier is an identifier of each face object identified, the face object and the face tracking identifier correspond to each other, and the human body tracking identifier and the face tracking identifier can be the same or different for the same human body, and the two are not related.

[0076] When one group of target hand raising prediction results and one group of target face prediction results are matched according to the target hand raising prediction result set and the target face prediction result set, it is judged whether the face object is within the human body object range, that is, whether the face frame 22 is within the human body frame 20, if yes, it is judged as matching, otherwise, it is judged as not matching.

[0077] The specific method is that, in order to improve the recognition accuracy, first, the human body frame 20 is expanded by M pixels to form a matching frame 24, the reference point C coordinate of the matching frame 24 is obtained according to the reference point A coordinate of the human body frame 20, the diagonal point C1 coordinate of the matching frame 24 is C1(x+W A +M, y+H A +M); secondly, the center point D coordinate D(x2, y2) of the human body object is calculated according to the reference point C coordinate and the diagonal point C1 coordinate of the matching frame; the center point E coordinate E(x3, y3) of the face object is calculated according to the reference point B coordinate B(x1, y1), height H B , and width W B of the face frame.

[0078] According to the center point coordinate D(x2, y2) of the human body object and the center point coordinate E(x3, y3) of the face object, the following is calculated:

[0079] |x2-x3|<W A / 2 (1)

[0080] |y2-y3|<H A / 2 (2)

[0081] If the formulas (1) and (2) are satisfied at the same time, it is determined that the face frame is in the human body frame, and it is determined that the matching condition is satisfied.

[0082] For a group of hand-raising prediction results satisfying the matching condition and the corresponding face prediction results, the corresponding human body tracking identifier and face tracking identifier are recorded. Subsequently, for a plurality of continuous to-be-identified images collected within a certain time, the corresponding hand-raising prediction results and face prediction results are determined according to the human body tracking identifier and the face tracking identifier whether the matching condition is satisfied; if the hand-raising prediction results and the face prediction results corresponding to the human body tracking identifier and the face tracking identifier in the plurality of continuous to-be-identified images satisfy the matching condition, it is determined that the human body object corresponding to the human body tracking identifier and the face tracking identifier is the identified hand-raising personnel.

[0083] Figure 3 The device structure diagram of the embodiment of the application is shown in the figure. As shown in the figure, the hand-raising behavior recognition device provided by the embodiment of the application comprises:

[0084] The hand-raising training module is used for identifying the to-be-identified image to obtain a target hand-raising prediction result set comprising at least one group of target hand-raising prediction results.

[0085] The to-be-identified image is identified by using the hand-raising training module to obtain a hand-raising prediction result set, the hand-raising prediction result set comprising at least one hand-raising prediction result of a human body object, each group of hand-raising prediction results comprising the reference point coordinate, height, width, recognition confidence of the human body object. According to the preset hand-raising behavior threshold value, when the recognition confidence is greater than or equal to the hand-raising behavior threshold value, the group of hand-raising prediction results corresponding to the recognition confidence is taken as the target hand-raising prediction result, and the target hand-raising prediction result set is formed by at least one group of target hand-raising prediction results.

[0086] The hand-raising behavior tracking module is used for labeling at least one hand-raising human body object according to the input target hand-raising prediction result set to generate the labeling information of each human body object.

[0087] The hand-raising behavior tracking module determines the number of the human body objects and the positions of each human body object according to the number of the target hand-raising prediction results in the target hand-raising prediction result set, and labels each human body object, including labeling the human body frame of each human body object and labeling the human body tracking identifier of each human body object, the human body tracking identifier corresponding to the human body object in one-to-one manner.

[0088] The face training module is configured to identify the to-be-identified image to obtain a target face prediction result set including at least one group of target face prediction results.

[0089] The face training module is configured to identify the to-be-identified image to obtain a target face prediction result set including at least one group of target face prediction results.

[0090] The face tracking module is configured to label at least one face object according to the input target face prediction result set to generate labeling information of each face object.

[0091] The face tracking module is configured to determine the number of the face objects and the positions of each face object according to the number of the target face prediction results in the target face prediction result set, and label each face object, including labeling the face frame of each face object and labeling the face tracking identifier of each face object, the face tracking identifier corresponding to the face object in one-to-one manner.

[0092] The matching module is configured to match the target hand-raising prediction result set and the target face prediction result set according to a matching condition.

[0093] The target hand-raising prediction result set includes at least one group of target hand-raising prediction results, and the target face prediction result set includes at least one group of target face prediction results. One group of target hand-raising prediction results is matched with each group of target face prediction results, and when one group of target hand-raising prediction results and one group of face prediction results satisfy the matching condition, the human body object corresponding to the target hand-raising prediction results and the face prediction results is determined as the identified hand-raising personnel. The matching condition is that whether the face object is within the human body object range, i.e., whether the face frame is within the human body frame, if yes, it is determined as matching, otherwise, it is determined as not matching.

[0094] The output module is configured to output a hand-raising behavior identification result for the target hand-raising prediction results and the corresponding target face prediction results satisfying the matching condition.

[0095] The hand-raising behavior recognition device further comprises:

[0096] An image processing module is configured to extract each frame of image from a video stream collected by the image collection device, and pre-process each frame of image, and take the pre-processed image as a to-be-recognized image.

[0097] The hand-raising behavior recognition device further comprises:

[0098] A recording module is configured to record the labeling information of the human object corresponding to the target hand-raising prediction result and the labeling information of the human face object corresponding to the target human face prediction result that meet the matching condition.

[0099] An output module is configured to, for the plurality of to-be-recognized images input successively, judge that each to-be-recognized image meets the matching condition according to the labeling information of the human object and the labeling information of the human face object recorded, and output the hand-raising recognition result.

[0100] In the embodiment of the present application, the recording module records the human tracking identifier corresponding to the target hand-raising prediction result and the human face tracking identifier corresponding to the target human face prediction result that meet the matching condition. The output module, for the plurality of to-be-recognized images input successively, judges that each to-be-recognized image meets the matching condition according to the human tracking identifier and the human face tracking identifier recorded, and outputs the hand-raising behavior recognition result.

[0101] The plurality of to-be-recognized images input successively can be a plurality of images extracted successively within a predetermined time.

[0102] In order to achieve the above object, an embodiment of a device for executing the hand-raising behavior recognition method is further provided in the present application. The device comprises:

[0103] One or more processors and a memory.

[0104] The device for executing the hand-raising behavior recognition method can further comprise an input device and an output device.

[0105] The processor, the memory, the input device and the output device can be connected through a bus or other means.

[0106] The memory is a non-volatile computer readable storage medium, which can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the hand-raising behavior recognition method in the embodiment of the present application. The processor executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory, that is, implements the hand-raising behavior recognition method of the method embodiment.

[0107] The memory can include a program storage area and a data storage area. The program storage area can store an operating system and application programs required for at least one function. The data storage area can store data created according to use of the device for recognizing the hand-raising behavior, and the like. In addition, the memory can include a high-speed random access memory, and can further include a non-volatile memory such as at least one of a magnetic disk storage device, a flash memory device, or other non-volatile solid state memory device. In some embodiments, the memory can optionally include a memory disposed remotely with respect to the processor, and these remote memories can be connected to the member user behavior monitoring device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0108] The input device can receive inputted digital or character information, and generate key signal inputs related to user settings and function controls of the device for recognizing the hand-raising behavior. The output device can include a display device such as a display screen.

[0109] The one or more modules are stored in the memory, and when executed by the one or more processors, perform the hand-raising behavior recognition method of any of the above-described method embodiments. The device for recognizing the hand-raising behavior according to the embodiments has the same or similar technical effects as the above-described method embodiments.

[0110] The embodiments of the present application also provide a non-transitory computer storage medium storing computer executable instructions, which can execute the processing method of the list item operation of any of the above-described method embodiments. The embodiments of the non-transitory computer storage medium have the same or similar technical effects as the above-described method embodiments.

[0111] Finally, it should be noted that a person of ordinary skill in the art can understand that all or part of the processes in the above-described embodiments can be completed by a computer program instructing related hardware. The program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-described embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like. The embodiments of the computer program have the same or similar technical effects as the above-described method embodiments.

[0112] In addition, typically, the apparatus, device, etc. described in the present disclosure can be various electronic terminal devices, such as a mobile phone, a personal digital assistant (PDA), a tablet computer (PAD), a smart television, etc., and can also be large terminal devices, such as a server, etc., and thus the protection scope of the present disclosure should not be limited to a certain type of apparatus, device. The client described in the present disclosure can be applied to any of the above electronic terminal devices in the form of electronic hardware, computer software, or a combination of both.

[0113] In addition, the method according to the present disclosure can also be implemented as a computer program executed by a CPU, which can be stored in a computer readable storage medium. When the computer program is executed by the CPU, the above-mentioned functions defined in the method of the present disclosure are performed.

[0114] In addition, the above-mentioned method steps and system units can also be implemented by a controller and a computer readable storage medium for storing a computer program that enables the controller to implement the above-mentioned steps or unit functions.

[0115] In addition, it should be understood that the computer readable storage medium (e.g., memory) described herein can be a volatile memory or a non-volatile memory, or can include both volatile memory and non-volatile memory. As an example and not a limitation, non-volatile memory can include read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM), which can serve as external cache memory. As an example and not a limitation, RAM can be obtained in a variety of forms, such as synchronous RAM (DRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM) and direct Rambus RAM (DRRAM). The storage device of the disclosed aspect is intended to include, but not limited to, these and other suitable types of memory.

[0116] The apparatus of the above-mentioned embodiments is used to implement the corresponding method in the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiments, which are not repeated here.

[0117] Those of ordinary skill in the art will realize that the foregoing discussion of any of the embodiments has been presented for the purpose of illustration and description and is not intended to limit the scope of the disclosure (including the claims) to the examples set forth in the description or illustration of specific embodiments. Further, the steps of any of the methods disclosed herein do not have to be performed in the precise order described. The steps of various embodiments can be performed in any order, unless otherwise specified or required by the circumstances. Other variations and modifications of the embodiments disclosed herein, in addition to those described herein, will be apparent to those of ordinary skill in the art from the foregoing description and accompanying drawings. Such variations and modifications are intended to fall within the scope of the disclosure. Accordingly, the disclosure is not limited to that precisely as shown and described.

[0118] In addition, to simplify the description and discussion, and so as not to obscure the application, well-known power / ground connections to integrated circuit (IC) chips and other components can or can not be shown in the provided figures. Furthermore, devices can be shown in block diagram form in order to avoid obscuring the application, and this also applies to similar block diagrams wherever they can be found in the present disclosure. In the description provided herein, numerous specific

[0119] Although the application has been described in conjunction with specific embodiments thereof, numerous alternatives, modifications, and variations will be readily apparent to those of ordinary skill in the art. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.

[0120] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations as falling within the scope of the broadest possible interpretation of the appended claims. Accordingly, any and all such alternatives, modifications, and variations should be included within the scope of the present application.

Claims

1. A hand-raising behavior recognition method, characterized by, The method comprises: inputting a to-be-recognized image; recognizing the to-be-recognized image by using a hand-raising training model to obtain a target hand-raising prediction result set comprising at least one group of target hand-raising prediction results; the target hand-raising prediction result set comprises a target hand-raising prediction result of at least one human body object, and the target hand-raising prediction result comprises reference point coordinates, height, width and recognition confidence of the human body object; inputting the target hand-raising prediction result set into a hand-raising behavior tracker, and respectively labeling at least one human body object that raises hands according to the target hand-raising prediction result set to generate labeling information of each human body object, wherein the labeling information of each human body object comprises a human body frame of each human body object; recognizing the to-be-recognized image by using a face training model to obtain a target face prediction result set comprising at least one group of target face prediction results; the target face prediction result set comprises a target face prediction result of at least one face object, and the target face prediction result comprises reference point coordinates, height, width and recognition confidence of the face object; inputting the target face prediction result set into a face tracker, and respectively labeling at least one face object according to the target face prediction result set to generate labeling information of each face object, wherein the labeling information of each face comprises a face frame of each face object; matching according to a matching condition based on the target hand-raising prediction result set and the target face prediction result set, comprising: for a group of target hand-raising prediction results and a group of target face prediction results that are matched: expanding the human body frame by a plurality of pixels to form a matching frame; determining reference point coordinates and diagonal point coordinates of the matching frame based on the reference point coordinates, height, width of the human body object and the expanded pixel value; determining the center point coordinates of the human body object based on the reference point coordinates and the diagonal point coordinates of the matching frame; calculating the center point coordinates of the face object based on the reference point coordinates, height and width of the face object; and judging whether the face frame is within the human body frame based on the positional relationship between the center point coordinates of the human body object and the center point coordinates of the face object, and if so, the matching condition is satisfied; outputting a hand-raising behavior recognition result for the target hand-raising prediction result and the corresponding target face prediction result that satisfy the matching condition.

2. The method of claim 1, wherein, The labeling information of each human body object comprises a human body tracking identifier of each human body object, and the human body tracking identifier is in one-to-one correspondence with the human body object.

3. The method of claim 1, wherein, The labeling information of each face comprises a face tracking identifier of each face object, and the face tracking identifier is in one-to-one correspondence with the face object.

4. The method of claim 1, wherein, Further comprising: for the target hand-raising prediction result and the corresponding target face prediction result that satisfy the matching condition, recording the labeling information of the human body object and the labeling information of the face object respectively, and for a plurality of to-be-recognized images that are continuously input within a predetermined time, judging whether the plurality of to-be-recognized images all satisfy the matching condition based on the recorded labeling information of the human body object and the labeling information of the face object, and if so, inputting a hand-raising behavior recognition result.

5. A hand-raising behavior recognition apparatus characterized by comprising: The method comprises: The hand-raising training module is configured to identify the to-be-identified image to obtain a target hand-raising prediction result set including at least one group of target hand-raising prediction results; The target hand-raising prediction result set includes a target hand-raising prediction result of at least one human object, and the target hand-raising prediction result includes reference point coordinates, height, width, and recognition confidence of the human object. The hand-raising behavior tracking module is configured to label at least one human object raising hands according to the input target hand-raising prediction result set to generate labeling information of each human object, and the labeling information of each human object includes a human body frame of each human object. The face training module is configured to identify the to-be-identified image to obtain a target face prediction result set including at least one group of target face prediction results. The face tracking module is configured to label at least one face object according to the input target face prediction result set to generate labeling information of each face object, and the labeling information of each face object includes a face frame of each face object. The matching module is configured to match the target hand-raising prediction result set and the target face prediction result set according to a matching condition, including: for a group of target hand-raising prediction results and a group of target face prediction results to be matched: expanding the human body frame by a plurality of pixels to form a matching frame; determining reference point coordinates and diagonal point coordinates of the matching frame according to the reference point coordinates, height, width of the human object, and the expanded pixel value; determining the center point coordinates of the human object according to the reference point coordinates and the diagonal point coordinates of the matching frame; calculating the center point coordinates of the face object according to the reference point coordinates, height, and width of the face object; and determining whether the face frame is within the human body frame according to the positional relationship between the center point coordinates of the human object and the center point coordinates of the face object, and if yes, the matching condition is met. The output module is configured to output a hand-raising behavior recognition result for the target hand-raising prediction result and the corresponding target face prediction result meeting the matching condition.

6. The apparatus of claim 5, wherein, The labeling information of each human object includes a human body tracking identifier of each human object, and the human body tracking identifier is in one-to-one correspondence with the human object.

7. The apparatus of claim 5, wherein, The labeling information of each face object includes a face tracking identifier of each face object, and the face tracking identifier is in one-to-one correspondence with the face object.

8. The apparatus of claim 5, wherein, The recording module is configured to record the labeling information of the human object corresponding to the target hand-raising prediction result meeting the matching condition and the labeling information of the face object corresponding to the target face prediction result. The output module is configured to output a hand-raising behavior recognition result when each to-be-identified image meets the matching condition according to the recorded labeling information of the human object and the labeling information of the face object for a plurality of to-be-identified images input successively. The processor executes the program to implement the method of any one of claims 1 to 4.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, ​

Citation Information

Patent Citations

  • Teaching behavior analysis method and device based on target tracking and attitude detection

    CN108846853A