Somatosensory recognition method, electronic equipment and storage medium
By using a human body recognition model to determine the target selection area and performing image scaling processing in motion-sensing games, the problems of user positioning restrictions and background interference are solved, achieving flexible and stable motion-sensing recognition effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-03
AI Technical Summary
Existing motion-sensing game recognition devices are limited by one-way interaction, requiring users to adjust their positions, and background interference can easily lead to recognition errors, affecting the flexibility and stability of the game process, especially in multi-player scenarios where performance requirements are high.
The human body recognition model determines the recognition area of the target participant in the target video image, and after scaling and displacement processing, it is input into the human body key point recognition model to eliminate background interference, reduce the input resolution requirement, and improve the convenience and accuracy of recognition.
It achieves flexible motion recognition without requiring users to adjust their positions, reduces model performance requirements, improves recognition stability and user coverage, and is suitable for multi-person scenarios and low-end devices.
Smart Images

Figure CN121600435A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video technology, and more particularly to a motion recognition method, electronic device, and storage medium. Background Technology
[0002] Currently, there are a large number of motion-sensing games available, such as motion-sensing games with the Switch controller, the Microsoft Kinect game console, and the "Everyday Jump Rope" app.
[0003] In related technologies, motion-sensing games often use cameras to capture and record human movements, and use machine learning models to input video footage captured by the camera, output identified human key points, and use the human key points and their changing positions, speeds, and other information to bind game objects and drive game characters on the screen.
[0004] However, devices that perform motion-sensing game recognition are limited by their one-way interaction method. The game design requires users to adjust their position so that their bodies are within the "recognition area" on the screen, allowing the device's camera to identify key body points and drive game objects. Furthermore, when the background of the recognition area contains many characters, it can easily lead to recognition errors and interfere with gameplay. Summary of the Invention
[0005] This application provides a motion-sensing recognition method, electronic device, and storage medium. It can directly use the target recognition area containing the features of the participant as the motion-sensing recognition region, improving the convenience, flexibility, and interactivity of motion-sensing recognition. It can also effectively eliminate interference from people in the background of video images, improving the stability and accuracy of motion-sensing recognition. Furthermore, it can reduce the input resolution requirements of the machine learning model (human key point recognition model) in the main motion-sensing recognition process, improving model efficiency while ensuring model accuracy, reducing model performance requirements, and increasing user coverage of motion-sensing recognition. The technical solution is as follows:
[0006] In a first aspect, embodiments of this application provide a motion recognition method, the method comprising:
[0007] Acquire a target video image; the target video image includes at least one feature of the person to be involved.
[0008] Based on the human body recognition model, the target recognition selection area corresponding to the target participant features in the target video image is determined; the target participant features are used to characterize the image information of the participant selected to participate in the somatosensory recognition or the image information of the participant in a specified posture from the at least one participant feature; the specified posture is the human body posture that the participant should make corresponding to the specified participant feature to participate in the somatosensory recognition.
[0009] The target region image corresponding to the above target recognition selection area is scaled and shifted to obtain the image to be recognized;
[0010] The above-mentioned image to be identified is input into the human key point recognition model, and the human key point recognition result corresponding to the above-mentioned image to be identified is output.
[0011] In this embodiment, on the one hand, based on the human body recognition model, the target recognition selection area corresponding to the features of the selected participants in the target video image who are participating in motion sensing or in a specified posture can be determined. Therefore, the target recognition selection area where the target participant features are located can be directly used as the area for motion sensing without the target participant features standing in the specified recognition area, thereby improving the convenience, flexibility and interactivity of motion sensing and solving the problem of motion sensing area limitation. On the other hand, by scaling and shifting the target area image corresponding to the target recognition selection area before inputting it into the human key point recognition model, the corresponding human key point recognition result can be obtained. By scaling the target area image corresponding to the target recognition selection area, the interference of people in the background of the video image can be effectively eliminated, improving the stability and accuracy of motion sensing. At the same time, the input resolution requirement of the machine learning model (human key point recognition model) in the main process of motion sensing can be reduced, improving the model running efficiency while ensuring the model running accuracy, reducing the model performance requirements, and increasing the user coverage of motion sensing.
[0012] In one possible implementation, the target video image includes features of multiple participants; the human body recognition model includes a multi-person human body recognition model.
[0013] The above-mentioned target recognition selection region, based on the human body recognition model, determines the target participant features in the target video image, including:
[0014] The target video image is input into the multi-person human body recognition model, and the human body recognition area information corresponding to the features of the multiple participants is output; the human body recognition area information includes the position and size information of the minimum bounding rectangle region corresponding to the feature of the participant in the target video image.
[0015] Based on the location and size information corresponding to each of the above-mentioned features of the participants, the human body recognition area corresponding to each of the above-mentioned features of the participants is displayed in the above-mentioned target video image;
[0016] The user inputs a recognition area selection operation; the recognition area selection operation is used to select the target human body recognition area information of the target participant feature for participating in the somatosensory recognition from the human body recognition areas corresponding to the multiple participant features;
[0017] In response to the above-mentioned identification area selection operation, the target identification selection area corresponding to the above-mentioned target participant characteristics is determined based on the above-mentioned target human body identification area information.
[0018] In one possible implementation, the target video image includes features of multiple participants; the human body recognition model includes a multi-person pose recognition model.
[0019] The above-mentioned target recognition selection region, based on the human body recognition model, determines the target participant features in the target video image, including:
[0020] The target video image is input into the multi-person posture recognition model, and the human body recognition area information and posture recognition results corresponding to the features of the multiple participants are output.
[0021] Based on the human body recognition area information corresponding to the target participant features, the target recognition selection area corresponding to the above-mentioned target participant features in the above-mentioned target video image is determined; the target participant features are used to characterize the image information of the participant whose posture recognition result is in a specified posture.
[0022] In one possible implementation, the aforementioned multi-person pose recognition model includes a first multi-person pose recognition model and / or a second multi-person pose recognition model; the first multi-person pose recognition model includes a multi-person human keypoint recognition network, used to identify the keypoint distance between the shoulder keypoints and leg keypoints of the human skeleton of each participant in the target video image corresponding to the features of the participants, and to determine whether the participants corresponding to each feature of the participants in the target video image are in a specified pose based on the keypoint distance, and is trained based on multiple human images with known human body region information and human body keypoint information; the second multi-person pose recognition model includes a multi-person human behavior pose recognition network, used to identify whether the participants corresponding to each feature of the participants in the target video image are in a specified pose, and is trained based on multiple human images with known human body region information and human body pose information; wherein, the human body image includes multiple human body features.
[0023] In one possible implementation, the number of the aforementioned target participant features in the target video image is multiple;
[0024] After determining the target recognition region corresponding to the features of the target participant in the target video image based on the human body recognition model, and before scaling and shifting the target region image corresponding to the target recognition region to obtain the image to be recognized, the method further includes:
[0025] Based on the target recognition selection area corresponding to the characteristics of each of the above-mentioned target participants, the corresponding target region image is extracted from the above-mentioned target video image;
[0026] The above-mentioned scaling and displacement processing of the target region image corresponding to the target recognition selection area to obtain the image to be recognized includes:
[0027] Based on the location of the target recognition selection area corresponding to the features of each of the aforementioned target participants, the target area images corresponding to the features of each of the aforementioned target participants are sequentially scaled and stitched together to obtain the image to be recognized corresponding to the aforementioned target video image.
[0028] In one possible implementation, after scaling and shifting the target region image corresponding to the target recognition selection area to obtain the image to be recognized, the method further includes:
[0029] The image to be identified is rendered and displayed; the image to be identified is updated as the target video image is updated.
[0030] In one possible implementation, before determining the target recognition selection area corresponding to the features of the target participant in the target video image based on the human body recognition model, the method further includes:
[0031] After the motion-sensing classroom application is launched, a question template is displayed;
[0032] Based on the above question template, the system receives the question text information input by the user and the number of participants in the question; the question text information includes the question content text and the question answer text; the number of the above target participant characteristics is less than or equal to the number of participants in the question;
[0033] After inputting the image to be recognized into the human key point recognition model and outputting the human key point recognition result corresponding to the image to be recognized, the method further includes:
[0034] Based on the above human body key point recognition results and the above question answer text, the target participant's target answer result corresponding to the above target participant characteristics is determined.
[0035] In this embodiment, by combining e-learning with the motion recognition provided in this embodiment, the interactive capabilities of electronic devices (interactive smart whiteboards) in the classroom are utilized to enhance the intelligence and engagement of the classroom. This allows for flexible and accurate motion recognition even in classroom spaces with many students and limited seating, lowering the barrier to entry for motion recognition and improving its practicality.
[0036] In one possible implementation, after inputting the image to be recognized into the human key point recognition model and outputting the human key point recognition result corresponding to the image to be recognized, the method further includes:
[0037] The movement of the target virtual object is driven by the above human key point recognition results; the above target virtual object is the virtual object bound to the target participant corresponding to the above target participant characteristics.
[0038] Secondly, embodiments of this application provide a motion-sensing recognition device, including:
[0039] An acquisition module is used to acquire a target video image; the target video image includes at least one feature of a person to be involved.
[0040] The first determining module is used to determine the target recognition selection area corresponding to the target participant features in the target video image based on the human body recognition model; the target participant features are used to characterize the image information of the participant selected to participate in the somatosensory recognition or the image information of the participant in a specified posture from the at least one participant feature; the specified posture is the human body posture that the participant should make corresponding to the specified participant feature to participate in the somatosensory recognition.
[0041] The scaling and displacement processing module is used to scale and displacement the target area image corresponding to the target recognition selection area to obtain the image to be recognized.
[0042] The key point recognition module is used to input the above-mentioned image to be recognized into the human key point recognition model and output the human key point recognition result corresponding to the above-mentioned image to be recognized.
[0043] In one possible implementation, the target video image includes features of multiple participants; the human body recognition model includes a multi-person human body recognition model.
[0044] The aforementioned first determining module includes:
[0045] The human body recognition unit is used to input the target video image into the multi-person human body recognition model and output the human body recognition area information corresponding to the features of the multiple participants; the human body recognition area information includes the position information and size information of the minimum bounding rectangle region corresponding to the feature of the participant in the target video image.
[0046] The display unit is used to display the human body recognition area corresponding to each of the multiple features of the participants in the target video image based on the position information and size information corresponding to each of the multiple features of the participants.
[0047] The receiving unit is used to receive the recognition area selection operation input by the user; the recognition area selection operation is used to select the target human body recognition area information of the target participant feature for participating in the somatosensory recognition from the human body recognition areas corresponding to the multiple participants' features.
[0048] The first determining unit is used to determine the target recognition selection area corresponding to the target participant characteristics based on the target human body recognition area information in response to the above-mentioned recognition area selection operation.
[0049] In one possible implementation, the target video image includes features of multiple participants; the human body recognition model includes a multi-person pose recognition model.
[0050] The aforementioned first determining module includes:
[0051] The posture recognition unit is used to input the target video image into the multi-person posture recognition model and output the human body recognition area information and posture recognition results corresponding to the features of the multiple participants.
[0052] The second determining unit is used to determine the target recognition selection area corresponding to the target participant features in the target video image based on the human body recognition area information corresponding to the target participant features; the target participant features are used to characterize the image information of the participant whose posture recognition result is in a specified posture.
[0053] In one possible implementation, the aforementioned multi-person pose recognition model includes a first multi-person pose recognition model and / or a second multi-person pose recognition model; the first multi-person pose recognition model includes a multi-person human keypoint recognition network, used to identify the keypoint distance between the shoulder keypoints and leg keypoints of the human skeleton of each participant in the target video image corresponding to the features of the participants, and to determine whether the participants corresponding to each feature of the participants in the target video image are in a specified pose based on the keypoint distance, and is trained based on multiple human images with known human body region information and human body keypoint information; the second multi-person pose recognition model includes a multi-person human behavior pose recognition network, used to identify whether the participants corresponding to each feature of the participants in the target video image are in a specified pose, and is trained based on multiple human images with known human body region information and human body pose information; wherein, the human body image includes multiple human body features.
[0054] In one possible implementation, the number of the aforementioned target participant features in the target video image is multiple;
[0055] The aforementioned motion recognition device also includes:
[0056] The cropping module is used to crop the corresponding target region image from the above target video image based on the target recognition selection area corresponding to the characteristics of each of the above target participants;
[0057] The aforementioned scaling and displacement processing module is specifically used for:
[0058] Based on the location of the target recognition selection area corresponding to the features of each of the aforementioned target participants, the target area images corresponding to the features of each of the aforementioned target participants are sequentially scaled and stitched together to obtain the image to be recognized corresponding to the aforementioned target video image.
[0059] In one possible implementation, the aforementioned motion recognition device further includes:
[0060] The rendering module is used to render and display the image to be identified; the image to be identified is updated as the target video image is updated.
[0061] In one possible implementation, the aforementioned motion recognition device further includes:
[0062] The display module is used to display the question template after the motion-sensing classroom application is launched;
[0063] The receiving module is used to receive the question text information and the number of participants in the question based on the above question template; the above question text information includes the question content text and the question answer text; the number of the above target participant features is less than or equal to the above number of participants in the question.
[0064] The aforementioned motion recognition device also includes:
[0065] The second determining module is used to determine the target answer result of the target participant corresponding to the above-mentioned target participant characteristics based on the above-mentioned human body key point recognition results and the above-mentioned question answer text.
[0066] In one possible implementation, the aforementioned motion-sensing recognition device further includes:
[0067] The driving module is used to drive the movement of the target virtual object based on the above human key point recognition results; the above target virtual object is the virtual object bound to the target participant corresponding to the above target participant characteristics.
[0068] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory; wherein the memory stores a computer program, the computer program being adapted to be loaded by the processor and execute the method steps provided by the first aspect of the embodiments of this application or any possible implementation thereof.
[0069] Fourthly, embodiments of this application provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps provided by the first aspect of the embodiments of this application or any possible implementation thereof.
[0070] It is understood that the motion recognition device provided in the second aspect, the electronic device provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the motion recognition method provided in the first aspect or any implementation of the first aspect. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the motion recognition method provided in the first aspect or any implementation of the first aspect, and will not be repeated here.
[0071] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0072] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0073] Figure 1 A schematic diagram of a motion-sensing recognition area provided in related technologies;
[0074] Figure 2 A schematic diagram illustrating an application scenario of a motion recognition method provided in an exemplary embodiment of this application;
[0075] Figure 3 A flowchart illustrating a motion recognition method provided for an exemplary embodiment of this application;
[0076] Figure 4 A schematic diagram illustrating the implementation process of determining a target recognition selection area, provided as an exemplary embodiment of this application;
[0077] Figure 5A A schematic diagram of a human body recognition area provided for an exemplary embodiment of this application;
[0078] Figure 5B A schematic diagram illustrating the selection of a target recognition region provided in an exemplary embodiment of this application;
[0079] Figure 6 A schematic diagram illustrating another implementation process for determining a target recognition selection area provided in an exemplary embodiment of this application;
[0080] Figure 7A A schematic diagram illustrating the selection of another target recognition region provided in an exemplary embodiment of this application;
[0081] Figure 7B A schematic diagram illustrating the selection of an image to be recognized, provided as an exemplary embodiment of this application;
[0082] Figure 8 A flowchart illustrating another motion recognition method provided as an exemplary embodiment of this application;
[0083] Figure 9 A schematic diagram showing a question template provided for an exemplary embodiment of this application;
[0084] Figure 10 A flowchart illustrating another motion recognition method provided as an exemplary embodiment of this application;
[0085] Figure 11 A schematic diagram of a target virtual object provided for an exemplary embodiment of this application;
[0086] Figure 12 A schematic diagram of the structure of a motion-sensing recognition device provided for an exemplary embodiment of this application;
[0087] Figure 13 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. Detailed Implementation
[0088] To make the features and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0089] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0090] In related technologies, motion-sensing games often use cameras to record human movements and employ machine learning models. The input is video footage captured by the camera, the output is identified human body key points, and these key points, along with information about their position and speed of change, are used to bind game objects and drive the game character on screen. However, in the aforementioned motion-sensing recognition schemes, users participating in motion-sensing recognition often need to adjust their position to place their bodies within the "recognition area" on the screen, allowing the camera of the electronic device performing the motion-sensing recognition to identify the participant's body key points to drive the game object. For example, but not limited to... Figure 1 As shown, when two people play a motion-sensing game, they often need to move their positions to stand within the fixed recognition areas 110 and 120 of the electronic device to achieve motion recognition. That is, when the participants' space is limited, such as in a classroom with many tables and chairs, they may not be able to stand completely within the recognition area, hindering the quick start of the game and affecting the flexibility of motion recognition. Furthermore, when there are many figures in the background of the recognition area, it can easily lead to motion recognition errors, interfering with the game. At the same time, the above solution requires high resolution for the model input; when the electronic device's performance is poor (e.g., single CPU), running a machine learning model may be insufficient to run a high-precision real-time multi-person motion recognition algorithm.
[0091] To address the problems existing in the above-mentioned related technologies, this application provides a motion recognition method. On the one hand, based on a human body recognition model, the target recognition selection area corresponding to the features of the selected participant in motion recognition or in a specified posture in the target video image can be determined. Therefore, the target participant does not need to stand in the specified recognition area; the target recognition selection area where the target participant's features are located can be directly used as the area for motion recognition, improving the convenience, flexibility, and interactivity of motion recognition and solving the problem of limited motion recognition area. On the other hand, by scaling and shifting the target area image corresponding to the target recognition selection area before inputting it into the human key point recognition model, the corresponding human key point recognition result is obtained. By scaling the target area image corresponding to the target recognition selection area, interference from people in the background of the video image can be effectively eliminated, improving the stability and accuracy of motion recognition. At the same time, the input resolution requirement of the machine learning model (human key point recognition model) in the main process of motion recognition can be reduced, improving the model's running efficiency while ensuring the model's running accuracy, reducing the model's performance requirements, and increasing the user coverage of motion recognition.
[0092] It is understood that the motion recognition method provided in this application can be applied not only to motion-sensing game scenarios, but also to classroom teaching scenarios or multi-player game scenarios with multiple viewers, etc. The embodiments of this application do not limit this.
[0093] For example, such as Figure 2 As shown, in a classroom teaching scenario, the motion recognition method provided in this application embodiment can be executed independently by the teaching terminal 200 in front of the classroom, or by the server corresponding to the application that needs to perform motion recognition installed in the teaching terminal 200, etc. This application embodiment does not limit this. After the application that needs to perform motion recognition in the teaching terminal 200 is turned on, it can, but is not limited to, first collect target video images within its visual range through its camera 210. The target video images include at least one feature of the person to be identified (e.g., image information corresponding to the smallest bounding rectangle area where the students in the classroom are located); then, based on the human body recognition model, the target recognition selection area corresponding to the target feature of the person to be identified in the target video image is determined. The target feature of the person to be identified is used to characterize the image information of the person to be identified or the image information of the person to be identified in a specified posture among at least one feature of the person to be identified. The specified posture is the human posture that the person to be identified should make according to the specified feature of the person to be identified. The target area image corresponding to the target recognition selection area is scaled and shifted to obtain the image to be identified; finally, the image to be identified is input into the human body key point recognition model, and the human body key point recognition result corresponding to the image to be identified is output. The target video image, target recognition selection area, image to be recognized, and human key point recognition results can all be displayed sequentially on the display screen 220 of the teaching terminal 200, thereby improving the user's sense of participation and interaction capabilities in motion recognition.
[0094] It is understood that the motion recognition method provided in this application embodiment can be executed by electronic devices such as terminals or servers. The terminals can be mobile phones, tablets, computers, etc. with motion recognition programs installed, and the servers can be hardware servers, virtual servers, cloud servers, etc. This application embodiment does not limit this.
[0095] Next, combine Figure 1 and Figure 2 Taking the motion sensing recognition performed by an electronic device as an example, this application introduces a motion sensing recognition method provided in an exemplary embodiment. Please refer to [link / reference needed] for details. Figure 3 This is a flowchart illustrating a motion recognition method provided in an exemplary embodiment of this application. Figure 3 As shown, this motion recognition method may include the following steps:
[0096] S301, acquire the target video image, the target video image includes at least one feature of the person to be involved.
[0097] Specifically, when the electronic device is turned on or the motion-sensing recognition program installed in the electronic device is activated, the electronic device can, but is not limited to, receive target video images sent by a camera connected to it, or directly acquire target video images within the visual range of motion-sensing recognition through its installed camera. The aforementioned target video images may include one or more features of the person to be involved, which is not limited in this embodiment of the application.
[0098] S302, Based on the human body recognition model, determine the target recognition selection area corresponding to the target participant features in the target video image. The target participant features are used to characterize the image information of the participant selected from at least one participant feature to participate in the somatosensory recognition or the image information of the participant in a specified posture.
[0099] Specifically, the specified posture refers to the human posture that the participant should adopt corresponding to the characteristics of the participant in the specified somatosensory recognition, such as, but not limited to, standing posture or hand-raising posture.
[0100] Understandably, the position of the target recognition selection area can be moved based on the positional movement of the target participant features in multiple consecutive frames of target video images, or it can be fixed at the position of the target recognition selection area corresponding to the first determined target participant features. This application embodiment does not limit this.
[0101] In some possible embodiments, the electronic device includes a display screen, and the target video image may, but is not limited to, include features of multiple participants; the human body recognition model may, but is not limited to, include a multi-person human body recognition model. Figure 4 As shown, the above-mentioned S302, based on the human body recognition model, determines the target recognition selection area corresponding to the characteristics of the target participant in the target video image, and the implementation process may include, but is not limited to, the following:
[0102] S401, input the target video image into the multi-person human body recognition model, and output the human body recognition area information corresponding to the features of multiple participants.
[0103] Specifically, the aforementioned multi-person human recognition model can be, but is not limited to, a YoloV8 model, a YoloV7 model, a YoloV5 model, an RTMP model, an RTMO model, etc. The aforementioned multi-person human recognition model can be trained on multiple human images with known human recognition region information, but is not limited to. The aforementioned human recognition region information can include, but is not limited to, the positional information (e.g., but not limited to, vertex corner coordinates) and size information (e.g., but not limited to, width and height) of the smallest bounding rectangle region (i.e., the human recognition region) corresponding to the identified features of the person to be identified (e.g., but not limited to, the human image information of the person to be identified) in the target video image.
[0104] S402, based on the location and size information corresponding to the features of multiple participants, displays the human recognition area corresponding to each of the features of multiple participants in the target video image.
[0105] For example, after determining the human body recognition area information corresponding to each of the multiple participants' features, the electronic device's display screen can show, based on the location and size information corresponding to each of the multiple participants' features, the following information: Figure 5A The target video image shown displays multiple human body recognition areas (i.e., blue selection boxes) corresponding to the features of the participants, allowing users of electronic devices to manually select the target recognition area corresponding to the feature of the participant to be included in this somatosensory recognition from multiple human body recognition areas identified by the multi-person human body recognition model.
[0106] S403, receive user input of recognition area selection operation, the recognition area selection operation is used to select the target human body recognition area information of the target participant feature to participate in the somatosensory recognition from the human body recognition areas corresponding to the features of multiple participants.
[0107] For example, after the display screen shows the human body recognition areas corresponding to the features of multiple participants in the target video image, the user can manually select the human body recognition area containing the features of the target participant to be included in the motion sensing recognition, for example, but not limited to clicking. Figure 5B The human body recognition area (i.e., the blue selection box) is shown in the lower right corner. Electronic devices can receive user input for the selection of this recognition area based on the display screen.
[0108] S404, in response to the recognition area selection operation, determines the target recognition selection area corresponding to the characteristics of the target participant based on the target human body recognition area information.
[0109] Specifically, the electronic device can respond to the user's input of a recognition area selection operation, determine the location of the selected target human recognition area (i.e., the human recognition area selected by the user) based on the information of the selected target human recognition area, and display the target recognition selection area corresponding to the characteristics of the target participant selected by the user on the display screen.
[0110] In some possible embodiments, the target video image may, but is not limited to, include features of multiple participants; the human body recognition model may, but is not limited to, include a multi-person pose recognition model. For example... Figure 6 As shown, the above-mentioned S302, based on the human body recognition model, can also include, but is not limited to, the following implementation process for determining the target recognition selection area corresponding to the characteristics of the target participant in the target video image:
[0111] S601, input the target video image into the multi-person posture recognition model, and output the human body recognition area information and posture recognition results corresponding to the features of multiple participants.
[0112] Specifically, the aforementioned posture recognition results are used to indicate whether a human body is in a specified posture. The aforementioned multi-person posture recognition model may include, but is not limited to, a first multi-person posture recognition model and / or a second multi-person posture recognition model; the aforementioned first multi-person posture recognition model may include, but is not limited to, a multi-person human keypoint recognition network. When the specified posture is a standing posture, the aforementioned first multi-person posture recognition model is used to identify the keypoint distance between the shoulder keypoint and leg keypoint of the human skeleton of each participant in the target video image, and to determine whether the participant corresponding to each participant feature in the target video image is in the specified posture (i.e., standing posture) based on the aforementioned keypoint distance. When the specified posture is a raised hand posture, the aforementioned first multi-person posture recognition model is used to identify the keypoint distance between the head keypoint and hand keypoint of the human body of each participant in the target video image, and to determine whether the participant corresponding to each participant feature in the target video image is in the specified posture (i.e., raised hand posture) based on the aforementioned keypoint distance. The aforementioned first multi-person posture recognition model is trained on multiple human images with known human body region information and human body keypoint information. The aforementioned second multi-person pose recognition model may include, but is not limited to, a multi-person human behavior pose recognition network, used to identify whether the participants in the target video image are in a specified pose (e.g., but not limited to, standing or raising a hand). This second multi-person pose recognition model may be trained on, but is not limited to, multiple human images based on known human region information and human pose information; the human pose information indicates whether the human body is in a specified pose. The aforementioned human images include multiple human features. The aforementioned first multi-person pose recognition model may be, but is not limited to, a multi-person human keypoint detection machine learning model trained and optimized based on a YOLOV8-pose network and a human keypoint dataset (multiple human images with known human region information and human keypoint information). The aforementioned second multi-person pose recognition model may be, but is not limited to, a multi-person human behavior pose detection machine learning model trained and optimized based on a YOLOV8 network and a human pose dataset (multiple human images based on known human region information and whether the human body is in a specified pose).
[0113] Understandably, in order to ensure the accuracy of motion recognition, the human body recognition area information and posture recognition results corresponding to the above-mentioned multiple participants' features can be obtained by combining the recognition results corresponding to the first multi-person posture recognition model and the second multi-person posture recognition model. For example, but not limited to, only when both the first multi-person posture recognition model and the second multi-person posture recognition model recognize that the participants corresponding to the participants' features are in the specified posture will the posture recognition result of the participants corresponding to the participants' features be determined to be in the specified posture.
[0114] S602, Based on the human body recognition area information corresponding to the characteristics of the target participant, determine the target recognition selection area corresponding to the characteristics of the target participant in the target video image.
[0115] Specifically, the aforementioned target participant features are used to characterize the image information of the participant in a specified posture, as indicated by the corresponding posture recognition result. Once the target participant features in a specified posture (e.g., but not limited to, the human image information corresponding to a standing person or a person raising their hand) are identified in the target video image, the human recognition area corresponding to these features can be directly determined as the target recognition selection area in the target video image. This allows for flexible determination of the target recognition selection area for the target participant features involved in motion sensing recognition through the recognition of multiple human postures. Motion sensing recognition can be achieved without the target participant corresponding to the motion sensing feature standing in the specified recognition area, thus improving the flexibility and practicality of motion sensing recognition.
[0116] Please continue to refer to the following. Figure 3 ,like Figure 3 As shown, in step S302 above, after determining the target recognition selection area corresponding to the characteristics of the target participant in the target video image based on the human body recognition model, the motion recognition method may further include the following steps:
[0117] S303, scale and shift the target area image corresponding to the target recognition selection area to obtain the image to be recognized.
[0118] Optionally, when there is only one target participant feature selected or in a specified posture in the target video image, the target area image corresponding to the target participant feature can be directly enlarged to a preset size to obtain the corresponding image to be recognized, and displayed in the center of the electronic device's display screen.
[0119] Optionally, when there are multiple target participant features selected or in a specified posture in the target video image, after determining the target recognition selection area corresponding to the target participant features in the target video image based on the human body recognition model in S302, before scaling and shifting the target area image corresponding to the target recognition selection area to obtain the image to be recognized in S303, the corresponding target area image can be extracted from the target video image based on the target recognition selection area corresponding to each target participant feature in the target video image, thereby eliminating the influence of other target participant areas and background areas outside the target participant human body area in the target video image on the target participant human body perception recognition, and improving the accuracy and performance of multi-person human body perception recognition. The process of scaling and shifting the target region image corresponding to the target recognition selection area to obtain the image to be recognized in S303 above can be, but is not limited to: based on the position (center point coordinates) of the target recognition selection area corresponding to the features of each target participant in the target video image, the target region images corresponding to the features of each target participant are scaled and stitched in order (for example, but not limited to, according to the center point coordinates of the target recognition selection area from left to right or from right to left) to obtain the image to be recognized corresponding to the target video image.
[0120] For example, when determining Figure 7A After selecting the target recognition regions corresponding to the features of the left and right target participants in the target video image, the corresponding target region images can be extracted from the target video image based on the target recognition regions corresponding to the features of the two target participants. These two target region images are then sequentially scaled and stitched together to obtain the desired result. Figure 7B The image to be identified is shown.
[0121] Furthermore, in step S303 above, after scaling and shifting the target area image corresponding to the target recognition selection area to obtain the image to be recognized, the electronic device can, but is not limited to, render the image to be recognized on the display screen, so that the target participants and viewers participating in the motion recognition can more intuitively see the target area image corresponding to the target participant, and understand the progress and status of the motion recognition in real time. The image to be recognized is updated as the acquired target video image is updated. That is, after the target recognition selection area is determined, the image to be recognized will be updated in real time according to the target area image corresponding to the target recognition selection area in the real-time acquired target video image, thereby realizing real-time motion recognition of the target participant's features within the target recognition selection area.
[0122] Understandably, the recognition frequency of the human key point recognition model in electronic devices should be greater than or equal to the acquisition frequency of the target video image, so as to avoid the problem of missed recognition in the acquired target video image and ensure the real-time performance and accuracy of motion sensing recognition.
[0123] S304: Input the image to be recognized into the human key point recognition model and output the human key point recognition result corresponding to the image to be recognized.
[0124] Specifically, the aforementioned human key point recognition model is used to perform human key point recognition on its input image to be recognized. It can, but is not limited to, output the target human body recognition box information and human key point recognition results within the target recognition selection area in the image to be recognized, thereby determining the target action or target posture corresponding to each target participant in the image to be recognized.
[0125] Understandably, the human keypoint recognition model can be, but is not limited to, the YoloV8-pose model, YoloV7 model, YoloV5 model, RTMP model, RTMO model, etc. Since the features of each target participant in the image to be recognized have undergone human body region scaling compared to the target video image, and the target participant's body occupies the main part of the image to be recognized, the aforementioned human keypoint recognition model, compared to the human recognition model in S302, does not require high-resolution input to recognize the key points of the human body. The input size of the model is a crucial factor determining the computational load and affecting the model's speed. For example, when the input size decreases from 640*640 to 320*320, theoretically, the model's computational speed can be more than doubled. That is, by scaling and shifting the target region image corresponding to the target recognition selection area of each target participant's features before using the human keypoint recognition model for human keypoint recognition, the motion recognition provided in this application embodiment can achieve efficient model output on most low-end models, improving the efficiency, real-time performance, practicality, and applicability of motion recognition.
[0126] In this embodiment, on the one hand, based on the human body recognition model, the target recognition selection area corresponding to the features of the selected participants in the target video image who are participating in motion sensing or in a specified posture can be determined. Therefore, the target participant whose features are located does not need to stand in the specified recognition area. The target recognition selection area can be directly used as the area for motion sensing, improving the convenience, flexibility and interactivity of motion sensing and solving the problem of limited motion sensing area. On the other hand, by scaling and shifting the target area image corresponding to the target recognition selection area before inputting it into the human key point recognition model, the corresponding human key point recognition result is obtained. By scaling the target area image corresponding to the target recognition selection area, the interference of people in the background of the video image can be effectively eliminated, improving the stability and accuracy of motion sensing. At the same time, the input resolution requirement of the machine learning model (human key point recognition model) in the main process of motion sensing can be reduced. While ensuring the accuracy of model operation, the efficiency of model operation is improved, the performance requirements of the model are reduced, and the user coverage of motion sensing is increased.
[0127] To enhance students' enthusiasm and participation in answering questions during class, the motion recognition method provided in this application can be applied to classroom teaching scenarios. Please refer to the following for details. Figure 8 This is a flowchart illustrating another motion recognition method provided in an exemplary embodiment of this application. Figure 8 As shown, in a classroom teaching scenario, this motion recognition method may include, but is not limited to, the following steps:
[0128] S801, after the motion-sensing classroom application is launched, the question template is displayed.
[0129] Specifically, the electronic devices used in classroom teaching (such as, but not limited to, tablets) are equipped with motion-sensing classroom applications. When teachers need to use these applications to interact with students, they can first trigger the application's launch via buttons or the screen on the electronic device. After the application is launched, it will display a settable question template on the screen (such as, but not limited to, […]). Figure 9 As shown), this allows teachers to quickly fill in the corresponding question content text and question answer text, as well as other question text information.
[0130] Optionally, when the teacher fills in a multiple-choice question, the teacher can also set different human answering actions for different options based on the question template, so that the electronic device can determine whether the target participant (student) has answered the question correctly based on the target answering action indicated in the human key point recognition results of the identified target participant characteristics (student answering the question) and the human answering action corresponding to the question answer text.
[0131] S802 receives the question text information and the number of participants in answering the question based on the question template.
[0132] Specifically, after displaying the question template on the screen, the electronic device can also, but is not limited to, receive user-inputted question text information and the number of participants based on the question template. The question text information may, but is not limited to, include the question content text and the question answer text. The number of the aforementioned target participant characteristics is less than or equal to the aforementioned number of participants, that is, the actual number of participants should be less than or equal to the number of participants set by the teacher.
[0133] Understandably, the text of the above-mentioned questions may be, but is not limited to, English words for body parts or English phrases or sentences for body movements, and the corresponding answer text may be, but is not limited to, the corresponding description text of body movements; the text of the above-mentioned questions may be, but is not limited to, multiple-choice question text, and the corresponding answer text may be, but is not limited to, the corresponding answer option text and the description text of the body movement corresponding to the answer, etc., and this application embodiment does not limit this.
[0134] S803, acquire the target video image, the target video image includes at least one feature of the person to be involved.
[0135] Specifically, S803 is the same as S301, and will not be repeated here.
[0136] S804, based on the human body recognition model, determine the target recognition selection area corresponding to the target participant features in the target video image. The target participant features are used to characterize the image information of the participant selected from at least one participant feature to participate in the somatosensory recognition or the image information of the participant in a specified posture.
[0137] Specifically, S804 is the same as S302, and will not be repeated here.
[0138] S805 scales and shifts the target region image corresponding to the target recognition selection area to obtain the image to be recognized.
[0139] Specifically, S805 is the same as S303, and will not be repeated here.
[0140] S806 inputs the image to be recognized into the human key point recognition model and outputs the human key point recognition result corresponding to the image to be recognized.
[0141] Specifically, S806 is the same as S304, and will not be repeated here.
[0142] S807, based on the human body key point recognition results and the question answer text, determines the target participant's target answer result corresponding to the target participant's characteristics.
[0143] Specifically, after obtaining the human body key point recognition results corresponding to the target participant's characteristics, i.e., the student participating in answering the question, the electronic device can, but is not limited to, first determine the target participant's target answering action (e.g., but not limited to, opening the mouth, blinking, shaking the head, etc.) based on the target participant's human body key point recognition results. Then, it matches the identified target participant's target answering action with the corresponding human body answering action described in the question answer text. If the matching degree exceeds a threshold (e.g., but not limited to 80% or 90%), the target participant's target answering result is determined to be correct. If the matching degree does not exceed the threshold, the target participant's target answering result is determined to be incorrect. Thus, by combining electronic teaching with the motion recognition provided in the embodiments of this application, and utilizing the interactive capabilities of the interactive smart whiteboard in the classroom, the intelligence and fun of the classroom are improved. This allows for flexible and accurate motion recognition even in classroom spaces with many people and a large number of tables and chairs, reducing the threshold for using motion recognition and improving its practicality.
[0144] To improve the flexibility and real-time performance of motion-sensing games, the motion recognition method provided in this application can also be applied to virtual game activity scenarios. Please refer to [link / reference needed] for details. Figure 10 This is a flowchart illustrating another motion recognition method provided in an exemplary embodiment of this application. Figure 10 As shown, in virtual game activity scenarios, this motion recognition method may include, but is not limited to, the following steps:
[0145] S1001, acquire the target video image, the target video image includes at least one feature of the person to be involved.
[0146] Specifically, S1001 is the same as S301, and will not be repeated here.
[0147] S1002, Based on the human body recognition model, determine the target recognition selection area corresponding to the target participant features in the target video image. The target participant features are used to characterize the image information of the participant selected from at least one participant feature to participate in the somatosensory recognition or the image information of the participant in a specified posture.
[0148] Specifically, S1002 is the same as S302, and will not be repeated here.
[0149] S1003, the target area image corresponding to the target recognition selection area is scaled and shifted to obtain the image to be recognized.
[0150] Specifically, S1003 is the same as S303, and will not be repeated here.
[0151] S1004: Input the image to be recognized into the human key point recognition model and output the human key point recognition result corresponding to the image to be recognized.
[0152] Specifically, S1004 is the same as S304, and will not be repeated here.
[0153] S1005, Drive the movement of the target virtual object based on the human key point recognition results. The target virtual object is the virtual object bound to the target participant whose characteristics correspond to the target participant.
[0154] Specifically, the aforementioned human keypoint recognition results may include, but are not limited to, the positional information of human keypoints corresponding to the characteristics of each target participant in the image to be recognized. After the motion-sensing game (motion-sensing recognition) is started and the human keypoint recognition results corresponding to the image to be recognized are obtained, the target virtual object corresponding to the target participant characteristic displayed on the screen can be driven to perform corresponding movements based on the positional offset and positional offset direction of the same human keypoints corresponding to the target participant characteristics between the current frame of the image to be recognized and the previous frame of the image to be recognized. The aforementioned target virtual object may include, but is not limited to, the game character or corresponding virtual image bound to the target participant corresponding to the target participant characteristic.
[0155] For example, such as Figure 11 As shown, when the electronic device identifies target recognition selection areas 1110 and 1120 corresponding to the characteristics of two target participants in a standing posture, and after scaling and shifting the target area images corresponding to each of target recognition selection areas 1110 and 1120, it inputs them into the human key point recognition model to obtain the human key point recognition results corresponding to each of the two target recognition selection areas. Based on the real-time human key point recognition results in target recognition selection area 1110, it can directly drive the target virtual object 1130 bound to the target participant corresponding to the target participant characteristics in target recognition selection area 1110 on the electronic device display screen to perform corresponding movements, and based on the real-time human key point recognition results in target recognition selection area 1120, it can drive the target virtual object 1140 bound to the target participant corresponding to the target participant characteristics in target recognition selection area 1120 on the electronic device display screen to perform corresponding movements.
[0156] Please refer to the following. Figure 12 This is a schematic diagram of the structure of a motion-sensing recognition device provided in an exemplary embodiment of this application. Figure 12 As shown, the motion recognition device 1200 includes:
[0157] The acquisition module 1210 is used to acquire a target video image; the target video image includes at least one feature of a person to be involved.
[0158] The first determining module 1220 is used to determine the target recognition selection area corresponding to the target participant features in the target video image based on the human body recognition model; the target participant features are used to characterize the image information of the participant selected to participate in the somatosensory recognition or the image information of the participant in a specified posture among the at least one participant features; the specified posture is the human body posture that the participant should make corresponding to the specified participant features to participate in the somatosensory recognition.
[0159] The scaling and displacement processing module 1230 is used to scale and displacement the target area image corresponding to the above target recognition selection area to obtain the image to be recognized.
[0160] The key point recognition module 1240 is used to input the above-mentioned image to be recognized into the human key point recognition model and output the human key point recognition result corresponding to the above-mentioned image to be recognized.
[0161] In one possible implementation, the target video image includes features of multiple participants; the human body recognition model includes a multi-person human body recognition model.
[0162] The aforementioned first determining module 1220 includes:
[0163] The human body recognition unit is used to input the target video image into the multi-person human body recognition model and output the human body recognition area information corresponding to the features of the multiple participants; the human body recognition area information includes the position information and size information of the minimum bounding rectangle region corresponding to the feature of the participant in the target video image.
[0164] The display unit is used to display the human body recognition area corresponding to each of the multiple features of the participants in the target video image based on the position information and size information corresponding to each of the multiple features of the participants.
[0165] The receiving unit is used to receive the recognition area selection operation input by the user; the recognition area selection operation is used to select the target human body recognition area information of the target participant feature for participating in the somatosensory recognition from the human body recognition areas corresponding to the multiple participants' features.
[0166] The first determining unit is used to determine the target recognition selection area corresponding to the target participant characteristics based on the target human body recognition area information in response to the above-mentioned recognition area selection operation.
[0167] In one possible implementation, the target video image includes features of multiple participants; the human body recognition model includes a multi-person pose recognition model.
[0168] The aforementioned first determining module 1220 includes:
[0169] The posture recognition unit is used to input the target video image into the multi-person posture recognition model and output the human body recognition area information and posture recognition results corresponding to the features of the multiple participants.
[0170] The second determining unit is used to determine the target recognition selection area corresponding to the target participant features in the target video image based on the human body recognition area information corresponding to the target participant features; the target participant features are used to characterize the image information of the participant whose posture recognition result is in a specified posture.
[0171] In one possible implementation, the aforementioned multi-person pose recognition model includes a first multi-person pose recognition model and / or a second multi-person pose recognition model; the first multi-person pose recognition model includes a multi-person human keypoint recognition network, used to identify the keypoint distance between the shoulder keypoints and leg keypoints of the human skeleton of each participant in the target video image corresponding to the features of the participants, and to determine whether the participants corresponding to each feature of the participants in the target video image are in a specified pose based on the keypoint distance, and is trained based on multiple human images with known human body region information and human body keypoint information; the second multi-person pose recognition model includes a multi-person human behavior pose recognition network, used to identify whether the participants corresponding to each feature of the participants in the target video image are in a specified pose, and is trained based on multiple human images with known human body region information and human body pose information; wherein, the human body image includes multiple human body features.
[0172] In one possible implementation, the number of target participant features in the target video image is multiple; the somatosensory recognition device 1200 further includes:
[0173] The cropping module is used to crop the corresponding target region image from the above target video image based on the target recognition selection area corresponding to the characteristics of each of the above target participants;
[0174] The aforementioned scaling and displacement processing module 1230 is specifically used for:
[0175] Based on the location of the target recognition selection area corresponding to the features of each of the aforementioned target participants, the target area images corresponding to the features of each of the aforementioned target participants are sequentially scaled and stitched together to obtain the image to be recognized corresponding to the aforementioned target video image.
[0176] In one possible implementation, the aforementioned motion recognition device 1200 further includes:
[0177] The rendering module is used to render and display the image to be identified; the image to be identified is updated as the target video image is updated.
[0178] In one possible implementation, the aforementioned motion recognition device 1200 further includes:
[0179] The display module is used to display the question template after the motion-sensing classroom application is launched;
[0180] The receiving module is used to receive the question text information and the number of participants in the question based on the above question template; the above question text information includes the question content text and the question answer text; the number of the above target participant features is less than or equal to the above number of participants in the question.
[0181] The aforementioned motion recognition device 1200 also includes:
[0182] The second determining module is used to determine the target answer result of the target participant corresponding to the above-mentioned target participant characteristics based on the above-mentioned human body key point recognition results and the above-mentioned question answer text.
[0183] In one possible implementation, the aforementioned motion recognition device 1200 further includes:
[0184] The driving module is used to drive the movement of the target virtual object based on the above human key point recognition results; the above target virtual object is the virtual object bound to the target participant corresponding to the above target participant characteristics.
[0185] The division of modules in the above-described motion recognition device is for illustrative purposes only. In other embodiments, the motion recognition device can be divided into different modules as needed to complete all or part of the functions of the above-described motion recognition device. The implementation of each module in the motion recognition device provided in this application embodiment can be in the form of a computer program. This computer program can run on electronic devices such as terminals or servers. The program modules constituted by this computer program can be stored in the memory of electronic devices such as terminals or servers. When the computer program is executed by a processor, it implements all or part of the steps of the motion recognition method described in the embodiments of this application.
[0186] Please see below. Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 13 As shown, the electronic device 1300 includes: at least one processor 1310, at least one communication bus 1320, user interface 1330, at least one network interface 1340, and memory 1350.
[0187] The communication bus 1320 can be used to realize the connection and communication of the above components.
[0188] The user interface 1330 may include a display screen and a camera. Optionally, the user interface 1330 may also include a standard wired interface and a wireless interface.
[0189] The network interface 1340 may include a Bluetooth module, a Near Field Communication (NFC) module, a Wireless Fidelity (Wi-Fi) module, etc.
[0190] The processor 1310 may include one or more processing cores. The processor 1310 connects to various parts of the electronic device 1300 using various interfaces and lines, and performs various functions and processes data of the routing electronic device 1300 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1350, and by calling data stored in the memory 1350. Optionally, the processor 1310 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 1310 may integrate one or a combination of several of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display on the screen; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 1310 and may be implemented as a separate chip.
[0191] The memory 1350 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 1350 may include a non-transitory computer-readable medium. The memory 1350 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1350 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as motion sensing, scaling and displacement processing, key point recognition, etc.), instructions for implementing the various method embodiments described above, etc.; the data storage area may store data involved in the various method embodiments described above, etc. The memory 1350 may also be at least one storage device located remotely from the aforementioned processor 1310. Figure 13 As shown, the memory 1350, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and program instructions.
[0192] Specifically, the processor 1310 can be used to call program instructions stored in the memory 1350 and specifically perform the following operations: acquire a target video image; the target video image includes at least one feature of a participant; based on a human body recognition model, determine the target recognition selection area corresponding to the target participant feature in the target video image; the target participant feature is used to characterize the image information of the participant selected from the at least one feature of the participant to participate in somatosensory recognition or the image information of the participant in a specified posture; the specified posture is the human body posture that the participant should make corresponding to the specified feature of the participant to participate in somatosensory recognition; scale and shift the target area image corresponding to the target recognition selection area to obtain the image to be recognized; input the image to be recognized into a human body key point recognition model and output the human body key point recognition result corresponding to the image to be recognized.
[0193] In some possible embodiments, the target video image includes multiple features of the participants; the human body recognition model includes a multi-person human body recognition model.
[0194] When the processor 1310 executes the above-mentioned determination of the target recognition selection area corresponding to the target participant feature in the target video image based on the human body recognition model, it is specifically used to perform the following: inputting the target video image into the above-mentioned multi-person human body recognition model, and outputting the human body recognition area information corresponding to each of the multiple participants' features; the human body recognition area information includes the position information and size information of the minimum bounding rectangle region corresponding to the participants' feature in the target video image; displaying the human body recognition areas corresponding to each of the multiple participants' features in the target video image based on the position information and size information corresponding to each of the multiple participants' features; receiving a recognition area selection operation input by the user; the recognition area selection operation is used to select the target human body recognition area information of the target participant feature participating in the somatosensory recognition from the human body recognition areas corresponding to each of the multiple participants' features; and in response to the recognition area selection operation, determining the target recognition selection area corresponding to the target participant feature based on the target human body recognition area information.
[0195] In some possible embodiments, the target video image includes multiple features of the participants; the human body recognition model includes a multi-person pose recognition model.
[0196] When the processor 1310 executes the above-mentioned human body recognition model to determine the target recognition selection area corresponding to the target participant features in the above-mentioned target video image, it is specifically used to perform the following: inputting the above-mentioned target video image into the above-mentioned multi-person posture recognition model, outputting the human body recognition area information and posture recognition results corresponding to the features of the multiple participants; determining the target recognition selection area corresponding to the target participant features in the above-mentioned target video image based on the human body recognition area information corresponding to the target participant features; the target participant features are used to characterize the image information of the participant who is in a specified posture according to the posture recognition result.
[0197] In some possible embodiments, the above-mentioned multi-person posture recognition model includes a first multi-person posture recognition model and / or a second multi-person posture recognition model; the first multi-person posture recognition model includes a multi-person human keypoint recognition network, used to identify the keypoint distance between the shoulder keypoint and leg keypoint of the human skeleton of each participant in the target video image corresponding to the features of the participants, and to determine whether the participants corresponding to each feature of the participants in the target video image are in a specified posture based on the keypoint distance, and is trained based on multiple human images with known human body region information and human body keypoint information; the second multi-person posture recognition model includes a multi-person human behavior posture recognition network, used to identify whether the participants corresponding to each feature of the participants in the target video image are in a specified posture, and is trained based on multiple human images with known human body region information and human body posture information; wherein, the human body image includes multiple human body features.
[0198] In some possible embodiments, the number of target participant features in the target video image is multiple; after the processor 1310 executes the above-mentioned human body recognition model to determine the target recognition selection area corresponding to the target participant features in the target video image, before executing the above-mentioned scaling and displacement processing of the target region image corresponding to the target recognition selection area to obtain the image to be recognized, it is also used to execute:
[0199] Based on the target recognition selection area corresponding to the characteristics of each of the above-mentioned target participants, the corresponding target region image is extracted from the above-mentioned target video image.
[0200] When the processor 1310 performs the scaling and displacement processing on the target region image corresponding to the target recognition selection area to obtain the image to be recognized, it specifically performs the following:
[0201] Based on the location of the target recognition selection area corresponding to the features of each of the aforementioned target participants, the target area images corresponding to the features of each of the aforementioned target participants are sequentially scaled and stitched together to obtain the image to be recognized corresponding to the aforementioned target video image.
[0202] In some possible embodiments, after the processor 1310 performs the scaling and displacement processing on the target region image corresponding to the target recognition selection area to obtain the image to be recognized, it is further configured to perform:
[0203] The image to be identified is rendered and displayed; the image to be identified is updated as the target video image is updated.
[0204] In some possible embodiments, before the processor 1310 executes the above-described human body recognition model to determine the target recognition selection area corresponding to the target participant features in the target video image, it is further configured to perform:
[0205] After the motion-sensing classroom application is launched, a question template is displayed; based on the question template, the user-inputted question text information and the number of participants are received; the question text information includes the question content text and the question answer text; the number of the target participant characteristics is less than or equal to the number of participants.
[0206] After the processor 1310 executes the above-mentioned input of the image to be identified into the human key point recognition model and outputs the human key point recognition result corresponding to the image to be identified, it is also used to execute: based on the above-mentioned human key point recognition result and the above-mentioned question answer text, determine the target answer result of the target participant corresponding to the target participant's characteristics.
[0207] In some possible embodiments, after the processor 1310 performs the above-described actions of inputting the image to be recognized into the human key point recognition model and outputting the human key point recognition result corresponding to the image to be recognized, it is further configured to perform:
[0208] The target virtual object is driven to move based on the above human key point recognition results; the above target virtual object is the virtual object bound to the target participant corresponding to the above target participant characteristics on the above display screen.
[0209] This application also provides a computer storage medium storing instructions that, when run on a computer or processor, cause the computer or processor to perform one or more steps of any of the above methods. If the constituent modules of the above-described motion-sensing recognition device are implemented as software functional units and sold or used as independent products, they can be stored in the storage medium.
[0210] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0211] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.
[0212] The embodiments described above are merely preferred embodiments of this application and are not intended to limit the scope of this application. Any modifications and improvements made by those skilled in the art to the technical solutions of this application without departing from the spirit of this application should fall within the protection scope defined by the claims of this application.
Claims
1. A motion-sensing recognition method, characterized in that, The method includes: Acquire a target video image; the target video image includes at least one feature of the person to be involved. Based on the human body recognition model, the target recognition selection area corresponding to the target participant features in the target video image is determined; the target participant features are used to characterize the image information of the participant selected to participate in the somatosensory recognition or the image information of the participant in a specified posture from the at least one participant feature; the specified posture is the human body posture that the participant corresponding to the specified participant feature should make. The target region image corresponding to the target recognition selection area is scaled and shifted to obtain the image to be recognized; The image to be identified is input into the human key point recognition model, and the human key point recognition result corresponding to the image to be identified is output.
2. The method as described in claim 1, characterized in that, The target video image includes features of multiple participants; the human body recognition model includes a multi-person human body recognition model. The step of determining the target recognition selection region corresponding to the features of the target participant in the target video image based on the human body recognition model includes: The target video image is input into the multi-person human body recognition model, and the human body recognition area information corresponding to the features of the multiple participants is output; the human body recognition area information includes the position information and size information of the minimum bounding rectangle region corresponding to the feature of the participant in the target video image. Based on the location and size information corresponding to each of the multiple features of the participants, the human body recognition area corresponding to each of the multiple features of the participants is displayed in the target video image; The system receives a user input for a recognition area selection operation; the recognition area selection operation is used to select the target human body recognition area information of the target participant feature for participating in the somatosensory recognition from the human body recognition areas corresponding to the multiple participant features; In response to the recognition area selection operation, the target recognition selection area corresponding to the target participant's characteristics is determined based on the target human body recognition area information.
3. The method as described in claim 1, characterized in that, The target video image includes features of multiple participants; the human body recognition model includes a multi-person posture recognition model. The step of determining the target recognition selection region corresponding to the features of the target participant in the target video image based on the human body recognition model includes: The target video image is input into the multi-person posture recognition model, and the human body recognition area information and posture recognition results corresponding to the features of the multiple participants are output. Based on the human body recognition area information corresponding to the target participant features, the target recognition selection area corresponding to the target participant features in the target video image is determined; the target participant features are used to characterize the image information of the participant whose posture recognition result is in a specified posture.
4. The method as described in claim 3, characterized in that, The multi-person posture recognition model includes a first multi-person posture recognition model and / or a second multi-person posture recognition model; The first multi-person pose recognition model includes a multi-person human key point recognition network, which is used to identify the key point distance between the shoulder key points and leg key points of the human skeleton of each participant in the target video image, and to determine whether the participant corresponding to each participant feature in the target video image is in a specified pose based on the key point distance. It is trained based on multiple human images with known human body region information and human body key point information. The second multi-person posture recognition model includes a multi-person human behavior posture recognition network, which is used to identify whether the participants corresponding to the features of each participant in the target video image are in a specified posture, and is trained based on multiple human images with known human body region information and human posture information. The human body image includes multiple human body features.
5. The method as described in claim 1, characterized in that, The number of target participant features in the target video image is multiple; After determining the target recognition selection area corresponding to the features of the target participant in the target video image based on the human body recognition model, and before scaling and shifting the target region image corresponding to the target recognition selection area to obtain the image to be recognized, the method further includes: Based on the target recognition selection area corresponding to the characteristics of each target participant, the corresponding target region image is extracted from the target video image; The step of scaling and shifting the target region image corresponding to the target recognition selection area to obtain the image to be recognized includes: Based on the position of the target recognition selection area corresponding to the features of each target participant, the target area images corresponding to the features of each target participant are sequentially scaled and stitched together to obtain the image to be recognized corresponding to the target video image.
6. The method as described in claim 5, characterized in that, After scaling and shifting the target region image corresponding to the target recognition selection area to obtain the image to be recognized, the method further includes: The image to be identified is rendered and displayed; the image to be identified is updated as the target video image is updated.
7. The method as described in claim 1, characterized in that, Before determining the target recognition selection area corresponding to the target participant features in the target video image based on the human body recognition model, the method further includes: After the motion-sensing classroom application is launched, a question template is displayed; The system receives user-inputted question text information and the number of participants in the question based on the question template; the question text information includes the question content text and the question answer text; the number of target participant features is less than or equal to the number of participants in the question. After inputting the image to be identified into the human key point recognition model and outputting the human key point recognition result corresponding to the image to be identified, the method further includes: Based on the human body key point recognition results and the question answer text, the target participant's target answer result corresponding to the target participant's characteristics is determined.
8. The method as described in claim 1, characterized in that, After inputting the image to be identified into the human key point recognition model and outputting the human key point recognition result corresponding to the image to be identified, the method further includes: The target virtual object is driven to move based on the human body key point recognition results; the target virtual object is the virtual object bound to the target participant corresponding to the target participant characteristics.
9. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the method steps as claimed in any one of claims 1-8.
10. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the method steps as claimed in any one of claims 1-8.