Pose recognition method and device, electronic equipment and storage medium
By selecting the appropriate algorithm model according to the working conditions for hand position recognition, the problem of recognition lag in complex working conditions in virtual reality equipment is solved, and efficient and accurate position recognition under different conditions is achieved.
Patent Information
- Application Number
- CN202410141219.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-01
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art In virtual reality equipment, hand position recognition is prone to loss of tracking and lag under complex working conditions, and the algorithm is highly complex, affecting the user experience.
Different algorithm models are selected for pose recognition according to different working conditions: multiple algorithm models are used to identify them step by step in simple working conditions, and a single algorithm model is used to identify them at one time in complex working conditions, and predict pose information from bottom to upward.
Under different working conditions, the position information can be quickly and accurately identified, reducing delay and improving user experience.
Smart Images

Figure CN120451252A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a posture recognition method, device, electronic device, and storage medium. Background Art
[0002] In terminals such as virtual reality devices, the terminal's camera can control the terminal's functions by capturing the user's movements. For example, it can capture the user's hand gestures to generate a virtual hand image in the virtual space created by the terminal, or turn functions on or off based on the user's hand gestures. Summary of the Invention
[0003] The present disclosure provides a posture recognition method, device, electronic device and storage medium.
[0004] The present disclosure adopts the following technical solutions.
[0005] In some embodiments, the present disclosure provides a posture recognition method, comprising:
[0006] Acquire a target image, and determine a target operating condition corresponding to the target image, wherein the operating conditions include: a first operating condition and a second operating condition; the complexity of the first operating condition is less than the complexity of the second operating condition;
[0007] In response to the target working condition being a first working condition, using a first algorithm to identify the position and pose information of the target object in the target image; or, in response to the target working condition being a second working condition, using a second algorithm to identify the position and pose information of the target object in the target image;
[0008] Among them, the first algorithm adopts at least two algorithm models to identify the posture information of the target object in steps, and the second algorithm adopts one algorithm model to identify the posture information of the target object in one step.
[0009] In some embodiments, the present disclosure provides a posture recognition device, comprising:
[0010] a determination unit, configured to acquire a target image and determine a target operating condition corresponding to the target image, wherein the operating conditions include: a first operating condition and a second operating condition; the complexity of the first operating condition is less than the complexity of the second operating condition;
[0011] an algorithm unit, configured to, in response to the target working condition being a first working condition, identify the position and pose information of the target object in the target image using a first algorithm; or, in response to the target working condition being a second working condition, identify the position and pose information of the target object in the target image using a second algorithm;
[0012] Among them, the first algorithm adopts at least two algorithm models to identify the posture information of the target object in steps, and the second algorithm adopts one algorithm model to identify the posture information of the target object in one step.
[0013] In some embodiments, the present disclosure provides an electronic device comprising: at least one memory and at least one processor;
[0014] The memory is used to store program codes, and the processor is used to call the program codes stored in the memory to execute the above method.
[0015] In some embodiments, the present disclosure provides a computer-readable storage medium for storing program code, which, when executed by a processor, prompts the processor to perform the above method.
[0016] The present disclosure provides a posture recognition method comprising: acquiring a target image, determining a target operating condition corresponding to the target image, wherein the operating conditions include: a first operating condition and a second operating condition; the complexity of the first operating condition is less than the complexity of the second operating condition; in response to the target operating condition being the first operating condition, using a first algorithm to identify the posture information of the target object in the target image; or, in response to the target operating condition being the second operating condition, using a second algorithm to identify the posture information of the target object in the target image; wherein the first algorithm uses at least two algorithm models to identify the posture information of the target object in steps, and the second algorithm uses one algorithm model to identify the posture information of the target object in one step. The method proposed in the present disclosure can reduce the delay in identifying posture information. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0018] Figure 1 4 is a flowchart of a posture recognition method according to an embodiment of the present disclosure.
[0019] Figure 2 Schematic diagram of a posture recognition method according to an embodiment of the present disclosure.
[0020] Figure 3 It is a schematic diagram of the algorithm model adopted in the second algorithm of the embodiment of the present disclosure.
[0021] Figure 4 Schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0023] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0024] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0026] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0027] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed sequentially and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modification of “one” mentioned in the present disclosure is illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as “one or more”.
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] The solution provided by the embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0034] The terminal usually has a camera. For example, the terminal is an extended reality device such as a virtual reality device, a mixed reality device, an augmented reality device, etc. The camera of the terminal can capture the user's body movements, such as capturing the user's hands, and perform specific functions based on the posture information of the user's body movements. Taking the terminal as an extended reality device as an example, it restores the hands in the real world to the virtual world through bare hand tracking. In some technologies, a top-down approach is adopted, using multiple models to first obtain the position of the hand in the captured image, and then use the local image of the hand to restore the hand posture. This method has problems such as easy tracking and loss of the hand when the hand moves quickly, and high algorithm complexity. Due to the loss of the hand, multiple initializations will occur, and problems such as model disappearance and freezing will occur. The business complexity is high, the computing power overhead is large under complex working conditions, and the situation under different working conditions is not taken into account.
[0035] like Figure 1 As shown, Figure 1 4 is a flowchart of a posture recognition method according to an embodiment of the present disclosure, which includes the following steps.
[0036] S11. Acquire a target image and determine a target working condition corresponding to the target image.
[0037] In some embodiments, the execution end of the present method may be a terminal, such as an extended reality device such as a virtual reality device, a mixed reality device, an augmented reality device, etc. The target image may be an image of the target object captured by a camera on the terminal (such as a binocular camera), or may be an image captured by a camera on a non-terminal and then transmitted to the terminal. The target image may be an image in a video stream captured by the camera, or the target image may be a virtual image used to generate the target object in a virtual space. The working condition corresponding to the target image is the target working condition, and the working conditions include: a first working condition and a second working condition according to different types; the complexity of the first working condition is less than the complexity of the second working condition. The difference in working conditions reflects the difficulty of identifying the posture information of the target object in the target image. Under different working conditions, the complexity of identifying the posture information of the target object in the target image is also different. Compared with the image of the second working condition, the image of the first working condition is more friendly to the algorithm and is easier to identify the posture information.
[0038] S12. In response to the target working condition being the first working condition, using a first algorithm to identify the position and posture information of the target object in the target image; or, in response to the target working condition being the second working condition, using a second algorithm to identify the position and posture information of the target object in the target image.
[0039] In some embodiments, the target object may be a limb of the human body, such as a hand of the human body, and the posture information of the target object may be a gesture of the hand of the human body. In the first algorithm, at least two algorithm models are used to identify the posture information of the target object in steps, and in the second algorithm, one algorithm model is used to identify the posture information of the target object in one step. In some embodiments, at least two algorithm models are used in the first algorithm, and only one algorithm model is used in the second algorithm. The algorithm model may be a trained neural network algorithm model, and one algorithm model in the second algorithm may implement the functions of multiple algorithm models in the first algorithm. In the first algorithm, a top-down approach may be adopted, with each algorithm model performing one step respectively, and the posture information of the target object is identified in sequence through different algorithm models. For example, a tracker is first used to determine the local position of the target object, and then the position of the key point is determined by the local position, and then the posture information is determined. This calculation method has relatively high accuracy and high speed in identifying the posture information of the target object when the working conditions are good when the target image is captured (for example, the target object moves slowly and the light is good). However, when the working conditions are poor, it is easy to lose the target object in tracking, causing the terminal to freeze in the process of generating the virtual image of the target object, reducing the user experience, and because the data is in the The transmission among multiple algorithm modules can easily cause noise amplification. Therefore, in some embodiments of the present disclosure, for the case where the target working condition is the second working condition, the second algorithm is used to identify the posture information of the target object in the target image. At this time, only one algorithm model is used to predict the posture information of the target object from the target image at one time. It can adopt a bottom-up approach. Compared with the first algorithm that adopts multiple algorithm models, there is no need to use a tracker or rely on other algorithm models, there will be no noise amplification, and the recognition speed and accuracy are also higher than the first algorithm. The second algorithm is more suitable for scenarios with complex working conditions than the first algorithm, and it can achieve higher robustness with lower complexity.
[0040] In some embodiments of the present disclosure, when identifying the posture information of the target object in the target image, different algorithms are adopted according to the different target disclosures of the target image. When the target disclosure is relatively simple, multiple algorithm models are adopted in steps, for example, in a top-down manner, so as to improve the accuracy of the identified posture information when the target disclosure is relatively simple. When the target disclosure is relatively complex, it is not appropriate to adopt multiple algorithm models. Therefore, one algorithm model is adopted to identify the posture information of the target object in the target image at one time, so as to ensure the recognition accuracy and speed under complex working conditions. In the embodiments of the present disclosure, different numbers of algorithm models are adopted under different working conditions, so as to ensure that more accurate posture information can be obtained in a shorter time under different working conditions, so as to ensure the user experience.
[0041] In some embodiments of the present disclosure, determining the target working condition corresponding to the target image includes: determining the target working condition based on one or more of the position of the target object in the target image, the lighting condition of the target image, the processing time of the previous image within a preset time period before the target image, and the motion information of the target object.
[0042] In some embodiments, the target working condition represents the difficulty of identifying the posture information of the target object from the target image. The identification of the posture information of the target object from the target image is affected by multiple factors. Specifically, when the position of the target object in the target image is in the middle of the target image, it is easier to identify than when it is located at the edge of the target image (the position of the target object in the target image can be predicted based on the recognition results of the previous image of the target object in one or more frames before the target image). The target image is easier to identify when it is illuminated sufficiently than when it is illuminated insufficiently or excessively. The preceding image can be an image taken within a preset time period before the target image is captured. The target image can be an image in a video stream. The preceding image and the target image are captured at similar times, so the position, movement, etc. of the target object displayed by the two images will not change significantly. The processing time of the preceding image can be used to assess the working condition of the target image to a certain extent. If the processing time of the preceding image is very long or very short, the target image is likely to be a relatively complex working condition or a very simple one. The processing time of the preceding image can refer to the time it takes to identify the posture information of the target object from the preceding image. In some embodiments, the motion information of the target object can be obtained through sensors on a device associated with the target object. The device associated with the target object can have sensors. For example, if the target object is the user's hand, the hand is holding a handle, and the handle can have sensors. Alternatively, the target object can be a strap worn on the wrist or forearm, and the strap has sensors. The hand's motion information can also be indirectly sensed through information from the sensors on the wrist or forearm. Alternatively, a light spot of visible or invisible light is displayed on the handle, and optical tracking technology is used to identify the light spot on the handle, thereby determining the handle's position. Based on this position, the hand is positioned and motion information is determined. Sensors, for example, can be gyroscopes, accelerometers, geomagnetic sensors, or velocity sensors. This motion information can be used to determine the target object's movement speed, thereby helping to determine the target working condition of the target image.
[0043] In some embodiments of the present disclosure, if one or more of the following conditions are met, the target operating condition corresponding to the target object is determined to be the second operating condition: the target object is located outside a preset area in the target image, the lighting conditions do not meet the preset requirements, the processing time of the previous image is greater than a preset threshold, and the motion information meets the preset standard.
[0044] In some embodiments, the preset region may be, for example, the central region of the target image, and the size of the preset region may occupy at least 50% of the target image. If the target object is located outside the preset region, it indicates that the target object is at the edge of the target image, which is not conducive to identifying the target object in the target image. Therefore, the target operating condition is considered to be the second operating condition. In some embodiments, the preset requirement may be a pre-set brightness range. When the lighting conditions do not meet the preset requirement, it indicates that the target image is too bright or too dark, which is not conducive to identifying the target object in the target image. Therefore, the target operating condition is considered to be the second operating condition. In some embodiments, the preset threshold may be greater than twice the average processing time of the first algorithm for images in the first operating condition. If the processing time of the previous image is greater than the preset threshold, it indicates that the operating condition of the previous image is relatively complex and is not suitable for the first algorithm. Because the time interval between the capture of the previous image and the target image is less than the preset time interval (the preset time interval may be less than 1 second, such as 0.1 seconds, 0.3 seconds, or 0.5 seconds), the target image has a high similarity with the previous image. In this case, the target image is likely to be in the second operating condition. If the motion information meets the preset criteria, for example, the motion information shows that the target object's moving speed is greater than the preset speed, it indicates that the target object is moving at a relatively fast speed. At this time, the first algorithm is likely to cause tracking loss, so it is identified as the second working condition.
[0045] In some embodiments of the present disclosure, a first algorithm is used to identify the posture information of a target object in a target image, including: determining the local position of the target object in the target image through a detection model; inputting the local position into a key point model to determine the key point position in the target object; and inputting the key point position into the posture model to obtain the posture information of the target object.
[0046] In some embodiments, such as Figure 2 As shown on the left, the portion enclosed by the dotted box on the left schematically shows the first algorithm. Figure 2 The tracker in the method can be used to track the target object, so that the multi-camera faces the target object, the detection model obtains the target image taken by the terminal camera (multi-camera), identifies the position of the target object in the target image as the local position, and inputs the local position into the key point model. The key point model can extract the area where the target object is located as the local image based on the local position, and then determine the key point position on the local image. The key point position can be, for example, a skeleton point or a top point on the target object. For example, when the target object is a hand, the key point position can include the skeleton point of the finger, the fingertip and other positions. After obtaining the key point position, the geometric position relationship between the key point positions can be used to determine the pose information of the target object ( Figure 2The pose information may include 6-DOF information. In some embodiments, the target object is a hand, and the pose information of the target object may be a gesture made by the hand. In some embodiments, the hand gesture may control the terminal to execute a function pre-assigned to the gesture.
[0047] In some embodiments of the present disclosure, the algorithm model in the second algorithm is a trained algorithm model, and the input of the algorithm model in the second algorithm during the training process includes: an image of the target object; the output of the algorithm model in the second algorithm during the training process includes: posture information of the target object; and the output of the algorithm model in the second algorithm during the training process also includes: one or more of the local position of the target object in the input image, the key point position of the target object in the input image, and the local posture information of the target object in the input image.
[0048] In some embodiments, after the second algorithm obtains the input image, it extracts features from it to obtain basic features, and then predicts the pose information based on the basic features. Figure 2 When the BottomUp model in the training phase is used, a pre-annotated image is used as input to the algorithm model, and the algorithm model will additionally calculate the local position of the target object in the input image, the key point position of the target object in the input image, and the local pose information of the target object in the input image (pose information composed of parts of the target object). The calculated output data will be compared with the annotation of the input image, and then the parameters of the algorithm model will be continuously adjusted to reduce the loss function. In this way, by additionally training and outputting the above data for model training during the training phase, the accuracy of the pose information output by the algorithm model during use can be improved. Figure 3 For example, the target object is a hand. During the training phase, the hand detection frame (local position), palm key points (key point positions), finger postures (local posture information) and palm postures (posture information of the target object) will be output.
[0049] In some embodiments of the present disclosure, using a second algorithm to identify the pose information of a target object in a target image includes: using an algorithm model in the second algorithm to output only the pose information of the target object in the target image. In some embodiments, when using the algorithm model in the second algorithm, unlike in the training phase, only the pose information of the target object is output, without outputting the local position, key point positions, and local pose information, thereby improving response speed during use and reducing computing power consumption.
[0050] In some embodiments of the present disclosure, when the target operating condition is the first operating condition, the accuracy of posture information recognition using the first algorithm is higher than that using the second algorithm. In the first operating condition, the first algorithm has higher accuracy and the operating condition is relatively simple, so the time consumption is not too long. Therefore, using the first algorithm can better ensure the user experience.
[0051] In some embodiments of the present disclosure, when the target working condition is the second working condition, the speed of identifying the posture information using the second algorithm is higher than the speed of identifying the posture information using the first algorithm. Under the second working condition, the working condition is more complex, and the accuracy of the first algorithm cannot be guaranteed. The accuracy of the posture information identified by the second algorithm is lower than that of the second algorithm, and it is more time-consuming and prone to losing the target object and causing lag. The second algorithm does not rely on other algorithm models, has no tracker, has low complexity, and is faster than the first algorithm under the second working condition. The accuracy of identifying the posture information is also sufficient. Especially under the working condition where the target object moves rapidly, the second algorithm can achieve better robustness with lower complexity and reduce latency.
[0052] In the embodiment of the present disclosure, the first algorithm or the second algorithm is adopted according to different working conditions. The second algorithm predicts the posture information of the target object at one time, improves the detection success rate and reduces the delay, and can be applied to virtual reality devices for user bare hand tracking.
[0053] Some embodiments of the present disclosure further provide a posture recognition device, including:
[0054] a determining unit, configured to determine, in response to acquiring a target image, a target operating condition corresponding to the target image, wherein the operating condition includes: a first operating condition and a second operating condition; and the complexity of the first operating condition is less than the complexity of the second operating condition;
[0055] an algorithm unit, configured to, in response to the target working condition being the first working condition, identify the position and pose information of the target object in the target image using a first algorithm; or, in response to the target working condition being the second working condition, identify the position and pose information of the target object in the target image using a second algorithm;
[0056] Among them, the first algorithm adopts at least two algorithm models to identify the posture information of the target object in steps, and the second algorithm adopts one algorithm model to identify the posture information of the target object in one step.
[0057] In some embodiments, determining the target working condition corresponding to the target image includes: determining the target working condition based on one or more of the position of the target object in the target image, the lighting conditions of the target image, the processing time of the previous image within a preset time period before the target image, and the motion information of the target object.
[0058] In some embodiments, if one or more of the following conditions are met, the target operating condition corresponding to the target object is determined to be the second operating condition: the target object is located outside the preset area in the target image, the lighting conditions do not meet the preset requirements, the processing time of the previous image is greater than a preset threshold, and the motion information meets the preset standard.
[0059] In some embodiments, using a first algorithm to identify the pose information of the target object in the target image includes: determining a local position of the target object in the target image using a detection model;
[0060] Inputting the local position into a key point model to determine the key point position in the target object;
[0061] The key point positions are input into a pose model to obtain the pose information of the target object.
[0062] In some embodiments, the algorithm model in the second algorithm is a trained algorithm model, and the input of the algorithm model in the second algorithm during the training process includes: an image of the target object;
[0063] The output of the algorithm model in the second algorithm during the training process includes: the posture information of the target object; and the output of the algorithm model in the second algorithm during the training process also includes: the local position of the target object in the input image, the key point position of the target object in the input image, and one or more of the local posture information of the target object in the input image.
[0064] In some embodiments, using a second algorithm to identify the position and posture information of the target object in the target image includes: using an algorithm model in the second algorithm to output only the position and posture information of the target object in the target image.
[0065] In some embodiments, at least one of the following conditions is met: when the target working condition is the first working condition, the accuracy of identifying the posture information using the first algorithm is higher than the accuracy of identifying the posture information using the second algorithm; when the target working condition is the second working condition, the speed of identifying the posture information using the second algorithm is higher than the speed of identifying the posture information using the first algorithm; and the target object is a human hand.
[0066] For the embodiments of the device, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separation modules may or may not be separate. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those of ordinary skill in the art can understand and implement them without paying creative work.
[0067] The method and apparatus of the present disclosure are described above based on the embodiments and application examples. In addition, the present disclosure also provides an electronic device and a computer-readable storage medium, which are described below.
[0068] Reference below Figure 4 , which shows a schematic diagram of the structure of an electronic device (e.g., a terminal device or server) 800 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in the figure is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0069] The electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0070] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows the electronic device 800 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0071] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0072] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0073] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0074] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0075] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method of the present disclosure.
[0076] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0078] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0079] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0080] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0081] According to one or more embodiments of the present disclosure, a posture recognition method is provided, comprising:
[0082] In response to acquiring a target image, determining a target operating condition corresponding to the target image, wherein the operating conditions include: a first operating condition and a second operating condition; the complexity of the first operating condition is less than the complexity of the second operating condition;
[0083] In response to the target working condition being a first working condition, using a first algorithm to identify the position and pose information of the target object in the target image; or, in response to the target working condition being a second working condition, using a second algorithm to identify the position and pose information of the target object in the target image;
[0084] Among them, the first algorithm adopts at least two algorithm models to identify the posture information of the target object in steps, and the second algorithm adopts one algorithm model to identify the posture information of the target object in one step.
[0085] According to one or more embodiments of the present disclosure, a posture recognition method is provided, which, in response to acquiring a target image, determines a target working condition corresponding to the target image, including:
[0086] The target working condition is determined based on one or more of the position of the target object in the target image, the lighting condition of the target image, the processing time of a previous image within a preset time period before the target image, and the motion information of the target object.
[0087] According to one or more embodiments of the present disclosure, a posture recognition method is provided, wherein if one or more of the following conditions are met, the target operating condition corresponding to the target object is determined to be the second operating condition:
[0088] The target object is located outside a preset area in the target image, the lighting condition does not meet a preset requirement, the processing time of the previous image is greater than a preset threshold, and the motion information meets a preset standard.
[0089] According to one or more embodiments of the present disclosure, a posture recognition method is provided, which uses a first algorithm to recognize posture information of a target object in a target image, including:
[0090] Determine the local position of the target object in the target image by using a detection model;
[0091] Inputting the local position into a key point model to determine the key point position in the target object;
[0092] The key point positions are input into a pose model to obtain the pose information of the target object.
[0093] According to one or more embodiments of the present disclosure, a posture recognition method is provided, wherein the algorithm model in the second algorithm is a trained algorithm model, and the input of the algorithm model in the second algorithm during the training process includes: an image of the target object;
[0094] The output of the algorithm model in the second algorithm during the training process includes: the posture information of the target object; and the output of the algorithm model in the second algorithm during the training process also includes: the local position of the target object in the input image, the key point position of the target object in the input image, and one or more of the local posture information of the target object in the input image.
[0095] According to one or more embodiments of the present disclosure, a posture recognition method is provided, which uses a second algorithm to recognize posture information of a target object in the target image, including:
[0096] The algorithm model in the second algorithm is used to output only the posture information of the target object in the target image.
[0097] According to one or more embodiments of the present disclosure, a posture recognition method is provided, which satisfies at least one of the following conditions:
[0098] When the target working condition is the first working condition, the accuracy of the posture information identified by the first algorithm is higher than the accuracy of the posture information identified by the second algorithm;
[0099] When the target working condition is the second working condition, the speed of recognizing the posture information using the second algorithm is higher than the speed of recognizing the posture information using the first algorithm;
[0100] The target object is a human hand.
[0101] According to one or more embodiments of the present disclosure, a posture recognition device is provided, comprising:
[0102] a determining unit, configured to determine, in response to acquiring a target image, a target operating condition corresponding to the target image, wherein the operating conditions include: a first operating condition and a second operating condition; the complexity of the first operating condition is less than the complexity of the second operating condition;
[0103] an algorithm unit, configured to, in response to the target working condition being a first working condition, identify the position and pose information of the target object in the target image using a first algorithm; or, in response to the target working condition being a second working condition, identify the position and pose information of the target object in the target image using a second algorithm;
[0104] Among them, the first algorithm adopts at least two algorithm models to identify the posture information of the target object in steps, and the second algorithm adopts one algorithm model to identify the posture information of the target object in one step.
[0105] According to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one memory and at least one processor;
[0106] The at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute any one of the above methods.
[0107] According to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium is used to store program code, and when the program code is executed by a processor, the processor is prompted to perform the above method.
[0108] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0109] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0110] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A posture recognition method, characterized in that: include: Acquire a target image, and determine a target operating condition corresponding to the target image, wherein the operating conditions include: a first operating condition and a second operating condition; the complexity of the first operating condition is less than the complexity of the second operating condition; In response to the target working condition being a first working condition, using a first algorithm to identify the position and pose information of the target object in the target image; or, in response to the target working condition being a second working condition, using a second algorithm to identify the position and pose information of the target object in the target image; Among them, the first algorithm adopts at least two algorithm models to identify the posture information of the target object in steps, and the second algorithm adopts one algorithm model to identify the posture information of the target object in one step.
2. The method according to claim 1, characterized in that Determining a target operating condition corresponding to the target image includes: The target working condition is determined based on one or more of the position of the target object in the target image, the lighting condition of the target image, the processing time of a previous image within a preset time period before the target image, and the motion information of the target object.
3. The method according to claim 2, characterized in that If one or more of the following conditions are met, it is determined that the target operating condition corresponding to the target object is the second operating condition: The target object is located outside a preset area in the target image, the lighting condition does not meet a preset requirement, the processing time of the previous image is greater than a preset threshold, and the motion information meets a preset standard.
4. The method according to claim 1, wherein Using a first algorithm to identify the pose information of the target object in the target image includes: Determine the local position of the target object in the target image by using a detection model; Inputting the local position into a key point model to determine the key point position in the target object; The key point positions are input into a pose model to obtain the pose information of the target object.
5. The method according to claim 1, wherein The algorithm model in the second algorithm is a trained algorithm model, and the input of the algorithm model in the second algorithm during the training process includes: an image of the target object; The output of the algorithm model in the second algorithm during the training process includes: the posture information of the target object; and the output of the algorithm model in the second algorithm during the training process also includes: the local position of the target object in the input image, the key point position of the target object in the input image, and one or more of the local posture information of the target object in the input image.
6. The method according to claim 5, characterized in that Using a second algorithm to identify the pose information of the target object in the target image includes: The algorithm model in the second algorithm is used to output only the posture information of the target object in the target image.
7. The method according to claim 1, characterized in that Satisfy at least one of the following: When the target working condition is the first working condition, the accuracy of the posture information identified by the first algorithm is higher than the accuracy of the posture information identified by the second algorithm; When the target working condition is the second working condition, the speed of recognizing the posture information using the second algorithm is higher than the speed of recognizing the posture information using the first algorithm; The target object is a human hand.
8. A posture recognition device, characterized in that: include: a determination unit, configured to acquire a target image and determine a target operating condition corresponding to the target image, wherein the operating conditions include: a first operating condition and a second operating condition; the complexity of the first operating condition is less than the complexity of the second operating condition; an algorithm unit, configured to, in response to the target working condition being a first working condition, identify the position and pose information of the target object in the target image using a first algorithm; or, in response to the target working condition being a second working condition, identify the position and pose information of the target object in the target image using a second algorithm; Among them, the first algorithm adopts at least two algorithm models to identify the posture information of the target object in steps, and the second algorithm adopts one algorithm model to identify the posture information of the target object in one step.
9. An electronic device comprising: at least one memory and at least one processor; The at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the method according to any one of claims 1 to 7. 10 . A computer-readable storage medium, configured to store program code, wherein when the program code is executed by a processor, the processor is prompted to execute the method according to claim 1 .