A method and device for recognizing a bone posture based on virtual-real interaction
By using a multi-lens camera to detect human images, generate 3D stereoscopic images, and switch virtual 3D coordinates, the problem of interaction between human skeletons and virtual characters in a 3D environment is solved, thus realizing interaction between the human body and virtual characters.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN YUNZHONGXIANG TECH CO LTD
- Filing Date
- 2023-09-25
- Publication Date
- 2026-04-10
AI Technical Summary
Current technology cannot achieve interaction between human skeletons and virtual characters in a three-dimensional environment.
By acquiring human images through a multi-lens camera, detecting two-dimensional skeletons and joints, generating three-dimensional images, and switching virtual three-dimensional coordinates according to action commands, the interaction between the human body and the virtual character can be realized.
It enables effective interaction between the human body and virtual characters in a three-dimensional environment, generates a virtual skeletal model, and can generate corresponding actions based on human body movements.
Smart Images

Figure CN117275039B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of skeleton posture recognition, and particularly relates to a skeleton posture recognition method and device based on virtual-real interaction. BACKGROUND
[0002] The so-called human skeleton posture recognition refers to matching abstract level features with a human model to obtain the posture of a target at different moments. The human skeleton posture recognition is a core problem of human motion capture. After recognizing the skeleton of a human body on the market, virtual interaction is not combined. In the prior art, virtual interaction tends to be interaction in a two-dimensional and fixed three-dimensional environment, and cannot interact with a virtual character in a three-dimensional environment according to the skeleton of a human body.
[0003] Therefore, it is urgent to provide a skeleton posture recognition method and device based on virtual-real interaction to solve the technical problem that a virtual character cannot be interacted with in a three-dimensional environment according to the skeleton of a human body in the prior art. SUMMARY
[0004] Therefore, it is urgent to provide a skeleton posture recognition method and device based on virtual-real interaction to solve the technical problem that a virtual character cannot be interacted with in a three-dimensional environment according to the skeleton of a human body in the prior art.
[0005] In one aspect, the present application provides a skeleton posture recognition method based on virtual-real interaction, comprising:
[0006] obtaining a preset number of human body images of a human body based on a multi-lens camera, detecting the preset number of human body images to obtain all two-dimensional skeletons and a preset number of joint nodes of the human body;
[0007] generating a three-dimensional image according to the preset number of human body images, and processing the three-dimensional image according to the all two-dimensional skeletons and the preset number of joint nodes to obtain virtual three-dimensional coordinates corresponding to each two-dimensional skeleton;
[0008] when an action instruction is received, determining a target two-dimensional skeleton according to the action instruction, and switching the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton to three-dimensional coordinates corresponding to the target action according to the action in the action instruction
[0009] In some possible implementation manners, the detecting the preset number of human body images to obtain the all two-dimensional skeletons and the preset number of joint nodes comprises:
[0010] detecting the preset number of human body images according to a preset human skeleton detection neural network to obtain at least one two-dimensional skeleton corresponding to each human body image;
[0011] The preset number of human body images are compared, and duplicate two-dimensional skeletons in the preset number of human body images are removed to obtain all two-dimensional skeletons of the human body;
[0012] According to the joints of each two-dimensional skeleton, a preset number of joint nodes of all two-dimensional skeletons are obtained.
[0013] In some possible implementation manners, the obtaining of the preset number of joint nodes of all two-dimensional skeletons according to the joints of each two-dimensional skeleton includes:
[0014] According to the joints of all two-dimensional skeletons, a corresponding joint node of each two-dimensional skeleton is obtained.
[0015] All joint nodes of all two-dimensional skeletons are screened to obtain a preset number of joint nodes.
[0016] In some possible implementation manners, the processing of the three-dimensional stereoscopic image according to the all two-dimensional skeletons and the preset number of joint nodes to obtain a virtual three-dimensional coordinate corresponding to each two-dimensional skeleton includes:
[0017] According to relative information of the multi-lens camera and the three-dimensional stereoscopic image, depth and position data of the three-dimensional stereoscopic image are obtained.
[0018] According to the all two-dimensional skeletons and the preset number of joint nodes, a human body contour of the three-dimensional stereoscopic image is detected to obtain a height of each two-dimensional skeleton on the three-dimensional stereoscopic image.
[0019] According to the height corresponding to each two-dimensional skeleton, the depth, and the position data, a virtual three-dimensional coordinate corresponding to each two-dimensional skeleton is generated.
[0020] In some possible implementation manners, after the processing of the three-dimensional stereoscopic image according to the all two-dimensional skeletons and the preset number of joint nodes to obtain a virtual three-dimensional coordinate corresponding to each two-dimensional skeleton, the method further includes:
[0021] Each two-dimensional skeleton is detected to obtain an identifier corresponding to each two-dimensional skeleton.
[0022] According to the binding of the identifier and a skeleton corresponding to the human body, the skeleton of the human body is detected.
[0023] In some possible implementation manners, when the action instruction is received, a target two-dimensional skeleton is determined according to the action instruction, including:
[0024] By detecting the human body, an action instruction of the human body is received; the action instruction includes a target skeleton that performs an action.
[0025] Based on the target skeleton of the action command, determine the target identifier that is bound to the target skeleton;
[0026] Based on the target identifier, the target two-dimensional skeleton in the virtual space is determined.
[0027] In some possible implementations, switching the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton to the three-dimensional coordinates corresponding to the target action according to the action in the action instruction includes:
[0028] Based on the action command, the target skeleton is determined to perform the target action;
[0029] Perform corresponding operations on the target 2D skeleton according to the target action, and switch the virtual 3D coordinates corresponding to the target 2D skeleton to the 3D coordinates corresponding to the target action.
[0030] In some possible implementations, generating a three-dimensional image based on the preset number of human body images includes:
[0031] The preset number of human body images are stitched together to obtain a three-dimensional image.
[0032] In some possible implementations, after determining that the target skeleton performs a target action based on the action instruction, the method further includes:
[0033] The target action is optimized to obtain the optimized target action;
[0034] Determine whether the amplitude of the optimized target action is greater than the amplitude threshold;
[0035] If not, wait for the next action command.
[0036] On the other hand, the present invention also provides a skeletal pose recognition device based on virtual-real interaction, comprising:
[0037] The image acquisition module is used to acquire a preset number of human body images based on a multi-lens camera, detect the preset number of human body images, and obtain all two-dimensional skeletons and a preset number of joints of the human body.
[0038] The coordinate determination module is used to generate a three-dimensional image based on the preset number of human images, and to process the three-dimensional image based on all the two-dimensional skeletons and the preset number of joints to obtain the virtual three-dimensional coordinates corresponding to each two-dimensional skeleton.
[0039] The action execution module is configured to, when receiving the action instruction, determine a target two-dimensional skeleton according to the action instruction, and switch the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton to the three-dimensional coordinates corresponding to the action according to the action in the action instruction.
[0040] The beneficial effects of the above embodiments are that the skeleton posture recognition method based on virtual-real interaction provided by the present application obtains all two-dimensional skeletons and a preset number of joint nodes by processing human body images, generates a three-dimensional image, and processes the three-dimensional image according to all two-dimensional skeletons and the preset number of joint nodes to obtain virtual three-dimensional coordinates corresponding to each two-dimensional skeleton. When an action instruction is received, a target two-dimensional skeleton is determined according to the action instruction, and the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton are switched to the three-dimensional coordinates corresponding to the action according to the action in the action instruction. The present application obtains all two-dimensional skeletons and a preset number of joint nodes by processing human body images, generates a three-dimensional image, and thus obtains a virtual skeleton model of a human body. Furthermore, an action instruction can be generated according to an action performed by the human body, so that the virtual skeleton model can perform a corresponding action according to the action instruction, and the purpose of interaction between a human body and a virtual character in a three-dimensional environment is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0042] Figure 1 An embodiment flow diagram of the skeleton posture recognition method based on virtual-real interaction provided by the present application;
[0043] Figure 2 An embodiment structure diagram of the skeleton posture recognition device based on virtual-real interaction provided by the present application;
[0044] Figure 3 An embodiment structure diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0045] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of the present application.
[0046] Some block diagrams shown in the drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor systems and / or microcontroller systems.
[0047] Reference to "an embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily mutually exclusive of other embodiments. It is explicitly and implicitly understood that the embodiments described herein can be combined with other embodiments.
[0048] The embodiments of the present application provide a virtual-real interaction-based skeleton posture recognition method and device, which are described below.
[0049] Figure 1 An embodiment flow diagram of a virtual-real interaction-based skeleton posture recognition method provided by the present application is shown in Figure 1 The virtual-real interaction-based skeleton posture recognition method includes:
[0050] S101, acquiring a preset number of human body images of a human body based on a multi-lens camera, detecting the preset number of human body images, and obtaining all two-dimensional skeletons and a preset number of joint nodes of the human body;
[0051] S102, generating a three-dimensional image based on the preset number of human body images, and processing the three-dimensional image based on all two-dimensional skeletons and the preset number of joint nodes to obtain virtual three-dimensional coordinates corresponding to each two-dimensional skeleton;
[0052] S103, when an action instruction is received, determining a target two-dimensional skeleton based on the action instruction, and switching the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton to three-dimensional coordinates corresponding to a target action based on the action in the action instruction.
[0053] Compared with the prior art, the virtual-real interaction-based skeleton posture recognition method provided by the application obtains a preset number of human body images of a human body based on a multi-lens camera, detects the preset number of human body images to obtain all two-dimensional skeletons and a preset number of joint nodes of the human body, generates a three-dimensional image based on the preset number of human body images, processes the three-dimensional image based on all two-dimensional skeletons and the preset number of joint nodes to obtain virtual three-dimensional coordinates corresponding to each two-dimensional skeleton, determines a target two-dimensional skeleton based on a motion instruction when the motion instruction is received, and switches the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton to three-dimensional coordinates corresponding to the motion based on the motion in the motion instruction. The application obtains all two-dimensional skeletons and a preset number of joint nodes by processing human body images and generates a three-dimensional image, so that a virtual skeleton model of the human body can be obtained. Further, a motion instruction can be generated based on a motion of the human body, so that the virtual skeleton model can perform a corresponding motion based on the motion instruction, and the purpose of interaction between the human body and a virtual character in a three-dimensional environment is achieved.
[0054] It should be noted that, in order to detect the human body image, the preset number of human body images need to be detected. In some embodiments of the application, step S101 comprises:
[0055] The preset number of human body images are detected based on the preset human skeleton detection neural network to obtain at least one two-dimensional skeleton corresponding to each human body image.
[0056] The preset number of human body images are compared to remove repeated two-dimensional skeletons in the preset number of human body images to obtain all two-dimensional skeletons of the human body.
[0057] The preset number of joint nodes of all two-dimensional skeletons are obtained based on joints of each two-dimensional skeleton.
[0058] It should be noted that the preset human skeleton detection neural network can be connected with T (T is a positive integer greater than or equal to 1) stages after a VGG-19 network, and each stage has a structure of two full convolutional networks. VGG (Visual Geometry Group) belongs to the Department of Science and Engineering of the University of Oxford, which has published a series of convolutional network models starting with VGG. The preset human skeleton detection neural network can also be other existing neural network models. The specific model creation method is not limited in the embodiments of the application.
[0059] In specific embodiments of the present application, the preset number of human body images can be input into the preset human body skeleton detection neural network, and the preset number of human body images can detect each human body image to obtain at least one two-dimensional skeleton contained in each human body image, so that all two-dimensional skeletons of the preset number of human body images can be obtained. However, because the preset number of human body images contains repeated picture content, the detected two-dimensional skeletons have repeated skeletons, and therefore, the preset number of human body images needs to be compared to delete the repeated two-dimensional skeletons among all two-dimensional skeletons of the preset number of human body images, and all two-dimensional skeletons containing all skeletons of the human body are obtained.
[0060] In some embodiments of the present application, according to the joints of each two-dimensional skeleton, a preset number of joint nodes of all two-dimensional skeletons are obtained, including:
[0061] According to the joints of all two-dimensional skeletons, the joint nodes corresponding to each two-dimensional skeleton are obtained.
[0062] All joint nodes of all two-dimensional skeletons are screened to obtain a preset number of joint nodes.
[0063] In specific embodiments of the present application, when the preset human body skeleton detection neural network detects the human body image, the joint nodes corresponding to each two-dimensional skeleton can also be detected, so that all joint nodes can be obtained. The joint nodes can be screened according to the screening requirements, so that the required preset number of joint nodes can be obtained. The specific screening requirements can be set by the staff according to their own experience and actual situation, and the embodiments of the present application are not limited herein.
[0064] In some embodiments of the present application, step S102 includes:
[0065] The preset number of human body images are image spliced to obtain a three-dimensional image.
[0066] In specific embodiments of the present application, the preset number of human body images can be image spliced to generate a three-dimensional image.
[0067] In some embodiments of the present application, step S102 includes:
[0068] According to the relative information of the multi-lens camera and the three-dimensional image, the depth and position data of the three-dimensional image are obtained.
[0069] According to all two-dimensional skeletons and the preset number of joint nodes, the human body contour of the three-dimensional image is detected to obtain the height of each two-dimensional skeleton on the three-dimensional image.
[0070] According to the height, depth and position data corresponding to each two-dimensional skeleton, a virtual three-dimensional coordinate corresponding to each two-dimensional skeleton is generated.
[0071] In specific embodiments of the present application, after generating the three-dimensional stereoscopic image, the depth of the three-dimensional stereoscopic image can be obtained according to the distance between the multi-lens camera for shooting the human body image and the three-dimensional stereoscopic image, the position data of the three-dimensional stereoscopic image can also be obtained according to the relative information between the multi-lens camera and the three-dimensional stereoscopic image, and the contour of the human body in the three-dimensional stereoscopic image can be restored according to all the two-dimensional skeletons and the preset number of joint nodes, so as to be consistent with the human body. The height corresponding to each two-dimensional skeleton can be obtained according to the position of each two-dimensional skeleton on the three-dimensional stereoscopic image, so that the height can be determined as the Z axis, the depth can be determined as the X axis, and the position data can be determined as the Y axis, thereby obtaining the virtual three-dimensional coordinates corresponding to each two-dimensional skeleton.
[0072] In some embodiments of the present application, step S102 further comprises:
[0073] detecting each two-dimensional skeleton to obtain an identifier corresponding to each two-dimensional skeleton;
[0074] binding the skeleton corresponding to the identifier to the human body, and detecting the skeleton of the human body.
[0075] In specific embodiments of the present application, an identifier can be set for each two-dimensional skeleton, which can be bound to the skeleton of the corresponding part of the human body, so that the human body can be detected.
[0076] In some embodiments of the present application, step S103 comprises:
[0077] receiving an action instruction of the human body by detecting the human body; the action instruction comprises a target skeleton for performing an action;
[0078] determining a target identifier bound to the target skeleton according to the target skeleton of the action instruction;
[0079] determining a target two-dimensional skeleton in the virtual space according to the target identifier.
[0080] In specific embodiments of the present application, when it is detected that the human body has performed an action, an action instruction issued by the device can be received, the action instruction can include a target skeleton corresponding to the action and a corresponding target action, so that the target identifier of the target skeleton for performing the action can be determined, and the target two-dimensional skeleton corresponding to the three-dimensional stereoscopic image in the virtual three-dimensional space can be determined according to the target identifier.
[0081] In some embodiments of the present application, step S103 comprises:
[0082] determining that the target skeleton performs a target action according to the action instruction;
[0083] According to the target action, corresponding operation is performed on the target two-dimensional skeleton, and the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton are switched to the three-dimensional coordinates corresponding to the target action.
[0084] In specific embodiments of the present application, after the target identifier is determined, corresponding operation can be performed on the target two-dimensional skeleton corresponding to the target identifier according to the target action of the target skeleton in the action instruction. After the action is performed, the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton can be switched to the virtual three-dimensional coordinates corresponding to the target action after the target action is completed, thereby updating the virtual three-dimensional coordinates. Thus, the three-dimensional image in the virtual three-dimensional space can be controlled according to the action of the human body, and virtual interaction between the human body and the three-dimensional image in the virtual three-dimensional space is realized.
[0085] In some embodiments of the present application, according to the action instruction, after the target skeleton performs the target action, the method further comprises:
[0086] optimizing the target action to obtain an optimized target action;
[0087] judging whether the action amplitude of the optimized target action is greater than an amplitude threshold value;
[0088] If not, the next action instruction is waited for.
[0089] In specific embodiments of the present application, when the action instruction is received, the target action in the action instruction can be optimized. The optimization can be processing the coordinates of the target action to obtain stable coordinates, so that the three-dimensional image performs the target action more smoothly. The optimized target action can also be detected to judge whether the target action is an action operated by the human body. The action amplitude of the target action can be judged. If it is greater than the amplitude threshold value, it is determined that the target action is an action operated by the human body. Then, the step of "according to the target action, corresponding operation is performed on the target two-dimensional skeleton, and the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton are switched to the three-dimensional coordinates corresponding to the target action" can be performed. If it is not greater than the amplitude threshold value, it means that the human body may have moved slightly. At this time, the three-dimensional image in the virtual three-dimensional space does not need to be operated. The step can be performed again when the next action instruction is received, thereby monitoring the action of the human body.
[0090] In order to better implement the skeleton posture recognition method based on virtual-real interaction in the embodiments of the present application, on the basis of the skeleton posture recognition method based on virtual-real interaction, correspondingly, the embodiments of the present application also provide a skeleton posture recognition device based on virtual-real interaction, as shown in Figure 2 The skeleton posture recognition device based on virtual-real interaction comprises:
[0091] The image acquisition module 201 is used to acquire a preset number of human images based on a multi-lens camera, detect the preset number of human images, and obtain all two-dimensional skeletons and a preset number of joints of the human body.
[0092] The coordinate determination module 202 is used to generate a three-dimensional image based on a preset number of human images, and to process the three-dimensional image based on all two-dimensional bones and a preset number of joints to obtain the virtual three-dimensional coordinates corresponding to each two-dimensional bone.
[0093] The motion execution module 203 is used to determine the target two-dimensional skeleton according to the motion command when a motion command is received; and to switch the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton to the three-dimensional coordinates corresponding to the target motion according to the motion command.
[0094] The skeletal posture recognition device based on virtual-real interaction provided in the above embodiments can realize the technical solutions described in the above embodiments of the skeletal posture recognition method based on virtual-real interaction. The specific implementation principles of each module or unit can be found in the corresponding content in the above embodiments of the skeletal posture recognition method based on virtual-real interaction, and will not be repeated here.
[0095] like Figure 3 As shown, the present invention also provides an electronic device 300. The electronic device 300 includes a processor 301, a memory 302, and a display 303. Figure 3 Only some components of the electronic device 300 are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0096] In some embodiments, memory 302 may be an internal storage unit of electronic device 300, such as a hard disk or memory of electronic device 300. In other embodiments, memory 302 may also be an external storage device of electronic device 300, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on electronic device 300.
[0097] Furthermore, the memory 302 may include both internal storage units of the electronic device 300 and external storage devices. The memory 302 is used to store application software and various types of data installed on the electronic device 300.
[0098] The processor 301 may, in some embodiments, be a central processing unit (CPU), a microprocessor, or other data processing chip, for running program codes stored in the memory 302 or processing data, such as the method for recognizing bone posture based on virtual-real interaction in the present application.
[0099] The display 303 may, in some embodiments, be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, or the like. The display 303 is used to display information of the electronic device 300 and to display a visualized user interface. The components 301-303 of the electronic device 300 communicate with each other through a system bus.
[0100] In some embodiments of the present application, when the processor 301 executes the program for recognizing bone posture based on virtual-real interaction in the memory 302, the following steps can be implemented:
[0101] A preset number of human body images of a human body are acquired based on a multi-lens camera, and all two-dimensional bones and a preset number of joint nodes of the human body are obtained by detecting the preset number of human body images;
[0102] A three-dimensional image is generated according to the preset number of human body images, and the three-dimensional image is processed according to all two-dimensional bones and the preset number of joint nodes, to obtain virtual three-dimensional coordinates corresponding to each two-dimensional bone;
[0103] When an action instruction is received, a target two-dimensional bone is determined according to the action instruction, and the virtual three-dimensional coordinates corresponding to the target two-dimensional bone are switched to three-dimensional coordinates corresponding to a target action according to an action in the action instruction.
[0104] It should be understood that, in addition to the above functions, the processor 301 may, when executing the program for recognizing bone posture based on virtual-real interaction in the memory 302, also implement other functions, which can be referred to the description of the corresponding method embodiments above.
[0105] Further, the type of the electronic device 300 is not limited, and the electronic device 300 can be a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop, or the like. Exemplary embodiments of the portable electronic device include, but are not limited to, a portable electronic device running an IOS, an android, a microsoft, or other operating system. The portable electronic device can also be other portable electronic devices, such as a laptop having a touch-sensitive surface (e.g., a touch panel), and the like. It should also be understood that in some other embodiments of the present application, the electronic device 300 can not be a portable electronic device, but a desktop computer having a touch-sensitive surface (e.g., a touch panel).
[0106] Accordingly, the embodiments of the present application also provide a computer readable storage medium for storing computer readable programs or instructions, which, when executed by a processor, can implement the method steps or functions of the virtual-real interaction based skeleton posture recognition method provided by the above embodiments.
[0107] Those skilled in the art can understand that all or part of the processes of the above-mentioned embodiments can be completed by a computer program instructing related hardware (such as a processor, a controller, etc.) to complete, and the computer program can be stored in a computer readable storage medium. The computer readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.
[0108] The above describes the virtual-real interaction based skeleton posture recognition method and device provided by the present application in detail, and the principle and implementation mode of the present application are described by applying specific examples. The above embodiment is only used to help understand the method and core idea of the present application; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range can be changed, and the above description should not be understood as limiting the present application.
Claims
1. A skeletal pose recognition method based on virtual-real interaction, characterized in that, include: A preset number of human images are acquired using a multi-lens camera, and the preset number of human images are detected to obtain all two-dimensional skeletons and a preset number of joints of the human body. Based on the preset number of human body images, a three-dimensional image is generated, and the three-dimensional image is processed according to all the two-dimensional skeletons and the preset number of joints to obtain the virtual three-dimensional coordinates corresponding to each two-dimensional skeleton; When an action command is received, the target two-dimensional skeleton is determined according to the action command; and according to the action in the action command, the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton are switched to the three-dimensional coordinates corresponding to the target action. After processing the three-dimensional image based on all the two-dimensional skeletons and the preset number of joints to obtain the virtual three-dimensional coordinates corresponding to each two-dimensional skeleton, the process further includes: Each two-dimensional skeleton is detected to obtain the corresponding identifier for each two-dimensional skeleton; The human body's skeleton is then detected by binding the identifier to the corresponding skeleton. When an action command is received, determining the target two-dimensional skeleton based on the action command includes: By detecting the human body, the system receives motion commands from the human body; the motion commands include the target skeleton to perform the motion. Based on the target skeleton of the action command, determine the target identifier that is bound to the target skeleton; Based on the target identifier, a target two-dimensional skeleton in the virtual space is determined; the step of processing the three-dimensional image based on all the two-dimensional skeletons and the preset number of joints to obtain the virtual three-dimensional coordinates corresponding to each two-dimensional skeleton includes: Based on the relative information between the multi-lens camera and the three-dimensional stereo image, the depth and position data of the three-dimensional stereo image are obtained; The human body contour of the three-dimensional image is detected and restored based on all the two-dimensional bones and the preset number of joints to obtain the height of each two-dimensional bone on the three-dimensional image. Based on the height, depth, and position data of each two-dimensional skeleton, generate virtual three-dimensional coordinates corresponding to each two-dimensional skeleton; The step of switching the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton to the three-dimensional coordinates corresponding to the target action according to the action in the action command includes: Based on the action command, the target skeleton is determined to perform the target action; Perform corresponding operations on the target 2D skeleton according to the target action, and switch the virtual 3D coordinates corresponding to the target 2D skeleton to the 3D coordinates corresponding to the target action.
2. The skeletal pose recognition method based on virtual-real interaction according to claim 1, characterized in that, The step of detecting the preset number of human images to obtain all two-dimensional skeletons and a preset number of joints of the human body includes: The preset number of human images are detected by a preset human skeleton detection neural network to obtain at least one two-dimensional skeleton corresponding to each human image. The preset number of human body images are compared, and duplicate two-dimensional skeletons are removed from the preset number of human body images to obtain all two-dimensional skeletons of the human body; Based on the joints of each two-dimensional bone, a preset number of joints for all two-dimensional bones are obtained.
3. The skeletal pose recognition method based on virtual-real interaction according to claim 2, characterized in that, The step of obtaining a preset number of joints for all two-dimensional bones based on the joints of each two-dimensional bone includes: Based on the joints of all the two-dimensional bones, obtain the joint points corresponding to each two-dimensional bone; All joints of all the two-dimensional skeletons are filtered to obtain a preset number of joints.
4. The skeletal pose recognition method based on virtual-real interaction according to claim 1, characterized in that, The step of generating a three-dimensional image based on the preset number of human body images includes: The preset number of human body images are stitched together to obtain a three-dimensional image.
5. The skeletal pose recognition method based on virtual-real interaction according to claim 1, characterized in that, After determining that the target skeleton performs a target action according to the action command, the process further includes: The target action is optimized to obtain the optimized target action; Determine whether the amplitude of the optimized target action is greater than the amplitude threshold; If not, wait for the next action command.
6. A skeletal pose recognition device based on virtual-real interaction, characterized in that, include: The image acquisition module is used to acquire a preset number of human body images based on a multi-lens camera, detect the preset number of human body images, and obtain all two-dimensional skeletons and a preset number of joints of the human body. The coordinate determination module is used to generate a three-dimensional image based on the preset number of human images, and to process the three-dimensional image based on all the two-dimensional skeletons and the preset number of joints to obtain the virtual three-dimensional coordinates corresponding to each two-dimensional skeleton. The motion execution module is used to determine the target two-dimensional skeleton according to the motion command when a motion command is received; and to switch the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton to the three-dimensional coordinates corresponding to the target motion according to the motion command. After processing the three-dimensional image based on all the two-dimensional skeletons and the preset number of joints to obtain the virtual three-dimensional coordinates corresponding to each two-dimensional skeleton, the process further includes: Each two-dimensional skeleton is detected to obtain the corresponding identifier for each two-dimensional skeleton; The human body's skeleton is then detected by binding the identifier to the corresponding skeleton. When an action command is received, determining the target two-dimensional skeleton based on the action command includes: By detecting the human body, the system receives motion commands from the human body; the motion commands include the target skeleton to perform the motion. Based on the target skeleton of the action command, determine the target identifier that is bound to the target skeleton; Based on the target identifier, a target two-dimensional skeleton in the virtual space is determined; the step of processing the three-dimensional image based on all the two-dimensional skeletons and the preset number of joints to obtain the virtual three-dimensional coordinates corresponding to each two-dimensional skeleton includes: Based on the relative information between the multi-lens camera and the three-dimensional stereo image, the depth and position data of the three-dimensional stereo image are obtained; The human body contour of the three-dimensional image is detected and restored based on all the two-dimensional bones and the preset number of joints to obtain the height of each two-dimensional bone on the three-dimensional image. Based on the height, depth, and position data of each two-dimensional skeleton, generate virtual three-dimensional coordinates corresponding to each two-dimensional skeleton; The step of switching the virtual three-dimensional coordinates corresponding to the target two-dimensional skeleton to the three-dimensional coordinates corresponding to the target action according to the action in the action command includes: Based on the action command, the target skeleton is determined to perform the target action; Perform corresponding operations on the target 2D skeleton according to the target action, and switch the virtual 3D coordinates corresponding to the target 2D skeleton to the 3D coordinates corresponding to the target action.
Citation Information
Patent Citations
Multi-person real-time three-dimensional motion recognition and evaluation method, device, equipment and medium
CN114694257A
Human body posture synchronous animation display method and device and automobile
CN115619914A