Pose recognition method and device, storage medium and electronic device
By simultaneously capturing depth information through electronic and external devices to perform 3D modeling and identify object pose, the high cost of infrared cameras is solved, enabling flexible and economical pose recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2022-11-17
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies that use infrared cameras and infrared reflective dots to identify object poses result in high costs for electronic devices, making them unsuitable for widespread application.
The system uses the built-in camera of an electronic device to capture images simultaneously with an external device. By analyzing the first and second images, depth information is obtained, and 3D modeling is performed to identify the pose of the object, thus avoiding the use of infrared cameras and infrared reflectors.
It reduces the cost of equipment for recognizing object poses, improves flexibility and applicability, can be applied in a variety of scenarios, and accurately identifies the spatial position and posture of objects through 3D modeling.
Smart Images

Figure CN115719376B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to a pose recognition method, apparatus, storage medium, and electronic device. Background Technology
[0002] An object's pose reflects its spatial position and orientation, and it is necessary to identify the object's pose in many aspects such as spatial positioning and 3D modeling.
[0003] Generally, by installing an infrared camera on an electronic device to emit infrared light onto an object, and setting infrared reflective points on the object to reflect the infrared light, the infrared camera determines the pose of the object by capturing the reflected infrared light.
[0004] However, installing infrared cameras and infrared reflectors would make electronic devices expensive, which would not help save on production costs. Summary of the Invention
[0005] This application provides a pose recognition method, apparatus, storage medium, and electronic device, which can save the cost required for recognizing object pose.
[0006] In a first aspect, embodiments of this application provide a pose recognition method applied to an electronic device, the method comprising:
[0007] Acquire a first image and a second image of the object to be identified simultaneously captured by an electronic device and an external device, with the first image and the second image being captured from different angles;
[0008] Based on the first and second images, obtain the depth information of the object to be identified;
[0009] Based on the depth information, a 3D model of the object to be identified is obtained;
[0010] The current pose of the object to be identified is determined based on the 3D model.
[0011] Secondly, embodiments of this application also provide a pose recognition device, comprising:
[0012] The image acquisition module is used to acquire a first image and a second image of the object to be identified simultaneously captured by an electronic device and an external device. The first image and the second image are captured from different angles.
[0013] The depth calculation module is used to obtain the depth information of the object to be identified based on the first image and the second image;
[0014] The 3D modeling module is used to perform 3D modeling of the object to be identified based on depth information, thereby obtaining a 3D model of the object to be identified.
[0015] The pose recognition module is used to determine the current pose of the object to be recognized based on the 3D model.
[0016] Thirdly, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when run on a computer, causes the computer to execute the pose recognition method provided in any embodiment of this application.
[0017] Fourthly, embodiments of this application also provide an electronic device, including a processor and a memory, the memory having a computer program, and the processor executing the pose recognition method as provided in any embodiment of this application by calling the computer program.
[0018] The technical solution provided in this application involves simultaneously capturing images of the object to be identified using both a built-in camera on an electronic device and a camera on an external device. The external device then transmits the captured second image to the electronic device for image processing. The electronic device analyzes the first image it captured and the second image captured by the external device to obtain depth information about the object to be identified. This depth information describes the relative position between the object and the electronic device. Subsequently, a 3D model is created based on the 2D image and depth information of the object to be identified, resulting in a 3D model of the object. The pose of the object can then be identified based on this 3D model. Compared to existing technologies, this method does not require an infrared camera or infrared reflective points on the object, thus reducing the equipment cost for identifying the object's pose. Furthermore, this method is flexible and easy to use, applicable to various situations for identifying and tracking the pose of objects. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of a pose recognition system provided in an embodiment of this application.
[0021] Figure 2 This is a flowchart illustrating the pose recognition method provided in an embodiment of this application.
[0022] Figure 3 This is a schematic diagram illustrating the method for determining the depth information of an object to be identified based on a first image and a second image in the pose recognition method provided in this application embodiment.
[0023] Figure 4This is a schematic diagram illustrating the correction of the initial image pose based on the current camera pose in the pose recognition method provided in this application embodiment.
[0024] Figure 5 This is a schematic diagram of the pose recognition device provided in the embodiments of this application.
[0025] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.
[0027] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0028] In related technologies, recognizing the pose of an object requires electronic devices to have relevant components to support pose recognition. However, this approach can result in expensive electronic devices or large electronic devices that are not easily used flexibly in various scenarios.
[0029] To address the technical problems involved in related technologies, embodiments of this application provide a pose recognition method, apparatus, storage medium, and electronic device. The subject executing the pose recognition method can be the pose recognition apparatus provided in this application embodiment, or an electronic device integrating the pose recognition apparatus, wherein the pose recognition apparatus can be implemented in hardware or software. The electronic device can be a smartphone, foldable phone, tablet computer, PDA, laptop computer, camera, camcorder, vehicle terminal, or other device with shooting capabilities.
[0030] Before introducing the pose recognition method provided in the embodiments of this application, a system is first provided here. For details, please refer to [link to relevant documentation]. Figure 1 , Figure 1This is a schematic diagram of a pose recognition system provided in an embodiment of this application. The pose recognition system includes two devices with imaging capabilities, one referred to as an electronic device and the other as an external device. The terms "electronic device" and "external device" are relative. For example, taking a mobile terminal and an in-vehicle terminal as examples, when the mobile terminal is used as the electronic device, the in-vehicle terminal is used as the external device, and vice versa. The following embodiments use a mobile terminal as the electronic device and an in-vehicle terminal as the external device to describe the pose recognition method provided in this application.
[0031] The mobile terminal and the vehicle-mounted terminal are communicatively connected and can transmit data. When implementing the pose recognition method provided in this application, the mobile terminal and the vehicle-mounted terminal can be controlled to simultaneously capture images of the object to be recognized, resulting in a first image captured by the mobile terminal and a second image captured by the vehicle-mounted terminal. Then, either the mobile terminal or the vehicle-mounted terminal is selected to process the first and second images to obtain the pose of the object to be recognized.
[0032] Specifically, this embodiment of the application also includes a bracket on the vehicle equipped with the aforementioned vehicle-mounted terminal. The bracket and the vehicle-mounted terminal are at a certain angle, which can be preset as a fixed angle. When the mobile terminal is mounted on the bracket, the mobile terminal and the vehicle-mounted terminal form a fixed shooting angle for synchronous shooting. Of course, the angle between the bracket and the vehicle-mounted terminal can be adjusted as needed. For example, given the initial angle of the bracket, the changed angle can be obtained from the changed angle and the initial angle during the adjustment process. Specifically, multiple angles can be preset for the bracket, allowing the user to selectively switch the bracket's angle during actual shooting. Understandably, when the angle between the bracket and the vehicle-mounted terminal is changeable, the shooting angle between the mobile terminal and the vehicle-mounted terminal changes accordingly, thereby selectively shooting objects inside the vehicle, improving both the shooting angle and the shooting flexibility.
[0033] The pose recognition method provided in this application is further described in the following embodiments. Please refer to... Figure 2 , Figure 2 This is a flowchart illustrating the pose recognition method provided in an embodiment of this application. The specific flow of the pose recognition method provided in this embodiment of the application can be as follows:
[0034] 101. Acquire a first image and a second image of the object to be identified simultaneously captured by an electronic device and an external device, wherein the first image and the second image are captured from different angles.
[0035] This can be achieved by setting electronic devices and external devices to capture images of the object to be identified at the same time. The image captured by the electronic device is called the first image, and the image captured by the external device at the same time is called the second image.
[0036] For example, there are several ways to set up synchronized shooting between an electronic device and an external device. One method is to have either the electronic device or the external device send a synchronized shooting signal so that both have the same frame rate and start time. Another method is to add timestamps to the first and second images captured by the electronic device and the external device respectively, thus enabling synchronized shooting. Since there are many ways to control synchronized shooting, they will not be listed here.
[0037] In this scenario, the electronic device and the external device are positioned at a shooting angle, and during the shooting process, the electronic device and the external device capture images of the object to be identified from different shooting angles.
[0038] In this embodiment, the object to be identified can be any person or object in the current shooting scene. It can also be a specific person or object, such as an electronic device or external device that can track and capture images of a person.
[0039] Taking the aforementioned in-vehicle terminal as an example, when taking photos via the in-vehicle terminal and mobile terminal, the shooting scenario is the content inside the vehicle. The content inside the vehicle can be referred to as the object to be identified, including the driver and passengers. The choice can be made based on actual needs and is not limited here.
[0040] 102. Based on the first image and the second image, obtain the depth information of the object to be identified.
[0041] The method of obtaining the depth information of the object to be identified based on the first image and the second image can be, for example, sending the first image and the second image to a server for processing to obtain the depth information of the object to be identified, or having an electronic device perform local processing to obtain the depth information of the object to be identified.
[0042] When photographing the object to be identified, there is a corresponding shooting angle between the electronic device and the external device, and this shooting angle is known. Given the known shooting angle, the depth information of the object to be identified can be obtained from the intersection point of the electronic device and the external device positioned on the object to be identified, and from the distance from the intersection point to the electronic device. The depth information of the object to be identified describes the distance between the object to be identified and the electronic device.
[0043] For example, the intersection of the electronic device and the external device on the object to be identified can be determined by analyzing the first image and the second image. Both the first image and the second image contain the object to be identified. By analyzing the pixel coordinates of the object to be identified on the first image and the second image, the intersection of the electronic device and the external device on the object to be identified can be determined, thereby obtaining the depth information of the object to be identified.
[0044] 103. Based on the depth information, perform 3D modeling of the object to be identified to obtain a 3D model of the object to be identified.
[0045] The first and second images can represent the position information of the object to be identified in two-dimensional space. By adding the depth information of the object to be identified and analyzing it, the position information of the object to be identified in three-dimensional space can be obtained. Then, a three-dimensional model of the object to be identified can be obtained through three-dimensional modeling.
[0046] For example, the position coordinates of the object to be identified in two-dimensional space can be obtained from the first image and the second image, and then the depth information of the object to be identified can be used as the position coordinates of the z-axis. Thus, based on the transformation relationship between the camera coordinate system and the world coordinate system, the position coordinates of the object to be identified in three-dimensional space can be obtained.
[0047] Specifically, a three-dimensional model of the object to be identified can be constructed based on its position coordinates in three-dimensional space. This three-dimensional model describes the positional distribution of the pixels that constitute the object in three-dimensional space.
[0048] 104. Determine the current pose of the object to be identified based on the 3D model.
[0049] After constructing a 3D model of the object to be identified, the current pose of the object can be obtained from the 3D model. The 3D model can reflect the spatial position, orientation, and posture of the object to be identified.
[0050] In practice, this application is not limited by the execution order of the described steps. Without causing conflicts, some steps may be performed in other orders or simultaneously.
[0051] As can be seen from the above, the pose recognition method provided in this application embodiment simultaneously captures images of the object to be recognized using an electronic device and an external device positioned at a shooting angle to obtain a first image and a second image of the object. Then, based on the first image, the second image, and the shooting angle, the position coordinates of the object in three-dimensional space are analyzed. A three-dimensional model of the object is then constructed based on these position coordinates to display its current pose. This method can identify the pose of the object economically and can be flexibly applied to various scenarios, such as recognizing the pose of objects inside a vehicle in a driving scenario. Furthermore, the pose recognition method provided in this application embodiment can also flexibly replace the electronic device to form a pose recognition system with external devices for recognizing the pose of the object, making its usage more flexible and its application more widespread.
[0052] In some embodiments, obtaining depth information of the object to be identified based on the first image and the second image includes:
[0053] Determine the first coordinates of the feature points of the object to be identified in the first image, and the second coordinates in the second image;
[0054] Based on the first coordinate, the second coordinate, and the shooting angle between the electronic device and the external device, the pixel depth of the feature points is determined by a triangulation algorithm.
[0055] Depth information of the object to be identified is determined based on the pixel depth of feature points.
[0056] Please see Figure 3 , Figure 3 This diagram illustrates the method for determining the depth information of an object to be identified based on a first image and a second image, as provided in the embodiments of this application. The left side of the diagram shows an electronic device and a first image captured by the electronic device, denoted by O1. The right side shows an external device and a second image captured by the external device, denoted by O2. The object to be identified has multiple feature points. Taking one feature point as an example, its first coordinate in the first image is denoted by P, and its second coordinate in the second image is denoted by P', where both the first and second coordinates indicate pixel positions. The shooting angle between the electronic device and the external device is known, denoted by R. Based on the first coordinate, the second coordinate, and the shooting angle, the pixel depth of the feature point is determined using a triangulation algorithm. The pixel depth of the feature point describes its spatial position, denoted by T.
[0057] The smallest unit of the feature points of the object to be identified can be a pixel. By analyzing each feature point of the object to be identified, the pixel depth of each feature point can be obtained, and the pixel depth of all feature points constitutes the depth information of the object to be identified.
[0058] In some embodiments, determining the current object pose of the object to be identified based on the 3D model includes:
[0059] Determine the initial image pose of the object to be identified based on the 3D model;
[0060] Obtain the historical object pose of the object to be identified;
[0061] The initial image pose is corrected based on the historical object pose to obtain the current object pose.
[0062] In this embodiment, the pose obtained from the 3D model is determined as the initial image pose. Then, the initial image pose is corrected by using the historical object pose to correct the initial image pose and obtain the current object pose.
[0063] For example, by continuously capturing images of the object to be identified using electronic devices and external devices, a sequence of first images and a sequence of second images can be obtained. A 3D model of the object to be identified at that capture moment can be constructed based on the first and second images obtained at the same capture time, and then the object's pose at that capture moment can be obtained based on the 3D model. The object poses obtained before the current capture moment are referred to as historical object poses.
[0064] In this embodiment, the object pose at the previous shooting moment can be obtained as the historical object pose, or two or more object poses within a preset time period before the current shooting moment can be obtained as the historical object pose.
[0065] There are several ways to correct the initial image pose based on the historical object pose. For example, one could infer the change in position and posture of the object to be identified at the current shooting moment based on the historical object pose, thus correcting the initial image pose. Another example is to correct the connection relationships between key points in the initial image pose based on the connection relationships between key points reflected in the historical object pose.
[0066] This embodiment corrects the current image pose by modifying it based on the historical object pose, thereby correcting the initial image pose and making the current object pose obtained after the correction process more accurate.
[0067] In some embodiments, the initial image pose is corrected based on the historical object pose to obtain the current object pose, including:
[0068] The initial image pose is corrected based on the historical object poses to obtain the current image pose;
[0069] Obtain the current camera pose of the electronic device;
[0070] The current image pose is corrected based on the current camera pose to obtain the current object pose.
[0071] In this embodiment, the pose after correcting the initial image pose based on the historical object pose is called the current image pose. Then, the current image pose is further corrected by combining the current camera pose of the electronic device.
[0072] Specifically, after obtaining the historical object pose of the object to be identified using the aforementioned method, the initial image pose is corrected based on the historical object pose to obtain the current image pose. Then, the current image pose is corrected by combining the current camera pose of the electronic device. Correcting the current image pose using the current camera pose reduces errors caused by tilting the camera when capturing images of the object to be identified.
[0073] For example, when an electronic device is tilted to take a picture of a driver, there is a shooting deviation angle between the electronic device and the driver, which makes the driver image displayed in the first image not an orthographic projection image. This will affect the accuracy of the 3D model when constructing a 3D model of the object to be identified based on the first image.
[0074] For example, when correcting the current image pose based on the current camera pose, the current image pose can be corrected based on the shooting deviation angle of the current camera pose relative to the plane where the object to be identified is located. For details, please refer to... Figure 4 , Figure 4 This diagram illustrates the correction of the initial image pose based on the current camera pose in the pose recognition method provided in this application. The shooting deviation angle between the electronic device and the object to be recognized is denoted as A. After obtaining the shooting deviation angle A based on the current camera pose, the orthographic projection of the object to be recognized can be obtained. The initial image pose is then corrected based on the orthographic projection. The correction method involves correcting the initial image pose based on the orthographic projection, ensuring that the projection of the current object coincides with the orthographic projection, thereby completing the correction process for the initial image pose.
[0075] This application embodiment corrects the current image pose by using the current camera pose of an electronic device, which enables affine transformation processing of the 3D model to correct the pose deviation of the object to be identified as represented by the 3D model, thereby obtaining a more accurate current object pose in terms of scale and projection viewpoint.
[0076] As mentioned above, when correcting the initial image pose using historical object poses, the PNP (Perspective-n-Point) method can also be used. The PNP method is a method for solving the motion of 3D to 2D point pairs, aiming to solve the pose of the camera coordinate system relative to the world coordinate system. For example, after obtaining the historical object pose and the initial image pose, the predicted camera pose of the electronic device can be estimated based on the two. Then, the preset camera pose and the actual current camera pose of the electronic device can be compared to correct the initial image pose by reducing the difference between the two.
[0077] Understandably, in the embodiments of this application, the initial image pose can be corrected based on at least one of the historical object pose and the current camera pose of the electronic device to obtain the current object pose. The specific number of the two and the order of correction when both are used are not limited here. Any method that can correct the initial image pose can be used in this embodiment and is within the protection scope claimed by this application.
[0078] In some embodiments, the initial image pose is corrected based on the historical object pose to obtain the current object pose, including:
[0079] Obtain the current and historical camera poses of the electronic device;
[0080] Estimate the predicted image pose of the object to be identified based on the current camera pose, historical camera poses, and historical object poses;
[0081] The predicted image pose and the initial image pose are fused to obtain the current object pose of the object to be identified.
[0082] In this embodiment, the predicted image pose of the object to be identified can also be estimated based on the current camera pose, historical camera pose, and historical object pose, and then the predicted image pose and the initial image pose can be fused to obtain the current object pose.
[0083] Among them, there is a pair of pose data at the same shooting time. The pair of pose data includes the historical camera pose and the historical object pose at the shooting time. When multiple pairs of historical camera poses and historical object poses are known, the predicted image pose of the object to be identified can be estimated based on the current camera pose.
[0084] The method of fusing the predicted image pose and the initial image pose can be, for example, taking the median value of the positional deviation between the same key point in the two, and then obtaining the current object pose based on the median value of multiple key points.
[0085] In this embodiment, the correlation between the historical camera pose and the historical object pose of the electronic device is combined to estimate the predicted image pose of the object to be identified at the current shooting time. The current object pose is then obtained based on the fusion data of the predicted image pose and the initial image pose. This method can reduce the amount of data when correcting the initial image pose, thereby quickly and accurately correcting the initial image pose.
[0086] In some embodiments, obtaining the current camera pose of the electronic device includes:
[0087] Acquire gyroscope data from electronic devices;
[0088] The current camera pose of the electronic device is calculated based on the gyroscope data.
[0089] The electronic device is also equipped with a gyroscope or inertial measurement unit, also known as an IMU. By detecting the gyroscope data of the electronic device, the current pose of the electronic device can be analyzed, and the current pose of the electronic device also reflects the current pose of the camera on it.
[0090] The gyroscope data includes the angular velocity and acceleration data of the three axes of the electronic device.
[0091] Specifically, the current camera pose of the electronic device can be calculated by integrating the gyroscope data of the electronic device.
[0092] Understandably, when an electronic device does not have a gyroscope or inertial measurement unit installed in it, such components can be installed on it as an external device to detect the gyroscope data of the electronic device.
[0093] In some embodiments, after determining the current object pose of the object to be identified based on the 3D model, the method further includes:
[0094] The current object pose is sent to the display device to drive the virtual image of the object to be identified in the virtual scene loaded on the display device.
[0095] In this embodiment, after obtaining the current object pose of the object to be identified, a virtual image can be driven based on the current object pose. The virtual image can be displayed on an electronic device, on the aforementioned external device, or on another external device, depending on actual needs.
[0096] The device that displays the virtual image is called a display device. The display device can load a virtual scene and display the virtual image in the virtual scene. The virtual image can be constructed based on a three-dimensional model.
[0097] In this embodiment, the virtual image is driven by the current object's pose, which enables remote interaction with the virtual image. Of course, the display device can also synchronize the interactive scene to other external devices for display, thereby realizing multi-terminal linkage.
[0098] As one implementation method, the pose of the virtual avatar can be synchronously changed based on the current object pose.
[0099] Specifically, based on the key points of the virtual image, corresponding key points can be matched from the current object's pose, and then the position of the corresponding key points of the virtual image can be adjusted according to the key point position indicated by the current object's pose, so that the pose of the virtual image and the pose of the object to be identified are synchronized.
[0100] These key points can include skeletal key points or limb key points, and the specific choice can be made according to actual needs.
[0101] As another implementation, the facial expression of the virtual avatar can be adjusted according to the current pose of the object, so as to render the face of the virtual avatar based on facial feature points.
[0102] Specifically, when the virtual avatar is a human figure, the key points of the virtual avatar's facial features can be determined, and these key points can be matched with the current object's pose. The virtual avatar's facial features are then adjusted based on the key point positions corresponding to the current object's pose. Alternatively, the facial orientation of the object to be identified can be determined based on the current object's pose, and the facial orientation of the virtual avatar can be adjusted synchronously. Then, the face of the virtual avatar's model is rendered using the facial information of the object's 3D model.
[0103] Understandably, in addition to the above-mentioned methods of driving virtual images, virtual objects in a virtual scene can also be generated based on the three-dimensional model of the object to be identified, and then sent to the display device and presented in the corresponding spatial position of the virtual scene.
[0104] To better understand the solution for driving virtual avatars provided in the embodiments of this application, two application scenarios are provided for explanation.
[0105] One application scenario is VR gaming. Users wear VR glasses within a game vehicle equipped with an onboard terminal. The user places their phone on a bracket within the vehicle. After selecting a game mode via their phone, both the phone and the onboard terminal capture images of the user, creating a 3D model. The phone then sends this 3D model to the VR glasses, displaying it within the corresponding game scene—for example, showing the user's image inside the vehicle. Alternatively, a virtual avatar already exists in the game scene. The phone analyzes the user's current pose and sends it to the VR glasses, which then synchronize the avatar's pose accordingly. Furthermore, the virtual avatar's face can be replaced with the user's 3D face, with the replaced face's pose synchronized with the user's current avatar pose. This approach enhances game immersion and provides a superior gaming experience.
[0106] Another application scenario is live streaming. For example, users can live stream from their vehicles and capture images of themselves using the vehicle's in-vehicle terminal and a mobile phone used for live streaming. The user's current pose can then be obtained through 3D modeling and pose correction. This pose can then drive the virtual avatar in the live stream on the mobile phone to make synchronous pose changes, such as changes in limbs, head orientation, and facial expressions. This allows for flexible live streaming and enables the virtual avatar's pose changes to be driven quickly and accurately.
[0107] In some embodiments, the external device includes an in-vehicle terminal, the object to be identified includes the driver of the vehicle where the in-vehicle terminal is located, and after determining the current object pose of the object to be identified based on the 3D model, the device further includes:
[0108] If the vehicle is in motion, the driver's driving behavior is identified based on the current object's pose, and the identification result is obtained.
[0109] Among these methods, identifying whether a vehicle is in motion can involve detecting whether the vehicle's drive mechanism is activated, detecting whether the onboard terminal is turned on and performing journey calculations, or detecting whether the vehicle is moving using sensors. Since there are multiple identification methods, they will not be listed here.
[0110] For example, during driving behavior recognition, the driver's posture is compared with the current object's posture to determine differences. The driver's posture is determined based on the driver's body type and the vehicle's driver's seat type, or it may be a preset standard driving posture. This driver's posture can verify whether the driver's driving behavior is abnormal.
[0111] Specifically, the driver's pose can be compared with the current object's pose to determine the deviation between the two. During pose comparison, an overall comparison can be performed to determine whether the fit between the two poses is small. If the fit is large, it indicates abnormal driver behavior; if it is small, it indicates normal driver behavior. The standard for evaluating the fit can be determined by pre-setting a fit deviation.
[0112] Of course, local comparisons can also be performed to determine the fit between the two postures. For example, comparisons can be made separately for the head, shoulders and neck, and back. Only when the fit of all three is low is the driver's driving behavior considered normal; otherwise, it is considered abnormal. Different weights can also be assigned to different body parts, such as a higher weight for the head and a lower weight for the back, thereby enabling more accurate identification of the user's driving behavior.
[0113] In some embodiments, after recognizing the driver's driving behavior based on the current object pose and obtaining the recognition result, the method further includes:
[0114] If the recognition result describes abnormal driving behavior of the driver, then the driver will be prompted to correct their posture based on the difference between the current object's posture and the driver's posture.
[0115] In this embodiment, when the driver's driving behavior is determined to be abnormal, a posture correction prompt is also given based on the difference information between the current object posture and the driving posture, thereby reminding the driver to adjust the driving posture in time to avoid traffic accidents.
[0116] Specifically, the difference between the two can indicate the driver's posture deviation and positional deviation relative to the driving posture, and then provide posture correction prompts to the driver based on the driving posture. For example, if the driver's head is turned to the right relative to the driving posture, the prompt can remind the driver to move their head to the left. Alternatively, the prompt can be based on the magnitude of the difference between the two postures, indicating the degree of correction. For example, if the driver's current posture is too far to the right, the prompt can remind the driver to significantly correct their head to the left.
[0117] As mentioned above, recognizing the driver's driving behavior based on the driver's current position and posture can not only ensure the driver's driving safety, but also correct the driver's driving posture, and even provide timely reminders when the driver is fatigued.
[0118] As can be seen from the above, the pose recognition method proposed in this embodiment of the invention can continuously identify the pose of the object to be identified, so as to drive the pose of the virtual image to change synchronously according to the real-time pose. It can also predict the user's behavior by tracking the change information of pose in real time, and then provide timely reminders based on the user's abnormal behavior. Moreover, when identifying the pose of the object to be identified, this application improves the flexibility of pose capture by economically combining the shooting functions of electronic devices and external devices. On the other hand, it also combines the pose of the electronic device's gyroscope data for real-time correction. This correction method has low computational load, is fast and accurate, and can achieve good results when applied to various scenarios, making it easy for users to use.
[0119] In one embodiment, a pose recognition device is also provided. Please refer to [link / reference]. Figure 5 , Figure 5 This is a schematic diagram of the structure of a pose recognition device 200 provided in an embodiment of this application. The pose recognition device 200 is applied to an electronic device and includes:
[0120] The image acquisition module 201 is used to acquire a first image and a second image of the object to be identified simultaneously captured by an electronic device and an external device. The first image and the second image are captured from different angles.
[0121] The depth calculation module 202 is used to obtain the depth information of the object to be identified based on the first image and the second image;
[0122] The 3D modeling module 203 is used to perform 3D modeling of the object to be identified based on depth information, so as to obtain a 3D model of the object to be identified.
[0123] The pose recognition module 204 is used to determine the current pose of the object to be recognized based on the 3D model.
[0124] In some embodiments, the pose recognition module 204 is further configured to:
[0125] Determine the initial image pose of the object to be identified based on the 3D model;
[0126] Obtain the historical object pose of the object to be identified;
[0127] The initial image pose is corrected based on the historical object pose to obtain the current object pose.
[0128] In some embodiments, the pose recognition module 204 is further configured to:
[0129] The initial image pose is corrected based on the historical object poses to obtain the current image pose;
[0130] Obtain the current camera pose of the electronic device;
[0131] The current image pose is corrected based on the current camera pose to obtain the current object pose.
[0132] In some embodiments, the pose recognition module 204 is further configured to:
[0133] Acquire gyroscope data from electronic devices;
[0134] The current camera pose of the electronic device is calculated based on the gyroscope data.
[0135] In some embodiments, the pose recognition device 200 further includes a data transmission module, used for:
[0136] The current object pose is sent to the display device to drive the virtual image of the object to be identified in the virtual scene loaded on the display device.
[0137] In some embodiments, the external device includes an in-vehicle terminal, the object to be identified includes the driver of the vehicle where the in-vehicle terminal is located, and the pose recognition module 204 is further configured to:
[0138] If the vehicle is in motion, the driver's driving behavior is identified based on the current object's pose, and the identification result is obtained.
[0139] In some embodiments, the pose recognition module 204 is further configured to:
[0140] If the recognition result describes abnormal driving behavior of the driver, then the driver will be prompted to correct their posture based on the difference between the current object's posture and the driver's posture.
[0141] It should be noted that the pose recognition device 200 provided in this application embodiment belongs to the same concept as the pose recognition method in the above embodiment. The pose recognition device 200 can implement any of the methods provided in the pose recognition method embodiment. For details of its implementation process, please refer to the pose recognition method embodiment, which will not be repeated here.
[0142] As can be seen from the above, the pose recognition device 200 proposed in this application can continuously recognize the pose of the object to be recognized, so as to drive the pose of the virtual image to change synchronously according to the real-time pose. It can also predict the user's behavior by tracking the change information of pose in real time, and then provide timely reminders based on the user's abnormal behavior. Moreover, when recognizing the pose of the object to be recognized, this application improves the flexibility of capturing the pose of the object to be recognized by economically combining the shooting functions of electronic devices and external devices. On the other hand, it also combines the pose of the electronic device's gyroscope data for real-time correction. This correction method has low computational load, is fast and accurate, and can achieve good results when applying pose recognition in various scenarios, making it convenient for users to use.
[0143] This application also provides an electronic device, which can be a smartphone, foldable phone, tablet computer, PDA, laptop computer, camera, camcorder, vehicle terminal, or other device with shooting capabilities. Figure 6 As shown, Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 300 includes a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, and a computer program stored in the memory 302 and executable on the processor. The processor 301 and the memory 302 are electrically connected. Those skilled in the art will understand that the electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0144] The processor 301 is the control center of the electronic device 300. It connects various parts of the electronic device 300 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 302, and calling data stored in the memory 302, it performs various functions of the electronic device 300 and processes data, thereby monitoring the electronic device 300 as a whole.
[0145] In this embodiment, the processor 301 in the electronic device 300 loads the instructions corresponding to the processes of one or more applications into the memory 302 according to the following steps, and the processor 301 runs the applications stored in the memory 302 to realize various functions:
[0146] Acquire a first image and a second image of the object to be identified simultaneously captured by an electronic device and an external device, with the first image and the second image being captured from different angles;
[0147] Based on the first and second images, obtain the depth information of the object to be identified;
[0148] Based on the depth information, a 3D model of the object to be identified is obtained;
[0149] The current pose of the object to be identified is determined based on the 3D model.
[0150] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0151] although Figure 6 As not shown in the diagram, the electronic device 300 may also include a camera, sensor, wireless fidelity module, Bluetooth module, etc., which will not be described in detail here.
[0152] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0153] As can be seen from the above, the electronic device provided in this embodiment can continuously identify the pose of the object to be identified, so as to drive the pose of the virtual image to change synchronously according to the real-time pose. It can also predict the user's behavior by tracking the change information of pose in real time, and then provide timely reminders based on the user's abnormal behavior. Moreover, when identifying the pose of the object to be identified, this application improves the flexibility of capturing the pose of the object by economically combining the shooting functions of the electronic device and the external device. On the other hand, it also combines the pose of the electronic device's gyroscope data for real-time correction. This correction method has low computational load, is fast and accurate, and can achieve good results when applying pose to various scenarios, making it convenient for users.
[0154] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0155] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of computer programs that can be loaded by a processor to execute steps in any of the pose recognition methods provided in embodiments of this application. For example, the computer program can execute the following steps:
[0156] Acquire a first image and a second image of the object to be identified simultaneously captured by an electronic device and an external device, with the first image and the second image being captured from different angles;
[0157] Based on the first and second images, obtain the depth information of the object to be identified;
[0158] Based on the depth information, a 3D model of the object to be identified is obtained;
[0159] The current pose of the object to be identified is determined based on the 3D model.
[0160] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0161] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk, or optical disk, etc. Since the computer program stored in the storage medium can execute the steps of any of the pose recognition methods provided in the embodiments of this application, it can achieve the beneficial effects that any of the pose recognition methods provided in the embodiments of this application can achieve, as detailed in the preceding embodiments, and will not be repeated here.
[0162] The foregoing has provided a detailed description of a pose recognition method, apparatus, medium, and electronic device provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A pose recognition method, applied to electronic devices, characterized in that, include: Acquire a first image and a second image of the object to be identified simultaneously captured by the electronic device and the external device, wherein the first image and the second image are captured from different angles; Based on the first image and the second image, obtain the depth information of the object to be identified; The step of obtaining the depth information of the object to be identified based on the first image and the second image includes: Determine the first coordinates of the feature points of the object to be identified in the first image, and the second coordinates in the second image; Based on the first coordinate, the second coordinate, and the shooting angle between the electronic device and the external device, the pixel depth of the feature points is determined by a triangulation algorithm. Determine the depth information of the object to be identified based on the pixel depth of feature points; Based on the depth information, a 3D model of the object to be identified is performed to obtain a 3D model of the object to be identified. The current object pose of the object to be identified is determined based on the 3D model; Determining the current pose of the object to be identified based on the 3D model includes: The initial image pose of the object to be identified is determined based on the three-dimensional model; Obtain the historical object pose of the object to be identified; The initial image pose is corrected based on the historical object pose to obtain the current object pose; The step of correcting the initial image pose based on the historical object pose to obtain the current object pose includes: The initial image pose is corrected based on the historical object pose to obtain the current image pose; Obtain the current camera pose of the electronic device; The current image pose is corrected based on the current camera pose to obtain the current object pose; The step of obtaining the current camera pose of the electronic device includes: Acquire gyroscope data from the electronic device, the gyroscope data including the angular velocity and acceleration data of the three axes of the electronic device; The current camera pose of the electronic device is calculated based on the gyroscope data, wherein the current camera pose is obtained by integrating the gyroscope data. After determining the current object pose of the object to be identified based on the 3D model, the method further includes: The current object pose is sent to the display device to drive the virtual image of the object to be identified in the virtual scene loaded on the display device.
2. The pose recognition method as described in claim 1, characterized in that, The external device includes an in-vehicle terminal, and the object to be identified includes the driver of the vehicle where the in-vehicle terminal is located. After determining the current object pose of the object to be identified based on the 3D model, the method further includes: If the vehicle is in motion, the driver's driving behavior is identified based on the current object pose, and the identification result is obtained.
3. The pose recognition method as described in claim 2, characterized in that, After identifying the driver's driving behavior based on the current object pose and obtaining the identification result, the method further includes: If the recognition result describes the driver's driving behavior as abnormal, then the driver is given a posture correction prompt based on the difference information between the current object posture and the driving posture.
4. A pose recognition device, characterized in that, include: The image acquisition module is used to acquire a first image and a second image of the object to be identified simultaneously captured by an electronic device and an external device, wherein the first image and the second image are captured from different angles; A depth calculation module is used to obtain the depth information of the object to be identified based on the first image and the second image; The step of obtaining the depth information of the object to be identified based on the first image and the second image includes: Determine the first coordinates of the feature points of the object to be identified in the first image, and the second coordinates in the second image; Based on the first coordinate, the second coordinate, and the shooting angle between the electronic device and the external device, the pixel depth of the feature points is determined by a triangulation algorithm. Determine the depth information of the object to be identified based on the pixel depth of feature points; A 3D modeling module is used to perform 3D modeling on the object to be identified based on the depth information, thereby obtaining a 3D model of the object to be identified. The pose recognition module is used to determine the current pose of the object to be recognized based on the 3D model. Determining the current pose of the object to be identified based on the 3D model includes: The initial image pose of the object to be identified is determined based on the three-dimensional model; Obtain the historical object pose of the object to be identified; The initial image pose is corrected based on the historical object pose to obtain the current object pose; The step of correcting the initial image pose based on the historical object pose to obtain the current object pose includes: The initial image pose is corrected based on the historical object pose to obtain the current image pose; Obtain the current camera pose of the electronic device; The current image pose is corrected based on the current camera pose to obtain the current object pose; The step of obtaining the current camera pose of the electronic device includes: Acquire gyroscope data from the electronic device, the gyroscope data including the angular velocity and acceleration data of the three axes of the electronic device; The current camera pose of the electronic device is calculated based on the gyroscope data, wherein the current camera pose is obtained by integrating the gyroscope data. After determining the current object pose of the object to be identified based on the 3D model, the method further includes: The current object pose is sent to the display device to drive the virtual image of the object to be identified in the virtual scene loaded on the display device.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run on a computer, it causes the computer to perform the pose recognition method as described in any one of claims 1 to 3.
6. An electronic device comprising a processor and a memory, the memory storing a computer program, characterized in that, The processor executes the pose recognition method as described in any one of claims 1 to 3 by invoking the computer program.