Height detection method, device and storage medium

By using electronic devices to capture images and create 3D facial models, the problem of equipment dependence and inconvenient operation in traditional height measurement methods has been solved, achieving highly accurate height measurement without the need for specialized equipment and complete images.

CN115885316BActive Publication Date: 2026-05-19HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2021-07-29
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Traditional height measurement methods require specialized equipment and are inconvenient to operate. Existing technologies rely on devices such as binocular cameras and depth cameras, and require capturing complete human images, which have limitations and insufficient accuracy.

Method used

By acquiring multiple video frames from the image acquisition components of electronic devices and performing semantic plane detection to determine ground information, and then performing face detection, the facial pose is determined using a 3D face model, and the height is calculated by combining the device pose, thus avoiding reliance on professional equipment and the capture of complete human images.

Benefits of technology

It enables height measurement without the need for specialized equipment and complete human images. It is easy to operate and highly accurate, and can determine height through facial recognition and 3D technology, improving measurement accuracy and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115885316B_ABST
    Figure CN115885316B_ABST
Patent Text Reader

Abstract

The application relates to a height detection method and device and a storage medium, wherein the method comprises the following steps: performing semantic plane detection on a plurality of video frames collected by an image collection component of an electronic device to determine ground information in the plurality of video frames; performing face detection on the plurality of video frames to determine a face region; determining a first face pose of a target object in the plurality of video frames according to the face region and a preset three-dimensional face model; and determining a first height of the target object according to the ground information, the first face pose and a device pose of the electronic device. The height detection method and device provided by the application can automatically detect the height of a target object without manual positioning, and are convenient and accurate to use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a height detection method, device and storage medium. Background Technology

[0002] Traditional height measurement methods usually require manual operation of professional instruments, such as height measuring instruments, which is not only inefficient but also inconvenient to carry and unsuitable for personal use.

[0003] With the development of image processing technology, depth images of the human target to be measured can be captured using specialized equipment such as binocular cameras and depth cameras. The height of the human target can then be measured by processing the depth images. However, this method is highly dependent on specialized equipment such as binocular cameras and depth cameras, and it also requires capturing a complete human image (i.e., a full-body image of the human target), which has certain limitations. Summary of the Invention

[0004] In view of this, a height detection method, device and storage medium are proposed.

[0005] In a first aspect, embodiments of this application provide a height detection method, the method comprising: performing semantic plane detection on multiple video frames acquired by an image acquisition component of an electronic device to determine ground information in the multiple video frames; performing face detection on the multiple video frames to determine face regions; determining a first face pose of a target object in the multiple video frames based on the face regions and a preset three-dimensional face model; and determining a first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device.

[0006] The embodiments of this application can perform semantic plane detection on multiple video frames acquired by the image acquisition component of an electronic device to determine ground information in the multiple video frames; simultaneously, face detection is performed on the multiple video frames to determine face regions, and based on the face regions and a preset 3D face model, the first face pose of the target object in the multiple video frames is determined; then, based on the ground information, the first face pose, and the device pose of the electronic device, the first height of the target object is determined. Thus, height detection not only does not rely on professional equipment (such as binocular cameras, depth cameras, etc.), but also can determine the face pose of the target object through face recognition and 3D face technology, and then determine the height of the target object through the face pose, device pose, and ground information. There is no need to manually locate the target object, nor is it necessary to capture a complete human image of the target object. The operation is convenient and highly accurate.

[0007] According to the first aspect, in a first possible implementation of the height detection method, the method further includes at least one of the following: prompting the user to take a picture of the ground if no ground information is detected within a preset time period; prompting the user to adjust the device pose if the pitch angle of the electronic device indicated by the device pose does not meet a first preset condition; prompting the user to adjust the device pose and / or change the face pose of the target object if the first face pose does not meet a second preset condition; or prompting the user to adjust the device pose if the face region does not meet a third preset condition.

[0008] The embodiments of this application can prompt the user when at least one of the following conditions is met: the ground information is not detected within a preset time period; the pitch angle of the electronic device indicated by the device pose does not meet a first preset condition; the first face pose does not meet a second preset condition; or the face region does not meet a third preset condition. For example, the user can be prompted to take a picture of the ground, adjust the device pose, or change the face pose of the target object, so that the user can make corresponding adjustments, thereby improving the accuracy of height detection.

[0009] According to the first aspect or the first possible implementation of the first aspect, in the second possible implementation of the height detection method, determining the first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device includes: determining the second height of the target object based on the ground information, the first face pose, and the device pose; and performing post-processing on the second height to obtain the first height, wherein the post-processing includes Kalman filtering.

[0010] The embodiments of this application can determine the second height of the target object based on ground information, the first face pose and the device pose, and perform post-processing such as Kalman filtering on the second height to obtain the first height of the target object, thereby improving the accuracy of height detection.

[0011] According to the first aspect, or the first possible implementation of the first aspect, or the second possible implementation of the first aspect, in a third possible implementation of the height detection method, the method further includes: displaying the first height on the display interface of the electronic device.

[0012] In the embodiments of this application, after determining the first height of the target object, the first height of the target object can be displayed on the display interface of the electronic device through animation, text, augmented reality (AR) and other means, thereby improving the user experience.

[0013] According to the first aspect, in a fourth possible implementation of the height detection method, the ground information is located in the world coordinate system, the first face pose is located in the camera coordinate system, and determining the first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device includes: adjusting the first face pose according to a preset interpupillary distance reference value to obtain a second face pose; performing coordinate transformation on the second face pose according to the device pose to obtain a third face pose of the target object, the third face pose being located in the world coordinate system; and determining the first height of the target object based on the third face pose and the ground information.

[0014] The embodiments of this application can adjust and transform the first face pose located in the camera coordinate system to obtain the third face pose located in the world coordinate system, and determine the first height of the target object based on the third face pose and ground information, thereby enabling the calculation of the first height of the target object in the world coordinate system and improving the accuracy of height detection.

[0015] According to the fourth possible implementation of the first aspect, in the fifth possible implementation of the height detection method, adjusting the first face pose according to a preset interpupillary distance reference value to obtain the second face pose includes: determining a face size transformation coefficient according to the preset interpupillary distance reference value and the interpupillary distance value in the first face pose; adjusting the first face pose according to the face size transformation coefficient to obtain the second face pose of the target object.

[0016] In the embodiments of this application, the face size transformation coefficient is determined by using the pupil distance reference value and the pupil distance value in the first face pose, and the first face pose is adjusted according to the face size transformation coefficient to obtain the second face pose of the target object, thereby obtaining the actual face size and pose of the target object in the camera coordinate system.

[0017] According to the fourth possible implementation of the first aspect, in the sixth possible implementation of the height detection method, determining the first height of the target object based on the third face pose and the ground information includes: determining the top position of the target object's head based on the third face pose; and determining the first height of the target object based on the top position of the head and the ground information.

[0018] The embodiments of this application determine the top position of the target object's head and determine the target object's first height based on the top position and ground information. This method is simple, fast, and can improve the accuracy of height detection.

[0019] Secondly, embodiments of this application provide a height detection device applied to an electronic device, comprising an image acquisition component for acquiring multiple video frames; and a processing component configured to: perform semantic plane detection on the multiple video frames to determine ground information in the multiple video frames; perform face detection on the multiple video frames to determine face regions; determine a first face pose of a target object in the multiple video frames based on the face image of the face region and a preset three-dimensional face model; and determine a first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device.

[0020] The embodiments of this application can perform semantic plane detection on multiple video frames acquired by the image acquisition component of an electronic device to determine ground information in the multiple video frames; simultaneously, face detection is performed on the multiple video frames to determine face regions, and based on the face regions and a preset 3D face model, the first face pose of the target object in the multiple video frames is determined; then, based on the ground information, the first face pose, and the device pose of the electronic device, the first height of the target object is determined. Thus, height detection not only does not rely on professional equipment (such as binocular cameras, depth cameras, etc.), but also can determine the face pose of the target object through face recognition and 3D face technology, and then determine the height of the target object through the face pose, device pose, and ground information. There is no need to manually locate the target object, nor is it necessary to capture a complete human image of the target object. The operation is convenient and highly accurate.

[0021] According to the second aspect, in a first possible implementation of the height detection device, the processing component is further configured to: prompt the user to photograph the ground if no ground information is detected within a preset time period; prompt the user to adjust the device pose if the pitch angle of the electronic device indicated by the device pose does not meet a first preset condition; prompt the user to adjust the device pose and / or change the face pose of the target object if the first face pose does not meet a second preset condition; or prompt the user to adjust the device pose if the face region does not meet a third preset condition.

[0022] The embodiments of this application can prompt the user when at least one of the following situations occurs: the ground information is not detected within a preset time period; the pitch angle of the electronic device indicated by the device pose does not meet a first preset condition; the first face pose does not meet a second preset condition; or the face region does not meet a third preset condition. For example, the user can be prompted to take a picture of the ground, adjust the device pose, or change the face pose of the target object, so that the user can make corresponding adjustments, thereby improving the accuracy of height detection.

[0023] According to the second aspect or the first possible implementation of the second aspect, in the second possible implementation of the height detection device, determining the first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device includes: determining the second height of the target object based on the ground information, the first face pose, and the device pose; and performing post-processing on the second height to obtain the first height, wherein the post-processing includes Kalman filtering.

[0024] The embodiments of this application can determine the second height of the target object based on ground information, the first face pose and the device pose, and perform post-processing such as Kalman filtering on the second height to obtain the first height of the target object, thereby improving the accuracy of height detection.

[0025] According to the second aspect, or the first possible implementation of the second aspect, or the second possible implementation of the second aspect, in a third possible implementation of the height detection device, the processing unit is further configured to display the first height on the display interface of the electronic device.

[0026] In the embodiments of this application, after determining the first height of the target object, the first height of the target object can be displayed on the display interface of the electronic device through animation, text, augmented reality (AR) and other means, thereby improving the user experience.

[0027] According to the second aspect, in a fourth possible implementation of the height detection device, the ground information is located in the world coordinate system, the first face pose is located in the camera coordinate system, and determining the first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device includes: adjusting the first face pose according to a preset interpupillary distance reference value to obtain a second face pose; performing coordinate transformation on the second face pose according to the device pose to obtain a third face pose of the target object, the third face pose being located in the world coordinate system; and determining the first height of the target object based on the third face pose and the ground information.

[0028] The embodiments of this application can adjust and transform the first face pose located in the camera coordinate system to obtain the third face pose located in the world coordinate system, and determine the first height of the target object based on the third face pose and ground information, thereby enabling the calculation of the first height of the target object in the world coordinate system and improving the accuracy of height detection.

[0029] According to the fourth possible implementation of the second aspect, in the fifth possible implementation of the height detection device, adjusting the first face pose according to a preset interpupillary distance reference value to obtain the second face pose includes: determining a face size transformation coefficient according to the preset interpupillary distance reference value and the interpupillary distance value in the first face pose; adjusting the first face pose according to the face size transformation coefficient to obtain the second face pose of the target object.

[0030] In the embodiments of this application, the face size transformation coefficient is determined by using the pupil distance reference value and the pupil distance value in the first face pose, and the first face pose is adjusted according to the face size transformation coefficient to obtain the second face pose of the target object, thereby obtaining the actual face size and pose of the target object in the camera coordinate system.

[0031] According to the fourth possible implementation of the second aspect, in the sixth possible implementation of the height detection device, determining the first height of the target object based on the third face pose and the ground information includes: determining the top position of the target object's head based on the third face pose; and determining the first height of the target object based on the top position of the head and the ground information.

[0032] The embodiments of this application determine the top position of the target object's head and determine the target object's first height based on the top position and ground information. This method is simple, fast, and can improve the accuracy of height detection.

[0033] Thirdly, embodiments of this application provide a height measurement device, comprising: an image acquisition component for acquiring multiple video frames; a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement one or more of the height detection methods described in the first aspect or multiple possible implementations of the first aspect when executing the instructions.

[0034] The embodiments of this application can perform semantic plane detection on multiple video frames acquired by the image acquisition component of an electronic device to determine ground information in the multiple video frames; simultaneously, face detection is performed on the multiple video frames to determine face regions, and based on the face regions and a preset 3D face model, the first face pose of the target object in the multiple video frames is determined; then, based on the ground information, the first face pose, and the device pose of the electronic device, the first height of the target object is determined. Thus, height detection not only does not rely on professional equipment (such as binocular cameras, depth cameras, etc.), but also can determine the face pose of the target object through face recognition and 3D face technology, and then determine the height of the target object through the face pose, device pose, and ground information. There is no need to manually locate the target object, nor is it necessary to capture a complete human image of the target object. The operation is convenient and highly accurate.

[0035] Fourthly, embodiments of this application provide a non-volatile computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement one or more of the height detection methods described in the first aspect or various possible implementations of the first aspect.

[0036] The embodiments of this application can perform semantic plane detection on multiple video frames acquired by the image acquisition component of an electronic device to determine ground information in the multiple video frames; simultaneously, face detection is performed on the multiple video frames to determine face regions, and based on the face regions and a preset 3D face model, the first face pose of the target object in the multiple video frames is determined; then, based on the ground information, the first face pose, and the device pose of the electronic device, the first height of the target object is determined. Thus, height detection not only does not rely on professional equipment (such as binocular cameras, depth cameras, etc.), but also can determine the face pose of the target object through face recognition and 3D face technology, and then determine the height of the target object through the face pose, device pose, and ground information. There is no need to manually locate the target object, nor is it necessary to capture a complete human image of the target object. The operation is convenient and highly accurate.

[0037] Fifthly, embodiments of this application provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in an electronic device, the processor in the electronic device executes one or more of the height detection methods described in the first aspect or various possible implementations of the first aspect.

[0038] The embodiments of this application can perform semantic plane detection on multiple video frames acquired by the image acquisition component of an electronic device to determine ground information in the multiple video frames; simultaneously, face detection is performed on the multiple video frames to determine face regions, and based on the face regions and a preset 3D face model, the first face pose of the target object in the multiple video frames is determined; then, based on the ground information, the first face pose, and the device pose of the electronic device, the first height of the target object is determined. Thus, height detection not only does not rely on professional equipment (such as binocular cameras, depth cameras, etc.), but also can determine the face pose of the target object through face recognition and 3D face technology, and then determine the height of the target object through the face pose, device pose, and ground information. There is no need to manually locate the target object, nor is it necessary to capture a complete human image of the target object. The operation is convenient and highly accurate.

[0039] These and other aspects of this application will become more apparent in the description of the following embodiments(s). Attached Figure Description

[0040] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this application together with the specification and serve to explain the principles of this application.

[0041] Figure 1 A schematic diagram of the structure of an electronic device according to an embodiment of this application is shown.

[0042] Figure 2 A software structure block diagram of an electronic device according to an embodiment of this application is shown.

[0043] Figure 3 A flowchart of a height detection method according to an embodiment of this application is shown.

[0044] Figure 4 A schematic diagram illustrating the detection process of ground information according to an embodiment of this application is shown.

[0045] Figure 5 This diagram illustrates the process of determining the first facial pose of a target object according to an embodiment of this application.

[0046] Figure 6 A schematic diagram showing the height of a target object according to an embodiment of this application is provided.

[0047] Figure 7 A schematic diagram illustrating the height detection process according to an embodiment of this application is shown.

[0048] Figure 8 A block diagram of a height detection device according to an embodiment of this application is shown. Detailed Implementation

[0049] Various exemplary embodiments, features, and aspects of this application will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0050] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0051] Furthermore, to better illustrate this application, numerous specific details are provided in the following detailed embodiments. Those skilled in the art should understand that this application can be implemented without certain specific details. In some instances, methods, means, components, and circuits well-known to those skilled in the art have not been described in detail in order to highlight the main points of this application.

[0052] In related technologies, measuring human height typically requires specialized equipment such as binocular cameras and depth cameras, which is highly dependent on the equipment and requires capturing complete images of the human body, thus having certain limitations. For example, some technical solutions use binocular cameras to capture scene images, obtain the image coordinates of the apex of the human head in the scene image, and obtain the depth information corresponding to the apex of the head generated by the binocular camera based on the image coordinates of the apex. Then, using the image coordinates and depth information of the apex of the head, the coordinates of the apex in the camera coordinate system are calculated. Finally, based on the coordinates of the apex in the camera coordinate system and the installation height, pitch angle, and tilt angle of the binocular camera, the height of the human target is measured.

[0053] This technical solution not only requires a binocular camera (i.e., it is equipment-dependent), but also necessitates a fixed camera pose and a known camera installation height, thus limiting its application scenarios. Furthermore, this technical solution requires capturing a complete image of the human body to achieve height measurement, which also presents certain limitations.

[0054] For example, some technical solutions can generate a dense semantic map based on simultaneous localization and mapping (SLAM) technology. Then, planar semantic detection can be achieved using this dense semantic map, and object height can be automatically identified based on the inherent relationships between semantic elements. Finally, based on ground extraction and focal target segmentation and projection, the object's length and width can be calculated, ultimately yielding the object's bounding box size (length, width, and height). Since the human body is a type of object in a generalized sense, this technical solution can be used to measure human height.

[0055] However, generating dense semantic maps based on semantic SLAM technology relies on a depth camera (i.e., it is device-dependent). Furthermore, while this technology works well for objects with surfaces parallel to the ground, the human body has a complex shape and no obvious flat surface on its head, resulting in lower measurement accuracy. In addition, this technology requires capturing the entire target object, including its top, to reconstruct a complete outline. For scenarios involving human height measurement, this necessitates the photographer capturing the target from a higher angle, which is inconvenient and has certain limitations.

[0056] With the development of artificial intelligence (AI), some technical solutions have adopted face-height models obtained through machine learning to measure human height. For example, a face classifier and a face-height model can be trained separately first. Then, the image of the human target to be measured is input into the face classifier for face detection to obtain the face image of the human target. This face image is then input into the face-height model for processing to obtain the height of the human target.

[0057] However, the core of this technical solution is the face height model. The face height model obtained through machine learning not only has poor interpretability and is highly dependent on training data, but also has difficulty in generalizing due to the different face height relationships among different ethnic groups, resulting in low accuracy of measurement results.

[0058] In other technical solutions, height is measured by manually manipulating an augmented reality (AR) ruler. For example, the spatial equation of the ground can be obtained using planar detection and SLAM technology. Then, the measuring person needs to position a virtual anchor point at the feet of the human target (i.e., the object being measured) and pull the virtual AR ruler upwards to the top of the head, stopping there. Then, the length of the AR ruler, i.e., the height of the human target, is obtained through a three-dimensional (3D) spatial coordinate system established by SLAM.

[0059] However, this technical solution not only requires manual intervention and has low measurement efficiency, but also involves clicking on a two-dimensional (2D) image and projecting it onto a 3D plane through ray projection. Due to object occlusion, manual operation errors, and other reasons, the virtual anchor point may appear to be positioned at the feet of the human target, but the actual position may be significantly different. In other words, the virtual anchor point is not accurately positioned, which leads to inaccurate height measurement results.

[0060] To address the aforementioned technical problems, this application provides a height detection method applicable to electronic devices. The height detection method of this application can perform semantic plane detection on multiple video frames acquired by the image acquisition component of an electronic device to determine ground information in the multiple video frames; simultaneously, it performs face detection on the multiple video frames to determine face regions; based on the face regions and a preset 3D face model, it determines the first face pose of the target object in the multiple video frames; and based on the ground information, the first face pose, and the device pose of the electronic device, it determines the first height of the target object.

[0061] This method of detecting the height of a target object does not rely on professional equipment (such as binocular cameras, depth cameras, etc.), and can determine the facial pose of the target object through facial recognition and 3D facial technology. Then, the height of the target object can be determined by facial pose, device pose and ground information. There is no need to manually locate the target object or take a complete human image of the target object. It is convenient to operate and highly accurate.

[0062] The electronic devices described in this application can be touchscreen or non-touchscreen. Touchscreen electronic devices can be controlled by clicking or swiping on the display screen using fingers, styluses, etc. Non-touchscreen electronic devices can be connected to input devices such as mice, keyboards, and touch panels for control.

[0063] Figure 1 This illustration shows a schematic diagram of an electronic device 100 according to an embodiment of this application. The electronic device 100 may include at least one of the following: mobile phone, foldable electronic device, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device, in-vehicle device, smart home device, or smart city device. This application does not impose any special limitations on the specific type of the electronic device 100.

[0064] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) connector 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0065] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0066] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors. The processor can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0067] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 may be a cache memory. This memory can store instructions or data that the processor 110 has used or that are used frequently. If the processor 110 needs to use the instruction or data, it can directly retrieve it from this memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0068] In some embodiments, the processor 110 may include one or more interfaces. These interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc. The processor 110 can connect to modules such as touch sensors, audio modules, wireless communication modules, displays, and cameras through at least one of these interfaces.

[0069] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0070] Electronic device 100 can implement display functions through a GPU, display screen 194, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0071] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or more display screens 194.

[0072] Electronic device 100 can perform functions such as taking pictures and videos through camera 193, ISP, video codec, GPU, display screen 194, application processor AP, neural network processor NPU, etc., that is, to realize image and video acquisition and other related functions.

[0073] The camera 193 can be used to acquire color image data of the subject. In some embodiments, the camera 193 can also be used to acquire depth data of the subject. That is, the camera in the electronic device 100 can be a regular camera that does not acquire depth data, such as a monocular camera, or a professional camera that can acquire depth data, such as a binocular camera or a depth camera. This application does not limit the specific type of camera 193.

[0074] The ISP (Image Signal Processor) can be used to process color image data captured by the camera. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's image sensor. The light signal is converted into an electrical signal, which is then transmitted to the ISP for processing, transforming it into a visible image. The ISP can also perform algorithmic optimizations on image noise, brightness, and skin tones. Furthermore, the ISP can optimize parameters such as exposure and color temperature for the shooting scene.

[0075] In some embodiments, the electronic device 100 may include one or more cameras 193. Specifically, the electronic device 100 may include one front-facing camera and at least one rear-facing camera. The front-facing camera is typically used to capture color image data of the photographer facing the display screen 194, while the rear-facing camera is used to capture color image data of the subject (such as a person, landscape, etc.) in front of the photographer.

[0076] In some embodiments, the CPU, GPU, or NPU in the processor 110 can process multiple video frames captured by the camera 193. Specifically, the processor 110 can detect multiple video frames captured by the image acquisition component (i.e., the camera 193) of the electronic device 100 to determine ground information in the multiple video frames; simultaneously, it can perform face detection on the multiple video frames to determine face regions, and determine the first face pose of the target object in the multiple video frames based on the face regions and a preset 3D face model; then, based on the ground information, the first face pose, and the device pose of the electronic device, it can determine the first height of the target object. In some embodiments, the first height of the target object can also be displayed on the display screen 194 of the electronic device 100.

[0077] The gyroscope sensor 180B in the electronic device 100 can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 around three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and controls the lens to move in the opposite direction to counteract the shake of the electronic device 100, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation, motion-sensing games, and other scenarios.

[0078] The accelerometer 180E in electronic device 100 can detect the magnitude of acceleration of electronic device 100 in various directions (generally the x, y, and z axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic device and applied to applications such as screen orientation switching and pedometers.

[0079] In some embodiments, components such as the gyroscope sensor 180B and the accelerometer sensor 180E of the electronic device 100 can constitute an inertial measurement unit (IMU) for measuring the device pose of the electronic device 100.

[0080] The touch sensor 180K in electronic device 100 is also called a "touch device". Touch sensor 180K can be disposed on display screen 194, and the touch sensor 180K and display screen 194 together form a touch screen, also called a "touchscreen". Touch sensor 180K is used to detect touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be disposed on the surface of electronic device 100, in a different location than display screen 194.

[0081] The buttons 190 in the electronic device 100 may include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch buttons. The electronic device 100 can receive button input and generate key signal inputs related to user settings and function control. For example, when taking photos or videos using the camera application (APP) of the electronic device 100, the camera APP can provide buttons such as "Start Photo / Video" and "End Video" for user operation.

[0082] The motor 191 in the electronic device 100 can generate vibration alerts. The motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to different touch operations applied to different applications (such as taking photos, playing audio, etc.). The motor 191 can also correspond to different vibration feedback effects for touch operations applied to different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, taking photos, and video recording) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0083] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.

[0084] Figure 2 A software structure block diagram of an electronic device 100 according to an embodiment of this application is shown.

[0085] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: the application layer, the application framework layer, the Android runtime (ART) and native C / C++ libraries, the Hardware Abstraction Layer (HAL), and the kernel layer.

[0086] The application layer can include a series of application packages.

[0087] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0088] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0089] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, resource manager, notification manager, activity manager, input manager, etc.

[0090] The window manager provides Window Manager Service (WMS), which can be used for window management, window animation management, surface management, and as a relay station for the input system.

[0091] Content providers store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, phone calls made and received, browsing history and bookmarks, phone books, etc.

[0092] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0093] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0094] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0095] The Activity Manager Service (AMS) can be used to start, switch, and schedule system components (such as activities, services, content providers, and broadcast receivers), as well as manage and schedule application processes.

[0096] The Input Manager Service (IMS) provides input management services, which can be used to manage system inputs such as touchscreen input, keypad input, and sensor input. IMS retrieves events from input device nodes and, through interaction with the WMS (Windows Management System), distributes these events to appropriate windows.

[0097] The Android runtime consists of the core libraries and the Android runtime itself. The Android runtime is responsible for converting source code into machine code. The Android runtime primarily employs ahead-of-time (AOT) compilation and just-in-time (JIT) compilation techniques.

[0098] The core library primarily provides basic Java class library functionalities, such as libraries for fundamental data structures, mathematics, I / O, tools, databases, and networking. It also provides APIs for users to develop Android applications.

[0099] The native C / C++ library can include multiple functional modules. Examples include: Surface Manager, Media Framework, libc, OpenGL ES, SQLite, and Webkit. The Surface Manager manages the display subsystem and provides 2D and 3D layer blending for multiple applications. The Media Framework supports playback and recording of various common audio and video formats, as well as still image files. The Media Library supports various audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG. OpenGL ES provides drawing and manipulation of 2D and 3D graphics in applications. SQLite provides a lightweight relational database for applications on electronic devices.

[0100] The Hardware Abstraction Layer (HAL) runs in user space, encapsulates kernel-level drivers, and provides calling interfaces to the upper layers.

[0101] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0102] The following describes the workflow of the software and hardware of the electronic device 100 by way of example, using a height detection scenario from an embodiment of this application.

[0103] Assuming height detection is implemented through a height app on an electronic device, when performing height detection, the user can touch the height app icon on the electronic device's screen. When the touch sensor 180K receives the touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, touch operation timestamp, etc.). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking a single touch operation as an example, and the control corresponding to the single touch operation being the height app icon, the height app calls the application framework layer's interface to launch the height app, and then calls the kernel layer to launch the camera driver, capturing multiple video frames through the camera 193, i.e., capturing a video stream through the camera 193. These multiple video frames may include the target object whose height is to be detected.

[0104] After acquiring multiple video frames, the electronic device 100 can perform ground detection, face detection and other related processing on the multiple video frames through the processor 110, thereby determining the height of the target object.

[0105] Figure 3 A flowchart illustrating a height detection method according to an embodiment of this application is shown. Figure 3 As shown, the height detection method includes: step S310, performing semantic plane detection on multiple video frames acquired by the image acquisition component of the electronic device to determine the ground information in the multiple video frames.

[0106] The image acquisition component can be a camera of an electronic device. This camera can be a regular camera that does not collect depth data, such as a monocular camera, or a professional camera that can collect depth data, such as a binocular camera or a depth camera.

[0107] When the image acquisition component is a standard camera, the multiple video frames captured by the image acquisition component are color (red, green, blue, RGB) video frames. Since multiple video frames can form a video stream, it can also be considered that the image acquisition component is capturing an RGB video stream. When the image acquisition component is a professional camera, the multiple video frames captured by the image acquisition component may include depth data in addition to RGB image data. It should be noted that this application does not limit the specific type of image acquisition component.

[0108] Multiple video frames (i.e., RGB video streams) can be acquired using the image acquisition component of an electronic device, and the acquired video frames can be processed by plane detection, semantic segmentation, and other methods to determine the ground information in the multiple video frames. The ground information can be represented by a plane equation in space, or by other methods; this application does not impose any limitations on this.

[0109] In one possible implementation, when determining ground information in multiple video frames, planar detection can first be performed on the multiple video frames acquired by the image acquisition unit to determine the position information of multiple planes in the multiple video frames. Optionally, SLAM technology can be used to determine the position information of multiple planes in multiple video frames.

[0110] For example, 3D information can be extracted from multiple video frames acquired by an image acquisition unit to obtain sparse point cloud data, while simultaneously determining the device pose of the electronic device when acquiring each video frame. Then, based on the device pose of the electronic device when acquiring each video frame, a plane fitting algorithm is used to perform plane fitting on the sparse point cloud data to obtain the position information of multiple planes in the multiple video frames. The position information of each plane can be represented by plane equations in space.

[0111] By extracting 3D information from multiple video frames, sparse point cloud data is obtained. Based on the device pose when the electronic device collects each video frame, plane fitting is performed on the sparse point cloud data to obtain the position information of multiple planes in multiple video frames. This not only improves processing efficiency but also improves the accuracy of the position information of each plane.

[0112] While determining the positional information of multiple planes in multiple video frames, semantic segmentation can also be performed on the multiple video frames acquired by the image acquisition unit to obtain the semantic segmentation results of each video frame. Specifically, for any video frame among multiple video frames, semantic recognition can be performed on the video frame to identify the category of objects in the video frame, such as the ground, table, wall, etc., and according to the identified object category, each pixel in the video frame is labeled to obtain the semantic segmentation result of the video frame.

[0113] Then, based on the positional information of multiple planes in multiple video frames and the semantic segmentation results of each video frame, semantic recognition can be performed on multiple planes in multiple video frames to obtain multiple semantic plane information, such as desktop, wall, and ground information. Then, ground information can be selected from multiple semantic plane information.

[0114] Figure 4 A schematic diagram illustrating the detection process of ground information according to an embodiment of this application is shown. Figure 4As shown, assuming that the height detection method of this application embodiment is implemented through an application called Height APP on an electronic device (e.g., a mobile phone), after the user opens the Height APP, the Height APP can acquire images through the image acquisition component (e.g., a camera) of the electronic device to obtain multiple video frames 410 (i.e., RGB video streams); then, three-dimensional information is extracted from the multiple video frames 410 to obtain sparse point cloud data 420, and the device pose 430 when the electronic device acquires the multiple video frames 410 is determined; and based on the device pose 430 when the electronic device acquires the multiple video frames 410, a plane fitting algorithm is used to perform plane fitting on the sparse point cloud data 420 to obtain the position information 440 of multiple planes in the multiple video frames 410.

[0115] While determining the position information 440 of multiple planes in multiple video frames 410, semantic segmentation 450 can also be performed on multiple video frames 410 to obtain semantic segmentation results 460; then, based on the position information 440 and semantic segmentation results 460 of multiple planes, semantic recognition can be performed on multiple planes in multiple video frames 410 to obtain multiple semantic plane information 470, and ground information 480 can be selected from the multiple semantic plane information 470, wherein the ground information 480 can be represented by plane equations in space.

[0116] The above description exemplifies the ground information detection process using only multiple video frames (i.e., RGB video streams) acquired by the image acquisition unit as input. In some embodiments, depth data acquired by the image acquisition unit and device pose information of the electronic device acquired by the inertial measurement unit (IMU) can also be used as input simultaneously to improve the accuracy of ground detection.

[0117] By performing planar detection and semantic segmentation on multiple video frames acquired by the image acquisition unit, electronic devices can automatically perceive the shooting scene and acquire multiple semantic planar information, thereby automatically recognizing ground information. This not only avoids manual operation by the user, such as manually selecting the ground, but also improves the accuracy of ground detection.

[0118] In one possible implementation, if no ground information is detected in multiple video frames within a preset time period (e.g., 5 seconds, 10 seconds), the user can be prompted to take a picture of the ground. For example, prompts such as "Please take a picture of the ground" or "No ground detected" can be broadcast to the user via voice, or the prompts can be displayed on the height app's screen using text, animation, or other means, allowing the user to adjust the shooting content in a timely manner, thereby improving the efficiency of ground detection.

[0119] It should be noted that those skilled in the art can set the content and manner of the prompting information when ground information is not detected in multiple video frames according to the actual situation, and this application does not impose any restrictions on this.

[0120] Step S320: Perform face detection on the multiple video frames to determine face regions. While determining the ground information in the multiple video frames, face detection can be performed on the multiple video frames through feature extraction, key point detection, and other methods. If a complete face is detected, face regions can be determined from the multiple video frames, and the object corresponding to the face region can be identified as the target object. There can be one or more face regions, and there can also be one or more target objects; this application does not impose any restrictions on either.

[0121] In one possible implementation, after determining the face region from multiple video frames, it can be determined whether the face region meets a third preset condition. This third preset condition is that the face region is located within a preset area of ​​the video frame in which it resides. The preset area can be the central area of ​​the video frame containing the face region. For example, the central area of ​​the video frame containing the face region can be set as an area centered on its center point, with an area equal to half its size. It should be noted that those skilled in the art can set the preset area of ​​the video frame containing the face region according to actual circumstances, and this application does not impose any restrictions on this.

[0122] If the face region does not meet the third preset condition, the user can be prompted to adjust the device position of the electronic device through voice broadcast, text display, animation display, etc., so that the face region meets the third preset condition, thereby making the face region located within the preset area of ​​the video frame, thus improving the accuracy of height detection.

[0123] Step S330: Determine the first facial pose of the target object in the multiple video frames based on the facial region and the preset 3D facial model. When determining the first facial pose of the target object in the multiple video frames, a 3D facial model of the target object can first be established using a pre-trained neural network based on the facial region and the preset 3D facial model (i.e., the average 3D facial model). For example, the facial region and the preset 3D facial model can be input into a pre-trained convolutional neural network (CNN) for registration to obtain the 3D facial model of the target object.

[0124] Then, based on the preset parameters of the 3D facial model, such as constraints on the facial structure (interpupillary distance, distance from the tip of the nose to the top of the head, etc.), the position and rotation information of the target object's 3D facial model relative to the image acquisition component of the electronic device can be determined. This position and rotation information of the target object's 3D facial model relative to the image acquisition component of the electronic device is then used to determine the first facial pose of the target object. The rotation information can be represented by pitch angle, roll angle, and yaw angle.

[0125] Figure 5 This diagram illustrates the process of determining the first facial pose of a target object according to an embodiment of this application. Figure 5 As shown, face detection 520 can be performed on multiple video frames 510 acquired by the image acquisition component of the electronic device to determine the face region 530; the face region 530 and the preset face 3D model 540 are input into the pre-trained convolutional neural network CNN 550 for registration to obtain the face 3D model 560 of the target object; according to the parameters of the preset face 3D model 540, the position information and rotation information of the face 3D model of the target object relative to the image acquisition component of the electronic device are determined, and the position information and rotation information are determined as the first face pose 570 of the target object.

[0126] In this way, the 3D model of the target object's face can be determined based on the face region and the preset 3D face model. Then, based on the parameters of the preset 3D face model, the first face pose of the target object can be determined. Thus, the 3D face reconstruction technology can be used to determine the first face pose of the target object, which can not only improve processing efficiency but also improve the accuracy of the first face pose of the target object, thereby improving the accuracy of height detection.

[0127] In one possible implementation, a neural network (such as a convolutional neural network CNN 550) can be pre-trained based on multiple sample face regions and a pre-defined 3D face model. For example, for any sample face region, the sample face region and the pre-defined 3D face model can be input into the neural network for registration to obtain a sample face 3D model; then, the sample face 3D model is reverse-rendered, that is, projected into a two-dimensional space to obtain a reverse-rendered image; the network loss of the neural network is determined based on the differences between each reverse-rendered image and the corresponding sample face region; and the network parameters of the neural network are adjusted based on this network loss.

[0128] Training ends when the neural network meets preset training termination conditions, resulting in a trained neural network. These termination conditions may include, for example, the neural network reaching a preset training epoch threshold, the network loss converging within a certain range, or the neural network passing validation on a validation set. Those skilled in the art can set the training termination conditions for the neural network according to actual circumstances, and this application does not impose any restrictions on this.

[0129] In one possible implementation, after determining the first face pose of the target object, it can be determined whether the first face pose meets a second preset condition. The second preset condition is that the pitch angle of the first face pose is within a preset second angle range, the roll angle of the first face pose is within a preset third angle range, and the yaw angle of the first face pose is within a preset fourth angle range. The second, third, and fourth angle ranges can be the same or different. It should be noted that those skilled in the art can set the specific values ​​of the second, third, and fourth angle ranges according to actual conditions, and this application does not impose any restrictions on this.

[0130] If the first facial pose of the target object does not meet the second preset condition, the user can be prompted to adjust the device pose of the electronic device and / or change the facial pose of the target object through voice broadcast, text display, animation display, etc., so that the first facial pose of the target object meets the second preset condition, thereby making the face of the target object face the image acquisition component of the electronic device. In other words, it can make the face in the video frame acquired by the image acquisition component the frontal face of the target object, thereby improving the accuracy of height detection.

[0131] Step S340: Determine the first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device. The device pose of the electronic device is the pose of the electronic device when it captures the video frame containing the face region.

[0132] In one possible implementation, before determining the first height of the target object, it can be determined whether the pitch angle of the electronic device indicated by the device pose of the electronic device meets a first preset condition. The first preset condition is that the pitch angle of the electronic device indicated by the device pose of the electronic device is within a preset first angle range.

[0133] If the pitch angle of the electronic device indicated by the device posture does not meet the first preset condition, the user can be prompted to adjust the device posture through voice broadcast, text display, animation display, etc., to avoid excessive upward or downward shooting during video frame acquisition, thereby improving the accuracy of height detection.

[0134] When determining the initial height of a target object, the coordinate system of the ground information and the target object's initial facial pose can be determined first. When using SLAM technology, the ground information is located in world coordinates. Based on data collected by the inertial measurement unit (IMU), the Y-axis of the world coordinate system can be defined as the vertical direction of the real world. Since physical information such as distance and object size in the world coordinate system is the same as in the real world, a connection can be established between the virtual world coordinate system and the real world. Therefore, the object size calculated in the world coordinate system is the actual size of the object in the real world.

[0135] The first facial pose of the target object is its facial pose relative to the image acquisition component of the electronic device, which is located in the camera coordinate system. In the camera coordinate system, the image acquisition component of the electronic device is located at the origin. From the perspective of the 3D model of the target object's face, the position of the image acquisition component is fixed. Furthermore, since there is no size comparison between the object in the camera coordinate system and the object in the real world, the facial size of the target object in the camera coordinate system is scaled down compared to the actual facial size in the real world. Therefore, facial size adjustment and coordinate system transformation are necessary.

[0136] In one possible implementation, when determining the first height of the target object, the first facial pose of the target object can be adjusted according to a preset interpupillary distance reference value to obtain the second facial pose of the target object. The first facial pose indicates the same face size as the target object's face size within the face region, while the second facial pose indicates the actual face size of the target object. The second facial pose is located in the camera coordinate system. In other words, the first facial pose of the target object can be adjusted in the camera coordinate system so that the adjusted second facial pose indicates the actual face size of the target object.

[0137] For example, the interpupillary distance value in the first face pose of the target object can be determined, and the face size transformation coefficient can be determined based on the preset interpupillary distance reference value and the interpupillary distance value in the first face pose; then, the first face pose can be adjusted according to the face size transformation coefficient to obtain the second face pose of the target object.

[0138] By using the interpupillary distance reference value and the interpupillary distance value in the first face pose, the face size transformation coefficient is determined. Based on the face size transformation coefficient, the first face pose is adjusted to obtain the second face pose of the target object, thereby obtaining the actual face size and pose of the target object in the camera coordinate system.

[0139] After obtaining the second face pose of the target object, the second face pose can be transformed into coordinates based on the device pose of the electronic device to obtain the third face pose of the target object, wherein the third face pose of the target object is located in the world coordinate system.

[0140] In one possible implementation, the second face pose P can be determined using the following formula (1). C Perform coordinate transformation to obtain the third face pose P of the target object. w :

[0141] P w =T -1 P C (1)

[0142] In formula (1), T represents the rigid body transformation matrix determined based on the device pose (R, t) of the electronic device. Where R represents the rotation matrix in the device pose of the electronic device, and t represents the translation matrix in the device pose of the electronic device.

[0143] Then, the first height of the target object can be determined based on the third-person face pose and ground information. In one possible implementation, the top position of the target object's head can be determined based on the third-person face pose, and then the first height of the target object can be determined based on the top position of the target object's head and ground information. For example, assuming the top position of the target object determined based on the third-person face pose is (x1, y1, z1), and the ground information is represented by the plane equation F = f(x, y, z) in space, the first distance L1 from (x1, y1, z1) to the ground F = f(x, y, z) can be calculated in the Y-axis direction, and this first distance L1 is determined as the first height L of the target object, i.e., L = L1.

[0144] By determining the top position of the target's head and using this position along with ground information, the target's initial height can be determined quickly and easily, improving the accuracy of height detection.

[0145] In one possible implementation, the nose tip position of the target object can be determined based on the third person's facial pose. Then, based on the nose tip position of the target object and the preset proportional relationship between the nose tip to the chin and the nose tip to the top of the head, the top position of the target object can be determined. Finally, based on the top position of the target object's head and the ground information, the first height of the target object can be determined.

[0146] For example, assuming the nose tip position of the target object determined based on the third person's face pose is (x2, y2, z2), and the ground information is represented by the plane equation F = f(x, y, z) in space, the head position (x3, y3, z3) of the target object can be determined based on the nose tip position (x2, y2, z2) and the preset proportional relationship between the nose tip to the chin and the nose tip to the top of the head. Then, the second distance L2 from (x3, y3, z3) to the ground F = f(x, y, z) is calculated in the Y-axis direction, and this second distance L2 is determined as the first height L of the target object, i.e., L = L2.

[0147] By determining the position of the target's nose tip and the proportional relationship between the nose tip to the chin and the nose tip to the top of the head, the target's head position can be determined. Based on the target's head position and ground information, the target's initial height can be determined, thereby improving the accuracy of height detection.

[0148] In one possible implementation, when there are multiple face regions corresponding to the target object, for any given face region, the second height of the target object can be determined using a method similar to the one described above, based on ground information from multiple video frames, the first face pose of the target object, and the device pose when the electronic device captured the video frame containing that face region. Then, Kalman filtering and averaging are performed on the multiple second heights to obtain the first height of the target object. This approach improves the accuracy of height detection.

[0149] In one possible implementation, after determining the target object's initial height, this initial height can also be displayed on the electronic device's display interface. For example, after determining the target object's initial height, it can be displayed on the electronic device's display interface using animation, text, augmented reality (AR), or other methods. When height detection is implemented through a height-themed app, the target object's initial height can be displayed on the app's display interface. The height app's display interface may include a real-time image interface of video frames captured by the electronic device's image acquisition component.

[0150] Figure 6 This diagram illustrates the display of the height of a target object according to an embodiment of this application. Figure 6As shown, the user uses a height-themed app on the electronic device 600 to detect the height of the target object 630. The display interface 610 of the height-themed app displays video frames captured in real time by the image acquisition component (not shown) of the electronic device 600. When the height-themed app detects the height of the target object 630, it can display the height at a preset position above the head of the target object 630 in the display interface 610 of the height-themed app by means of an augmented reality icon 620. The displayed information can be "Height: 175CM".

[0151] The above Figure 6 The method of displaying height is illustrated using only one target object as an example. It should be noted that the heights of multiple target objects can also be displayed in the same way. Those skilled in the art can further configure the display method and position of the target object's height according to actual circumstances; this application does not impose any limitations on this.

[0152] The height detection method according to the embodiments of this application can perform semantic plane detection on multiple video frames acquired by the image acquisition component of an electronic device to determine ground information in the multiple video frames; simultaneously, it can perform face detection on the multiple video frames to determine face regions, and determine the first face pose of the target object in the multiple video frames based on the face regions and a preset 3D face model; then, based on the ground information, the first face pose, and the device pose of the electronic device, it can determine the first height of the target object. Thus, the height detection not only does not rely on professional equipment (such as binocular cameras, depth cameras, etc.), but can also determine the face pose of the target object through face recognition and 3D face technology, and then determine the height of the target object through the face pose, device pose, and ground information. It does not require manual positioning of the target object, nor does it require taking a complete human image of the target object. It is convenient to operate and highly accurate.

[0153] Figure 7 A schematic diagram illustrating the height detection process according to an embodiment of this application is shown. Figure 7 As shown, assuming a user performs height detection on a target object using a height-themed app running on an electronic device, step S701 is executed when the user opens the height app. The height app acquires multiple video frames (i.e., a video stream) through the image acquisition component of the electronic device. Optionally, the image acquisition component can continuously acquire the video stream during the height detection process.

[0154] When semantic plane detection is implemented using SLAM technology, in step S702, it can be determined whether SLAM is successfully initialized. If SLAM is not successfully initialized, the user is prompted to move the electronic device, and step S701 is re-executed. If SLAM is successfully initialized, step S703 is executed to perform semantic plane detection on multiple video frames, and in step S704, it is determined whether ground information is detected within a preset time period.

[0155] If no ground information is detected within the preset time period, the user is prompted to take a picture of the ground, and step S701 continues. If ground information is detected within the preset time period, step S705 is executed to perform face detection on multiple video frames, determine the face region, and in step S706, it is determined whether the face region meets the third preset condition. The third preset condition is that the face region is located within the preset area of ​​the video frame in which it is located.

[0156] If the face region does not meet the third preset condition, the user is prompted to adjust the device pose of the electronic device, and step S701 is executed again; if the face region meets the third preset condition, step S707 is executed, and the first face pose of the target object in multiple video frames is determined according to the face region and the preset face 3D model. In step S708, it is determined whether the first face pose meets the second preset condition. The second preset condition is that the pitch angle of the first face pose is within the preset second angle range, the roll angle of the first face pose is within the preset third angle range, and the yaw angle of the first face pose is within the preset fourth angle range.

[0157] If the first face pose does not meet the second preset condition, the user is prompted to adjust the device pose of the electronic device and / or change the face pose of the target object, and step S701 is executed again; if the first face pose meets the second preset condition, step S709 is executed to determine whether the pitch angle of the electronic device indicated in the device pose of the electronic device meets the first preset condition. The first preset condition is that the pitch angle of the electronic device indicated in the device pose of the electronic device is within a preset first angle range.

[0158] If the pitch angle of the electronic device indicated in the device pose does not meet the first preset condition, the user is prompted to adjust the device pose of the electronic device, and step S701 is executed again; if the pitch angle of the electronic device indicated in the device pose meets the first preset condition, step S710 is executed, and the first height of the target object is determined based on the ground information, the first face pose, and the device pose of the electronic device; then step S711 is executed, and the first height of the target object is displayed on the display interface of the height APP through augmented reality (AR) or other methods.

[0159] The height detection method of this application, through SLAM technology and semantic segmentation, can automatically identify ground information and automatically detect the height of target objects without manual operation (such as manual selection or marking of target objects). Furthermore, it can simultaneously detect the height of multiple target objects, thereby simplifying the height detection process and improving efficiency. In addition, the embodiments of this application, by acquiring three-dimensional information through SLAM technology, can also avoid contact between electronic devices and the human body of the target object, ensuring safety and reliability.

[0160] The height detection method in this application is based on multiple video frames captured by a common camera (e.g., a monocular camera) to detect height, without the need for specialized equipment such as depth cameras, thus reducing equipment dependence. In some embodiments, users can perform height detection using handheld devices (e.g., mobile phones, smartwatches, etc.). Furthermore, the embodiments of this application obtain the facial pose of the target object quickly and accurately through face recognition and 3D face reconstruction technology, which not only improves the accuracy of height detection but is also applicable to scenarios involving target object movement and changes in shooting angle.

[0161] Figure 8 A block diagram of a height detection device according to an embodiment of this application is shown. Figure 8 As shown, the height detection device is applied to an electronic device. The height detection device includes: an image acquisition unit 810 for acquiring multiple video frames; and a processing unit 820 configured to: perform semantic plane detection on the multiple video frames to determine ground information in the multiple video frames; perform face detection on the multiple video frames to determine face regions; determine the first face pose of a target object in the multiple video frames based on the face image of the face region and a preset 3D face model; and determine the first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device.

[0162] In one possible implementation, the processing unit is further configured to: prompt the user to take a picture of the ground if no ground information is detected within a preset time period; prompt the user to adjust the device pose if the pitch angle of the electronic device indicated by the device pose does not meet a first preset condition; prompt the user to adjust the device pose and / or change the face pose of the target object if the first face pose does not meet a second preset condition; or prompt the user to adjust the device pose if the face region does not meet a third preset condition.

[0163] In one possible implementation, determining the first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device includes: determining the second height of the target object based on the ground information, the first face pose, and the device pose; and performing post-processing on the second height to obtain the first height, wherein the post-processing includes Kalman filtering.

[0164] In one possible implementation, the processing unit is further configured to display the first height on the display interface of the electronic device.

[0165] An embodiment of this application provides a height detection device, including: an image acquisition component for acquiring multiple video frames; a processor and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions.

[0166] Embodiments of this application provide a non-volatile computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the above-described method.

[0167] Embodiments of this application provide a computer program product including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.

[0168] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital video disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing.

[0169] The computer-readable program instructions or code described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0170] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as "C" or similar languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information from computer-readable program instructions. These electronic circuits can execute computer-readable program instructions to implement various aspects of this application.

[0171] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0172] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0173] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0174] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.

[0175] It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented using hardware (such as circuits or ASICs (Application Specific Integrated Circuits)) that performs the corresponding function or action, or using a combination of hardware and software, such as firmware.

[0176] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, disclosure, and appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0177] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A height detection method, characterized in that, include: Semantic plane detection is performed on multiple video frames acquired by the image acquisition component of an electronic device to determine ground information in the multiple video frames, wherein the multiple video frames are color RGB video frames. Face detection is performed on the multiple video frames to determine the face regions; Based on the face region and the preset 3D face model, the first face pose of the target object in the multiple video frames is determined; The first height of the target object is determined based on the ground information, the first face pose, and the device pose of the electronic device. The method also includes: If semantic synchronous localization and mapping SLAM fails to initialize, the user is prompted to move the electronic device.

2. The method according to claim 1, characterized in that, The method further includes at least one of the following: If the ground information is not detected within a preset time period, the user is prompted to take a picture of the ground. If the pitch angle of the electronic device indicated by the device pose does not meet the first preset condition, the user is prompted to adjust the device pose. If the first face pose does not meet the second preset condition, the user is prompted to adjust the device pose and / or change the face pose of the target object. or If the face area does not meet the third preset condition, the user is prompted to adjust the device pose.

3. The method according to claim 1 or 2, characterized in that, Determining the first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device includes: Based on the ground information, the first face pose, and the device pose, the second height of the target object is determined; The second height is post-processed to obtain the first height, and the post-processing includes Kalman filtering.

4. The method according to any one of claims 1-3, characterized in that, The method further includes: The first height is displayed on the display interface of the electronic device.

5. A height detection device, characterized in that, Applied to electronic devices, including: An image acquisition component is used to acquire multiple video frames, wherein the multiple video frames are color RGB video frames; The processing unit is configured as follows: Perform semantic plane detection on the multiple video frames to determine the ground information in the multiple video frames; Face detection is performed on the multiple video frames to determine the face regions; Based on the face image of the face region and the preset 3D face model, determine the first face pose of the target object in the multiple video frames; Based on the ground information, the first face pose, and the device pose of the electronic device, the first height of the target object is determined: The processing unit is further configured to: If semantic synchronous localization and mapping SLAM fails to initialize, the user is prompted to move the electronic device.

6. The apparatus according to claim 5, characterized in that, The processing unit is also configured to be at least one of the following: If the ground information is not detected within a preset time period, the user is prompted to take a picture of the ground. If the pitch angle of the electronic device indicated by the device pose does not meet the first preset condition, the user is prompted to adjust the device pose. If the first face pose does not meet the second preset condition, the user is prompted to adjust the device pose and / or change the face pose of the target object. or If the face area does not meet the third preset condition, the user is prompted to adjust the device pose.

7. The apparatus according to claim 5 or 6, characterized in that, Determining the first height of the target object based on the ground information, the first face pose, and the device pose of the electronic device includes: Based on the ground information, the first face pose, and the device pose, the second height of the target object is determined; The second height is post-processed to obtain the first height, and the post-processing includes Kalman filtering.

8. The apparatus according to any one of claims 5-7, characterized in that, The processing unit is further configured to: The first height is displayed on the display interface of the electronic device.

9. A height detection device, characterized in that, include: Image acquisition unit, used to acquire multiple video frames; processor; Memory used to store processor-executable instructions; The processor is configured to implement the method of any one of claims 1-4 when executing the instructions.

10. A non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1-4.

11. A computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, wherein when the computer-readable code is executed in an electronic device, a processor in the electronic device performs the method of any one of claims 1-4.