Display method and related device
Patent Information
- Application Number
- CN202611132101.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-28
- Publication Date
- 2026-09-25
AI Technical Summary
然而,相关技术中普遍采用静态、人工主导的方式:距离调整依赖人工试错,无法基于用户的身高或肢体比例提供量化引导;环境光变化易造成画面局部过曝或过暗,导致肢体轮廓、体表细节等关键特征识别误差率升高;设备摆放高度或角度存在偏差时,缺乏动态变焦及空间视角校准机制,动作捕捉准确率下降,且无实时变焦补偿能力,相关技术缺乏基于用户个体特征、环境光照变化和设备姿态偏差的动态自适应调节能力
[0011]本公开实施例提供的显示方法及相关设备,通过获取检测期间的原始图像、用于表征设备与用户之间空间关系是否符合第一预定条件的第一状态数据以及用于表征环境光照是否符合第二预定条件的第二状态数据,再根据所述第一状态数据对所述原始图像或用户进行距离引导操作、视角矫正操作或变焦操作中的至少一个,并根据所述第二状态数据对原始图像的环境光照状态进行调整,之后基于所述第一调整操作和/或所述第二调整操作调整后的所述原始图像进行显示,从而可以在用户动作检测过程中动态自适应地优化拍摄参数与空间关系,提升动作捕捉的准确性和效率。
Smart Images

Figure CN122824975A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a display method and related equipment. Background Technology
[0002] With the improvement of mobile terminal computing power and the maturity of computer vision technology, vision-based user action detection systems are widely used in scenarios such as skills assessment, medical rehabilitation, and somatosensory interaction. The real-time acquisition and analysis of user limb movements and appearance features exhibit characteristics of strong real-time performance, high precision, and strong environmental dependence, which places higher demands on the system's adaptive adjustment capabilities.
[0003] In related technologies, to accurately capture users' body posture and movement details, it is usually necessary to ensure that the camera's angle of view, clarity, and distance from the user are appropriate. However, these technologies generally adopt a static, manually-driven approach: distance adjustment relies on manual trial and error, and cannot provide quantitative guidance based on the user's height or body proportions; changes in ambient light can easily cause local overexposure or underexposure of the image, leading to an increased error rate in recognizing key features such as body contours and surface details; when there are deviations in the device's placement height or angle, there is a lack of dynamic zoom and spatial perspective calibration mechanisms, resulting in a decrease in motion capture accuracy, and there is no real-time zoom compensation capability. These technologies lack the ability to dynamically and adaptively adjust based on individual user characteristics, changes in ambient light, and device posture deviations. Summary of the Invention
[0004] In view of this, the purpose of this disclosure is to provide a display method and related equipment to solve or partially solve the above-mentioned problems to a certain extent.
[0005] In a first aspect, this disclosure provides a display method, comprising:
[0006] Acquire raw images, first state data, and second state data during the detection process; the first state data is used to characterize whether the spatial relationship between the device and the user meets a first predetermined condition; the second state data is used to characterize whether the ambient lighting meets a second predetermined condition. Based on the first state data, a first adjustment operation is performed on the original image or the user; the first adjustment operation includes at least one of performing a distance guidance operation on the user, a perspective correction operation on the original image, or a zoom operation on the original image. Based on the second state data, a second adjustment operation is performed on the original image to adjust the original image from the current ambient lighting state to the target lighting state; The original image is displayed based on the adjustment performed by the first adjustment operation and / or the second adjustment operation.
[0007] A second aspect of this disclosure provides a display device, comprising: The acquisition module is configured to acquire raw images, first state data, and second state data during the detection period; the first state data is used to characterize whether the spatial relationship between the device and the user meets a first predetermined condition; the second state data is used to characterize whether the ambient lighting meets a second predetermined condition. The first adjustment module is configured to perform a first adjustment operation on the original image or the user based on the first state data; the first adjustment operation includes at least one of performing a distance guidance operation on the user, a perspective correction operation on the original image, or a zoom operation on the original image. The second adjustment module is configured to perform a second adjustment operation on the original image based on the second state data, so as to adjust the original image from the current ambient lighting state to the target lighting state. The display module is configured to display the original image after adjustment based on the first adjustment operation and / or the second adjustment operation.
[0008] A third aspect of this disclosure provides a computer device including one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, and the one or more programs include instructions for performing the method of the first aspect.
[0009] A fourth aspect of this disclosure provides a non-volatile computer-readable storage medium comprising a computer program that, when executed by one or more processors, causes the one or more processors to perform the method described in the first aspect.
[0010] A fifth aspect of this disclosure provides a computer program product comprising one or more computer programs that, when executed by one or more processors, implement the method as described in the first aspect.
[0011] The display method and related devices provided in this disclosure acquire an original image during detection, first state data characterizing whether the spatial relationship between the device and the user meets a first predetermined condition, and second state data characterizing whether the ambient light meets a second predetermined condition. Then, based on the first state data, at least one of a distance guidance operation, a perspective correction operation, or a zoom operation is performed on the original image or the user. The ambient light state of the original image is adjusted based on the second state data. Finally, the original image is displayed based on the original image adjusted by the first adjustment operation and / or the second adjustment operation. This allows for dynamic and adaptive optimization of shooting parameters and spatial relationships during user motion detection, thereby improving the accuracy and efficiency of motion capture. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A schematic diagram of an exemplary system provided by an embodiment of this disclosure is shown; Figure 2 A flowchart illustrating an exemplary display method provided in an embodiment of this disclosure is shown; Figure 3 A schematic diagram of a virtual footprint positioning frame according to an embodiment of the present disclosure is shown; Figure 4 A schematic diagram of the mobile phone tilt angle according to an embodiment of the present disclosure is shown; Figure 5 A schematic diagram of the hardware structure of an exemplary display device provided in an embodiment of this disclosure is shown; Figure 6 A schematic diagram of the hardware structure of an exemplary computer device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0015] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar words used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0016] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.
[0017] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.
[0018] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0019] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0020] Figure 1 A schematic diagram of an exemplary system 100 provided in this disclosure embodiment is shown. This exemplary system 100 may include a terminal device 110, a server 130, and a display device 160. Optionally, software and / or an application 120 (hereinafter referred to as application 120) may be installed on the terminal device 110. A user 140 may interact with the application 120 via the terminal device 110 and / or an attachment device to the terminal device 110.
[0021] In some embodiments, application 120 may be downloaded and installed on terminal device 110. In some embodiments, application 120 may also be accessed in other ways, such as via a web page. Figure 1 In system 100, in response to application 120 being launched, terminal device 110 can display the interface of application 120.
[0022] In some embodiments, terminal device 110 can communicate with server 130 through network 150 to provide services to application 120. Network 150 can be a wired network or a wireless network. In some cases, intermediate devices or network nodes such as routers and switches can be further configured in network 150, and terminal device 110 communicates with server 130 through these intermediate devices or network nodes, which can also be considered as communicating through network 150.
[0023] Terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 can also support any type of user-facing interface (such as "wearable" circuitry). Application 120 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc. Server 130 can be a server providing various services, such as a backend server supporting various applications or software displayed on terminal device 110. Server 130 here can be hardware or software. When it is hardware, it can be implemented as a distributed server cluster consisting of multiple servers or as a single server. When these are software applications, they can be implemented as multiple software programs or software modules (e.g., to provide distributed services) or as a single software program or software module. No specific limitations are made here.
[0024] In some cases, the display device 160 can be independent of the terminal device 110 and the server 130. In other cases, if the terminal device 110 and / or the server 130 have display capabilities, the display device 160 can be integrated with the terminal device 110 and / or the server 130.
[0025] It should be understood that the structure and function of the various elements in system 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0026] In some exemplary scenarios, the terminal device 110 or the server 130 can execute the display method provided in the embodiments of this application to automatically optimize the image acquisition quality and display it during detection. However, the inventors of this disclosure have found that in related technologies, it is usually necessary to rely on users to manually adjust device parameters, repeatedly try and fail at positioning, and use fixed exposure strategies, which leads to low detection efficiency, poor accuracy, and situations such as false detection and missed detection.
[0027] Furthermore, the display systems of related technologies typically rely on manual adjustment of shooting distance, manual auxiliary perspective correction, and fixed exposure strategies to cope with changes in ambient light. When dealing with users with different individual characteristics (such as height and body proportions), they cannot provide quantitative guidance and dynamic calibration based on spatial relationship anomalies (such as improper distance or perspective deviation). In scenes with changing light or local overexposure / underexposure, the error rate of key feature recognition is high, and the capture of body movement details is inaccurate. When there is a deviation in the device placement angle, the lack of dynamic zoom and spatial perspective calibration leads to a decrease in motion capture accuracy. On the other hand, existing systems provide users with a single guidance method. Different detection stages have different sensitivities to spatial relationships and lighting conditions. Existing systems lack a dynamic weight adjustment mechanism based on detection steps and lack intelligent sorting processing of the causal relationship between spatial relationship anomalies and ambient lighting anomalies, resulting in unreasonable adjustment operation order and affecting the overall detection and final display effect.
[0028] In summary, the challenges of image quality optimization during the detection process are not a matter of adjusting a single parameter, but rather a result of multiple factors, including individual feature differences, dynamic changes in ambient light, spatial perspective distortion, multimodal interactive guidance, and the specificity of the detection process. Systematic improvements in these fundamental capabilities are urgently needed.
[0029] In view of this, a first aspect of the present disclosure provides a display method that can solve or partially solve the above-mentioned problems to a certain extent.
[0030] Figure 2 A schematic flowchart of an exemplary display method 200 provided in an embodiment of this disclosure is shown. This display method 200 can be used to automatically optimize image acquisition quality during detection. Optionally, this display method 200 can be... Figure 1 The display device 160 can be implemented, or it can be made by Figure 1 The system 100 is implemented jointly by servers 130 or various entities through interactive means.
[0031] like Figure 2 As shown, the display method 200 may further include the following steps.
[0032] In step 201, the original image, first state data, and second state data during the detection period are acquired; the first state data is used to characterize whether the spatial relationship between the device and the user meets the first predetermined condition; the second state data is used to characterize whether the ambient lighting meets the second predetermined condition.
[0033] In this step, the detection period can refer to the time period required for limb and body shape detection during a skills talent interview, the time period for limb function assessment in the field of medical rehabilitation, or the time period required for capturing user movements during motion-sensing games and virtual reality (VR) interactions. During this detection period, the terminal device 110 (e.g., a mobile phone or tablet) can acquire the user's video stream in real time through its camera and obtain raw image frames from it.
[0034] Optionally, the first state data is used to characterize whether the spatial relationship between the device and the user meets a first predetermined condition. The first predetermined condition may include, but is not limited to, the distance between the device and the user being outside the optimal shooting range, or the device having a pitch or yaw deviation relative to the user's viewing angle. This first state data can be collected by a time-of-flight sensor (ToF), depth camera, gyroscope, or accelerometer built into the terminal device. For example, the user's height data and the actual distance between the device and the user can be obtained through a ToF sensor, or the tilt angle of the device can be obtained through a gyroscope, thereby determining whether the spatial relationship meets the first predetermined condition.
[0035] Optionally, the second state data is used to characterize whether the ambient lighting meets a second predetermined condition. The second predetermined condition may include overall ambient light being too dark or too bright, or local areas in the original image (such as key detection areas like the user's limbs or face) being overexposed or underexposed. This second state data can be obtained by acquiring ambient light intensity data through the ambient light sensor of the terminal device, or by obtaining local light intensity data through image analysis of the original image (e.g., brightness histogram analysis, local area grayscale value statistics), thereby determining whether the current lighting meets the second predetermined condition.
[0036] In some embodiments, the original image, first state data, and second state data can be continuously and in real-time acquired during detection to provide input for subsequent automatic adjustment operations. The frequency of acquiring the above data can be dynamically configured according to the requirements of the detection process; for example, a higher sampling frequency can be used in stages with rapid changes in motion, while a lower sampling frequency can be used in static detection stages.
[0037] In step 202, based on the first state data, a first adjustment operation is performed on the original image or the user; the first adjustment operation includes at least one of performing a distance guidance operation on the user, a perspective correction operation on the original image, or a zoom operation on the original image.
[0038] In this step, one or more first adjustment operations can be performed on the original image or the user based on the first state data obtained in step 201. The first state data characterizes whether the spatial relationship between the device and the user meets a first predetermined condition, such as abnormal distance or abnormal viewpoint. These abnormalities can cause the user's body proportions, limb integrity, or pose clarity in the original image to fail to meet detection requirements. Therefore, the purpose of the first adjustment operation is to actively correct or compensate for these spatial relationship abnormalities, so that subsequent detection can be performed based on images of higher quality.
[0039] Optionally, the distance guidance operation is user-oriented and used to handle situations where the actual distance between the device and the user is inconsistent with the ideal shooting distance. When the first state data indicates that the current distance is too close or too far, this disclosure can prompt the user to move forward, backward, left, or right by generating a visual positioning marker (e.g., rendering a virtual reference frame in the display interface) or outputting voice guidance, thereby allowing the user's position to fall within the optimal imaging range. This operation does not directly change the image data, but rather guides the user to actively adjust their spatial position through interactive means.
[0040] Optionally, the perspective correction operation targets the original image and is used to address the perspective deviation of the device relative to the user. For example, when the device is placed on a table or the ground, excessive elevation or depression angles may cause perspective distortion of the user in the image (such as legs appearing too short or head too large). In this case, based on the tilt angle information in the first state data, the user can be guided to manually straighten the device through voice or visual prompts. Alternatively, if the deviation is small, geometric transformations (such as perspective transformations) can be applied to the original image to correct the distortion, making the corrected image more consistent with the observation effect of a frontal view.
[0041] Optionally, the zoom operation is applied to the original image to handle scenarios where distance deviations cannot be quickly eliminated through simple guidance. For example, when a user cannot move to the optimal distance due to environmental limitations, digital zoom processing can be performed on the original image based on the ratio of the current distance to the ideal distance. This involves scaling and cropping the image to ensure that the user's size proportions in the image meet the requirements of the detection window without changing the user's actual position. The zoom operation can be used in conjunction with other adjustment operations as a supplement to or alternative to distance guidance.
[0042] In some cases, the first adjustment operation may include any one of distance guidance, perspective correction, and zoom operations, or two or all three simultaneously. The specific operations selected, and whether they are performed on the user or the original image, can be determined based on the specific type and severity of the spatial relationship reflected in the first state data, whether it meets the first predetermined condition, and the image quality requirements of the current detection stage. For example, when only distance deviation exists and the perspective is normal, only distance guidance can be performed; when both distance deviation and perspective deviation exist, multiple operations can be performed in parallel or sequentially. Through the above automatic adjustment, this disclosure can dynamically maintain a suitable spatial relationship between the user and the device during the detection process, reducing human intervention and improving detection efficiency and accuracy.
[0043] In step 203, based on the second state data, a second adjustment operation is performed on the original image to adjust the original image from the current ambient lighting state to the target lighting state.
[0044] In this step, a second adjustment operation can be automatically performed based on the second state data obtained in step 201. The second state data indicates whether the ambient lighting meets a second predetermined condition, such as overall ambient light being too dark or too bright, or local areas in the original image being overexposed or too dark. These anomalies can cause the contrast, brightness, or detail clarity of the original image to fail to meet detection requirements, thereby affecting the accuracy of subsequent recognition of areas such as the user's limbs and face. Therefore, the purpose of the second adjustment operation is to compensate for ambient light in the original image, that is, to adjust it from the current ambient lighting state to a suitable target lighting state for detection, in order to improve the visual quality and information recognizability of the image.
[0045] Optionally, the ambient light compensation can be applied to the entire original image to change the overall brightness state. For example, when the second state data indicates that the ambient light intensity is lower than a preset threshold, the overall brightness of the image can be increased; when the ambient light intensity is too high, the overall brightness can be appropriately reduced so that the image grayscale distribution falls within a suitable range.
[0046] Optionally, the ambient light compensation can also be performed on local areas of the original image to adjust the lighting conditions of those local areas. For example, if key areas such as the user's arm or face are overexposed in the original image, while the background area is properly lit, brightness attenuation or contrast adjustment can be applied only to that local area without affecting other areas. Through local compensation, details of key detection areas can be highlighted without compromising the overall image's naturalness.
[0047] Optionally, based on the type of illumination mismatch reflected in the second-state data, an appropriate strategy can be selected from multiple preset lighting compensation strategies to perform ambient light compensation. Each strategy corresponds to a set of adjustment parameters from the current state to the target state. Through the above adjustments, this disclosure can dynamically improve the illumination conditions of the original image during the detection process, realizing the transformation from the current ambient light state to the target illumination state, and providing clearer image evidence for subsequent detection.
[0048] In step 204, the original image is displayed based on the adjustments made by the first adjustment operation and / or the second adjustment operation.
[0049] In this step, the original image, after being processed by the aforementioned first adjustment operation and / or second adjustment operation, can be presented on the display interface of the terminal device.
[0050] Optionally, when only the first adjustment operation is performed, the original image processed by the first adjustment operation is displayed; when only the second adjustment operation is performed, the original image processed by the second adjustment operation is displayed. When both the first and second adjustment operations are performed, they can be applied sequentially to the same frame of the original image. Specifically, the first adjustment operation can be performed first to perform distance guidance, viewpoint correction, and / or zoom processing on the original image, followed by the second adjustment operation for ambient light compensation; alternatively, the second adjustment operation can be performed first for ambient light compensation, followed by the first adjustment operation. When neither is performed, the original image is displayed directly. Regardless of the scenario, the final displayed image is a single frame, not multiple different display results.
[0051] After the adjustments in steps 202 and 203, the imaging quality problems caused by spatial discrepancies or ambient lighting in the original image have been improved or eliminated. At this point, displaying the adjusted original image can provide users with a clearer, more proportionate, and moderately bright detection screen.
[0052] Optionally, the display can be performed in real time, meaning it is displayed immediately after each frame of the original image is adjusted, allowing the user to observe the current detection screen and the effects of various adjustment operations in real time. For example, when the system performs distance guidance and renders a virtual footprint positioning frame, the user can simultaneously see the positioning frame and their adjusted self-image on the display interface, thus completing the positioning adjustment according to the guidance.
[0053] Optionally, the display can also serve as the output of the final detection result. That is, after the detection process is completed, the adjusted original image (e.g., an image containing the correct pose and good lighting) is saved for subsequent detection. By displaying based on the adjusted original image, this disclosure ensures that the user sees an optimized image, thereby improving the interactive experience and the reliability of the detection results.
[0054] In some embodiments, the method 200 may further include the following steps: During the detection process, the first state data and the second state data are continuously acquired, and the first adjustment operation and the second adjustment operation are repeatedly executed.
[0055] In this step, the first and second adjustment operations performed in the aforementioned steps are not one-time operations, but are repeatedly performed throughout the detection period.
[0056] Optionally, during the detection process, first state data and second state data can be continuously acquired, that is, real-time monitoring of changes in the spatial relationship between the device and the user and changes in ambient lighting conditions.
[0057] Optionally, whenever an abnormality in spatial relationships or ambient lighting is detected, the first adjustment operation and / or the second adjustment operation can be repeated based on the latest acquired state data to ensure that the original image is always in the best acquisition quality throughout the detection process.
[0058] Through continuous monitoring and repeated adjustments, this disclosure can adaptively maintain a suitable spatial relationship between the device and the user, as well as good lighting conditions, in a dynamically changing detection environment, thereby providing stable and reliable image data for subsequent detection.
[0059] In some embodiments, the distance guidance operation in step 202 above may further include the following steps: Optionally, the first state data includes the user's height data during the detection period and the actual distance between the terminal device and the user.
[0060] In step 2021, the target shooting distance is calculated based on the height data, and distance guidance data is generated based on the deviation between the target shooting distance and the actual distance.
[0061] In this step, a suitable distance range can be determined based on the user's height data. This range ensures that the user's entire body or upper body appears in a suitable proportion within the device's shooting window. Specifically, the user's height data can be collected in real time using a front-facing camera combined with a depth sensor. A skeletal keypoint detection algorithm is then used to identify the coordinates of multiple key points from the user's head to feet, calculating the height proportion (e.g., the proportion of the legs to the whole body). Then, based on a pre-defined height-distance mapping model, the target shooting distance for the current user can be calculated by substituting the height data. This mapping model can be a linear regression model trained on diverse body type datasets, capable of outputting a highly accurate recommended distance range.
[0062] Then, the target shooting distance is compared with the actual distance obtained in step 201, and the difference between the two is calculated. This difference can reflect the degree of deviation of the user's current position from the ideal position. For example, a large difference indicates that the user is standing too close or too far away.
[0063] Based on this deviation, distance guidance data can be generated. This distance guidance data can include the direction of the deviation (too close or too far) and the magnitude of the deviation, so that it can be used to generate visual positioning markers or voice prompts later. In this way, this disclosure can dynamically determine the target shooting distance based on the user's individual height characteristics and body proportions, and quantify the deviation of the user's current position, providing a quantitative basis for subsequent guidance operations.
[0064] In step 2022, based on the distance guidance data, a virtual footprint positioning frame is rendered in the display interface of the terminal device; and / or, Based on the distance guidance data, a first voice command is output to guide the user to move to the target location corresponding to the target shooting distance.
[0065] In this step, based on the distance guidance data generated in step 2021, positioning guidance can be provided to the user through visual and / or auditory means, so that the user can move quickly and accurately to the target shooting distance.
[0066] Figure 3 A schematic diagram of a virtual footprint positioning frame according to an embodiment of the present disclosure is shown.
[0067] Optionally, such as Figure 3 As shown, a virtual footprint positioning frame can be rendered on the terminal device's display interface based on the distance guidance data. This virtual footprint positioning frame can be represented as a pair of semi-transparent footprint graphics, a set of colored outlines, or other visual elements with positional indication, rendered at a specific location on the display interface. The size of the positioning frame can be dynamically adjusted according to the user's height ratio, for example, set to a certain proportion of the viewport width. The position of the positioning frame can be calculated through spatial mapping: determining the coordinates of the optimal standing center point based on the viewport center. By observing the virtual footprint positioning frame, the user can adjust their stance by moving their feet to align with the markings within the frame. The position of the virtual footprint positioning frame can be pre-determined based on the ideal imaging area corresponding to the target shooting distance, or its rendering position on the interface can be dynamically adjusted based on the magnitude and direction of the deviation in the distance guidance data. Furthermore, for different detection actions (e.g., standing, squatting, arm extension, etc.), the position of the virtual footprint positioning frame 301 can be dynamically adjusted accordingly to ensure that the user's limb movements always remain within the effective range of the detection viewport. Through this augmented reality-style visual guidance, the user can intuitively understand the direction and range of movement required.
[0068] Optionally, based on the distance guidance data, a first voice command can be output to guide the user to move to the target position corresponding to the target shooting distance. This first voice command may include a specific direction of movement (such as forward, backward, left, or right) and a distance to move, for example, "Please move forward a short distance" or "Please move backward half a step." Through voice guidance, users can adjust their position without looking at the screen, making it suitable for scenarios where it is inconvenient to view the screen or where vision is limited.
[0069] In some embodiments, the two guidance methods described above can be used selectively or simultaneously. For example, when the environment is noisy, visual guidance can be primarily provided using a virtual footprint positioning frame; when the user's attention is distracted, both voice commands and visual markers can be used for dual guidance. Through the above guidance operations, this disclosure can effectively shorten the user's positioning adjustment time and improve detection efficiency.
[0070] In some embodiments, the method 200 may further include the following steps: In step 205, a virtual footprint positioning frame is rendered in the display interface of the terminal device; the virtual footprint positioning frame is used to guide the user through the distance.
[0071] In this step, a virtual footprint positioning frame can be generated on the terminal device's display interface. This virtual footprint positioning frame is a visual implementation of distance guidance operations, used to provide users with a positioning reference marker. By observing this positioning frame, users can intuitively understand the target location they need to move to, thus facilitating the completion of the distance guidance operation. This positioning frame can be represented as a footprint-shaped graphic, a semi-transparent outline, or other visual elements with clear location indication.
[0072] Optionally, the virtual footprint positioning frame is used to guide the user's distance movement. By observing the positioning frame, the user can move their feet to align with the markers within the frame, thus adjusting their position and bringing the actual distance within the target shooting range. Simultaneously, the positioning frame can be used in conjunction with a first voice command: the voice tells the user the direction of movement, while the positioning frame provides the specific target position; both work together to guide the user.
[0073] Through the above rendering, this disclosure can provide users with intuitive and accurate positioning guidance, reduce the difficulty and time of distance adjustment, and improve detection efficiency.
[0074] In step 206, obtain the action type required for the current detection step and / or the user's leg proportion.
[0075] In this step, we can further obtain auxiliary information related to the detection process in order to make adaptive adjustments to the aforementioned guiding elements such as the virtual footprint positioning frame.
[0076] Optionally, the action type required for the current detection stage can be acquired. The action type may include different detection actions such as standing, squatting, extending arms, and turning. Different action types have different requirements regarding the user's position and limb coverage within the frame.
[0077] Optionally, the user's leg proportions can be obtained. The leg proportions can be calculated from the original image using a skeletal keypoint detection algorithm, such as identifying the coordinates of key points like the hip, knee, and ankle, thereby determining the proportion of the legs to the total body height. This proportion can be used to more accurately locate the rendering position of the virtual footprint positioning box on the display interface.
[0078] By obtaining the above information, this disclosure can provide a basis for subsequently adjusting the rendering position and size of the virtual footprint positioning frame.
[0079] In step 207, adjustment data for the virtual footprint positioning frame is generated based on the action type and / or the leg proportion; the adjustment data is used to indicate that the virtual footprint positioning frame is adjusted from its current state to a target state that matches the action type and / or leg proportion.
[0080] In this step, based on the action type and / or leg proportions obtained in step 206, relevant data for adjusting the virtual footprint positioning frame can be generated. This adjustment data is used to instruct the virtual footprint positioning frame to be adjusted from its current state (i.e., its position and / or size before adjustment) to a target state (i.e., its position and / or size after adjustment) that matches the action type and / or leg proportions. This adjustment data may include the rendering position of the positioning frame in the display interface (e.g., horizontal coordinate offset, vertical coordinate offset) and the dimensions of the positioning frame (e.g., width, height, or proportion of the viewport).
[0081] Optionally, when the action type required for the current detection stage is obtained, the adjustment data of the virtual footprint positioning frame can be determined based on that action type. For example, if the current detection stage requires the user to perform a standing full-body detection, the virtual footprint positioning frame can be rendered in the center of the bottom of the screen so that the user's entire body can fall completely into the viewport after standing firmly on both feet. If the current detection stage requires the user to perform a squatting detection, since the user's center of gravity shifts downward during the squatting action, the position of the virtual footprint positioning frame can be adjusted accordingly to the lower part of the screen, or the vertical size of the positioning frame can be appropriately reduced to match the actual position of the feet in the screen when squatting. If the current detection stage requires the user to perform an arm extension detection, the positioning frame can be moved upward or its size adjusted while keeping the foot position unchanged to ensure that the user's extended arms do not exceed the upper boundary of the screen. By dynamically generating adjustment data based on the action type, the virtual footprint positioning frame can be made more suitable for the actual needs of different detection actions.
[0082] Optionally, when the user's leg proportions are obtained, adjustment data for the virtual footprint positioning frame can be generated based on these proportions. Leg proportions refer to the percentage of a user's leg length to their total height. Different users have significantly different leg proportions, which affects their optimal center position in the image. For example, for users with larger leg proportions, their feet are positioned lower than their heads, so the rendering position of the virtual footprint positioning frame can be appropriately shifted downwards in the image; while for users with smaller leg proportions, the positioning frame can be moved relatively upwards. By fine-tuning the vertical position of the positioning frame based on the leg proportions, users with different body proportions can maintain a relatively consistent proportion of their entire body or upper body in the image after standing according to the positioning frame.
[0083] In some embodiments, adjustment data can be generated simultaneously based on the action type and leg proportions. For example, in the squat detection process, the position of the positioning frame is first lowered by a baseline offset based on the action type, and then the offset is personalized based on the user's leg proportions, so that the final generated virtual footprint positioning frame can not only meet the spatial requirements of the detected action, but also adapt to the individual body characteristics of different users.
[0084] By generating the aforementioned adjustment data, this disclosure can provide a quantitative basis for subsequently adjusting the rendering position and / or size of the virtual footprint positioning frame, thereby achieving more accurate and personalized positioning guidance.
[0085] In step 208, the rendering position and / or size of the virtual footprint positioning frame in the display interface of the terminal device are adjusted according to the adjustment data.
[0086] In this step, the rendering position and / or size of the virtual footprint positioning frame in the display interface can be adjusted based on the adjustment data generated in step 207, so that the positioning frame can adapt to the needs of the current detection process and the individual characteristics of the user.
[0087] Optionally, the rendering position of the virtual footprint positioning box can be adjusted based on the adjustment data. For example, if the adjustment data indicates that the positioning box needs to be shifted downwards on the screen, the vertical coordinate of the positioning box can be decreased (assuming the origin is at the top left corner), making the positioning box closer to the bottom of the screen; if the adjustment data indicates that the positioning box needs to be moved upwards, the vertical coordinate can be increased. The magnitude of the position adjustment can be determined based on the type of action (such as shifting it downwards by a baseline amount when squatting) or the proportion of the legs (such as shifting it further downwards when the leg proportion is large).
[0088] Optionally, the size of the virtual footprint positioning frame can be adjusted based on the adjustment data. For example, if the adjustment data indicates that a larger standing area is needed for the current detection stage (such as standing posture detection), the width and height of the positioning frame can be appropriately enlarged; if a more compact standing position is needed (such as squatting posture detection), the positioning frame can be reduced accordingly. The size can also be fine-tuned according to the user's leg proportions so that the positioning frame visually matches the size of the user's feet.
[0089] Through the above adjustments, the virtual footprint positioning frame can always present as a suitable guide marker under different detection stages and different user conditions, thereby helping users to complete the positioning more accurately.
[0090] In some embodiments, the zoom operation in step 202 above may further include the following steps: In step 2023, in response to the deviation between the actual distance and the target shooting distance lasting for a period of time exceeding a preset duration, the digital zoom ratio is determined based on the deviation.
[0091] In this step, it can be determined whether the user can actively adjust their position. It should be noted that this step is not necessarily executed after step 2022. In other words, the distance guidance operation (step 2022) and the zoom operation (step 2023) can be two independent or parallel adjustment strategies. Regardless of whether the virtual footprint positioning frame rendering or voice command output in step 2022 is performed, as long as the deviation between the actual distance and the target shooting distance persists and exceeds a preset duration, this disclosure can initiate step 2023 to determine the digital zoom ratio based on the deviation.
[0092] Specifically, a preset duration can be set (e.g., several seconds, which can be dynamically configured according to the real-time requirements of the detection process). When the deviation between the actual distance and the target shooting distance persists, and the duration of this deviation exceeds the preset duration, it indicates that the user may have failed to eliminate the distance deviation through physical movement within a reasonable time due to reasons such as limited space, inconvenient physical movement, or difficulty in receiving or understanding guidance information.
[0093] At this point, this disclosure can determine the digital zoom ratio based on the magnitude of the current deviation. Generally, the larger the deviation, the farther the user's actual position deviates from the ideal position, and the greater the required zoom ratio; the smaller the deviation, the smaller the zoom ratio can be used. The digital zoom ratio can be a magnification greater than 1, used to zoom in when the user is far away; or it can be a reduction magnification less than 1, used to expand the field of view when the user is close. By determining a suitable digital zoom ratio, a basis is provided for subsequent zoom processing of the original image, thereby ensuring that the user's image size in the image meets the detection requirements without changing the user's actual position. This method is particularly convenient for users who have difficulty adjusting their position using conventional interaction methods due to physical limitations or cognitive abilities, allowing them to complete the detection without forced movement and eliminating additional difficulties caused by interaction barriers.
[0094] In step 2024, the original image is zoomed according to the digital zoom ratio.
[0095] In this step, the actual zoom operation can be performed on the original image based on the digital zoom ratio determined in step 2023. Zoom processing can involve scaling the image to change the aspect ratio of the user's image within the frame.
[0096] Specifically, when the digital zoom ratio is greater than 1, the original image can be magnified, resulting in a correspondingly larger image of the user. Conversely, when the digital zoom ratio is less than 1, the original image can be reduced, resulting in a correspondingly smaller image of the user. Zoom processing can employ interpolation algorithms (such as bilinear or bicubic interpolation) to generate the scaled image, preserving image smoothness and detail clarity. Simultaneously, to compensate for potential sharpness loss due to magnification, the zoomed image can be sharpened (e.g., by analyzing high-frequency components). If the sharpness does not meet requirements, deep learning image super-resolution algorithms can be applied to enhance the image and improve detail.
[0097] Through the above zoom processing, when the user's actual position does not reach the target shooting distance, the imaging ratio imbalance caused by the distance deviation can be compensated at the image level, so that subsequent detection can be carried out based on the user image of appropriate size.
[0098] In some embodiments, the perspective correction operation in step 202 above may further include the following steps: Optionally, the first state data includes the viewing angle deviation data of the terminal device relative to the user.
[0099] In step 2025, in response to the viewpoint deviation data being greater than a preset angle threshold, a second voice command is output and a straightening prompt is generated to guide the user to manually adjust the placement angle of the terminal device.
[0100] Figure 4 A schematic diagram of the mobile phone tilt angle according to an embodiment of the present disclosure is shown.
[0101] like Figure 4 As shown, when a terminal device (such as a mobile phone) is placed on a table or the ground for full-body detection, ideally, the device's pitch angle should be within a suitable range, such as... Figure 4 As shown in the left image, taking a full-body standing or squatting pose with a mobile phone placed on the ground as an example, when the phone is placed upright, the ground horizon is located at the bottom of the viewport in the image, the user's feet are fully visible, and the full-body pose is clearly presented in the detection viewport.
[0102] However, when the phone is tilted at a large angle (e.g., the bottom of the phone is raised by a stone, or the top of the phone is tilted backward), the ground level will shift upward to the top of the screen, such as... Figure 4 As shown in the right image, even if the user moves backward, their feet cannot enter the frame, and their entire body cannot be captured. Similarly, if the phone is tilted too far downwards (e.g., the top of the phone is tilted downwards), the ground horizon will shift downwards or even out of the frame, causing the user's head to extend beyond the top edge of the image, again preventing a full-body shot.
[0103] In this step, the three-axis rotational angular velocity and linear acceleration of the device can be collected in real time using the gyroscope and accelerometer built into the terminal device. A fusion algorithm (such as Kalman filtering) is then used to calculate the device's tilt angle, thus obtaining the viewing angle deviation data. When the detected viewing angle deviation data exceeds a preset angle threshold (e.g., a pitch deviation exceeding 10 degrees), it indicates that the current viewing angle cannot be fully corrected by the software algorithm, or that the corrected image will lose crucial limb information such as feet or head due to over-cropping. At this point, a second voice command can be output, such as "Please straighten your phone so that the camera is facing upwards and vertically positioned" or "Please align the bottom of your phone with the horizontal line on the ground." Simultaneously, a straightening prompt is generated on the display interface, such as overlaying a highlighted horizontal reference line and dynamically displaying the current horizontal line's position offset, guiding the user to rotate the device until the reference line coincides with the actual horizontal line on the ground.
[0104] Through this combination of voice and visual guidance, users can quickly and intuitively adjust the device to a suitable shooting angle, ensuring the effectiveness of full-body detection. For example, when a user hears the voice prompt and sees a red horizontal line appear at the top of the screen, they can simply adjust the phone's angle to move the line down to the bottom of the screen to correct the viewing angle. This method is particularly suitable for users unfamiliar with device operation, lowering the barrier to entry.
[0105] In step 2026, in response to the viewpoint deviation data being less than or equal to a preset angle threshold, the original image is subjected to perspective transformation processing to generate a viewpoint-corrected image.
[0106] In this step, when the viewing angle deviation data is within a relatively small range, the viewing angle can be automatically corrected by the image processing algorithm without the user having to manually adjust the device.
[0107] Optionally, the viewing angle deviation data may include the pitch angle deviation and / or yaw angle deviation of the terminal device. The preset angle thresholds can be set according to actual detection needs; for example, the pitch angle deviation threshold can be set to 10 degrees, and the yaw angle deviation threshold can be set to 8 degrees. When the detected viewing angle deviation data is less than or equal to the corresponding preset threshold, it indicates that the perspective distortion of the image is relatively mild. In this case, it can be directly corrected through perspective transformation, saving user operation time while ensuring the correction effect.
[0108] For example, suppose a user places their phone on the ground for a full-body standing posture detection. Due to the uneven ground, the phone has a slight tilt angle (e.g., 5 degrees). In this case, the user in the image will exhibit a slight trapezoidal distortion, appearing wider at the bottom and narrower at the top, meaning the feet appear larger and the head smaller, resulting in a slightly distorted overall proportion. Since 5 degrees is less than the preset threshold of 10 degrees, this disclosure does not require the user to manually straighten the phone; instead, it directly performs perspective transformation on the original image. Specifically, edge detection algorithms can be used to extract scene features and lines (such as horizontal lines on the ground or vertical lines on the walls), and then line detection methods can be applied to identify the horizontal reference lines in the image. Next, four key points in the image (e.g., the user's feet, head, and the edges of both sides of the body) are identified. By calculating the transformation relationship from the distorted trapezoid to a standard rectangle, the image is remapped to generate a perspective-corrected image. In the corrected image, the proportions of the user's feet and head are restored to normal, and the overall posture is closer to the realistic visual effect.
[0109] In addition, during the correction process, to avoid black areas without pixels or distortion at the edges of the corrected image, an adaptive cropping algorithm can be applied: first calculate the effective bounding rectangle of the corrected image, and then crop away the excess blank or distorted parts at the edges (e.g., remove a small number of pixels from the edges) to ensure that the user's limb movements are fully presented in the detection window.
[0110] For example, when detecting a user's upper limb extension movements, the phone may have a small yaw angle deviation, meaning it rotates about 4 degrees to the left or right in the horizontal direction. This causes one arm to appear shorter than the other in the image, affecting the judgment of the symmetry of the arm extensions. Since 4 degrees is less than a preset threshold (e.g., 8 degrees), this disclosure can also perform horizontal perspective correction on the image through perspective transformation. After correction, the visual lengths of the left and right arms in the image tend to be consistent, thus ensuring the accuracy of subsequent detection.
[0111] For example, when a user is seated for upper body detection, the phone may have a slight downward tilt deviation (e.g., 3 degrees), causing the user's head to appear slightly larger and the shoulders narrower in the image. Since the deviation is small, perspective transformation processing can be activated to correct the downward view to a frontal view, restoring the proportions of the head and shoulders to normal.
[0112] In some embodiments, to improve the real-time performance of perspective transformation processing, the graphics processing unit (GPU) of a mobile device can be used to accelerate the image transformation algorithm in hardware. For example, by integrating a lightweight mobile inference framework (such as TensorFlow Lite) to handle the remapping calculations in perspective transformation, the processing latency of a single frame image can be controlled within a very short time (e.g., less than 50 milliseconds), thereby ensuring a real-time interactive experience during the detection process.
[0113] Optionally, after perspective transformation and cropping, the sharpness of the image after perspective correction can be evaluated. For example, the sharpness of the image can be evaluated by calculating the gradient distribution of the image's grayscale changes. When the calculated sharpness value is lower than a preset sharpness threshold, it indicates that the corrected image may have edge blurring issues. In this case, further sharpening filtering (such as an unsharpening mask algorithm) can be applied. The basic principle of this algorithm is to subtract the blurred version from the original image to obtain the difference in edge details, and then add this difference back to the original image according to a certain proportion, thereby enhancing the edge and texture details in the image. Through this processing, the user's limb contours, joints, and other key parts can be more clearly distinguished in the image, which is beneficial for subsequent action recognition and feature detection.
[0114] By employing the aforementioned perspective correction technique, users are no longer required to physically adjust the phone's angle multiple times, significantly reducing the number of manual interventions (e.g., from an average of 3-5 times to 1 time), thereby improving detection efficiency and user experience. Simultaneously, the improved image quality after correction also leads to a significant increase in the accuracy of subsequent motion capture. This solution is particularly suitable for users unfamiliar with or with limited understanding of device operation, avoiding the confusion and frustration caused by repeated adjustments.
[0115] In some embodiments, the second adjustment operation in step 203 above may further include the following steps: Optionally, the second state data includes ambient light intensity data.
[0116] In step 2031, based on the preset intensity range to which the ambient light intensity data belongs, a first target strategy is determined from multiple preset light compensation strategies, and the original image is adjusted based on the first target strategy so that the original image is adjusted from the current brightness state to the target brightness state.
[0117] In this step, the ambient light intensity data can be range-determined based on the pre-defined light intensity range. The corresponding overall compensation strategy can then be selected based on the determination result. In this way, the brightness, contrast, exposure and other parameters of the entire original image can be uniformly adjusted to achieve the transformation of the original image from the current brightness state to the target brightness state.
[0118] Optionally, the preset intensity range can be divided according to the actual detection scenario and the light sensitivity of the terminal device. For example, the light intensity can be divided into three ranges: low light range (e.g., ambient light intensity below 100 lux), normal range (e.g., 100 to 800 lux), and overexposure range (e.g., above 800 lux). Each range corresponds to a different compensation strategy, and each strategy includes a set of parameters to adjust the image from the current brightness state to the target brightness state.
[0119] For example, in a certain detection scenario, the ambient light sensor of the terminal device collects light intensity in real time at a sampling frequency of 10Hz. When it detects that the current ambient light intensity is only 50 lux, which is lower than the preset minimum shooting light intensity requirement of 100 lux, it determines that the current shooting environment is too dark, that is, the current brightness state of the original image is too dark. At this time, the selected first target strategy may include the following operations: First, increase the brightness of the terminal device screen to 80% as an auxiliary light source to provide supplementary lighting for the user's face and limbs; second, increase the camera's sensitivity from the default ISO100 to ISO800 to enhance the image sensor's sensitivity to light; at the same time, extend the exposure time from the default 1 / 60 second to 1 / 30 second, so that the photosensitive element can accumulate more photons over a longer period of time. Furthermore, the exposure value (EV) can be optimized using the logarithmic formula for exposure value. This involves calculating the corresponding exposure compensation value based on the adjustment of ISO and exposure time, ensuring that the average brightness of the adjusted image falls within a suitable grayscale range (e.g., between 100 and 200 grayscale values). This adjusts the original image from its overly dark state to a target state of normal brightness. Through this overall adjustment, an image that was originally too dark due to insufficient light can become bright and clear, making details such as the user's limb contours and key detection areas, which were previously hidden in the shadows, visible.
[0120] For example, when the ambient light intensity data falls within the normal range (e.g., 100 to 800 lux), it is determined that the current lighting conditions are basically suitable, meaning the current brightness state is close to the target state. In this case, the first target strategy could be to turn off auxiliary lighting, use an automatic exposure algorithm, and only perform slight contrast stretching or color balance adjustments when necessary to maintain the natural look of the image, avoid unnecessary computational overhead, and keep the original image near the target brightness state.
[0121] For example, when the ambient light intensity data falls into the overexposure range (e.g., above 800 lux), the ambient light is determined to be too strong, which can easily lead to excessive overall brightness and loss of detail in the image; that is, the current brightness state of the original image is overexposed. In this case, the first target strategy may include: reducing the ISO to a lower level (e.g., ISO 100), shortening the exposure time, activating high dynamic range compositing mode (i.e., continuously acquiring and merging multiple frames of images with different exposures), or applying local tone mapping technology to suppress brightness overflow in highlight areas, so that the overall brightness of the image falls back from the overexposed state to the target state, and the overall brightness dynamic range of the image falls back to an appropriate level, preserving more detail in both bright and dark areas.
[0122] In some embodiments, the preset intensity range and its corresponding strategy parameters can be dynamically configurable. For example, the range boundaries and strategy parameters can be adjusted according to the characteristics of the terminal device's camera module and the specific requirements of the detection scene (such as the need for high contrast to highlight limb edges), thereby flexibly defining the standard for the target brightness state. Through the above processing, this disclosure can automatically adjust the original image as a whole under different lighting conditions, realizing the transformation from the current brightness state to the target brightness state, providing basic quality assurance for subsequent local lighting compensation or feature detection.
[0123] In some embodiments, the second adjustment operation in step 203 above may further include the following steps: In step 2032, at least one local region image of the user is obtained from the original image after the second adjustment operation.
[0124] In this step, one or more local area images of the user can be extracted from the image obtained after overall lighting compensation in step 2031, so that targeted lighting analysis and adjustment can be performed on these areas in the future.
[0125] Optionally, the local area may include the user's face, upper limbs (such as arms and palms), lower limbs (such as legs and feet), or torso. Different detection stages may focus on different local areas. For example, when detecting finger grasping ability, the hand area needs to be focused on; when detecting whole-body coordination, multiple limb areas may need to be monitored simultaneously.
[0126] For example, in the original image after overall adjustment, the user's overall brightness has been restored to a suitable range, but the arm area may be locally overexposed due to direct illumination from ambient light, resulting in the loss of skin texture or limb contour details in that area. In this case, a local image of the arm area can be extracted from the entire frame image using image segmentation algorithms (such as deep learning-based human keypoint detection models or traditional skin color segmentation methods).
[0127] For example, in steps that require detecting a user's hair color or head features, the user's head area can be segmented from the overall image to analyze whether the lighting in that area is uniform and whether there are areas that are too dark or too exposed.
[0128] For example, when assessing a user's lower limb flexibility (such as squatting), local images of the left and right legs can be obtained separately to assess bilateral symmetry and movement consistency.
[0129] In some embodiments, multiple local region images can be obtained from the overall adjusted original image. For example, multiple regions such as the head, torso, and limbs can be extracted simultaneously to form a set of local images. These local regions can be automatically located and cropped using a preset human body part detection model, or the type of region to be extracted can be dynamically selected based on the configuration information of the current detection stage.
[0130] Through the above operations, this disclosure can separate different body parts of the user from the overall image, providing input data for subsequent individual light intensity detection and refined light compensation of each local area.
[0131] In step 2033, local light intensity data corresponding to the local region image is obtained.
[0132] In this step, based on the local area image extracted in step 2032, the lighting conditions of that area can be independently analyzed and quantified to obtain the light intensity data of that local area. This data can reflect whether there are problems such as local overexposure, local underexposure, or uneven lighting in that area.
[0133] Optionally, local light intensity data can be obtained by analyzing the pixel grayscale distribution in a local area of the image. For example, the average grayscale value of all pixels in the local area can be calculated as an overall brightness index of the area; alternatively, a grayscale histogram of the area can be calculated to observe whether the distribution range of grayscale values is concentrated in bright areas (indicating overexposure) or dark areas (indicating underexposure).
[0134] Optionally, the local light intensity data may also include an indicator of the uniformity of illumination within the area. For example, the local area can be divided into multiple sub-blocks, the average brightness of each sub-block can be calculated, and then the variance or range of brightness between the sub-blocks can be statistically analyzed. If the variance is too large, it indicates that there is uneven brightness within the area.
[0135] For example, in step 2032, an image of the user's arm region was extracted. Analyzing this arm region revealed an average grayscale value of 220 (assuming 0 represents pure black and 255 pure white), significantly higher than the normal range (e.g., 100 to 180), indicating localized overexposure in this area. Furthermore, observing the grayscale histogram revealed a large concentration of pixels in the high grayscale range, while the low grayscale range contained almost no pixels, further confirming the overexposure phenomenon.
[0136] For example, an image of the user's facial region was extracted, and its average grayscale value was calculated to be 60, which is far below the normal range, indicating that the area is locally too dark. Further analysis divided the region into two sub-blocks, the left face and the right face. The average grayscale value of the left face was 55, and that of the right face was 65. The difference was small, indicating that the uniformity of illumination was acceptable, but the overall brightness was insufficient.
[0137] For example, when detecting a user's leg movements, two local images of the left and right legs were extracted. The average brightness of the two regions was calculated separately: 120 for the left leg and 125 for the right leg, both within the normal range and with little difference, indicating that the lighting conditions in the leg areas were good and no additional compensation was needed.
[0138] In some embodiments, local light intensity data can be obtained by performing more refined statistical analysis on local image regions, such as calculating the median gray level, the mode gray level, or extracting the area ratio of highlight and shadow regions in the image. This data can serve as the basis for subsequent determination of whether local lighting adjustment is needed and which adjustment strategy to choose.
[0139] Through the above operations, this disclosure can obtain the lighting conditions information of different body parts, thereby providing accurate input for subsequent targeted local lighting adjustments.
[0140] In step 2034, in response to the fact that the local area still has local overexposure or local underexposure, a second target strategy is determined from multiple preset light compensation strategies based on the local light intensity data, and the local area is individually adjusted based on the second target strategy so that the local area is adjusted from the current local brightness state to the target local brightness state.
[0141] In this step, based on the local light intensity data obtained in step 2033, it is determined whether there are any lighting abnormalities such as overexposure or underexposure in the local area. If so, according to the specific lighting conditions of the area, an applicable second target strategy is selected from multiple preset lighting compensation strategies. Then, independent lighting adjustment is performed only on the local area to adjust the local area from the current local brightness state to the target local brightness state without affecting other areas in the image.
[0142] Optionally, determining whether a local area is overexposed or underexposed can be based on a preset grayscale threshold. For example, when the average grayscale value of a local area is higher than a certain upper threshold (e.g., grayscale value 200), it is determined to be locally overexposed, meaning the current local brightness state is too bright; when the average grayscale value is lower than a certain lower threshold (e.g., grayscale value 80), it is determined to be locally underexposed, meaning the current local brightness state is too dark. A more accurate determination can also be made by combining the distribution characteristics of the grayscale histogram; for example, overexposed areas show obvious peaks in the high grayscale range while lacking in the low grayscale range, while underexposed areas show the opposite.
[0143] Optionally, multiple preset lighting compensation strategies may include: local brightness attenuation strategy (for overexposed areas), local brightness enhancement strategy (for underexposed areas), local contrast stretching strategy, local gamma correction strategy, local tone mapping strategy, etc. Each strategy can be set with different adjustment intensity parameters, such as light, medium, and heavy adjustment, corresponding to different target local brightness states. Based on the degree of overexposure or underexposure reflected in the local lighting intensity data, a second target strategy with the corresponding intensity can be selected to adjust the local area from its current state to the desired target state.
[0144] For example, suppose a local area image of the user's arm is extracted in step 2032, and step 2033 calculates the average grayscale value of this area to be 230. The grayscale histogram shows that the pixels are concentrated in the high-brightness area, indicating local overexposure. In this case, the second target strategy determined from multiple preset strategies can be a "local brightness attenuation strategy." The intensity is set to "medium," meaning that the brightness of the pixels in the arm area is appropriately reduced (e.g., the grayscale value of each pixel is reduced by a certain proportion), while maintaining the original texture contrast within the area. This adjusts the local area from its current overexposed brightness state to a target brightness state with clear details, allowing the previously lost skin texture or limb edges to reappear. This adjustment only applies to the arm area; other areas such as the background or face are unaffected.
[0145] For example, if the average grayscale value of a local area of a user's face is 55, it is determined to be locally too dark. In this case, the secondary objective strategy could be a combination of a "local brightness enhancement strategy" and a "local contrast stretching strategy." First, the overall brightness of the facial area is increased to restore the average grayscale value to a normal range (e.g., 120-150). Then, the contrast of this area is moderately stretched to make the facial features clearer, thus achieving a transition from an overly dark state to a normal brightness target state. After adjustment, facial features that were originally hidden in the shadows become visible, while other parts of the image (such as the background and clothing) remain unchanged.
[0146] For example, when detecting a user's leg movements, it was found that the average grayscale value of a local area on the left leg was 130 (normal), but there was a small overexposed area (grayscale value as high as 240) inside this area due to reflections, while the grayscale value of the outer shadow area was only 60. In this case, it can be determined that this local area is both overexposed and underexposed, and the lighting is extremely uneven. The corresponding second target strategy could be a "local tone mapping strategy," which involves high dynamic range compression processing on this area: compressing the highlights and raising the shadows, adjusting the entire local area from its current uneven brightness state to a target state of uniform brightness, ultimately making the light distribution of the entire area more uniform. After adjustment, the leg contour and texture details are clearly visible throughout the entire area.
[0147] In some embodiments, adjusting the lighting of a local area individually can be achieved using a mask-based processing approach: First, a mask of the same size as the original image is generated, with only the pixels corresponding to the area to be adjusted having non-zero values. Then, the adjustment algorithm is applied to the entire image, but the algorithm only updates the corresponding pixels based on the mask's position, leaving other areas unchanged. Alternatively, the extracted local area image can be processed directly, and then the processed area image can be pasted back into the corresponding position of the original image, while ensuring smooth edge blending to avoid obvious stitching artifacts. Regardless of the method used, the core objective is to adjust the local area from its current brightness state to the target brightness state.
[0148] By adjusting the local lighting as described above, this disclosure can accurately correct the problem of overexposure or underexposure in key detection areas caused by ambient light, and realize the conversion of the brightness state of the local area from the current to the target, without destroying the natural appearance of the overall image, thereby providing higher quality image basis for subsequent feature recognition and action detection.
[0149] Furthermore, through the combination of overall and local adjustments, the overall adjustment first performs uniform processing on the entire image. However, due to the influence of ambient light (such as direct sunlight through a window or local indoor light sources), key detection areas may still be overexposed or underexposed after the overall adjustment, and the overall adjustment cannot accurately repair such local problems. This disclosure further performs local adjustments to achieve the transformation of the brightness state of local areas from the current state to the target state, without compromising the overall natural appearance of the image, thereby providing higher-quality image data for subsequent feature recognition and action detection.
[0150] In some embodiments, step 204 above, which displays the original image adjusted by the first adjustment operation and / or the second adjustment operation, may further include the following steps: In step 2041, a causal relationship is determined between the types of spatial relationships that meet the first predetermined conditions and the types of ambient lighting that meet the second predetermined conditions.
[0151] In this step, correlation analysis can be performed on various types that may occur during the detection process. First, the internal causal relationships between different types of spatial relationships that meet the first predetermined condition are analyzed. Then, the causal relationships between different types of spatial relationships that meet the first predetermined condition and different types of ambient lighting that meet the second predetermined condition are further analyzed. By identifying these causal relationships, the execution order of subsequent adjustment operations can be arranged more rationally, avoiding repeated anomalies or adjustment failures due to improper adjustment order.
[0152] Optionally, the types of spatial relationships that meet the first predetermined conditions may include situations such as the distance between the equipment and the user being too close or too far, or the equipment's pitch or yaw angle deviation being too large. There may be a certain causal relationship between these types.
[0153] For example, when the device has a significant pitch deviation (such as an excessively high pitch angle), the user might subconsciously move backward to try and get their feet into the frame in order to ensure their entire body is in the shot. However, due to the excessive pitch angle, this movement is often inefficient or even ineffective, and may even further widen the discrepancy between the actual distance and the target shooting distance, causing or exacerbating distance anomalies. In this case, the viewpoint type is the cause, and the distance anomaly is the consequence. If the viewpoint is not corrected first and the user is blindly guided to move their position, the user may repeatedly adjust but still fail to meet the detection requirements, increasing user frustration and preparation time.
[0154] For example, when the distance between the device and the user is too close, although the anomaly of the viewpoint type itself is not directly caused by the anomaly of the distance type, the close distance may restrict the user's physical movement space. This could cause certain parts of the user's body (such as hands or head) to extend beyond the edge of the frame when performing certain actions (such as extending an arm or squatting). If the device itself already has a slight viewpoint deviation, this deviation will be amplified by the close distance, making the image distortion more obvious. Therefore, the anomaly of the distance type can exacerbate the impact of the anomaly of the viewpoint type on the detection results.
[0155] After analyzing the causal relationships between the various types of spatial relationships that meet the first predetermined condition, we can further analyze the causal relationships between these types and the various types of ambient lighting that meet the second predetermined condition.
[0156] For example, when the device is too close to the user, the user's body may block the main light source in the environment (such as a ceiling light or window), causing the user to appear in partial shadows or overall darkness in the image. In this case, there is a causal relationship between distance type and lighting type: the abnormal distance type is the antecedent, and the abnormal lighting type is the consequence. If ambient light compensation is applied directly without first adjusting the distance, it may result in over-compensation or under-compensation because the obstruction is not eliminated; conversely, if the user is guided to move to a suitable distance first, the lighting conditions will naturally improve after the obstruction is removed, and little or no light compensation may be needed.
[0157] For example, when the device's tilt angle is too large (e.g., the phone is tilted upwards), the camera's orientation may cause the lens to be directly pointed at a strong light source (such as a ceiling light or sunlight outside a window), resulting in localized overexposure or overall excessive brightness in the image. In this case, the abnormal viewing angle is the cause of the abnormal lighting. If the exposure parameters are blindly reduced without first correcting the viewing angle, other areas of the image may become too dark; however, by guiding the user to straighten the device, the lens is no longer directly facing the light source, the lighting conditions return to normal, and the overexposure problem is alleviated.
[0158] For example, when distance-related anomalies and viewpoint-related anomalies coexist, ambient lighting anomalies may be caused by one or both of them. For instance, a user being too close to the light source might block it, creating a shadow, while a viewpoint deviation might cause the lens to be directly facing another strong light source, resulting in overexposure. In this situation, the image may simultaneously contain both localized underexposure and localized overexposure. Following the causal order, the anomaly that fundamentally alleviates the lighting problem should be addressed first (e.g., correcting the viewpoint to resolve overexposure, then adjusting the distance to resolve shadow blocking).
[0159] In some embodiments, causal relationships between different types can be determined by preset causal rules or causal models trained based on historical data. For example, it can be predefined that if the device pitch angle deviation exceeds a certain threshold, and the deviation between the actual distance and the target shooting distance continues to increase, then the anomaly of the viewpoint type leads to the anomaly of the distance type; if the actual distance is less than a certain threshold of the target shooting distance and the shadow area of the user in the image exceeds a certain proportion, then the anomaly of the distance type leads to the anomaly of the lighting type.
[0160] By determining the above-mentioned multi-level causal relationships, this disclosure can identify the causal chains between various types of spatial relationships that meet the first predetermined condition, as well as the causal relationships between these types and various types of ambient lighting that meet the second predetermined condition. This lays the foundation for subsequent adjustment operations based on the principle that the preceding type takes precedence over the following type in the causal relationship, avoids ineffective or repeated adjustments, and improves the overall adjustment efficiency.
[0161] In step 2042, in response to the existence of multiple types with causal relationships, the first adjustment operation and / or the second adjustment operation corresponding to each type are executed sequentially in the order that the type as the cause takes precedence over the type as the result in the causal relationship.
[0162] In this step, when step 2041 determines that multiple types have a causal relationship, the type acting as the cause can be processed first, followed by the type acting as the result, according to the order in the causal chain. Here, "the type acting as the cause" refers to the type that leads to the occurrence of another type in the causal relationship, and "the type acting as the result" refers to the type that is caused by another type in the causal relationship. The order between the two is determined by causal logic. By performing the adjustment operation in this order, the problem can be eliminated at its root or the complexity of subsequent adjustments can be reduced, avoiding poor adjustment results or the need for repeated adjustments due to improper order.
[0163] Optionally, multiple types with causal relationships can all belong to types whose spatial relationships meet the first predetermined condition (e.g., viewpoint type causes distance type anomaly), or they can simultaneously include spatial relationship types and ambient lighting types (e.g., distance type causes ambient lighting type anomaly, or viewpoint type causes ambient lighting type anomaly). In either case, the corresponding adjustment operations are performed sequentially in the order of priority for the type as the cause and subsequent type as the result.
[0164] For example, suppose step 2041 determines that the viewpoint type (excessive device pitch angle deviation) is the cause of the distance type anomaly (the user's actual distance deviates from the target shooting distance). That is, due to the phone's tilted angle, even if the user moves backward, they cannot get their feet into the frame; instead, they may move further away. In this case, the causal chain is: viewpoint type (cause) → distance type (result). According to this step, the adjustment operation corresponding to the viewpoint type should be performed first (e.g., outputting a voice command and generating a ground level alignment prompt in step 2025 to guide the user to manually adjust the device's placement angle). After the viewpoint is corrected, the distance type should be evaluated to see if it still exists. Usually, after viewpoint correction, the user can fall within the target shooting distance range without significant movement. At this point, performing distance guidance operations (such as rendering a virtual footprint positioning frame or outputting a voice command) will be more effective.
[0165] For example, suppose step 2041 determines that the distance type (device too close to user) is the cause of the abnormal ambient lighting type (user's body blocking the light source, creating a local shadow). The causal chain is: distance type (cause) → ambient lighting type (result). According to this step, the adjustment operation corresponding to the distance type should be performed first (e.g., distance guidance operation or zoom operation). After the user moves to a suitable distance, the obstruction is removed, and the local shadow caused by the obstruction may disappear naturally or be significantly reduced. At this time, ambient light compensation operation (e.g., overall adjustment or local light adjustment) can be performed as needed, or even unnecessary compensation can be avoided.
[0166] In some embodiments, step 2042 above, which sequentially executes the first adjustment operation and / or the second adjustment operation corresponding to each type, may further include the following steps: the types of spatial relationships that meet the first predetermined conditions include viewpoint types and distance types, and the types of ambient lighting that meet the second predetermined conditions include lighting types; the causal relationship includes the viewpoint type as a cause type leading to the distance type as a result type, and the distance type as a cause type leading to the lighting type as a result type; the sequential execution of the first adjustment operation and / or the second adjustment operation corresponding to each type includes: sequentially executing the first adjustment operation corresponding to the viewpoint type, the first adjustment operation corresponding to the distance type, and the second adjustment operation corresponding to the lighting type.
[0167] In this step, when three or more types form a causal chain, such as viewpoint type (cause) causing an anomaly in distance type (intermediate type), and distance type causing an anomaly in ambient lighting type (result), the corresponding adjustment operations should be performed sequentially in the order of viewpoint type → distance type → ambient lighting type.
[0168] Specifically, the first adjustment operation corresponding to the viewing angle type is executed. The viewing angle type involves the pitch or yaw angle deviation of the terminal device relative to the user. Correspondingly, the adjustment operation may include: when the viewing angle deviation data is greater than a preset angle threshold, outputting a voice command and generating a straightening prompt to guide the user to manually adjust the placement angle of the terminal device; when the viewing angle deviation data is less than or equal to the preset angle threshold, performing perspective transformation processing on the original image to generate a perspective-corrected image. Through the above operations, the perspective distortion problem caused by the device angle deviation is eliminated or alleviated.
[0169] Then, the first adjustment operation corresponding to the distance type is executed. The distance type involves the deviation between the actual distance between the terminal device and the user and the target shooting distance. Correspondingly, the adjustment operation may include: calculating the target shooting distance based on the user's height data, and generating distance guidance data based on the deviation between the target shooting distance and the actual distance; rendering a virtual footprint positioning frame on the terminal device's display interface according to the distance guidance data, and / or outputting a first voice command to guide the user to move to the target position corresponding to the target shooting distance. If the user fails to move to the position within a preset time, the digital zoom ratio can also be determined based on the deviation, and the original image can be zoomed. Through the above operations, the problem of user image ratio imbalance caused by improper positioning is eliminated or alleviated.
[0170] Finally, a second adjustment operation corresponding to the lighting type is performed. The lighting type relates to whether the ambient lighting meets a second predetermined condition. Correspondingly, the adjustment operation may include: determining a first target strategy from multiple preset lighting compensation strategies based on the preset intensity range to which the ambient light intensity data belongs, and adjusting the original image as a whole to adjust the original image from the current brightness state to the target brightness state; furthermore, if local areas are still overexposed or underexposed, the lighting of the local areas can be further adjusted individually to adjust the local areas from the current local brightness state to the target local brightness state. Through the above operations, the problems of unsuitable image brightness or loss of detail caused by abnormal ambient lighting are eliminated or alleviated.
[0171] First, the device's viewing angle is corrected, then the user's position is guided, and finally, ambient light compensation is performed depending on the lighting conditions. This progressive adjustment avoids ineffective or repetitive adjustments due to unresolved anomalies of the underlying cause. For example, if distance guidance or zooming is performed first, image distortion caused by device viewing angle deviation will still exist, resulting in distorted image proportions and limited subsequent adjustments. However, if the viewing angle is corrected first and then the position is guided, the user can fall within the target shooting distance range without significant movement after the viewing angle is corrected, making distance guidance more effective. Similarly, after both viewing angle and distance are adjusted, lighting compensation is performed on the image with optimized spatial relationships, resulting in more accurate compensation without over-compensation. Through this progressive adjustment, this disclosure eliminates problems at their root, reduces unnecessary repetitive adjustments, and improves overall adjustment efficiency.
[0172] In some embodiments, if there is a parallel relationship between multiple types (i.e., no direct causal relationship), the adjustment operations can be performed in parallel, or the execution order can be determined according to a preset priority (e.g., based on the degree of impact on detection quality). This step only constrains the case where a causal relationship exists, and does not particularly limit the case where there is no causal relationship.
[0173] By performing the adjustment operations in the order of causality as described above, this disclosure can improve adjustment efficiency, reduce unnecessary repetitive adjustments, make the user experience smoother, and at the same time ensure that the quality of the detected image is fundamentally improved.
[0174] In step 2043, the adjusted original image is displayed.
[0175] In this step, after determining the causal relationship and performing the corresponding adjustment operations for each type in step 2042, the original image processed by the above adjustment operations is presented on the display interface of the terminal device.
[0176] Optionally, when only the first adjustment operation corresponding to the spatial relationship type (e.g., viewpoint correction or zoom) is performed in step 2042 without performing the second adjustment operation corresponding to the ambient lighting type, this step displays the original image processed by the first adjustment operation. When only the second adjustment operation corresponding to the ambient lighting type is performed in step 2042 without performing the first adjustment operation, this step displays the original image processed by the second adjustment operation. When multiple adjustment operations corresponding to different types are performed sequentially according to the causal chain in step 2042, this step displays the original image processed by the aforementioned multiple adjustment operations sequentially. Ultimately, the same frame image is always displayed, rather than multiple different display results.
[0177] Optionally, the display can be performed in real time, meaning it is displayed immediately after each frame of the original image is adjusted, allowing the user to observe the current detection screen and the effects of various adjustment operations in real time. For example, when the system performs distance guidance and renders a virtual footprint positioning frame, the user can simultaneously see the positioning frame and their adjusted self-image on the display interface, thus completing the positioning adjustment according to the guidance.
[0178] Optionally, the display can also serve as the output of the final detection result, that is, after the detection process is completed, the adjusted original image (e.g., an image containing the correct pose and good lighting) is saved for subsequent detection.
[0179] By performing the adjustment operations in the order of cause and effect as described above and then displaying the result, this disclosure ensures that users see an optimized image, thereby improving the interactive experience and the reliability of the test results.
[0180] In some embodiments, the method 200 may further include the following steps: In step 211, the current detection step identifier is obtained; the detection step identifier is used to characterize the type of action to be detected in the current detection process.
[0181] In this step, a marker can be obtained to identify the current detection stage, namely the detection step identifier. This identifier reflects the specific action type that the user needs to perform during the current detection process, such as standing, squatting, extending arms, turning around, or grasping with fingers. Different action types have different requirements for image acquisition quality. For example, full-body standing posture detection requires the user's entire body to be fully captured in the image, while arm extension detection focuses more on the imaging proportion and edge sharpness of the upper limb area.
[0182] Optionally, the detection step identifier can be a preset number, name, or code, corresponding to different stages in the interview process. For example, in a skills talent AI interview system, the interview process can include multiple consecutive detection steps: the first step is full-body standing posture detection, the second step is squatting and standing motion detection, the third step is arm extension detection, and the fourth step is finger dexterity detection. Each step corresponds to a unique detection step identifier (such as step number 1, 2, 3, 4). Before or at the start of each detection step, the system can read the current detection step identifier from the process control module or configuration file.
[0183] For example, when the detection step is labeled "Full-body standing posture detection," it means the user needs to be captured in a standing posture so the system can assess the integrity of their limbs and overall coordination. In this case, the requirements for shooting distance and angle are relatively strict, ensuring the entire body, from head to toe, is within the frame. When the detection step is labeled "Squatting and standing detection," it means the user needs to complete a squatting and standing sequence. In this case, in addition to the full body being captured, special attention needs to be paid to the range of motion of the knee and hip joints; therefore, there are specific requirements for the clarity and proportion of the lower limbs in the image.
[0184] For example, when the detection step is labeled "arm extension detection," the user needs to extend both arms to the sides or forward to assess the flexibility and symmetry of the upper limbs. In this case, the image needs to fully encompass both arms, and the arms should occupy a sufficient pixel area in the image to identify joint angles and movement trajectories. When the detection step is labeled "finger grip detection," the user may need to demonstrate opening and closing or gripping actions. In this case, the image needs to show a close-up or magnified view of the hand area, placing higher demands on resolution and local illumination uniformity.
[0185] In some embodiments, the detection step identifier can be obtained through a predefined interview script or dynamic configuration. For example, a system administrator or interviewer can set different combinations of detection items according to job requirements. The system executes these items sequentially according to the order of the combination, and passes the corresponding detection step identifier to the corresponding image acquisition and adjustment module at the beginning of each stage.
[0186] By obtaining the detection step identifier, this disclosure can know the specific requirements of the current detection step for image quality, thereby providing a basis for dynamically adjusting the execution priority or weight of the first adjustment operation and the second adjustment operation according to the identifier.
[0187] In step 212, according to the detection step identifier, corresponding weights are configured for the first adjustment operation and the second adjustment operation, respectively.
[0188] In this step, based on the detection step identifier obtained in step 211 (i.e., the type of action to be detected in the current detection process), corresponding weights can be assigned to the first adjustment operation (including at least one of distance guidance operation, viewpoint correction operation, and zoom operation) and the second adjustment operation (ambient light compensation operation). The weight value can reflect the importance or priority of the adjustment operation in the current detection stage. The higher the weight, the greater the impact of the operation on the current detection quality, and it may need to be executed first or given more attention in resource allocation.
[0189] Optionally, the weights can be preset values (e.g., decimals between 0 and 1, or integers between 1 and 10), or multi-dimensional comprehensive scores. Different detection step identifiers correspond to different weight configuration schemes. For example, a mapping table can be pre-established to associate each detection step identifier with the first adjustment operation weight and the second adjustment operation weight.
[0190] For example, when the detection step is labeled "full-body standing posture detection," this step requires the user's entire body, from head to toe, to be clearly and completely presented in the image. In this case, spatial relationships (distance and viewing angle) have a critical impact on image quality: too close a distance will result in cropped feet, and a deviation in viewing angle will cause image distortion. In contrast, the impact of ambient lighting is relatively minor (as long as the brightness is basically uniform and the outlines are clear). Therefore, a higher weight (e.g., 0.8) can be assigned to the first adjustment operation, and a lower weight (e.g., 0.2) to the second adjustment operation. This means that in this detection step, the system should prioritize ensuring the correctness of spatial relationships, and only compensate for lighting when necessary.
[0191] For example, when the detection step is labeled "arm extension detection," this step requires a focus on observing the user's upper limb extension range and symmetry. In this case, distance and viewing angle are equally important (the arm needs to be fully visible in the frame without distortion), but local lighting conditions can significantly impact the recognition of hand details (such as whether the fingers are separated). Therefore, a moderate weight (e.g., 0.6) can be assigned to the first adjustment operation, and a moderate weight (e.g., 0.4) to the second adjustment operation, achieving a relatively balanced approach.
[0192] For example, when the detection step is labeled "local feature detection" (such as hair color, specific skin areas), this step has extremely high requirements for lighting uniformity and detail clarity, while the requirements for full-body distance and viewing angle are relatively lenient (only local areas need to be in the frame). In this case, a lower weight (e.g., 0.2) can be assigned to the first adjustment operation, and a higher weight (e.g., 0.8) can be assigned to the second adjustment operation, prioritizing ambient light compensation and local lighting optimization to ensure clear details in key areas.
[0193] In some embodiments, the weight configuration can be further refined to the individual sub-operations within the first adjustment operation. For example, in "full-body standing posture detection", the weight of the distance guidance operation can be higher than that of the viewpoint correction operation; in "squatting and standing detection", the weight of the viewpoint correction operation may be even higher, because the device pitch angle deviation during squatting and standing will seriously affect the capture of the squatting posture.
[0194] By dynamically configuring weights based on the detection step identifiers, this disclosure enables adjustment operations to better align with the actual needs of the current detection process, avoiding the waste of resources on unnecessary adjustments while ensuring that critical adjustment operations are prioritized.
[0195] In step 212, the adjustment operation to be performed first is determined according to the weight.
[0196] In this step, the first adjustment operation (including at least one of distance guidance operation, viewpoint correction operation, and zoom operation) and the second adjustment operation (ambient light compensation operation) can be prioritized based on the weights configured in step 211. This determines which adjustment operation or set of adjustment operations should be executed first in the current detection process. The larger the weight value, the higher the importance of the adjustment operation in the current detection process, and the more it should be executed first; the smaller the weight value, the more it can be postponed or temporarily not executed when resources are limited.
[0197] Optionally, a weight threshold can be set. Only when the weight exceeds the threshold will the corresponding adjustment operation be included in the priority execution queue. If multiple operations exceed the threshold, they are executed sequentially in descending order of weight. Alternatively, the weight values of the first adjustment operation and the second adjustment operation can be directly compared, and the operation with the higher weight is selected as the priority adjustment operation.
[0198] For example, continuing from the example in step 211, assuming the current detection step is labeled "full-body standing posture detection," according to the pre-configuration, the weight of the first adjustment operation is 0.8, and the weight of the second adjustment operation is 0.2. In this case, the weight of the first adjustment operation is significantly higher than that of the second adjustment operation. Therefore, this step can determine to prioritize the execution of the first adjustment operation, that is, to prioritize handling abnormal spatial relationships between the device and the user (such as improper distance or viewing angle deviation). After the spatial relationship is adjusted to a suitable range, the second adjustment operation is then performed for ambient light compensation, depending on the situation. If the overall brightness of the original image already meets the requirements after the spatial relationship adjustment, the second adjustment operation can even be skipped, thereby saving computational resources.
[0199] For example, the current detection step is labeled "local feature detection," which has extremely high requirements for illumination uniformity and detail clarity. The pre-configured weight for the first adjustment operation is 0.2, and the weight for the second adjustment operation is 0.8. In this case, the weight of the second adjustment operation is significantly higher than that of the first. Therefore, this step can prioritize the second adjustment operation, that is, prioritize compensating for ambient light (especially for individually adjusting overexposed or underexposed areas), ensuring that details in key detection areas are clearly visible. After the light compensation is completed, the spatial relationship is then evaluated to determine if adjustments are needed (e.g., whether the distance is appropriate, whether there is a viewing angle deviation). If the spatial relationship basically meets the detection requirements, the first adjustment operation can be temporarily postponed.
[0200] For example, when the weights of the first and second adjustment operations are equal or close (e.g., both 0.5), it can be determined that they are executed in parallel, that is, spatial relationship adjustment and ambient light compensation are performed simultaneously or alternately. Alternatively, the priority can be dynamically determined based on the severity of the current anomaly: if the spatial relationship deviation exceeds a certain threshold, the first adjustment operation is executed first; if the lighting anomaly exceeds a certain threshold, the second adjustment operation is executed first.
[0201] In some embodiments, once the priority adjustment operation is determined, the system can prioritize allocating computing and interaction resources to that operation during subsequent continuous monitoring (e.g., detecting the corresponding status data more frequently and triggering adjustment instructions more promptly), while operations with lower weight can have their detection frequency appropriately reduced or be triggered only when necessary.
[0202] By determining the priority of adjustment operations based on weights, this disclosure enables adjustment strategies to better align with the core needs of the current detection process, avoiding wasting time and resources on unimportant adjustments, thereby improving overall detection efficiency and user experience.
[0203] In another feasible embodiment, the display method provided in this disclosure can also be applied to limb function assessment and rehabilitation training monitoring scenarios. Taking the limb function assessment of post-stroke patients as an example, users (e.g., patients undergoing rehabilitation) capture and analyze the range of motion of the upper limbs through a mobile terminal device. First, image quality optimization is completed according to the distance guidance operation (calculating the target shooting distance based on height and generating a virtual footprint positioning frame and voice commands), perspective correction operation (graded processing: voice guidance to straighten the device when there is a large deviation, and automatic correction of perspective transformation when there is a small deviation), and ambient light compensation operation (overall brightness adjustment and local overexposure / underexposure correction) in the aforementioned embodiments, to ensure that the range of motion of the patient's upper body and arms can be clearly and completely captured in the detection window.
[0204] After image quality optimization, video streams of patients performing specific movements, such as elbow flexion and extension or shoulder abduction, are acquired. Upper limb keypoint coordinates, including the shoulder, elbow, and wrist, are extracted in real-time using a skeletal keypoint detection algorithm (e.g., a lightweight mobile model based on MediaPipe). Based on these coordinates, joint range of motion is calculated: the elbow flexion and extension angle is obtained by the angle between the shoulder-elbow vector and the elbow-wrist vector, specifically calculated by dividing the dot product of the two vectors by their magnitudes and then taking the inverse cosine. Further analysis of the angle's change over time is performed to extract the maximum range of motion, the time required to reach the maximum angle, and the smoothness of the angle change (quantified by calculating the variance of the first derivative of the angle sequence; a smaller variance indicates smoother movement). All outputs are objective numerical values and do not make any judgments regarding disease status or rehabilitation effects. These values can be exported for reference by rehabilitation therapists or patients.
[0205] In another application, for users with movement disorders (such as Parkinson's patients), the method disclosed herein can be used to monitor the degree of involuntary tremors or bradykinesia. After the aforementioned image quality optimization, the patient is guided to perform standard movements, such as extending the arm forward and holding it still for several seconds, or tapping the fingers at a fixed rhythm. For tremor monitoring, coordinate sequences of key hand points (such as fingertips or wrist joints) are extracted at a high sampling rate (e.g., 60 frames per second), and time-frequency analysis is performed on the horizontal or vertical displacement: the displacement signal is de-trended and then converted to the frequency domain using a Fast Fourier Transform, the power spectral density is calculated, and the peak frequency is identified; this peak frequency is the main periodic frequency of the tremor. For bradykinesia monitoring, the joint angular velocity or displacement velocity when the patient performs the finger-tapping action is analyzed. The motion velocity curve is obtained by dividing the Euclidean distance of the key point coordinates in adjacent frames by the time interval, and then the average, maximum, and standard deviation of the velocity are statistically analyzed. The above numerical indicators are output, without including disease severity grading or diagnostic conclusions.
[0206] In another application example, the method disclosed herein can be used for monitoring activity patterns during postoperative recovery. Taking a patient after knee replacement surgery as an example, after image quality optimization, the coordinates of the hip, knee, and ankle joints are identified through skeletal key point detection, and the knee flexion angle (the angle between the hip-knee vector and the knee-ankle vector) is calculated. An ideal reference angle trajectory can be generated (based on general rehabilitation guidelines or customized according to the patient's preoperative assessment data). The actual angle trajectory is dynamically time-normalized and compared with the reference trajectory, outputting the cumulative value of the angle deviation and the duration exceeding the preset safe angle range. If the patient's movement angle exceeds the safe threshold (e.g., the flexion angle exceeds 120 degrees), a voice or visual prompt "Please reduce the squatting range" is provided, serving as a real-time reminder. This prompt is only based on the preset safe range and does not involve diagnosis of the recovery status or prognosis.
[0207] This medical scenario embodiment shares the same core image quality optimization steps as the aforementioned interview scenario embodiment. The difference lies in the medical scenario's focus on extracting quantitative indicators of movement (such as joint angles, tremor frequency, movement smoothness, and velocity curves) and providing real-time movement standardization reminders. All of these indicators are objective numerical outputs and do not contain any diagnostic or treatment conclusions. Those skilled in the art will understand that the method disclosed herein can also be applied to other scenarios requiring dynamic adjustment of shooting parameters and extraction of limb movement parameters, such as sports training assistance, motor skill learning, and dance posture correction; these will not be elaborated upon here.
[0208] In another embodiment, the display method provided in this disclosure can also be applied to motion-sensing games and virtual reality interaction scenarios. Taking motion-sensing games (such as dance or fitness games) as an example, users use mobile terminal devices to capture body movements in real time and control game characters. First, image quality optimization is completed according to the distance guidance operation (calculating the target shooting distance based on the user's height and generating a virtual footprint positioning frame and voice commands), perspective correction operation (graded processing: voice guidance to straighten the device when there is a large deviation, and automatic correction of perspective transformation when there is a small deviation), and ambient light compensation operation (overall brightness adjustment and local overexposure / underexposure correction) in the aforementioned embodiments, ensuring that the user's whole-body body movements can fall clearly and completely into the detection window, providing high-quality image input for subsequent action recognition.
[0209] After image quality optimization, the system acquires the user's real-time video stream at a higher frame rate (e.g., 30 or 60 frames per second). Full-body keypoint coordinates, including those for the head, shoulders, elbows, wrists, hips, knees, and ankles, are extracted in real-time using skeletal keypoint detection algorithms (e.g., mobile models based on BlazePose or MoveNet). Based on these coordinates, the matching degree between the user's current pose and the game's preset action template is calculated. Specifically, the spatial sequence of the user's keypoints is aligned with the reference sequence of the preset template action, for example, by calculating the similarity of their pose vectors using dynamic time warping or cosine similarity. The matching degree can be quantified as a score between 0 and 100, with higher scores indicating a closer resemblance between the user's action and the standard action.
[0210] For example, in dance games, users are required to complete a series of continuous arm swings and body twists. The system calculates in real time the change in the angle between the user's wrist coordinates and shoulder coordinates, as well as the relative rotation angle between the hip and shoulder, in each frame, comparing these motion trajectories frame by frame with a standard dance template. If the user's arm swing amplitude reaches more than 80% of the template requirement and the timing of the movements is basically in line with the music beat, the matching score is high (e.g., 85 points); if the user's arms are not fully extended or the movement is significantly delayed, the score is lowered accordingly. The game interface provides real-time visual feedback (such as displaying ratings like "Perfect," "Good," and "Miss") and score rewards based on the matching score, encouraging users to adjust their movements.
[0211] In virtual reality (VR) interaction scenarios, taking virtual assembly or virtual object manipulation as an example, users interact with objects in the virtual environment through hand gestures (e.g., grasping, rotating, and placing parts). After image quality optimization, 21 coordinate points of the finger joints are extracted using a hand keypoint detection algorithm (e.g., MediaPipe Hands). Based on these coordinates, the finger bending angle (e.g., the pinching angle between the thumb and forefinger), palm posture (3D spatial orientation), and the distance between the fingertip and the virtual object's collider are calculated. When the user makes a pinching gesture and the distance between the fingertip and the virtual part is less than a preset threshold, a grasping event is triggered; when the user rotates their wrist, the rotation angle of the virtual part is updated in real time based on the change in the Euler angle of the wrist joint. The interaction response latency is controlled to a low level (e.g., within 50 milliseconds) by optimizing the image processing pipeline to ensure immersion.
[0212] Furthermore, in motion-sensing games, if the system detects that part of the user's limbs (such as hands or feet) extend beyond the detection window boundary, it can dynamically adjust the zoom level or output voice prompts (such as "Please take a step back and let your arm fully enter the frame") to guide the user back to the optimal shooting range. This is consistent with the zoom and distance guidance operations in the aforementioned embodiments. Simultaneously, for game scenes in low-light environments (such as a nighttime living room), light compensation can improve screen brightness and optimize local contrast, ensuring clear edges of the user's limbs and avoiding misjudgments of movement.
[0213] This entertainment scenario embodiment shares the same core image quality optimization steps as the aforementioned interview scenario embodiment. The difference lies in that the entertainment scenario focuses on real-time motion matching scoring, gesture recognition and interactive event triggering, as well as gamified visual and auditory feedback. All of the above processing is based on objective data-driven interactive responses and does not involve the retention or analysis of user identity features or privacy information. Those skilled in the art will understand that the methods disclosed herein can also be applied to other scenarios requiring dynamic adjustment of shooting parameters and real-time motion capture, such as fitness guidance, dance instruction, and remote motion collaborative training; these will not be elaborated upon here.
[0214] As can be seen from the above embodiments, the display method provided by this disclosure continuously acquires first state data for characterizing whether the spatial relationship between the device and the user meets the first predetermined condition and second state data for characterizing whether the ambient light meets the second predetermined condition during the detection period, and automatically executes a first adjustment operation (including distance guidance operation, viewpoint correction operation and / or zoom operation) and a second adjustment operation (ambient light compensation) accordingly, which can dynamically optimize the acquisition quality of the original image in detection scenarios such as skills talent interviews.
[0215] Specifically, this disclosure dynamically calculates the target shooting distance based on the user's height data and leg proportions, and generates a virtual footprint positioning frame and / or voice commands, achieving personalized, contactless positioning guidance. When the user cannot complete the positioning adjustment within the preset time, intelligent zoom automatically compensates for distance deviations, which is especially convenient for users who have difficulty with conventional interactive operations due to physical or cognitive reasons. Regarding device placement angle deviations, this disclosure adopts a tiered processing mechanism: for large deviations, voice and visual prompts based on the ground level guide the user to manually correct the position; for small deviations, perspective transformation automatically corrects image distortion, and sharpening is combined with clarity assessment to ensure image quality. In terms of lighting compensation, this disclosure first adjusts the image as a whole according to the ambient light intensity range, and then performs individual brightness attenuation, brightness enhancement, or tone mapping for abnormal lighting in key local areas (such as the arms and face), effectively avoiding the impact of overall adjustment on local details. Furthermore, this disclosure identifies the causal relationships between various types of spatial relationship anomalies and between them and environmental lighting anomalies, and performs adjustment operations in the order of prior anomalies over subsequent anomalies to avoid ineffective repeated adjustments; at the same time, it dynamically configures the weights of the first and second adjustment operations according to the action type corresponding to the current detection step (such as full-body standing posture, arm extension, local feature detection), and determines the priority adjustment strategy accordingly, so that the adjustment operations are more in line with the core needs of different detection stages.
[0216] Through the aforementioned multi-dimensional and adaptive linkage adjustments, this disclosure significantly reduces the number and difficulty of manual user intervention, shortens detection preparation time, and improves the clarity and accuracy of limb contours and key detection areas in images. Thus, in scenarios requiring dynamic optimization of shooting parameters, such as skills talent interviews, medical rehabilitation limb assessments, motion-sensing games, and VR interactions, it effectively ensures detection efficiency and user experience.
[0217] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0218] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0219] Based on the same inventive concept, corresponding to any of the above-described embodiments, this disclosure also provides a display device.
[0220] refer to Figure 5 The display device includes: The acquisition module 501 is configured to acquire the original image, first state data and second state data during the detection period; the first state data is used to characterize whether the spatial relationship between the device and the user meets a first predetermined condition; the second state data is used to characterize whether the ambient lighting meets a second predetermined condition. The first adjustment module 502 is configured to perform a first adjustment operation on the original image or the user based on the first state data; the first adjustment operation includes at least one of performing a distance guidance operation on the user, a perspective correction operation on the original image, or a zoom operation on the original image. The second adjustment module 503 is configured to perform a second adjustment operation on the original image to adjust the original image from the current ambient lighting state to the target lighting state. Display module 504 is configured to display the original image after adjustment based on the first adjustment operation and / or the second adjustment operation.
[0221] The first state data includes the user's height data during the detection period and the actual distance between the terminal device and the user; In this exemplary embodiment, the first adjustment module 502 is specifically configured as follows: The target shooting distance is calculated based on the height data, and distance guidance data is generated based on the deviation between the target shooting distance and the actual distance. Based on the distance guidance data, render a virtual footprint positioning frame in the display interface of the terminal device; and / or, Based on the distance guidance data, a first voice command is output to guide the user to move to the target location corresponding to the target shooting distance.
[0222] In this exemplary embodiment, the first adjustment module 502 is specifically configured as follows: If the deviation between the actual distance and the target shooting distance lasts for a preset duration, the digital zoom ratio is determined based on the deviation. The original image is zoomed according to the digital zoom ratio.
[0223] In this exemplary embodiment, the display device further includes: The rendering module is configured to render a virtual footprint positioning frame in the display interface of the terminal device; the virtual footprint positioning frame is used to guide the user through the distance operation. The detection phase acquisition module is configured to acquire the action type required for the current detection phase and / or the user's leg proportion; The adjustment data generation module is configured to generate adjustment data for the virtual footprint positioning frame based on the action type and / or the leg proportion; the adjustment data is used to indicate adjusting the virtual footprint positioning frame from its current state to a target state that matches the action type and / or leg proportion. The positioning frame adjustment module is configured to adjust the rendering position and / or size of the virtual footprint positioning frame in the display interface of the terminal device according to the adjustment data.
[0224] Wherein, the first state data includes viewing angle deviation data of the terminal device relative to the user; in this exemplary embodiment, the first adjustment module 502 is specifically configured as follows: In response to the viewpoint deviation data being greater than a preset angle threshold, a second voice command is output and a straightening prompt is generated to guide the user to manually adjust the placement angle of the terminal device; In response to the viewpoint deviation data being less than or equal to a preset angle threshold, the original image is subjected to perspective transformation processing to generate a viewpoint-corrected image.
[0225] The second state data includes ambient light intensity data; In this exemplary embodiment, the second adjustment module 503 is specifically configured as follows: Based on the preset intensity range to which the ambient light intensity data belongs, a first target strategy is determined from multiple preset light compensation strategies, and the original image is adjusted based on the first target strategy so that the original image is adjusted from the current brightness state to the target brightness state.
[0226] In this exemplary embodiment, the display device further includes: The local region image acquisition module is configured to acquire at least one local region image of the user from the original image adjusted by the second adjustment operation; The local light intensity acquisition module is configured to acquire local light intensity data corresponding to the local region image; The light adjustment module is configured to, in response to the presence of local overexposure or local underexposure in the local area, determine a second target strategy from multiple preset light compensation strategies based on the local light intensity data, and perform individual light adjustment on the local area based on the second target strategy, so that the local area is adjusted from the current local brightness state to the target local brightness state.
[0227] In this exemplary embodiment, the display device further includes: The causal relationship determination module is configured to determine the causal relationship between each type of spatial relationship that meets the first predetermined condition and each type of ambient lighting that meets the second predetermined condition; The first execution module is configured to, in response to the existence of multiple types with causal relationships, sequentially execute the adjustment operations corresponding to each type in the order that the type acting as the cause takes precedence over the type acting as the result in the causal relationship.
[0228] In this exemplary embodiment, the display device further includes: The detection step identifier acquisition module is configured to acquire the current detection step identifier; the detection step identifier is used to characterize the type of action to be detected in the current detection process. The weight allocation module is configured to assign corresponding weights to the first adjustment operation and the second adjustment operation respectively, based on the detection step identifier. The second execution module is configured to determine the adjustment operation to be executed first based on the weight.
[0229] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0230] The apparatus of the above embodiments is used to implement the corresponding display method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0231] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the display method described in any of the above embodiments.
[0232] Figure 6 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0233] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0234] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0235] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0236] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0237] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0238] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0239] The electronic devices described above are used to implement the corresponding display methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0240] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the display method as described in any of the above embodiments.
[0241] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0242] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the display method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0243] Based on the same inventive concept, corresponding to the display method described in any of the above embodiments, this disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the display method. Corresponding to the execution entity for each step in each embodiment of the display method, the processor executing the corresponding step may belong to the corresponding execution entity.
[0244] The computer program products of the above embodiments are used to cause the computer and / or the processor to perform the display method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0245] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0246] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0247] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0248] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A display method, comprising: Acquire the raw image, first state data, and second state data during the detection period; The first state data is used to characterize whether the spatial relationship between the device and the user meets the first predetermined condition; The second state data is used to characterize whether the ambient lighting meets the second predetermined condition; Based on the first state data, a first adjustment operation is performed on the original image or the user; the first adjustment operation includes at least one of performing a distance guidance operation on the user, a perspective correction operation on the original image, or a zoom operation on the original image. Based on the second state data, a second adjustment operation is performed on the original image to adjust the original image from the current ambient lighting state to the target lighting state; The original image is displayed based on the adjustment performed by the first adjustment operation and / or the second adjustment operation.
2. The method as described in claim 1, wherein, The first status data includes the user's height data during the detection period and the actual distance between the terminal device and the user; The distance guidance operation includes: The target shooting distance is calculated based on the height data, and distance guidance data is generated based on the deviation between the target shooting distance and the actual distance. Based on the distance guidance data, render a virtual footprint positioning frame in the display interface of the terminal device; and / or, Based on the distance guidance data, a first voice command is output to guide the user to move to the target location corresponding to the target shooting distance.
3. The method of claim 2, wherein, The zoom operation includes: If the deviation between the actual distance and the target shooting distance lasts for a preset duration, the digital zoom ratio is determined based on the deviation. The original image is zoomed according to the digital zoom ratio.
4. The method of claim 1, wherein, The method further includes: A virtual footprint positioning frame is rendered in the display interface of the terminal device; the virtual footprint positioning frame is used to guide the user through the distance operation. Obtain the type of action required for the current detection step and / or the user's leg proportions; Based on the action type and / or the leg proportion, adjustment data for the virtual footprint positioning frame is generated; the adjustment data is used to indicate adjusting the virtual footprint positioning frame from its current state to a target state that matches the action type and / or leg proportion; Based on the adjustment data, adjust the rendering position and / or size of the virtual footprint positioning frame in the display interface of the terminal device.
5. The method of claim 1, wherein, The first state data includes the viewing angle deviation data of the terminal device relative to the user; The perspective correction operation includes: In response to the viewpoint deviation data being greater than a preset angle threshold, a second voice command is output and a straightening prompt is generated to guide the user to manually adjust the placement angle of the terminal device; In response to the viewpoint deviation data being less than or equal to a preset angle threshold, the original image is subjected to perspective transformation processing to generate a viewpoint-corrected image.
6. The method of claim 1, wherein, The second state data includes ambient light intensity data; The second adjustment operation includes: Based on the preset intensity range to which the ambient light intensity data belongs, a first target strategy is determined from multiple preset light compensation strategies, and the original image is adjusted based on the first target strategy so that the original image is adjusted from the current brightness state to the target brightness state.
7. The method of claim 6, wherein, The method further includes: Obtain at least one local region image of the user from the original image after the second adjustment operation; Obtain the local light intensity data corresponding to the local region image; In response to the presence of local overexposure or local underexposure in the local area, a second target strategy is determined from multiple preset light compensation strategies based on the local light intensity data, and the local area is individually adjusted based on the second target strategy to adjust the local area from the current local brightness state to the target local brightness state.
8. The method of claim 1, wherein, Displaying the original image adjusted based on the first adjustment operation and the second adjustment operation includes: Determine the causal relationship between each type of spatial relationship that meets the first predetermined condition and each type of ambient lighting that meets the second predetermined condition; In response to the existence of multiple types with causal relationships, the first adjustment operation and / or the second adjustment operation corresponding to each type are executed sequentially in the order that the type as the cause takes precedence over the type as the result in the causal relationship. The adjusted original image is displayed.
9. The method of claim 8, wherein, The spatial relationships that meet the first predetermined conditions include viewpoint types and distance types; the ambient lighting that meets the second predetermined conditions includes lighting types; the causal relationship includes the viewpoint type as a cause type leading to the distance type as a result type, and the distance type as a cause type leading to the lighting type as a result type. The sequential execution of the first adjustment operation and / or the second adjustment operation corresponding to each type includes: The first adjustment operation corresponding to the viewpoint type, the first adjustment operation corresponding to the distance type, and the second adjustment operation corresponding to the lighting type are executed sequentially.
10. The method of claim 1, wherein, The method further includes: Obtain the current detection step identifier; the detection step identifier is used to characterize the type of action to be detected in the current detection process. Based on the detection step identifier, corresponding weights are configured for the first adjustment operation and the second adjustment operation, respectively. Based on the weights, the adjustment operations to be performed first are determined.
11. A display device, comprising: The acquisition module is configured to acquire the raw image, first state data and second state data during the detection period; The first state data is used to characterize whether the spatial relationship between the device and the user meets the first predetermined condition; The second state data is used to characterize whether the ambient lighting meets the second predetermined condition; The first adjustment module is configured to perform a first adjustment operation on the original image or the user based on the first state data; the first adjustment operation includes at least one of performing a distance guidance operation on the user, a perspective correction operation on the original image, or a zoom operation on the original image. The second adjustment module is configured to perform a second adjustment operation on the original image based on the second state data, so as to adjust the original image from the current ambient lighting state to the target lighting state. The display module is configured to display the original image after adjustment based on the first adjustment operation and / or the second adjustment operation.
12. A computer device comprising one or more processors, a memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, the one or more programs comprising instructions for performing the method of any one of claims 1 to 10.
13. A non-volatile computer-readable storage medium comprising a computer program, which, when executed by one or more processors, causes the one or more processors to perform the method of any one of claims 1 to 10.
14. A computer program product comprising one or more computer programs that, when executed by one or more processors, implement the method as described in any one of claims 1 to 10.