Shooting control method, device, extended reality device and computer-readable storage medium

By obtaining user gaze point information in extended real-life devices and generating annotation boxes, the problems of low scene selection accuracy and poor picture clarity in the prior art are solved, and higher scene selection accuracy and shooting clarity are achieved, while reducing the occupation of equipment resources.

CN119211511BActive Publication Date: 2025-05-27FALCON INNOVATIONS TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411671948.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-05-27
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

The existing extended reality equipment has low accuracy in scene selection and poor clarity in shooting, which cannot meet the needs of users, affects the user experience, and occupies the equipment's computing and storage resources, increasing the equipment's power consumption.

Method used

By obtaining the user's gaze point information, a label box is generated on the virtual display screen based on the gaze point information, a focus of the real scene picture is represented, and a target image is generated in the shooting control command.

Benefits of technology

It has achieved improvements in the accuracy of scene selection, improved shooting clarity, saved the computing resources of the equipment, reduced the energy consumption of the equipment, and improved the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119211511B_ABST
    Figure CN119211511B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a shooting control method, apparatus, extended reality device, and computer-readable storage medium. The method includes: obtaining gaze point information of a user, where the gaze point information is used to represent a real-world scene image that the user gazes at through a physical display screen of the extended reality device; generating a marking frame on a virtual display screen according to the gaze point information, where the marking frame is used to represent the focus of the real-world scene image; and in response to a shooting control instruction, shooting the real-world scene image within the marking frame to generate a target image. The shooting control method provided by the embodiments of the present application is applied to an extended reality device, which can improve the accuracy of scene selection during shooting of the extended reality device, and thus improve the shooting clarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of extended reality display, and particularly to a shooting control method, apparatus, extended reality device, and computer-readable storage medium. Background Art

[0002] Extended Reality (XR) technology enables users to interact with virtual and real worlds by superimposing virtual objects, images, videos, or other digital content on the real world. Wearable XR terminal devices represented by smart glasses are considered the best implementation carriers of "XR + AI" technology, integrating rich functional applications such as communication, music, photography, navigation, translation, health detection, etc.

[0003] When a user takes a photo using an existing extended reality device, for example, using an AR glasses, most of the time, it is by means of a camera embedded in the frame. As the glasses move, environmental pictures within a fixed viewing angle range are obtained in real time and presented in an alternative frame of the virtual screen through extended reality display. Then, the user determines the target image in the alternative frame of the virtual screen, and finally, a shooting operation is performed.

[0004] However, the scene selection accuracy of the existing shooting method is relatively low, and the clarity of the captured picture is poor, which cannot meet the user's needs and thus affects the user experience. On the other hand, after focusing, the real scene to be photographed is captured and pushed to the device side for preview in the preview frame, which occupies the computing and storage resources of the device and increases the power consumption of the device. Summary of the Invention

[0005] Embodiments of the present application provide a shooting control method, apparatus, extended reality device, and computer-readable storage medium, which can improve the accuracy of scene selection during shooting of extended reality devices, and thus improve the shooting clarity.

[0006] In a first aspect, embodiments of the present application provide a shooting control method applied to an extended reality device. The method includes:

[0007] Obtaining the gaze point information of the user, where the gaze point information is used to represent the real scene picture that the user gazes through the physical display screen of the extended reality device;

[0008] Generating a marking frame on the virtual display screen according to the gaze point information, where the marking frame is used to represent the focus of the real scene picture;

[0009] In response to a shooting control instruction, shooting the real scene picture within the marking frame to generate a target image.

[0010] Optionally, in some embodiments of the present application, the obtaining of the user's fixation point information includes:

[0011] Obtaining the user's eye image information through the gaze tracking module of the extended reality device;

[0012] Determining the pupil center position and the eye rotation angle according to the eye image information;

[0013] Determining the fixation point information according to the pupil center position and the eye rotation angle.

[0014] Optionally, in some embodiments of the present application, generating a marking frame on the virtual display screen according to the fixation point information includes:

[0015] Determining the display position parameters of the marking frame according to the first pose relationship between the fixation point information and the extended reality device;

[0016] Determining the display size dimension parameters of the marking frame according to the shooting parameters of the shooting module;

[0017] Generating the marking frame according to the display position parameters and the display size dimension parameters.

[0018] Optionally, in some embodiments of the present application, the generating of the marking frame according to the display position parameters and the display size dimension parameters further includes:

[0019] Generating a first marking frame according to the display position parameters and the display size dimension parameters;

[0020] When there is a target object in the real scene picture represented in the first marking frame, obtaining the contour information of the target object;

[0021] Generating a second marking frame in the first marking frame according to the contour information;

[0022] Responding to a selection control instruction, determining the marking frame from the first marking frame and the second marking frame;

[0023] Wherein, the first marking frame and the second marking frame have different presentation forms.

[0024] Optionally, in some embodiments of the present application, after determining the marking frame from the first marking frame and the second marking frame in response to a selection control instruction, the method further includes:

[0025] Adjusting the shooting parameters of the shooting module according to the marking frame.

[0026] Optionally, in some embodiments of the present application, generating the annotation box at the display position according to the annotation range further includes:

[0027] Determining the eye position of the user through the gaze tracking module of the extended reality device, and obtaining the module position of the shooting module of the extended reality device;

[0028] Determining a second pose relationship according to the eye position and the module position;

[0029] Performing calculation processing on the original photo taken by the shooting module according to the second pose relationship to obtain the target image corresponding to the real scene picture.

[0030] Optionally, in some embodiments of the present application, the responding to the shooting control instruction to shoot the real scene picture within the annotation box to generate a target image includes:

[0031] Receiving a shooting control instruction sent by the user;

[0032] Performing shooting processing on the real scene picture according to the shooting control instruction to generate the target image.

[0033] In a second aspect, an embodiment of the present application further provides a shooting control device applied to an extended reality device. The device includes:

[0034] An acquisition module, configured to acquire gaze point information of a user, where the gaze point information is used to represent a real scene picture that the user gazes through the physical display screen of the extended reality device;

[0035] A processing module, configured to generate an annotation box on a virtual display screen according to the gaze point information, where the annotation box is used to represent the focus of the real scene picture;

[0036] The processing module is further configured to respond to a shooting control instruction to shoot the real scene picture within the annotation box to generate a target image.

[0037] Optionally, in some embodiments of the present application, the acquisition module is configured to:

[0038] Obtaining eye image information of the user through the gaze tracking module of the extended reality device;

[0039] Determining the pupil center position and the eye rotation angle according to the eye image information;

[0040] Determining the gaze point information according to the pupil center position and the eye rotation angle.

[0041] Optionally, in some embodiments of the present application, the processing module is configured to:

[0042] Determine the display position parameters of the annotation box according to the relationship between the gaze point information and the first pose of the extended reality device;

[0043] Determine the display size dimension parameters of the annotation box according to the shooting parameters of the shooting module;

[0044] Generate the annotation box according to the display position parameters and the display size dimension parameters.

[0045] Optionally, in some embodiments of the present application, the processing module is further configured to:

[0046] Generate a first annotation box according to the display position parameters and the display size dimension parameters;

[0047] When there is a target object in the real scene picture represented in the first annotation box, obtain the contour information of the target object;

[0048] Generate a second annotation box in the first annotation box according to the contour information;

[0049] In response to a selection control instruction, determine the annotation box from the first annotation box and the second annotation box;

[0050] Wherein, the first annotation box and the second annotation box have different presentation forms.

[0051] Optionally, in some embodiments of the present application, the processing module is further configured to:

[0052] Adjust the shooting parameters of the shooting module according to the annotation box.

[0053] Optionally, in some embodiments of the present application, the processing module is further configured to:

[0054] Determine the eye position of the user through the gaze tracking module of the extended reality device, and obtain the module position of the shooting module of the extended reality device;

[0055] Determine a second pose relationship according to the eye position and the module position;

[0056] Perform calculation processing on the original photo taken by the shooting module according to the second pose relationship to obtain the target image corresponding to the real scene picture.

[0057] Optionally, in some embodiments of the present application, the processing module is further configured to:

[0058] Receive a shooting control instruction sent by the user;

[0059] Perform shooting processing on the real - world scene image according to the shooting control instruction to generate the target image.

[0060] In a fourth aspect, an embodiment of the present application further provides an extended reality device, and the device includes:

[0061] A gaze - tracking module, configured to obtain the gaze point information of the user, where the gaze point information is used to represent the real - world scene image that the user gazes at through the physical display screen of the extended reality device;

[0062] A processing module, configured to generate a marking frame on the virtual display screen according to the gaze point information, where the marking frame is used to represent the focus of the real - world scene image;

[0063] A shooting module, configured to respond to a shooting control instruction and shoot the real - world scene image within the marking frame to generate a target image.

[0064] In a fifth aspect, an embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer - readable storage medium. The processor of the computer device reads the computer instructions from the computer - readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various optional implementation manners of the embodiments of the present application.

[0065] In summary, in the embodiment of the present application, by obtaining the gaze point information of the user, where the gaze point information is used to represent the real - world scene image that the user gazes at through the physical display screen of the extended reality device; generating a marking frame on the virtual display screen according to the gaze point information, where the marking frame is used to represent the focus of the real - world scene image; and responding to a shooting control instruction to shoot the real - world scene image within the marking frame to generate a target image, the technical solution directly frames and shoots in the real - world scene within the user's true field of view according to the user's gaze point information, achieving the technical effects of improving the accuracy of scene selection and the clarity of shooting. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the technical solutions in the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those skilled in the art can obtain other drawings according to these drawings without creative efforts.

[0067] Figure 1 It is a schematic diagram of the scenario of the shooting control method provided by the embodiment of the present application;

[0068] Figure 2It is a schematic flowchart of the shooting control method provided by the embodiments of the present application;

[0069] Figure 3 It is a schematic structural diagram of an extended reality device corresponding to the shooting control method provided by the embodiments of the present application;

[0070] Figure 4 It is a schematic overall flowchart of the shooting control method provided by the embodiments of the present application;

[0071] Figure 5 It is a schematic structural diagram of the shooting control device provided by the embodiments of the present application;

[0072] Figure 6 It is a schematic structural diagram of the extended reality device provided by the embodiments of the present application. Detailed implementation manners

[0073] Next, the technical solutions in the present application will be clearly and completely described in conjunction with the accompanying drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0074] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the above features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0075] In this application, the term "exemplary" is used to mean "serving as an example, illustration, or instance". Any embodiment described as "exemplary" in this application is not necessarily to be construed as more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the present invention. In the following description, details are set forth for the purpose of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed in this application.

[0076] First, the terms related to this application are explained as follows:

[0077] Extended Reality: Extended Reality, XR, is a technology that combines virtual information with real-world scenarios to create an enhanced perceptual environment.

[0078] Extended Reality device: Used to fuse virtual content with the real world to provide an enhanced visual experience. These devices typically employ head-mounted displays (HMDs), smart glasses, or other forms of wearable devices.

[0079] The embodiments of this application provide a shooting control method, apparatus, extended reality device, and computer-readable storage medium. Specifically, the embodiments of this application provide a shooting control apparatus applicable to the shooting control method, and the shooting control apparatus includes an extended reality device and the main control device of the extended reality device.

[0080] In the prior art, with the rapid development and increasing maturity of extended reality display technology, the combination of extended reality display and artificial intelligence large models endows extended reality devices with rich and diverse functional applications. More and more wearable extended reality devices (such as VR headsets, AR glasses, etc.) have been successively launched on the market, and first-person perspective intelligent photography is an important one of these functions.

[0081] However, existing extended reality devices such as wearable smart glasses mainly rely on cameras embedded in the devices to obtain environmental pictures within a fixed viewing angle range relative to the user's perspective in real time as the device moves. The camera takes pictures of the environment in real time and streams the images to the display optical engine of the extended reality, which is presented in an alternative frame (preview frame) of the virtual screen for the user to preview through extended reality display. Then, the user determines the target image in the alternative frame of the virtual screen, and finally performs a shooting operation on the target image.

[0082] However, this shooting method still cannot meet the user's need to directly frame within the real field of view and directly shoot a clear picture within the ideal framing area, which affects the user experience of the AR glasses. Moreover, the existing shooting method also needs to stream the picture to be shot to the alternative frame after shooting, occupying the computing resources and storage resources of the device, and at the same time increasing the power consumption of the device.

[0083] Therefore, the shooting methods of existing extended reality devices have various problems such as low scene selection accuracy, poor clarity of the captured picture, inability to meet the user's needs, and thus affecting the user experience.

[0084] The embodiments of the present application provide a shooting control method, apparatus, extended reality device, and computer-readable storage medium. The technical solution adopted is to obtain the gaze point information of the user, where the gaze point information is used to represent the real-world scene picture that the user gazes through the physical display screen of the extended reality device; generate a marking frame on the virtual display screen according to the gaze point information, where the marking frame is used to represent the focus of the real-world scene picture; and in response to a shooting control instruction, shoot the real-world scene picture within the marking frame to generate a target image.

[0085] In summary, the shooting control method in the embodiments of the present application can directly frame and shoot in the real-world scene within the user's real field of view according to the user's gaze point information, achieving the technical effects of improving the scene selection accuracy and the shooting clarity. At the same time, it is not necessary to first shoot the real scene to be shot and stream it to the preview frame of the device for the user to preview before the formal shooting, saving the computing resources of the device and reducing the device power consumption.

[0086] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments.

[0087] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the scene of the shooting control method provided by the embodiments of the present application. The shooting control system may include an extended reality device 100 and a main control device 200. The extended reality device 100 and the main control device 200 can be communicatively connected in any way, including but not limited to signal communication through electrical circuits and communication through wireless signals. The wireless signals can be computer network communications such as the TCP / IP protocol suite (TCP / IP Protocol Suite, TCP / IP) and the User Datagram Protocol (UDP). The extended reality device 100 can receive control signals on a remote control or a control panel. The extended reality device 100 can also receive instruction information sent by the main control device 200. The extended reality device 100 can perform corresponding operations according to the corresponding instruction information, such as the shooting control method in the present application.

[0088] In the embodiments of the present application, the extended reality device 100 includes, but is not limited to, a head-mounted display (HMD), smart glasses, or other forms of wearable devices, etc.

[0089] Those skilled in the art can understand that Figure 1 the application environment shown is only one application scenario of the solution of the present application, and does not constitute a limitation on the application scenario of the solution of the present application. Other application environments may also include more or fewer extended reality devices than Figure 1 shown, for example Figure 1 only 1 extended reality device is shown herein, and specifically it is not limited here.

[0090] In addition, as Figure 1 shown, the main control device 200 may include any hardware device such as a CPU or a single-chip microcomputer embedded inside the extended reality device 100 that can perform data processing and instruction sending. Specifically, it is not limited here. The main control device 200 may be any hardware device such as a CPU or a single-chip microcomputer embedded inside other wearable devices such as mobile phones, bracelets, iPads, and wristbands that can perform data processing and instruction sending. Specifically, it is not limited here.

[0091] It should be noted that Figure 1 the scene schematic diagram of the shooting control system shown is only an example. The shooting control system and scene described in the embodiments of the present application are for more clearly explaining the technical solution of the embodiments of the present application, and do not constitute a limitation on the technical solution provided by the embodiments of the present application. Those skilled in the art know that with the evolution of the shooting control system and the emergence of new service scenarios, the technical solution provided by the embodiments of the present application is equally applicable to similar technical problems.

[0092] Specifically, please refer to Figure 2 , Figure 2 which is a schematic flowchart of the extended reality device provided by the embodiments of the present application for executing the shooting control method. Among them, the specific execution process of the extended reality device for executing the shooting control method is as follows:

[0093] S201: Obtain the gaze point information of the user, where the gaze point information is used to represent the real scene picture that the user gazes through the physical display screen of the extended reality device.

[0094] As an optional embodiment, as Figure 3Schematic diagram of the structure of the extended reality device. The shooting control method provided in this application is applicable to the extended reality device. Taking the AR glasses as an example, in addition to the temple and the lens frame of the traditional glasses, the AR glasses also have a physical display screen (an optical display module including an optical combiner), a shooting module, a gaze tracking module, and a computing and processing module.

[0095] It should be noted that the installation positions of the above-mentioned various modules in Figure 3 are only for examples and can be changed according to the specific structure of the extended reality device.

[0096] Optionally, the optical display module includes an ultra-small optical engine and an optical coupler. The optical engine can be based on micro-led or micro-oled, and the optical coupler can be based on a waveguide or a semi-transmissive semi-reflective lens. In this embodiment, the optical combiner is arranged on the spectacle lens, that is, all or part of the spectacle lens. Therefore, the real-world scene picture that the user gazes at through the physical display screen of the extended reality device is the real-world scene picture that the user gazes at through the spectacle lens.

[0097] Optionally, the photographing module is an image sensor mainly composed of a camera. The photographing module can take pictures of the external environment that the user wants to photograph, and the camera faces the outside of the extended reality device. The extended reality device can also be equipped with other cameras, such as cameras for gaze tracking, cameras for gesture tracking, and cameras for lower body posture capture.

[0098] Optionally, the gaze tracking module can be an eye tracker, which can directly output the position of the user's fixation point; it can also be a module composed of cameras corresponding to the left and right eyes respectively plus a computing processor that executes a specific algorithm program. The cameras are used to take pictures of the user's left and right eye images, and through the operation and processing of the user's eye physiological characteristic data, the position of the user's fixation point is finally obtained.

[0099] It should be noted that the gaze tracking module is one of the core components of the extended reality device for executing the shooting control method, and it is responsible for capturing the user's eye movement information in real time. This module usually consists of one or more high-precision cameras and can be installed inside the AR glasses, directly aiming at the user's eyes.

[0100] Optionally, the computing and processing module is a chip integrating one or more processing units such as CPU, GPU, NPU, DSP, etc.

[0101] It should be noted that the extended reality device in the embodiments of the present application mainly includes the above modules, and may also include an acoustic module, a communication module, a sensor module, etc. The acoustic module may include a speaker and a microphone, and the communication module may be Bluetooth communication, 5G communication, etc., so as to operate or control the extended reality device in various ways. The sensor module may include GPS, IMU, gyroscope, etc., which can be used to achieve tracking and positioning.

[0102] In some embodiments of the present application, when the camera function is in the on state, the computing and processing module can obtain the accurate position of the user's fixation point at this time from the gaze tracking module.

[0103] Optionally, the shooting control method provided in the embodiments of the present application further includes: obtaining the eye image information of the user through the gaze tracking module of the extended reality device; determining the pupil center position and the eye rotation angle according to the eye image information; and determining the fixation point information according to the pupil center position and the eye rotation angle.

[0104] As an optional embodiment, as Figure 4 shown in the schematic diagram of the overall process of the shooting control method, before obtaining the user's fixation point information, the user can activate the shooting function of the AR glasses through a voice command, such as: start shooting, or a gesture action, such as: draw a shooting symbol in the air. After receiving the start command, the computing and processing module immediately controls the relevant application program to enter the standby state, loads the relevant shooting module and initializes the parameters of the gaze tracking module, such as: waking up the eye tracker, waking up the camera for shooting, etc., to ensure that the system can quickly respond to the user's next operation.

[0105] Optionally, after entering the shooting standby state, the gaze tracking module is used to monitor the user's gaze position in real time, and the computing and processing module is ready to receive and process the fixation point information from the gaze tracking module at any time, so as to prepare for the subsequent shooting process.

[0106] As an optional embodiment, after receiving the fixation point information of the gaze tracking module, the fixation point position of the user can be calculated through a relevant fixation point algorithm, and this position is the point that the user is focusing on in the field of view at this time.

[0107] Specifically, still as Figure 4As shown, after the user gazes at a certain position for a certain period of time, for example: 5 seconds, the eye tracking module obtains the user's eye image information. After the camera of the eye tracking module captures the user's eye image, through a series of calculations and processes, data such as the pupil center position and the eye rotation angle of the user can be obtained. Subsequently, using data such as the pupil center position and the eye rotation angle, relevant fixation point algorithms are adopted, such as: the pupil center - corneal reflection method or the 3D eye model, etc., to calculate the user's fixation point, that is, the specific position in the real scene that the user is currently focusing on through the physical display screen of the extended reality device.

[0108] It should be noted that the calculated user fixation point data can be coordinates in a three - dimensional space, used to represent the projection position of the user's line of sight in the real world. After the calculation and processing module receives these fixation point data, it can be used for subsequent image processing and annotation box positioning processes, that is, a border of the image to be captured can be marked in the user's real field of view with the fixation point as the center.

[0109] S202: Generate an annotation box on the virtual display screen according to the fixation point information, where the annotation box is used to represent the focus of the real - scene picture.

[0110] In the embodiment of the present application, the virtual display screen is opposite to the physical display screen. It refers to the virtual display screen generated in front of the user's eyes after the light emitted by the light engine undergoes multiple propagations and conversions through multiple optical components. The user can simultaneously observe the real field of view through this physical display screen, that is, the virtual display screen is located in the user's real field of view, and various virtual information can be superimposed and displayed on the virtual display screen. The focus of the real - scene picture refers to the geometric center point of the real - scene picture corresponding to the image to be taken.

[0111] It should be noted that the extended reality device is equipped with an optical display module, including: an ultra - small light engine and an optical coupler. The ultra - small light engine usually includes a micro - display screen and a lens made of liquid crystal or organic light - emitting diode (OLED) technology. The optical coupler can be a waveguide sheet. Taking smart glasses as an example, the waveguide sheet can be part or all of the spectacle lens. They are placed in the extended reality device, and the light emitted by the light engine is propagated and converted through optical components, and finally a virtual display screen in front of the user's eyes is formed.

[0112] In the embodiment of the present application, after the calculation and processing module comprehensively processes data such as the fixation point position, the internal and external parameters of the shooting module, and the internal and external parameters of the virtual screen, the position of the annotation box on the virtual display screen is obtained.

[0113] In the embodiments of the present application, the focus of the real-world scene image to be captured is determined based on the fixation point information. Generally, the focus is the geometric center point of the annotation box to be generated. After determining the focus position, the range contour of the real-world scene image to be captured can be determined by combining the internal and external parameters of the imaging module, and the geometric size and shape of the virtual annotation box can be determined based on the range contour, and the annotation box is displayed on the virtual display screen.

[0114] Optionally, the imaging control method provided by the embodiments of the present application further includes: determining the display position parameter of the annotation box according to the relationship between the fixation point information and the first pose of the extended reality device; determining the display size parameter of the annotation box according to the imaging parameters of the imaging module; and generating the annotation box according to the display position parameter and the display size parameter.

[0115] In the embodiments of the present application, after the computing and processing module receives the user's fixation point data, it will calculate the size and shape of the annotation box in combination with the imaging parameters of the imaging module.

[0116] Specifically, the real-world position of the real-world scene image to be captured is determined according to the fixation point information, and the pose relationship between the real-world position and the extended reality device is determined; the pose relationship between the virtual image displayed on the virtual screen and the extended reality device can also be determined by combining the internal and external parameters of the imaging module and the internal and external parameters of the virtual screen; according to the above pose relationship, the display position parameter of the annotation box on the virtual display screen can be determined; and then according to the imaging parameters of the imaging module of the extended reality device (for example: focal length, aperture, field of view angle, lens distortion coefficient, etc.), the display size of the annotation box is determined, and the annotation box is generated according to the display size and the display position parameter.

[0117] It should be noted that the imaging parameters include the first parameter information (internal parameters, for example: focal length, aperture, field of view angle, lens distortion coefficient, etc.) and the second parameter information (external parameters, for example: the installation position and angle of the imaging module relative to the glasses) of the imaging module. These imaging parameters are used to accurately project the fixation point from the three-dimensional space onto the two-dimensional plane of the captured image, and determine the size and shape of the captured photo corresponding to the selected real scene. The computing and processing module can determine the display position parameter of the annotation box according to the fixation point data (fixation point information) and the pose relationship, and further determine the display size according to the first parameter information, so as to generate the annotation box according to the display position parameter and the display size. The pose relationship can include a position relationship and / or a posture relationship, etc., for example: relative position, relative posture, etc.

[0118] In the embodiments of the present application, the display position data (display position parameters and display size) of the annotation box can be determined only based on the first parameter information and the pose relationship. First, determine the real position corresponding to the focus position in the real scene, and determine the device position of the extended reality device according to the sensor module, so as to determine the relative pose relationship between the area to be extended and the extended reality device based on the focus real position and the device position. Secondly, obtain the display parameters of the virtual display screen relative to the extended reality device. The display parameters include the conversion matrix relationship between the coordinate system of the virtual display screen and the coordinate system of the extended reality device. After determining the relative pose of the extended reality area relative to the extended reality device and the conversion matrix relationship between the coordinate system of the virtual display screen and the coordinate system of the extended reality device, the position data of the virtual display can be calculated, and the extended reality device can generate an annotation box at the specified correct position (the geometric center of the annotation box coincides with the fixation point, i.e., the focus of the captured photo, which is the correct position) according to the display data.

[0119] It should be noted that the display parameters of the virtual display screen relative to the extended reality device are determined by the optical hardware parameters of the extended reality device, that is, the first parameter information, which is generally obtained through pre-calibration and stored in the storage chip built into the extended reality device.

[0120] As a possible embodiment, the calculation and processing module can also determine the display position parameters of the annotation box according to the fixation point data (fixation point information) and the pose relationship, then adjust the display position parameters according to the second parameter information, and further determine the display size according to the first parameter information, so as to generate an annotation box according to the adjusted display position parameters and display size. This enables the shooting angle of the shooting module to be adjusted to be consistent with the human eye observation angle before officially capturing an image.

[0121] Optionally, the shooting control method provided in the embodiments of the present application further includes: determining the eye position of the user through the gaze tracking module of the extended reality device, and obtaining the module position of the shooting module of the extended reality device; determining the second pose relationship according to the eye position and the module position; performing calculation and processing on the original photo captured by the shooting module according to the second pose relationship to obtain a target image corresponding to the real scene picture.

[0122] In the embodiments of the present application, since there is a difference between the shooting module of the extended display device and the human eye position, therefore, the display position data (display position parameters and display size) of the annotation box can also be determined by combining the first parameter information, the second parameter information and the pose relationship.

[0123] Specifically, first, determine the real position corresponding to the focus position in the real scene, and determine the device position of the extended reality device according to the sensor module, so as to determine the relative pose relationship between the area to be extended and the extended reality device based on the focus real position and the device position; secondly, obtain the display parameters of the virtual display screen relative to the extended reality device, and the display parameters include the conversion matrix relationship between the coordinate system of the virtual display screen and the coordinate system of the extended reality device. After determining the relative pose of the extended reality area relative to the extended reality device and determining the conversion matrix relationship between the coordinate system of the virtual display screen and the coordinate system of the extended reality device, the position data of the virtual display can be calculated; further, the eye position of the user and the module position of the shooting module can also be obtained through the eye tracking module to obtain the relative pose relationship between the user's eyes and the shooting module; thus, the display position data is adjusted according to this pose relationship, and a marking frame is displayed according to the display size determined by the shooting module at the adjusted display position data, so that the shooting picture of the shooting module is consistent with the picture observed by the human eye.

[0124] It should be noted that the second parameter information is mainly used to adjust the display position of the marking frame to make the shooting picture of the shooting module consistent with the picture observed by the human eye. The effect of ensuring the picture consistency can also be achieved by processing the captured image with devices such as a computer.

[0125] Optionally, still as Figure 4 shown, the calculation and processing module first determines the display position parameters of the fixation point in the image according to the internal parameter data of the camera and the fixation point data of the user. Then, considering the external parameter data of the camera, the calculation and processing module adjusts the position of the fixation point according to the installation position and angle of the shooting module relative to the glasses to ensure that it is consistent with the fixation point in the real world. The adjustment process can adopt complex projection transformation algorithms. Through these algorithms, the system can accurately map the user's fixation point from the actual scene to the image captured by the camera, so that the image content obtained by the shooting module is consistent with the content of the real scene observed by the human eye. Similarly, the shooting angle of the image captured by the shooting module can also be made consistent with the angle of the real scene observed by the human eye.

[0126] As a possible embodiment, a telescopic component can also be set for the shooting module so that when the image is officially shot, the shooting module, the human eye, and the angle of the real scene to be shot are in the same straight line, so that the angle of the shot image is consistent with the angle of the real scene observed by the human eye.

[0127] As a possible embodiment, after the image is officially shot, the calculation and processing module can also adjust the captured image according to the first parameter information, the second parameter information, and the pose relationship to make the angle of the captured image consistent with the angle of the real scene observed by the human eye.

[0128] Before displaying the real - world scene image on the virtual display screen according to the fixation point information, optionally, the shooting control method provided by the embodiments of the present application further includes: generating a first annotation box according to the display position parameter and the display size parameter; when there is a target object in the real - world scene image represented within the first annotation box, obtaining the contour information of the target object; generating a second annotation box within the first annotation box according to the contour information; and in response to a selection control instruction, determining an annotation box from the first annotation box and the second annotation box.

[0129] In the embodiments of the present application, the forms of the first annotation box and the second annotation box are different. For example, the first annotation box can be displayed by a solid line, and the second annotation box can be displayed by a dashed line.

[0130] Specifically, the embodiments of the present application can further optimize the position and size of the annotation box according to the image features corresponding to the fixation point position, such as the contour of the target object obtained from the edge detection result or the shape of a specific object, or regenerate a second annotation box corresponding to the contour of the target object or the object shape within the first annotation box, and ensure that the second annotation box accurately covers the area of interest of the user.

[0131] Optionally, for example, if the target object existing in the observed real - world scene image is a water cup, first generate a first annotation box according to the fixation point position, further determine that the contour of the water cup is the contour of the target object, generate a second annotation box according to the contour of the water cup, and display the second annotation box within the first annotation box.

[0132] It should be noted that the target object can be one or more, the positional relationship of the target objects can be an overlapping relationship, a partially overlapping relationship, a juxtaposed relationship, etc. The first annotation box and the second annotation box can be displayed simultaneously or only one of them can be displayed. The number of the second annotation boxes can also be multiple, which is not specifically limited in the present application.

[0133] Optionally, after generating the second annotation box, the user can select the annotation box in various ways and confirm the final annotation box according to the selection result, so that the shooting module can shoot the image within the annotation box.

[0134] It should be noted that the user can implement the selection operation of the annotation box through gestures, voice, or other devices connected to the extended reality device.

[0135] After generating the second annotation box within the first annotation box according to the annotation range, optionally, the shooting control method provided by the embodiments of the present application further includes: adjusting the shooting parameters of the shooting module according to the annotation box.

[0136] Optionally, if the annotation box selected by the user is the second annotation box, that is, the image that the user wants to capture is the image within the second annotation box. For example, if the user only wants to capture an image of a water cup, the shooting parameters of the shooting module can be dynamically adjusted according to the size and range of the second annotation box, so that the obtained image is clearer.

[0137] Optionally, if the annotation box selected by the user is the first annotation box, that is, the image that the user wants to capture is the image within the first annotation box. For example, if the user wants to capture an image of a water cup and the environment around the water cup, the shooting parameters of the shooting module can be dynamically adjusted according to the size and range of the first annotation box (since the first annotation box is generated based on the shooting parameters of the shooting module, it is also possible not to make adjustments), so that the obtained image is clearer.

[0138] Optionally, the calculation and processing module can also be dynamically adjusted to cope with the slight movement of the user's head or the slight deviation of the line of sight. The annotation box can track the user's fixation point in real time and maintain the correct position in the image.

[0139] In the embodiment of the present application, after the annotation box is generated, the calculation and processing module will directly superimpose and display the annotation box on the user's virtual display screen. The annotation box contains the target image corresponding to the real scene picture. The calculation and processing module enters a waiting state at this time, waiting for the user's shooting confirmation instruction.

[0140] It should be noted that the display of the annotation box is implemented through the optical display module. The annotation box can also cover the area where the user is looking in a semi-transparent manner to ensure that it does not interfere with the user's observation of the real scene.

[0141] Through the embodiment of the present application, the user can accurately select the real scene position and range of the photo to be captured in real time in the real field of view in a natural interaction manner, preview the photo to be captured within the preview area of the annotation box in advance to determine whether to capture. The scene selection is accurate, the captured photo is clearer, and there is no need to continue the photo preview link in the alternative area or preview area, reducing resource occupancy.

[0142] In the embodiment of the present application, after the annotation box is generated, the calculation and processing module can project the initial image corresponding to the real scene picture in the annotation box onto the annotation box on the virtual display screen, so that the user can preview the initial image without taking a photo.

[0143] In the embodiment of the present application, after the annotation box is generated, the shooting parameters of the shooting module can be dynamically adjusted according to the annotation range of the annotation box, such as: focal length, aperture, field of view angle, and lens distortion coefficient, etc.; so that the shooting parameters conform to the size of the annotation box range, achieving the technical effect of obtaining a clearer shooting image.

[0144] In an embodiment of the present application, after the annotation box is generated, the computing and processing module can project the real scene picture in the annotation box into the annotation box on the virtual display screen, and generate a corresponding virtual image based on the initial image, add the virtual image to the initial image, generate a virtual and virtual reality combined image, and display it in the annotation box.

[0145] S203: In response to the shooting control instruction, the real scene picture within the marked frame is shot to generate a target image.

[0146] Optionally, the shooting control method provided in the embodiment of the present application also includes: receiving a shooting control instruction sent by a user; shooting and processing the real scene picture according to the shooting control instruction to generate a target image.

[0147] In an embodiment of the present application, the user can confirm the shooting in different ways, for example: continue to look at the marked box area for a period of time (such as 3 seconds), then the computing and processing module will automatically confirm that the user wants to shoot the area.

[0148] Optionally, the user can also explicitly issue a shooting command through voice commands, such as: take a photo.

[0149] Optionally, the user can also confirm the shooting through gestures.

[0150] It should be noted that during the confirmation process, the gaze tracking module will continue to track the user's gaze position to ensure that the annotation box can adjust its position at any time to adapt to the user's possible gaze movement.

[0151] As an optional embodiment, after the user confirms taking a photo by gaze, voice or gesture, the computing processing module immediately controls the shooting module to perform the shooting operation. The camera in the shooting module will quickly focus on the area covered by the marked box, perform necessary optimization such as automatic focus and exposure adjustment, and ensure that the captured image quality meets expectations.

[0152] Optionally, after shooting, the generated high-resolution image will first be stored in the local memory of the AR glasses for the user to view or edit immediately. At the same time, according to the user's settings, the image can be automatically uploaded to the cloud storage or transferred to the user's other devices, such as smartphones, tablets, etc., through the built-in communication module, such as Bluetooth, Wi-Fi or 5G communication, for backup and sharing.

[0153] Through the embodiments of the present application, the gaze point information of the user is acquired, where the gaze point information is used to represent the real - world scene picture that the user gazes at through the physical display screen of the extended reality device; a marking frame is generated on the virtual display screen according to the gaze point information, where the marking frame is used to represent the focus of the real - world scene picture; in response to a shooting control instruction, the real - world scene picture within the marking frame is shot to generate a target image. According to the user's gaze point information, the real - world scene within the user's true field of view is directly framed and shot, achieving the technical effects of improving the scene - selection accuracy and the shooting clarity.

[0154] To facilitate better implementation of the shooting control method of the present application, the present application also provides a shooting control device based on the above - mentioned shooting control method. The meanings of the nouns are the same as those in the above - mentioned shooting control method, and the specific implementation details can be referred to the descriptions in the method embodiments.

[0155] Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of the shooting control device provided by the embodiments of the present application, where the shooting control device 500 is applied to an extended reality device, and specifically as follows:

[0156] An acquisition module 501, configured to acquire the gaze point information of the user, where the gaze point information is used to represent the real - world scene picture that the user gazes at through the physical display screen of the extended reality device;

[0157] A processing module 502, configured to generate a marking frame on the virtual display screen according to the gaze point information, where the marking frame is used to represent the focus of the real - world scene picture;

[0158] The processing module 502 is further configured to, in response to a shooting control instruction, shoot the real - world scene picture within the marking frame to generate a target image.

[0159] Optionally, in some embodiments of the present application, the acquisition module 501 is configured to:

[0160] Acquire the eye image information of the user through the gaze - tracking module of the extended reality device;

[0161] Determine the pupil center position and the eyeball rotation angle according to the eye image information;

[0162] Determine the gaze point information according to the pupil center position and the eyeball rotation angle.

[0163] Optionally, in some embodiments of the present application, the processing module 502 is configured to:

[0164] Determine the display position parameters of the marking frame according to the first - pose relationship between the gaze point information and the extended reality device;

[0165] Determine the display size parameter of the annotation box according to the shooting parameters of the shooting module;

[0166] Generate the annotation box according to the display position parameter and the display size parameter.

[0167] Optionally, in some embodiments of the present application, the processing module 502 is further configured to:

[0168] Generate a first annotation box according to the display position parameter and the display size parameter;

[0169] When there is a target object in the real scene picture represented in the first annotation box, obtain the contour information of the target object;

[0170] Generate a second annotation box in the first annotation box according to the contour information;

[0171] In response to the selection control instruction, determine the annotation box from the first annotation box and the second annotation box;

[0172] Wherein, the first annotation box and the second annotation box have different presentation forms.

[0173] Optionally, in some embodiments of the present application, the processing module 502 is further configured to:

[0174] Adjust the shooting parameters of the shooting module according to the annotation box.

[0175] Optionally, in some embodiments of the present application, the processing module 502 is further configured to:

[0176] Determine the eye position of the user through the gaze tracking module of the extended reality device, and obtain the module position of the shooting module of the extended reality device;

[0177] Determine the second pose relationship according to the eye position and the module position;

[0178] Perform calculation processing on the original photo taken by the shooting module according to the second pose relationship to obtain the target image corresponding to the real scene picture.

[0179] Optionally, in some embodiments of the present application, the processing module 502 is further configured to:

[0180] Receive the shooting control instruction sent by the user;

[0181] Perform shooting processing on the real scene picture according to the shooting control instruction to generate the target image.

[0182] In the embodiment of the present application, the acquisition module 501 first acquires the gaze point information of the user, where the gaze point information is used to represent the real - world scene image that the user gazes at through the physical display screen of the extended reality device. Then, the processing module 502 generates a marking frame on the virtual display screen according to the gaze point information, where the marking frame is used to represent the focus of the real - world scene image. Then, the processing module 502 responds to the shooting control instruction and shoots the real - world scene image within the marking frame to generate a target image.

[0183] Among them, in the embodiment of the present application, the technical solution of acquiring the gaze point information of the user, where the gaze point information is used to represent the real - world scene image that the user gazes at through the physical display screen of the extended reality device; generating a marking frame on the virtual display screen according to the gaze point information, where the marking frame is used to represent the focus of the real - world scene image; and responding to the shooting control instruction to shoot the real - world scene image within the marking frame to generate a target image, directly frames and shoots in the real - world scene within the user's true field of view according to the user's gaze point information, achieving the technical effects of improving the accuracy of scene selection and the clarity of shooting.

[0184] In addition, the present application also provides an extended reality device, as Figure 6 shown, which shows the structural schematic diagram of the extended reality device involved in the present application. Specifically:

[0185] The extended reality device may include components such as a processor 601 with one or more processing cores, a memory 602 with one or more computer - readable storage media, a power supply 603, an input unit 604, and a shooting module 605. Those skilled in the art can understand that Figure 6 the structure of the extended reality device shown in

[0186] does not constitute a limitation on the extended reality device. It may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0187] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the extended reality device. In addition, the memory 602 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.

[0188] The extended reality device further includes a power supply 603 for powering each component. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 603 may also include any components such as one or more DC or AC power supplies, a recharge system, a power device debugging circuit, a power converter or inverter, and a power status indicator.

[0189] The extended reality device may further include an input unit 604, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0190] The extended reality device may further include a shooting module 605, which can be used to respond to a shooting control instruction and perform a shooting operation.

[0191] Although not shown, the extended reality device may further include a display unit, a gaze tracking module, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 601 in the extended reality device will load the executable files corresponding to the processes of one or more application programs into the memory 602 according to the following instructions, and the processor 601 will run the application programs stored in the memory 602, so as to implement the steps in any one of the shooting control methods provided in the embodiments of the present application.

[0192] Through the embodiments of the present application, the method includes obtaining the gaze point information of the user, where the gaze point information is used to represent the real - world scene image that the user gazes at through the physical display screen of the extended reality device; generating a marking frame on the virtual display screen according to the gaze point information, where the marking frame is used to represent the focus of the real - world scene image; and in response to a shooting control instruction, shooting the real - world scene image within the marking frame to generate a target image. According to the user's gaze point information, the method directly frames and shoots in the real - world scene within the user's true field of view, achieving the technical effects of improving the scene - selection accuracy and the shooting clarity.

[0193] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.

[0194] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer - readable storage medium and loaded and executed by a processor.

[0195] Therefore, the present application provides a computer - readable storage medium on which a computer program is stored. The computer program can be loaded by a processor to execute the steps in any of the shooting control methods provided by the present application.

[0196] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.

[0197] Among them, the computer - readable storage medium may include: read - only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0198] Since the instructions stored in the computer - readable storage medium can execute the steps in any of the shooting control methods provided by the present application, the beneficial effects that can be achieved by any of the shooting control methods provided by the present application can be realized. For details, reference may be made to the previous embodiments and will not be elaborated here.

[0199] The above has introduced in detail a shooting control method, device, extended reality device, and computer - readable storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, based on the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A shooting control method, characterized in that: Applied to an extended reality device, the method comprises: Acquire user's gaze point information, wherein the gaze point information is used to represent the real scene image that the user is gazing at through the physical display screen of the extended reality device; Generate a label frame on the virtual display screen according to the gaze point information, wherein the label frame is used to represent the focus of the real scene picture; the virtual display screen is generated in front of the human eye after the light emitted by the optical machine of the physical display screen enters the human eye after multiple propagation conversions; the virtual display screen is located in the real field of vision of the person; the virtual display screen is used to superimpose and display various virtual information; In response to a shooting control instruction, shooting the real scene picture in the marked frame to generate a target image; The step of generating a label frame on a virtual display screen according to the gaze point information includes: Determining the actual position of the real scene image to be captured according to the gaze point information; Determining a posture relationship between the real position and the extended reality device; Determine the positional relationship between the virtual image displayed on the virtual display screen and the extended reality device according to the internal and external parameters of the shooting module of the extended reality device and the internal and external parameters of the virtual display screen; Determine display position parameters of the annotation frame on the virtual display screen according to the posture relationship between the real position and the extended reality device, and the posture relationship between the virtual picture displayed on the virtual display screen and the extended reality device; Determining the display size of the annotation frame according to the shooting parameters of the shooting module; The annotation frame is generated according to the display size and the display position parameter.

2. The method according to claim 1, characterized in that The obtaining of the user's gaze point information includes: Acquiring eye image information of the user through the eye tracking module of the extended reality device; Determine the pupil center position and eyeball rotation angle according to the eye image information; The gaze point information is determined according to the pupil center position and the eyeball rotation angle.

3. The method according to claim 1, characterized in that The step of generating the annotation frame according to the display size and the display position parameter includes: Generate a first annotation frame according to the display position parameter and the display size; When a target object exists in the real scene picture represented in the first annotation frame, obtaining contour information of the target object; generating a second annotation frame within the first annotation frame according to the contour information; In response to a selection control instruction, determining the annotation box from the first annotation box and the second annotation box; The first annotation box and the second annotation box have different presentation forms.

4. The method according to claim 3, characterized in that After determining the annotation box from the first annotation box and the second annotation box in response to the selection control instruction, the method further includes: Adjust the shooting parameters of the shooting module according to the marking frame.

5. The method according to claim 1, characterized in that: In response to the shooting control instruction, shooting the real scene picture in the marked frame to generate a target image includes: Determine the eye position of the user through the sight tracking module of the extended reality device, and obtain the module position of the shooting module of the extended reality device; Determining a second posture relationship according to the eyeball position and the module position; The original photo taken by the shooting module is calculated and processed according to the second posture relationship to obtain the target image corresponding to the real scene picture.

6. The method according to claim 1, characterized in that The step of photographing the real scene image within the marked frame to generate a target image in response to the photographing control instruction includes: Receive shooting control instructions sent by users; The real scene picture is photographed and processed according to the shooting control instruction to generate the target image.

7. A shooting control device, characterized in that: Applied to an extended reality device, the device comprises: An acquisition module, used to acquire gaze point information of the user, wherein the gaze point information is used to represent a real scene image that the user is gazing at through a physical display screen of the extended reality device; A processing module, used for generating a labeling frame on a virtual display screen according to the gaze point information, wherein the labeling frame is used to represent the focus of the real scene picture; the virtual display screen is generated in front of the human eye after the light emitted by the optical machine of the physical display screen enters the human eye after multiple propagation conversions; the virtual display screen is located in the real field of vision of the person; the virtual display screen is used to superimpose and display various virtual information; The processing module is further used to respond to a shooting control instruction to shoot the real scene picture in the marked frame to generate a target image; Wherein, the processing module is specifically used for: Determining the actual position of the real scene image to be captured according to the gaze point information; Determining a posture relationship between the real position and the extended reality device; Determine the positional relationship between the virtual image displayed on the virtual display screen and the extended reality device according to the internal and external parameters of the shooting module of the extended reality device and the internal and external parameters of the virtual display screen; Determine display position parameters of the annotation frame on the virtual display screen according to the posture relationship between the real position and the extended reality device, and the posture relationship between the virtual picture displayed on the virtual display screen and the extended reality device; Determining the display size of the annotation frame according to the shooting parameters of the shooting module; The annotation frame is generated according to the display size and the display position parameter.

8. An extended reality device, characterized in that: The device comprises: An eye tracking module, used to obtain the user's gaze point information, wherein the gaze point information is used to represent the real scene image that the user is gazing at through the physical display screen of the extended reality device; A processing module is used to generate a label frame on a virtual display screen according to the gaze point information, wherein the label frame is used to represent the focus of the real scene picture; the virtual display screen is generated in front of the human eye after the light emitted by the optical machine of the physical display screen enters the human eye after multiple propagation conversions; the virtual display screen is located in the real field of vision of the person; the virtual display screen is used to superimpose and display various virtual information; A shooting module, used for shooting the real scene picture in the marked frame to generate a target image in response to a shooting control instruction; Wherein, the processing module is specifically used for: Determining the actual position of the real scene image to be captured according to the gaze point information; Determining a posture relationship between the real position and the extended reality device; Determine the positional relationship between the virtual image displayed on the virtual display screen and the extended reality device according to the internal and external parameters of the shooting module of the extended reality device and the internal and external parameters of the virtual display screen; Determine display position parameters of the annotation frame on the virtual display screen according to the posture relationship between the real position and the extended reality device, and the posture relationship between the virtual picture displayed on the virtual display screen and the extended reality device; Determining the display size of the annotation frame according to the shooting parameters of the shooting module; The annotation frame is generated according to the display size and the display position parameter.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the shooting control method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Shooting control method and device, electronic equipment and storage medium

    CN114079729A

  • Gaze tracking system

    US20120290401A1