Augmented reality method and apparatus, electronic device, and storage medium

By obtaining video data and determining the three-dimensional information of the real scene, responding to area selection triggers, processing and rendering the target image, the problem of single content and interaction of the existing augmented reality method is solved, and the augmented reality effect with flexible interaction and diverse content is achieved, improving the user experience.

WO2025113515A1PCT designated stage expired Publication Date: 2025-06-05BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2024/135040
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-01
Filing Date
2024-11-27
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The existing augmented reality methods are relatively single in terms of content and interaction, and cannot meet the diverse needs of users.

Method used

By obtaining video data and determining the three-dimensional information of the real scene corresponding to the video data, responding to the area selection trigger, determining the target area and processing the target image, and finally rendering the target image to the corresponding spatial location to achieve augmented reality effects with flexible interaction and diverse content.

Benefits of technology

Improve user experience, realize augmented reality effects with flexible interaction and diversified content, so that the target image can present a visual effect of time and space interlacing in real scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135040_05062025_PF_FP_ABST
    Figure CN2024135040_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in embodiments of the present disclosure are an augmented reality method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring video data, and determining three-dimensional information of a real scene corresponding to the video data; in response to an area selection trigger, determining at least a local area from the current video frame of the video data as a first target area; on the basis of the three-dimensional information of the real scene, determining the first spatial position in the real scene corresponding to an image in the first target area; processing the image in the first target area to obtain a target image; and rendering the target image at the first spatial position, so that a mapped image of the target image is presented in a target video frame, wherein the target video frame comprises a video frame temporally located after the current video frame and comprises the first spatial position.
Need to check novelty before this filing date? Find Prior Art

Description

Augmented reality method, device, electronic device and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed on December 1, 2023, with application number 202311642626.8 and invention name “An augmented reality method, device, electronic device and storage medium”. The entire contents of that application are incorporated by reference into this application. Technical Field

[0003] The embodiments of the present disclosure relate to the field of computer technology, and more particularly to an augmented reality method, device, electronic device, and storage medium. Background Art

[0004] Existing augmented reality methods can place three-dimensional models in space, but the content and interaction are relatively simple and cannot meet user needs. Summary of the Invention

[0005] The embodiments of the present disclosure provide an augmented reality method, device, electronic device, and storage medium, which have flexible interaction and diverse content, and can improve user experience.

[0006] In a first aspect, an embodiment of the present disclosure provides an augmented reality method, comprising:

[0007] Acquiring video data and determining three-dimensional information of a real scene corresponding to the video data;

[0008] In response to a region selection trigger, determining at least a local region from a current video frame of the video data as a first target region;

[0009] determining, based on the three-dimensional information of the real scene, a first spatial position in the real scene corresponding to the image within the first target area;

[0010] Processing the image within the first target area to obtain a target image;

[0011] The target image is rendered at the first spatial position so that a mapped image of the target image is presented in a target video frame; wherein the target video frame includes a video frame that is temporally subsequent to the current video frame and includes the first spatial position.

[0012] In a second aspect, an embodiment of the present disclosure further provides an augmented reality device, including:

[0013] A video acquisition module, configured to acquire video data and determine three-dimensional information of a real scene corresponding to the video data;

[0014] a region selection module, configured to determine, in response to a region selection trigger, at least a local region from a current video frame of the video data as a first target region;

[0015] a spatial position determination module, configured to determine a first spatial position in the real scene corresponding to the image within the first target area based on the three-dimensional information of the real scene;

[0016] An image processing module, configured to process the image within the first target area to obtain a target image;

[0017] An augmented reality module is used to render the target image at the first spatial position so that a mapping image of the target image is presented in a target video frame; wherein the target video frame includes a video frame that is located after the current video frame in time sequence and includes the first spatial position.

[0018] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:

[0019] one or more processors;

[0020] a storage device for storing one or more programs,

[0021] When the one or more programs are executed by the one or more processors, the one or more processors implement the augmented reality method as described in any of the embodiments of the present disclosure.

[0022] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the augmented reality method as described in any one of the embodiments of the present disclosure.

[0023] The technical solution of the embodiment of the present disclosure obtains video data and determines three-dimensional information of a real scene corresponding to the video data; in response to a region selection trigger, determines at least a local region from a current video frame of the video data as a first target region; determines a first spatial position in the real scene corresponding to an image within the first target region based on the three-dimensional information of the real scene; processes the image within the first target region to obtain a target image; renders the target image at the first spatial position so that a mapping image of the target image is presented in the target video frame; wherein the target video frame includes a video frame that is located after the current video frame in time sequence and includes the first spatial position. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0025] FIG1 is a schematic diagram of a flow chart of an augmented reality method provided by an embodiment of the present disclosure;

[0026] FIG2 is a schematic diagram of a first target area in an augmented reality method provided by an embodiment of the present disclosure;

[0027] FIG3 is a block diagram of an augmented reality method provided by an embodiment of the present disclosure;

[0028] FIG4 is a schematic diagram of an interface for selecting a region in an augmented reality method provided by an embodiment of the present disclosure;

[0029] FIG5 is a schematic diagram of a flow chart of an augmented reality method provided by an embodiment of the present disclosure;

[0030] FIG6 is a schematic diagram of an interface of an augmented reality method provided by an embodiment of the present disclosure;

[0031] FIG7 is a schematic structural diagram of an augmented reality device provided by an embodiment of the present disclosure;

[0032] FIG8 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0033] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0034] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0035] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0036] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0037] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0038] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0039] Figure 1 is a schematic flow diagram of an augmented reality method provided by an embodiment of the present disclosure. This embodiment of the present disclosure is applicable to augmented reality scenarios, such as augmented reality during video capture. The method can be performed by an augmented reality device, which can be implemented in software and / or hardware and can be configured in an electronic device, such as a mobile phone or computer.

[0040] As shown in FIG1 , the augmented reality method provided in this embodiment may include:

[0041] S110 : Acquire video data, and determine three-dimensional information of a real scene corresponding to the video data.

[0042] In the embodiments of the present disclosure, obtaining video data may include, for example, real-time acquisition of video data, or reading stored video data from a preset storage location. The scene in the video data is a real scene, and the augmented reality method provided by the present disclosure can enhance the visual effect of the real scene in the video data.

[0043] Specifically, the real scene can be reconstructed from the video frames in the video data based on existing methods to obtain three-dimensional information of the real scene.

[0044] For example, existing pose calculation methods, such as Structure from Motion (SfM), can be used to determine the camera parameters corresponding to each sample image (e.g., camera rotation and translation parameters). Based on the camera parameters, the pixel coordinates in each video frame can be projected into the world coordinate system to obtain sparse point cloud data of the real scene. Furthermore, the world coordinates of the pixels can be used to construct a Signed Distance Field (SDF) model. The three-dimensional information of the real scene can be determined based on the portion of the signed distance value output by the constructed signed distance field model where the signed distance value is zero.

[0045] For example, after determining the sparse point cloud, the Multiple View Stereo (MVS) algorithm can be used to estimate the depth image based on camera parameters, sparse point cloud and video frames. The depth image can then be fused and reconstructed to obtain a dense point cloud model to determine the three-dimensional information of the real scene.

[0046] In addition, other existing methods for reconstructing 3D information of a real scene from video frames can also be applied here, and are not exhaustively listed here. Furthermore, when reading video data, predetermined 3D information of a real scene can also be read.

[0047] S120 : In response to a region selection trigger, determine at least a local region from a current video frame of the video data as a first target region.

[0048] In the disclosed embodiments, video data can be loaded onto a display interface for display, and user interaction can be performed through the display interface to receive a region selection trigger. Furthermore, when capturing video data in real time, a region selection trigger can also be received when a preset gesture or posture is detected. Furthermore, other methods for receiving a region selection trigger can also be applied herein, and these are not exhaustive.

[0049] The augmented reality device can determine a local area or the entire image from the current video frame as a first target area based on a region selection trigger. When a local area is determined as the first target area, the first target area can include at least one local area, and each local area is independent, i.e., there is no connectivity between the local areas.

[0050] S130: Determine a first spatial position corresponding to the image in the first target area in the real scene according to the three-dimensional information of the real scene.

[0051] In the disclosed embodiment, after determining the first target area, the area corresponding to the first target area in the real scene can be determined based on the conversion relationship from pixel coordinates to real-world coordinates. Subsequently, based on the determined three-dimensional information of the real scene, three-dimensional information of the area corresponding to the first target area in the real scene can be determined, and the first spatial position can be determined based on this three-dimensional information.

[0052] Among them, determining the first spatial position based on the three-dimensional information of the area corresponding to the first target area in the real scene may include: determining the target point based on the three-dimensional information of the area corresponding to the first target area in the real scene; and determining the spatial plane based on the target point as the first spatial position.

[0053] For example, if the area corresponding to the first target area in the real scene is a planar area, the point within the planar area is the target point, and the planar area can be used as the spatial plane. For another example, if the area corresponding to the first target area in the real scene is a non-planar area, the point closest to the camera (i.e., the point with the smallest depth) can be used as the target point; the spatial plane is determined based on the camera viewing direction and the target point, such that the spatial plane contains the target point and is perpendicular to the camera viewing direction.

[0054] S140: Process the image in the first target area to obtain a target image.

[0055] In the disclosed embodiments, the image within the first target area can be processed based on existing image processing methods to obtain a target image. Existing image processing methods may include, but are not limited to, adding filters, stylization processing, and the like. Furthermore, when the image processing method includes stylization processing, scene depth information and / or character pose information of the image within the first target area can be first determined, and then, using a pre-trained generative model, stylization processing can be performed on the image within the first target area based on the depth information and / or pose information to obtain a target image similar to the image within the first target area.

[0056] By generating a target image similar to the image within the first target area, the overall scene after the target image is rendered to the first spatial location remains consistent with the original real scene, but the style of the scene at the first spatial location differs from the real scene, creating a visual effect of temporal and spatial interweaving, enhancing the user experience. For example, if the image within the first target area is a corridor, the generated target image will typically also be a corridor, but the style of the corridor in the target image differs from the real scene. This creates the visual effect of a corridor in another world within the real world, enhancing the user experience.

[0057] It is understandable that, after determining the first target area, there is no strict temporal relationship between determining the first spatial position based on step S130 and determining the target image based on step S140 .

[0058] S150, rendering the target image at the first spatial position so that a mapping image of the target image is presented in the first target video frame; wherein the first target video frame includes a video frame that is located after the current video frame in time sequence and includes the first spatial position.

[0059] In an embodiment of the present disclosure, the target image can be rendered to a first spatial position of a real scene; accordingly, based on the reverse process of "determining the first spatial position corresponding to the image within the first target area", that is, according to the conversion relationship from real-world coordinates to pixel coordinates, the target image rendered at the first spatial position can be reversely mapped to the video frame to obtain a mapped image of the target image.

[0060] Among them, the first target video frame may include the current video frame, and the video frame containing the first spatial position after the current video frame. Since the "conversion relationship from pixel coordinates to real-world coordinates" is inverse to the "conversion relationship from real-world coordinates to pixel coordinates", the area where the mapping image is presented in the current video frame is the first target area, that is, the mapping image of the target image is presented in the first target area of ​​the current video frame, so that the video frame image can be enhanced. For example, when the target image is an image of a different style (such as a cyberpunk style), a spatially interlaced visual effect can be generated; for example, when the target image is an image of different filters (such as filters of light at different time points), a temporally interlaced visual effect can be generated.

[0061] In addition, since the target image is rendered at the first spatial position of the real scene, when the video frame contains the first spatial position again after the perspective is moved, the target image at the first spatial position can still be mapped to the video frame again, thereby achieving the effect that the target image will continue to be mounted at a specific position in the real scene at any moving perspective, which can enhance the sense of reality.

[0062] In some optional implementations, processing the image within the first target area may include: performing stylized processing on the image of the first target area in at least two styles; correspondingly, rendering the target image at the first spatial position may include: cyclically rendering the target images of at least two styles at the first spatial position.

[0063] In these optional implementations, after performing at least two stylized processing on the image of the first target area, at least two styles of target images can be obtained. In this case, the at least two target images can be cyclically rendered at the first spatial location to dynamically change the image at the first spatial location. Accordingly, during the cyclic rendering of the target image at the first spatial location, if consecutive video frames contain the first spatial location, different mapped images may be presented in the consecutive video frames, thereby achieving a multi-dimensional spatiotemporal visual effect and further enhancing the user experience.

[0064] In some optional implementations, the first target area may include at least two; accordingly, rendering the target image at the first spatial position may include: rendering each target image at the corresponding first spatial position respectively.

[0065] For example, Figure 2 is a schematic diagram of a first target area in an augmented reality method provided by an embodiment of the present disclosure. Referring to Figure 2 , three posters on a wall can be simultaneously identified as first target areas. Furthermore, the first spatial positions corresponding to the three first target areas can be determined separately, and the images within the three target areas can be processed separately to obtain target images. Finally, each target image can be rendered at the corresponding first spatial position, thereby improving the efficiency of augmented reality and further enhancing the user experience.

[0066] In some optional implementations, after rendering the target image at the first spatial position, the method may further include: adding a target special effect to the target image in response to a special effect adding operation.

[0067] The special effect adding operation may be received through a display interface or other means. According to the special effect adding operation, a target special effect is added to the target image, which may include, for example, adding visual special effects and / or audio special effects to the target image and / or the video frame corresponding to the target image. The visual special effects may include, for example, the presentation of the edge lines of the target image (such as a bright edge effect), the presentation of the content of the target image (such as a blinds effect), and the transition of the current video frame before and after the mapping image of the target image is presented. The audio special effects may include, for example, the audio special effects of the current video frame before and after the mapping image of the target image is presented.

[0068] In these optional implementations, by adding target special effects to the target image, the visual presentation effect can be further enriched and the user experience can be improved.

[0069] In some optional implementations, when video data is collected in real time, after the target image is rendered at the first spatial position, the method may further include: continuing to collect video data.

[0070] In these optional implementations, after the current video frame is rendered during the capture process, video data can continue to be captured, and the rendered target image will continue to be presented at the first spatial position in the real scene. Furthermore, during subsequent video data capture, the aforementioned steps can be repeated, including selecting the first target area, determining the first spatial position and target image corresponding to the first target area, and rendering the new target image, until video capture is completed, resulting in augmented reality video data.

[0071] For example, Figure 3 is a block diagram of an augmented reality method provided by an embodiment of the present disclosure. Referring to Figure 3, video data can first be presented through a display interface, and three-dimensional information of a real scene can be determined. Then, a first target area can be determined through interaction with the display interface, and an image of the first target area can be processed to obtain a target image. Simultaneously, a first spatial position corresponding to the first target area can be determined based on the three-dimensional information of the real scene. Finally, the target image can be rendered at the first spatial position, and a mapped image of the target image can be presented in the current video frame.

[0072] The technical solution of the embodiment of the present disclosure obtains video data and determines three-dimensional information of a real scene corresponding to the video data; in response to a region selection trigger, determines at least a local region from a current video frame of the video data as a first target region; determines a first spatial position in the real scene corresponding to an image within the first target region based on the three-dimensional information of the real scene; processes the image within the first target region to obtain a target image; and renders the target image at the first spatial position so that in video frames containing the first spatial position after the current video frame, the region corresponding to the first spatial position presents the target image.

[0073] By selecting an area in a video frame corresponding to a real scene, generating a target image based on the image within the area, and rendering the target image to the spatial location corresponding to the selected area in the real scene, it is possible to add processed image content to the real scene, achieving a spatially interlaced visual effect. The target image can also be presented when the video frame contains the same spatial location again, enhancing the realism of the effect. Because the selection area is flexible and different target images can be generated for different areas, it enables flexible interaction and diverse content in augmented reality, improving the user experience.

[0074] The embodiments of this disclosure can be combined with the various optional solutions in the augmented reality method provided in the above embodiments. The augmented reality method provided in this embodiment describes in detail the steps for determining the first target area. Determining the first target area by detecting predicted gestures and / or displaying interface interactions can achieve selection flexibility and improve user experience.

[0075] FIG4 is a schematic diagram of an interface for region selection in an augmented reality method provided by an embodiment of the present disclosure. As shown in FIG4 , in the augmented reality method provided by this embodiment, in response to a region selection trigger, determining at least a local region from the current video frame of the video data as a first target region may include at least one of the following:

[0076] In response to detecting a preset gesture in the video data, at least a local area corresponding to the preset gesture is determined from the current video frame as a first target area; in response to an area selection operation of the current video frame in the video data, at least a local area corresponding to the area selection operation is determined from the current video frame as the first target area.

[0077] In the disclosed embodiments, at least one preset gesture can be pre-set, and a correspondence between each preset gesture and a region type can be configured. When a preset gesture is detected in video data, it can be considered that a region selection trigger has been received. Based on the correspondence, at least a partial region corresponding to the preset gesture can be determined from the current video frame as the first target region.

[0078] For example, in Figure 4(a), if the preset gesture includes a snap gesture, and the snap gesture corresponds to the entire video frame area, after the snap gesture is detected in the video data, the entire current video frame can be used as the first target area. For another example, if the preset gesture includes a clap gesture, and the clap gesture corresponds to a poster area, after the clap gesture is detected in the video data, the poster area in the current video frame can be used as the first target area. Furthermore, if at least one local area is determined in the current video frame, the first target area can be selected from the at least one local area based on a user selection operation.

[0079] In the embodiment of the present disclosure, the video data can also be loaded into the display interface for display, and the display interface can be used to interact with the user to receive a region selection trigger, wherein the region selection trigger can include at least one step of region selection operation.

[0080] For example, Figure 4(b) shows that the interface can deploy a first control for automatic region segmentation, and in response to the triggering of the first control, the current video frame of the video data can be automatically segmented to obtain at least one segmented region; and the user input for the selection operation of the segmented region can be received to determine at least a local region as the first target region.

[0081] For example, the display interface may deploy a second control for manual area segmentation, and may enter a manual segmentation interface in response to the triggering of the second control; receive segmentation operations input by the user (such as operations such as painting areas) in the manual segmentation interface, and determine at least a local area as the first target area based on the segmentation operation, etc.

[0082] The technical solution of the embodiment of the present disclosure describes in detail the steps of determining the first target area. By detecting the predicted gesture and / or displaying the interface interaction to determine the first target area, flexibility in selection can be achieved and the user experience can be improved. The augmented reality method provided in the embodiment of the present disclosure and the augmented reality method provided in the above embodiment belong to the same public concept. The technical details not fully described in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and in the above embodiment.

[0083] The various optional solutions in the augmented reality method provided in the embodiments of the present disclosure and the above embodiments can be combined. In the augmented reality method provided in this embodiment, when the first target area includes an instance segmentation area, the target image can be rendered to a first spatial position and then to other spatial positions, thereby being able to present multiple processed instances to enrich the presentation effect.

[0084] FIG5 is a flow chart of an augmented reality method provided by an embodiment of the present disclosure. As shown in FIG5 , the augmented reality method provided by this embodiment may include:

[0085] S510: Acquire video data, and determine three-dimensional information of a real scene corresponding to the video data.

[0086] S520 : In response to a region selection trigger, determine at least a local region from a current video frame of the video data as a first target region; wherein the first target region includes an instance segmentation region.

[0087] S530: Determine a first spatial position corresponding to the image in the first target area in the real scene according to the three-dimensional information of the real scene.

[0088] S540: Process the image in the first target area to obtain a target image.

[0089] There is no strict timing restriction for step S530 and step S540 .

[0090] S550: Render the target image at the first spatial position, so that the target image is presented in the area corresponding to the first spatial position in the video frames after the current video frame that include the first spatial position.

[0091] S560: In response to a repeated rendering trigger, determine a second target area from a current video frame of the video data; wherein the first target area is different from the second target area.

[0092] S570: Determine a second spatial position in the real scene corresponding to the image in the second target area according to the three-dimensional information of the real scene.

[0093] S580, rendering the target image at the second spatial position so that a mapping image of the target image is presented in a second target video frame; wherein the second target video frame includes a video frame that is located after the current video frame in time sequence and includes the second spatial position.

[0094] In the disclosed embodiments, the instance segmentation region may include, for example, a human body region or an object region. A repeat rendering trigger may be received via a display interface of the video data. For example, a mapping image of a target image in a current video frame may be provided with a repeat rendering control. When the control is triggered, a repeat rendering operation may be received.

[0095] In response to a repeated rendering trigger, at least one local area can be determined from the current video frame of the video data as the second target area (for example, the area after the mapped image is dragged can be used as the second target area); wherein the second target area is different from the first target area; based on the three-dimensional information of the real scene, the second spatial position corresponding to the image within each second target area in the real scene can be determined; and the target image can be rendered at each second spatial position.

[0096] Exemplarily, FIG6 is an interface diagram of an augmented reality method provided by an embodiment of the present disclosure. In FIG6, the instance segmented area is a human body area; processing the human body area may include, for example, changing the human body area, changing the human body area into an image of another style (for example, anime style), etc., to obtain a target image. As shown in FIG6, when the target image is rendered at the first spatial position corresponding to the human body area, a control for repeated rendering (for example, a "+1" control associated with the mapping image of the target image in the current video frame) may also be triggered, and the target image may be dragged to the second target area, and the target image may be rendered to the second spatial position corresponding to the image in the second target area. In the current video frame and subsequent video frames, in the case of containing the first spatial position and / or the second spatial position, the area corresponding to the first spatial position and / or the second spatial position in the video frame may present a mapping image of the target image.

[0097] The technical solution of the embodiment of the present disclosure, when the first target area includes an instance segmentation area, can render the target image to the first spatial position and then to other spatial positions, thereby presenting multiple processed instances to enrich the presentation effect. The augmented reality method provided by the embodiment of the present disclosure and the augmented reality method provided by the above embodiment belong to the same disclosed concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.

[0098] Figure 7 is a schematic diagram of the structure of an augmented reality device provided by an embodiment of the present disclosure. The augmented reality device provided by this embodiment is applicable to augmented reality scenarios, for example, to augmented reality scenarios during video shooting.

[0099] As shown in FIG7 , the augmented reality device provided by the embodiment of the present disclosure may include:

[0100] The video acquisition module 710 is used to acquire video data and determine three-dimensional information of a real scene corresponding to the video data;

[0101] a region selection module 720 for determining at least a local region from a current video frame of the video data as a first target region in response to a region selection trigger;

[0102] A spatial position determination module 730 is configured to determine a first spatial position corresponding to an image within a first target area in the real scene based on three-dimensional information of the real scene;

[0103] An image processing module 740 is configured to process the image within the first target area to obtain a target image;

[0104] The augmented reality module 750 is used to render the target image at a first spatial position so that a mapping image of the target image is presented in a first target video frame; wherein the first target video frame includes a video frame that is located after the current video frame in time sequence and includes the first spatial position.

[0105] In some optional implementations, the region selection module may perform at least one of the following:

[0106] In response to detecting a preset gesture in the video data, determining at least a local area corresponding to the preset gesture from the current video frame as a first target area;

[0107] In response to a region selection operation of a current video frame in the video data, at least a local region corresponding to the region selection operation is determined from the current video frame as a first target region.

[0108] In some optional implementations, the image processing module is configured to perform stylization processing of the image of the first target area in at least two styles;

[0109] Accordingly, the augmented reality module can be used to cyclically render target images of at least two styles at the first spatial position.

[0110] In some optional implementations, the first target area includes at least two;

[0111] Correspondingly, the augmented reality module can be used to render each target image at the corresponding first spatial position.

[0112] In some optional implementations, when the first target region includes an instance segmentation region, after rendering the target image at the first spatial location, the region selection module may be configured to determine, in response to a repeated rendering trigger, a second target region from a current video frame of the video data; wherein the first target region is different from the second target region;

[0113] A spatial position determination module may be configured to determine a second spatial position corresponding to an image within a second target area in the real scene based on three-dimensional information of the real scene;

[0114] The augmented reality module can be used to render the target image at a second spatial position so that a mapping image of the target image is presented in a second target video frame; wherein the second target video frame includes a video frame that is located after the current video frame in time sequence and includes the second spatial position.

[0115] In some optional implementations, the augmented reality module is further configured to:

[0116] After the target image is rendered at the first spatial position, a target special effect is added to the target image in response to a special effect adding operation.

[0117] In some optional implementations, the video acquisition module is used to collect video data in real time;

[0118] Correspondingly, after the augmented reality module renders the target image at the first spatial position, the video acquisition module is further configured to continue collecting video data.

[0119] The augmented reality device provided by the embodiments of the present disclosure can execute the augmented reality method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0120] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0121] Reference is now made to FIG8 , which illustrates a schematic diagram of the structure of an electronic device (e.g., a terminal device or server in FIG8 ) 800 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG8 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.

[0122] As shown in Figure 8, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0123] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although FIG8 shows the electronic device 800 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0124] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the augmented reality method of the embodiment of the present disclosure are performed.

[0125] The electronic device provided by the embodiment of the present disclosure and the augmented reality method provided by the above embodiment belong to the same disclosed concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0126] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the augmented reality method provided by the above embodiment is implemented.

[0127] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0128] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0129] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0130] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0131] Acquire video data and determine three-dimensional information of a real scene corresponding to the video data; in response to a region selection trigger, determine at least a local region from a current video frame of the video data as a first target region; determine a first spatial position in the real scene corresponding to an image within the first target region based on the three-dimensional information of the real scene; process the image within the first target region to obtain a target image; render the target image at the first spatial position so that a mapping image of the target image is presented in a first target video frame; wherein the first target video frame includes a video frame that is located after the current video frame in time sequence and includes the first spatial position.

[0132] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0134] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the names of the units and modules do not, in certain circumstances, limit the units and modules themselves.

[0135] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and the like.

[0136] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0137] According to one or more embodiments of the present disclosure, there is provided an augmented reality method, the method comprising:

[0138] Acquiring video data and determining three-dimensional information of a real scene corresponding to the video data;

[0139] In response to a region selection trigger, determining at least a local region from a current video frame of the video data as a first target region;

[0140] determining, based on the three-dimensional information of the real scene, a first spatial position in the real scene corresponding to the image within the first target area;

[0141] Processing the image within the first target area to obtain a target image;

[0142] The target image is rendered at the first spatial position so that a mapped image of the target image is presented in a first target video frame; wherein the first target video frame includes a video frame that is temporally located after the current video frame and includes the first spatial position.

[0143] According to one or more embodiments of the present disclosure, there is provided an augmented reality method, further comprising:

[0144] In some optional implementations, in response to the region selection trigger, determining at least a local region from the current video frame of the video data as the first target region includes at least one of the following:

[0145] In response to detecting a preset gesture in the video data, determining at least a local area corresponding to the preset gesture from the current video frame as a first target area;

[0146] In response to a region selection operation of a current video frame in the video data, at least a local region corresponding to the region selection operation is determined from the current video frame as a first target region.

[0147] According to one or more embodiments of the present disclosure, there is provided an augmented reality method, further comprising:

[0148] In some optional implementations, the processing the image in the first target area includes: performing stylization processing of the image in the first target area in at least two styles;

[0149] Correspondingly, rendering the target image at the first spatial position includes: cyclically rendering the target images of the at least two styles at the first spatial position.

[0150] According to one or more embodiments of the present disclosure, there is provided an augmented reality method, further comprising:

[0151] In some optional implementations, the first target area includes at least two;

[0152] Correspondingly, rendering the target image at the first spatial position includes: rendering each target image at a corresponding first spatial position respectively.

[0153] According to one or more embodiments of the present disclosure, there is provided an augmented reality method, further comprising:

[0154] In some optional implementations, when the first target region includes an instance segmentation region, after rendering the target image at the first spatial position, the method further includes:

[0155] In response to a repeat rendering trigger, determining a second target area from a current video frame of the video data; wherein the first target area is different from the second target area;

[0156] determining, based on the three-dimensional information of the real scene, a second spatial position in the real scene corresponding to the image within the second target area;

[0157] The target image is rendered at the second spatial position so that a mapped image of the target image is presented in a second target video frame; wherein the second target video frame includes a video frame that is temporally located after the current video frame and includes the second spatial position.

[0158] According to one or more embodiments of the present disclosure, there is provided an augmented reality method, further comprising:

[0159] In some optional implementations, after rendering the target image at the first spatial position, the method further includes:

[0160] In response to the special effect adding operation, a target special effect is added to the target image.

[0161] According to one or more embodiments of the present disclosure, there is provided an augmented reality method, further comprising:

[0162] In some optional implementations, the acquiring video data includes: collecting video data in real time;

[0163] Correspondingly, after rendering the target image at the first spatial position, the method further includes: continuing to collect video data.

[0164] According to one or more embodiments of the present disclosure, an augmented reality device is provided, the device comprising:

[0165] A video acquisition module, configured to acquire video data and determine three-dimensional information of a real scene corresponding to the video data;

[0166] a region selection module, configured to determine, in response to a region selection trigger, at least a local region from a current video frame of the video data as a first target region;

[0167] a spatial position determination module, configured to determine a first spatial position in the real scene corresponding to the image within the first target area based on the three-dimensional information of the real scene;

[0168] An image processing module, configured to process the image within the first target area to obtain a target image;

[0169] An augmented reality module is used to render the target image at the first spatial position so that a mapping image of the target image is presented in a target video frame; wherein the target video frame includes a video frame that is located after the current video frame in time sequence and includes the first spatial position.

[0170] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0171] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0172] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An augmented reality method, comprising: Acquire video data, and determine three-dimensional information of a real scene corresponding to the video data; In response to a region selection trigger, determining at least a local region from a current video frame of the video data as a first target region; Determining, according to the three-dimensional information of the real scene, a first spatial position in the real scene corresponding to the image in the first target area; Processing the image in the first target area to obtain a target image; The target image is rendered at the first spatial position so that a mapped image of the target image is presented in a first target video frame; wherein the first target video frame includes a video frame that is temporally located after the current video frame and includes the first spatial position.

2. The method according to claim 1, wherein: In response to the region selection trigger, determining at least a local region from the current video frame of the video data as the first target region includes at least one of the following: In response to detecting a preset gesture in the video data, determining at least a local area corresponding to the preset gesture from a current video frame as a first target area; In response to a region selection operation of a current video frame in the video data, at least a local region corresponding to the region selection operation is determined from the current video frame as a first target region.

3. The method according to claim 1, wherein: The processing of the image in the first target area includes: performing stylization processing of the image in the first target area in at least two styles; Correspondingly, rendering the target image at the first spatial position includes: cyclically rendering the target images of the at least two styles at the first spatial position.

4. The method according to claim 1, wherein: The first target area includes at least two; Correspondingly, rendering the target image at the first spatial position includes: rendering each of the target images at a corresponding first spatial position respectively.

5. The method according to claim 1, wherein: In a case where the first target region includes an instance segmentation region, after rendering the target image at the first spatial position, the method further includes: In response to a repeat rendering trigger, determining a second target area from a current video frame of the video data; wherein the first target area is different from the second target area; Determining, according to the three-dimensional information of the real scene, a second spatial position in the real scene corresponding to the image in the second target area; The target image is rendered at the second spatial position so that a mapped image of the target image is presented in a second target video frame; wherein the second target video frame includes a video frame that is temporally located after the current video frame and includes the second spatial position.

6. The method according to claim 1, wherein: After rendering the target image at the first spatial position, the method further includes: In response to the special effect adding operation, a target special effect is added to the target image.

7. The method according to any one of claims 1 to 6, wherein: The obtaining of video data includes: collecting video data in real time; Correspondingly, after rendering the target image at the first spatial position, the method further includes: continuing to collect video data.

8. An augmented reality device, comprising: A video acquisition module, used to acquire video data and determine three-dimensional information of a real scene corresponding to the video data; A region selection module, configured to determine at least a local region from a current video frame of the video data as a first target region in response to a region selection trigger; A spatial position determination module, configured to determine a first spatial position in the real scene corresponding to the image in the first target area according to the three-dimensional information of the real scene; An image processing module, used for processing the image in the first target area to obtain a target image; An augmented reality module is used to render the target image at the first spatial position so that a mapping image of the target image is presented in a target video frame; wherein the target video frame includes a video frame that is located after the current video frame in time sequence and includes the first spatial position.

9. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the augmented reality method as described in any one of claims 1 to 7.

10. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to perform the augmented reality method according to any one of claims 1 to 7 when executed by a computer processor.

Citation Information

Patent Citations

  • Augmented reality scene processing method and device, terminal device and system for realization of augmented reality scene

    CN107481327A

  • Display method and device based on augmented reality, equipment and storage medium

    CN112672185A

  • Image processing method and device, electronic equipment and storage medium

    CN114841984A

  • Augmented reality live broadcast method for projecting real person anchor image to remote real environment

    CN116801037A

  • Augmented reality creation using a real scene

    US20130307875A1

Cited By

  • Airport simulation scene and video fusion method for abnormal flyer identification

    CN121545104A