Image processing method and device of virtual scene, electronic equipment and storage medium

By generating and offsetting the left and right eye views, the problem of high computational resource consumption for naked-eye 3D effects in virtual scenes is solved, achieving efficient stereoscopic visual effects.

CN122120426APending Publication Date: 2026-05-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2024-11-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies consume high computational resources when generating naked-eye 3D effects for virtual scenes, which affects display efficiency.

Method used

By generating a first left view from the left eye perspective and a first right view from the right eye perspective, and offsetting them according to the offset value, a second left view and a second right view are obtained. These views are used to achieve a naked-eye 3D effect without wearing special glasses.

Benefits of technology

It saves computing resources needed to achieve naked-eye 3D effects, ensuring the accuracy and display efficiency of stereoscopic visual effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120426A_ABST
    Figure CN122120426A_ABST
Patent Text Reader

Abstract

The application provides an image processing method and device for a virtual scene, an electronic device and a storage medium. The method comprises: generating a first left view at a left eye perspective and a first right view at a right eye perspective for an initial image of a virtual scene; obtaining a left view offset value at the left eye perspective and a right view offset value at the right eye perspective; offsetting the first left view according to the left view offset value to obtain a second left view, and offsetting the first right view according to the right view offset value to obtain a second right view; and displaying the virtual scene in a naked-eye three-dimensional mode based on the second left view and the second right view. According to the application, the required computing resources for realizing the naked-eye three-dimensional effect of the virtual scene can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to computer technology, and more particularly to an image processing method, apparatus, electronic device, and storage medium for virtual scenes. Background Technology

[0002] Humans have a certain degree of parallax between their two eyes, causing the images seen by the left and right eyes to be slightly different. The brain processes these differences to perceive depth and three-dimensional space. Naked-eye 3D (3D) technology generates left and right parallax maps, sends these maps to a special program and a 3D display device, and the display device outputs left and right eye maps separately according to the observer's position, so that each eye sees different content, thus producing a stereoscopic effect.

[0003] In related technologies, for real-world scenes, two cameras can be used to capture images to obtain left and right disparity maps, and then a naked-eye 3D effect can be constructed based on these maps. For virtual scenes, a virtual camera can capture left and right eye images; however, the above methods consume significant computing resources, affecting the display efficiency of naked-eye 3D effects in virtual scenes. Summary of the Invention

[0004] This application provides an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product for virtual scenes, which can save the computing resources required to achieve naked-eye 3D effects in virtual scenes.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] This application provides an image processing method for a virtual scene, the method comprising:

[0007] For the initial image of the virtual scene, generate a first left view from the left eye perspective and a first right view from the right eye perspective;

[0008] Obtain the left view offset value from the left eye perspective and the right view offset value from the right eye perspective;

[0009] The first left view is offset according to the left view offset value to obtain the second left view, and the first right view is offset according to the right view offset value to obtain the second right view;

[0010] The virtual scene is displayed in naked-eye 3D mode based on the second left view and the second right view.

[0011] This application provides an image processing method for a virtual scene, the method comprising:

[0012] The virtual scene is displayed in a two-dimensional plane mode in the human-computer interaction interface;

[0013] In response to the switching operation of the two-dimensional plane mode, the virtual scene is displayed in a naked-eye three-dimensional mode, wherein the naked-eye three-dimensional mode is implemented by the image processing method of the virtual scene described in the embodiments of this application.

[0014] This application provides an image processing apparatus for a virtual scene, comprising:

[0015] The rendering module is used to generate a first left view from the left eye perspective and a first right view from the right eye perspective based on the initial image of the virtual scene.

[0016] The parameter acquisition module is used to acquire the left view offset value under the left eye view and the right view offset value under the right eye view.

[0017] The rendering module is further configured to offset the first left view according to the left view offset value to obtain the second left view, and to offset the first right view according to the right view offset value to obtain the second right view;

[0018] The display module is used to display the virtual scene in naked-eye 3D mode based on the second left view and the second right view.

[0019] This application provides an image processing apparatus for a virtual scene, comprising:

[0020] The display module is used to display virtual scenes in a two-dimensional plane mode in the human-computer interaction interface;

[0021] The display module is configured to display the virtual scene in a naked-eye 3D mode in response to a switching operation of the two-dimensional plane mode, wherein the naked-eye 3D mode is implemented by the image processing method of the virtual scene described in the embodiments of this application.

[0022] This application provides an electronic device, the electronic device comprising:

[0023] Memory is used to store executable instructions or computer programs.

[0024] The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the image processing method for virtual scenes provided in the embodiments of this application.

[0025] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs, which, when executed by a processor, implement the image processing method for a virtual scene provided in this application.

[0026] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the image processing method for virtual scenes provided in this application.

[0027] The embodiments of this application have the following beneficial effects:

[0028] For an initial image of a virtual scene, a first left view from the left-eye perspective and a first right view from the right-eye perspective are generated. Compared to related technologies that add a virtual camera to the virtual scene to obtain a disparity map or estimate depth to create a naked-eye 3D effect, generating images based on the initial image saves the computational resources required to achieve a naked-eye 3D effect. The first left view is offset according to the left view offset value to obtain a second left view, and the first right view is offset according to the right view offset value to obtain a second right view. By creating left-eye and right-eye views and appropriately offsetting them, the stereoscopic effect seen by the human eye can be simulated, thereby achieving a naked-eye 3D effect. Performing view generation and view offsetting in stages for either the left or right view allows for more precise control over the generation and offsetting process of each view, ensuring the accuracy of the final stereoscopic visual effect. Attached Figure Description

[0029] Figure 1A This is a schematic diagram of the first application mode of the image processing method for virtual scenes provided in the embodiments of this application;

[0030] Figure 1B This is a schematic diagram of the second application mode of the image processing method for virtual scenes provided in the embodiments of this application;

[0031] Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0032] Figure 3A This is a first flowchart illustrating the image processing method for a virtual scene provided in this application embodiment;

[0033] Figure 3B This is a second flowchart illustrating the image processing method for a virtual scene provided in this application embodiment;

[0034] Figure 3C This is a schematic diagram of the third process of the image processing method for a virtual scene provided in the embodiments of this application;

[0035] Figure 3D This is a schematic diagram of the fourth process of the image processing method for virtual scenes provided in the embodiments of this application;

[0036] Figure 4This is a schematic diagram of the fifth process of the image processing method for virtual scenes provided in the embodiments of this application;

[0037] Figure 5 This is a schematic diagram illustrating the principle of the naked-eye 3D mode provided in the embodiments of this application;

[0038] Figure 6 This is a schematic diagram illustrating the principle of the image processing method for virtual scenes provided in the embodiments of this application;

[0039] Figure 7A This is a first schematic diagram of the human-computer interaction interface provided in the embodiments of this application;

[0040] Figure 7B This is a second schematic diagram of the human-computer interaction interface provided in the embodiments of this application;

[0041] Figure 7C This is a third schematic diagram of the human-computer interaction interface provided in the embodiments of this application;

[0042] Figure 8A This is a schematic diagram of the component tree structure provided in the embodiments of this application;

[0043] Figure 8B This is a schematic diagram of the structure of the graphic layer tree provided in the embodiments of this application.

[0044] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0046] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0047] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0048] In this application, the facial (or other biometric) recognition technology (e.g., collecting images of a user's face or eyes) involved should be used in specific products or technologies in accordance with relevant laws and regulations when the above embodiments of this application are applied. Before collecting facial information, the information processing rules should be communicated and the individual consent of the target should be obtained. Facial information should be processed in strict accordance with the requirements of laws and regulations and personal information processing rules, and technical measures should be taken to ensure the security of relevant data.

[0049] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0051] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0052] 1) Virtual scenes utilize the scene output by the device that is different from the real world. Visual perception of the virtual scene can be formed with the naked eye or with the assistance of the device. For example, two-dimensional images are output through a display screen, and three-dimensional images are output through stereoscopic display technologies such as stereoscopic projection, virtual reality and augmented reality. In addition, various possible hardware can be used to form various perceptions that simulate the real world, such as auditory perception, tactile perception, olfactory perception and motion perception.

[0053] 2) Human-Computer Interaction Interface: An interface used to provide human-computer interaction functions and display information flow. Examples include Graphical User Interface (GUI) displays, Augmented Reality (AR) interfaces, Virtual Reality (VR) interfaces, Voice User Interface (VUI) interfaces, Interactive Projection Interfaces (using projection technology to display information on a flat surface), Eye-tracking interfaces (interfaces controlled by detecting the user's gaze), Holographic interfaces (three-dimensional holograms formed by projecting images using holographic projection technology, allowing viewing of stereoscopic images without special glasses), Multimodal interfaces (interfaces combining multiple interaction methods, such as tactile, visual, and auditory interaction), and Brain-Machine Interface (BMI) interfaces.

[0054] 3) Naked-eye 3D mode: A display technology that enables the perception of 3D effects without wearing any special glasses. It is a general term for technologies that achieve stereoscopic vision effects without the aid of external tools such as polarized glasses.

[0055] 4) Widget Tree: This is a data structure used to describe the user interface of an application. This data structure stores rendered content. The widget tree is a configuration data structure; its creation is very lightweight, and it is rebuilt during page refreshes.

[0056] 5) Image segmentation models: These are models in computer vision used to segment images into multiple parts or objects. They can identify and separate different regions in an image, each region potentially representing a single object or a group of objects with similar features. Common image segmentation models include: thresholding models, edge detection models, deep learning segmentation models (e.g., Convolutional Neural Networks (CNNs) using deep learning techniques), Fully Convolutional Networks (FCNs), and U-Nets. Image segmentation models are widely used in medical image analysis, object detection, image compression, and image editing. The choice of a suitable segmentation model typically depends on the specific application scenario, image type, and required segmentation accuracy.

[0057] In related technologies, glasses-free 3D technology generates left and right disparity maps, sends these maps to a special program and a 3D display device, and the display device outputs left and right eye maps according to the observer's position, so that each eye sees different content, thus producing a stereoscopic effect. For real-world scenes, two cameras can be used to capture images to obtain left and right disparity maps, and then a glasses-free 3D effect can be constructed based on these maps. For virtual scenes, left and right eye maps can be captured using a virtual camera; however, the above methods consume high computational resources, affecting the display efficiency of glasses-free 3D effects in virtual scenes. Depth estimation can also be performed on images to generate glasses-free 3D images, but this still suffers from high computational resource consumption, stuttering, and frame drops, affecting the display quality of glasses-free 3D.

[0058] This application provides an image processing method for virtual scenes, an image processing device for virtual scenes, an electronic device, a computer-readable storage medium, and a computer program product, which can save the computing resources required to achieve naked-eye 3D effects in virtual scenes.

[0059] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. These electronic devices can be implemented as terminal devices, such as laptops, tablets, desktop computers, set-top boxes, smart TVs, in-vehicle terminals, virtual reality (VR) devices, augmented reality (AR) devices, and other various types of terminals. They can also be implemented as servers. The following will describe exemplary applications when the electronic device is implemented as a terminal device or a server.

[0060] In one implementation scenario, see Figure 1A , Figure 1A This is a schematic diagram of the first application mode of the image processing method for virtual scenes provided in this application embodiment. It is applicable to some application modes that can complete the relevant data calculation of virtual scenes by relying entirely on the graphics processing hardware computing power of the terminal device 400, such as stand-alone / offline games, and complete the output of virtual scenes through various types of terminal devices 400 such as smartphones, tablets and virtual reality / augmented reality devices.

[0061] As an example, types of graphics processing hardware include central processing units (CPUs) and graphics processing units (GPUs).

[0062] When visual perception of a virtual scene is formed, the terminal device 400 calculates the data required for display through graphics computing hardware, and completes the loading, parsing and rendering of the display data. The graphics output hardware outputs video frames that can form visual perception of the virtual scene. For example, two-dimensional video frames are displayed on the display screen of a smartphone, or video frames that achieve a three-dimensional display effect are projected onto the lenses of augmented reality / virtual reality glasses. In addition, in order to enrich the perceptual effect, the terminal device 400 can also use different hardware to form one or more of auditory perception, tactile perception, motion perception and taste perception.

[0063] As an example, terminal device 400 runs a client (e.g., a standalone game application). During the operation of the client, it outputs a virtual scene for role-playing. The virtual scene can be an environment for game characters to interact with, such as plains, streets, valleys, etc., for game characters to fight. The first virtual object can be a game character controlled by the user, that is, the first virtual object is controlled by the real user and will move in the virtual scene in response to the real user's operation on the controller (e.g., touch screen, voice switch, keyboard, mouse, and joystick). For example, when the real user moves the joystick to the right, the first virtual object will move to the right in the virtual scene. It can also remain stationary, jump, and be controlled to shoot. Alternatively, terminal device 400 runs a video player that can play videos of the virtual scene in naked-eye 3D mode.

[0064] For example, the virtual scene can be a game virtual scene, and the user can be a player. The following will illustrate this with the example above.

[0065] For example, the human-computer interaction interface in terminal device 400 displays a virtual scene in a two-dimensional planar mode, responding to a user's trigger operation on terminal device 400 or a trigger operation from a control device associated with terminal device (e.g., a game controller, mouse, motion sensing device, or keyboard). The trigger operation is used to trigger a display mode switch. Terminal device 400 runs the image processing method for the virtual scene provided in this application embodiment to obtain a second left view and a second right view that are offset from the initial image. Terminal device 400 switches the display mode of the virtual scene from a two-dimensional planar mode to a naked-eye 3D mode, and then screen 101 in terminal device 400 switches to screen 102. In naked-eye 3D mode, terminal device 400 achieves a naked-eye 3D effect based on the second left view and the second right view.

[0066] The display of terminal device 400 is a display with naked-eye 3D mode functionality. This naked-eye 3D mode can be achieved through either a slit-type liquid crystal grating or a lenticular lens. Slit-type liquid crystal grating technology involves adding a slit-like grating in front of the display screen. Based on this slit-like grating, when the image for the left eye is displayed on the LCD screen, opaque stripes block the right eye; similarly, when the image for the right eye is displayed, opaque stripes block the left eye. By separating the visible images for the left and right eyes, the viewer sees a 3D image. Lens technology works by using the refraction principle of a lens to project corresponding pixels for the left and right eyes separately, achieving image separation. The biggest advantage of slit-type grating technology compared to lenticular lens technology is that the lens does not block light, resulting in a significant improvement in brightness.

[0067] Figure 1B This is a schematic diagram of the second application mode of the image processing method for virtual scenes provided in the embodiments of this application.

[0068] In the Figure 1B Before proceeding, let's first introduce the game modes involved in the terminal device and server collaborative implementation scheme. This scheme primarily involves two game modes: local game mode and cloud game mode. In local game mode, the terminal device and server collaboratively run the game processing logic. The player's input commands on the terminal device are partly processed by the terminal device's game logic, and partly by the server. Furthermore, the server-side game logic processing is often more complex and requires more computing power. In cloud game mode, the server handles all game logic processing, and the cloud server renders the game scene data into audio and video streams, which are then transmitted to the terminal device for display over the network. The terminal device only needs basic streaming media playback capabilities and the ability to receive player commands and send them to the server.

[0069] In another implementation scenario, see Figure 1B , Figure 1B This is a schematic diagram of the application mode of the image processing method for virtual scenes provided in the embodiments of this application. It is applied to terminal device 400 and server 200, and is suitable for application modes that rely on the computing power of server 200 to complete virtual scene calculation and output virtual scene on terminal device 400.

[0070] Taking the visual perception of forming a virtual scene as an example, server 200 calculates display data related to the virtual scene (such as scene data) and sends it to terminal device 400 via network 300. Terminal device 400 relies on graphics computing hardware to load, parse, and render the calculated display data, and relies on graphics output hardware to output the virtual scene to form visual perception. For example, it can display two-dimensional video frames on the display screen of a smartphone, or project video frames to achieve a three-dimensional display effect on the lenses of augmented reality / virtual reality glasses. As for the perception of the form of the virtual scene, it can be understood that it can be achieved with the help of the corresponding hardware output of terminal device 400, such as using a microphone to form auditory perception, using a vibrator to form tactile perception, and so on.

[0071] As an example, terminal device 400 runs a client (e.g., a web-based game application). During client operation, it outputs a virtual scene for role-playing. This virtual scene can be an environment for game characters to interact with, such as plains, streets, valleys, etc., for game characters to battle in. The first virtual object can be a game character controlled by the user; that is, the first virtual object is controlled by the real user and will move in the virtual scene in response to the real user's actions on a controller (e.g., a touchscreen, voice-activated switch, keyboard, mouse, and joystick). For example, when the real user moves the joystick to the right, the first virtual object will move to the right in the virtual scene. It can also remain stationary, jump, and be controlled to perform shooting operations. Alternatively, terminal device 400 runs a video player that can play videos of the virtual scene in a naked-eye 3D mode.

[0072] For example, the virtual scene can be a game virtual scene, server 200 can be a game platform server, and database 500 can be a game database that stores data such as game virtual scene data and player account data. The user can be a player. The following explanation is based on the above example.

[0073] For example, server 200 runs a game process and sends the game screen corresponding to the first virtual object to terminal device 400. The human-computer interaction interface on terminal device 400 displays the virtual scene in a two-dimensional plane mode. In response to a user's trigger operation on terminal device 400 or a trigger operation of a control device associated with terminal device (e.g., game controller, mouse, motion sensing device, or keyboard), terminal device 400 sends a display mode switching request to server 200 via the network. Server 200 obtains the image processing method for the virtual scene provided in this embodiment according to the display mode switching request, and obtains a second left view and a second right view that are offset from the initial image. Server 200 sends the second left view and the second right view to terminal device 400 via network 300. Terminal device 400 switches the display mode of the virtual scene from two-dimensional plane mode to naked-eye 3D mode, and then screen 101 on terminal device 400 switches to screen 102. In naked-eye 3D mode, terminal device 400 achieves a naked-eye 3D effect based on the second left view and the second right view.

[0074] In some embodiments, the image processing method for virtual scenes provided in this application can also be applied to the following scenarios:

[0075] 1. Video Application: The image processing method for virtual scenes provided in this application embodiment processes video frame images to form left and right views of the video frame images, and forms a naked-eye 3D effect based on the left and right views of the video frame images, thereby enhancing the stereoscopic and realistic feel of the video content.

[0076] 2. Advertising display: The image processing method of the virtual scene provided in this application embodiment is used to process the advertising image to form left and right views, and to form a naked-eye three-dimensional effect based on the left and right views of the advertising image, so as to provide a better advertising display effect, enhance the realism and three-dimensionality of the advertising image, and thus improve the recommendation efficiency of the advertisement.

[0077] In some embodiments, the terminal device or server can implement the image processing method for the virtual scene provided in this application by running a computer program. For example, the computer executable instructions can be microprogram-level commands, machine instructions, or software instructions. The computer program can be a native program or software module in the operating system; it can be a native application (APP), that is, a program that needs to be installed in the operating system to run, such as a game APP or an instant messaging APP; or it can be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to the browser environment to run. In summary, the above-mentioned computer executable instructions can be any form of instructions, and the above-mentioned computer program can be any form of application, module, or plugin.

[0078] Taking a computer program as an example, in actual implementation, the terminal device 400 has an application that supports virtual scenes installed and running. This application can be any of the following: a first-person shooter (FPS) game, a third-person shooter game, a virtual reality application, a 3D map application, or a multiplayer survival game. Users use the terminal device 400 to manipulate virtual objects located in the virtual scene, and these activities include, but are not limited to: adjusting body posture, crawling, walking, running, riding, jumping, driving, picking up items, shooting, attacking, throwing, and constructing virtual buildings—at least one of these. Illustratively, the virtual object can be a virtual character, such as a realistic or anime character.

[0079] This application embodiment can be implemented using database technology. A database, simply put, can be viewed as an electronic filing cabinet storing electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, capable of being shared by multiple users, having minimal redundancy, and being independent of application programs.

[0080] A Database Management System (DBMS) is a computer software system designed to manage databases, generally possessing basic functions such as storage, retrieval, security, and backup. DBMSs can be classified according to the database model they support, such as relational or XML (Extensible Markup Language); or according to the type of computer they support, such as server clusters or mobile devices; or according to the query language used, such as Structured Query Language (SQL) or XQuery; or according to performance priorities, such as maximum scale or maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, simultaneously supporting multiple query languages.

[0081] In some embodiments, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Electronic devices can be smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, etc., but are not limited to these. Terminal devices and servers can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.

[0082] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may be... Figure 1A or Figure 1B Terminal device 400, Figure 2 The terminal device 400 shown includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.

[0083] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0084] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0085] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0086] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0087] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0088] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0089] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0090] Presentation module 453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with user interface 430;

[0091] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.

[0092] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 An image processing apparatus 455 for a virtual scene stored in memory 450 is shown. This apparatus can be software in the form of programs and plugins, and includes the following software modules: a rendering module 4551, a parameter acquisition module 4552, and a display module 4553. These modules are logically linked and can therefore be arbitrarily combined or further separated according to the functions they implement. Figure 2 For ease of explanation, all the above modules are shown at once, but this should not be interpreted as excluding the implementation of the image processing device 455 in the virtual scene, which may only include the display module 4551. The functions of each module will be described below.

[0093] The image processing method for virtual scenes provided in this application will be described in conjunction with exemplary applications and implementations of the terminal devices provided in the embodiments of this application.

[0094] The following describes the image processing method for virtual scenes provided in the embodiments of this application. As mentioned above, the electronic device implementing the image processing method for virtual scenes in the embodiments of this application can be a terminal device or a server, or a combination of both. Therefore, the executing entity of each step will not be described again below.

[0095] It should be noted that the image processing examples below are illustrated using the naked-eye 3D mode applied in a game application. Based on the understanding of the following, those skilled in the art can apply the image processing method for virtual scenes provided in the embodiments of this application to the processing of other application scenarios. The virtual scene can be a virtual scene in a video or advertising image.

[0096] See Figure 3A , Figure 3A This is a first flowchart illustrating the image processing method for a virtual scene provided in this application embodiment, which will be combined with... Figure 3A The steps shown are explained. Figure 3A The entity responsible for executing the steps is Figure 1A or Figure 1B Terminal device 400.

[0097] In step 301, for the initial image of the virtual scene, a first left view from the left eye perspective and a first right view from the right eye perspective are generated.

[0098] To facilitate understanding, the application scenarios in the embodiments of this application are explained. Human eyes exhibit parallax, which is the directional difference produced when observing the same target from two points at a certain distance. Based on parallax, the images seen by the left and right eyes are slightly different, and the brain perceives depth and three-dimensional space by processing the differences between the left and right eye images. Naked-eye 3D mode is a display technology that enables the perception of 3D effects without wearing any special glasses. To achieve naked-eye 3D mode, a parallax is required between the left and right views, and these views are displayed using specific devices so that the left view can only be observed by the left eye, and the right view only by the right eye, allowing the viewer to perceive a sense of depth in the objects displayed on the screen.

[0099] For example, a virtual scene can be a game virtual scene, which can run through applications such as game applications and video players, or it can run based on an embedded application. The initial image of the virtual scene can include at least one of the following: environment scene materials of the virtual scene, UI components of the virtual scene, and virtual objects in the virtual scene. The virtual scene can be a three-dimensional virtual scene or a two-dimensional virtual scene. For a two-dimensional virtual scene, converting the initial image of the two-dimensional virtual scene into left and right views and implementing a naked-eye 3D effect based on the left and right views can improve the visual effect of the two-dimensional virtual scene.

[0100] For 3D virtual scenes, objects that are inherently 3D within the virtual scene do not need to be drawn; only the 2D content within the virtual scene needs to be drawn, thus saving computational resources required to create a naked-eye 3D mode. For example, if the virtual scene is a 3D game scene, where virtual objects and scene assets are 3D, and the UI components are 2D, and the initial image is an image of the UI components, then only the left and right views of the UI components need to be drawn.

[0101] The principles are essentially the same for left-eye and right-eye viewing angles. Taking the left-eye viewing angle as an example, it is the angle formed by light rays drawn from opposite ends (top and bottom, or left and right) of the screen at the optical center of the left eye when the left eye views the display device. For example, light rays drawn from the left and right edges of the screen form an angle at the left eye, or the angle drawn from two points symmetrically positioned around the center of the screen. (Reference) Figure 6 , Figure 6 This is a schematic diagram illustrating the principle of the image processing method for a virtual scene provided in this application embodiment. When a user observes the screen of a display device, the light rays on the screen intersect at the user's left eye, forming a left-eye view V1, and the light rays on the screen intersect at the user's right eye, forming a right-eye view V2. In this application embodiment, the left view from the left-eye perspective is an image of the virtual scene displayed on the screen as observed from the position of the left eye. The left-eye view contains visual information seen from the left eye, and this visual information differs from the visual information seen from the right eye.

[0102] The first left view from the left eye perspective and the first right view from the right eye perspective can be generated in the following way: based on the parallax corresponding to the left eye perspective and the parallax corresponding to the right eye perspective, the initial image is scaled, translated, and transformed to form the first left view from the left eye perspective and the first right view from the right eye perspective.

[0103] In some embodiments, reference Figure 3B , Figure 3B This is a schematic diagram of the second process of the image processing method for virtual scenes provided in the embodiments of this application. Figure 3AStep 301 in the process can be achieved through Figure 3B Steps 3011 to 3014 are implemented, and the details are explained below.

[0104] In step 3011, the size of the initial image of the virtual scene is reduced according to the pre-configured reduction ratio to obtain the first reduced image.

[0105] For example, the pre-configured scaling ratio can be set according to the actual needs of the terminal device's screen size. The size of the first scaled-down image is smaller than the size of the initial image. The scaling process can be implemented as follows: determine the size of the first scaled-down image based on the pre-configured scaling ratio; construct a grid corresponding to the first scaled-down image based on the size of the first scaled-down image; perform interpolation processing on the pixels of the initial image according to the size of the first scaled-down image (e.g., nearest neighbor interpolation, weighted interpolation) to form new pixel values; and fill each new pixel value into the grid corresponding to the first scaled-down image to form the first scaled-down image.

[0106] For example, in this embodiment of the application, the length of the first reduced image is the same as the initial image. That is, in the vertical direction, the size of the initial image is maintained, and in the horizontal direction, the initial image is reduced to form the first reduced image. The first reduced image is used as the basis for generating the first left view and the first right view. The first left view and the first right view are used for offsetting, and the second left view and the second right view obtained based on the offset are used to form a naked-eye 3D effect. The naked-eye 3D effect is achieved by displaying the second left view and the second right view in different areas of the screen, so that the viewer's left eye only sees the left view and the right eye only sees the right view. The sizes of the first left view and the first right view, the second left view and the second right view are the same, and the size of the display area of ​​the second left view and the second right view is the same as the initial image. Since the sizes of the second left view and the second right view are the same, the pre-configured reduction ratio of the first reduced image is set to 0.5, that is, in the horizontal direction, the width of the initial image is reduced to half of the original width to form the first reduced image.

[0107] In some embodiments, step 3011 can be implemented as follows: multiply the transformation matrix corresponding to the pre-configured scaling ratio with the initial drawing transformation matrix of the virtual scene to obtain a first drawing transformation matrix, wherein the first drawing transformation matrix is ​​used to perform a size transformation operation; perform a size transformation operation on each grid in the initial image based on the first drawing transformation matrix to obtain a first scaled-down image, wherein the length of the first scaled-down image is the same as that of the initial image, and the width of the first scaled-down image is smaller than that of the initial image.

[0108] For example, to facilitate understanding, the following explains the principle of scaling and translation based on the transformation matrix in image processing. The parameters of the transformation matrix are explained in conjunction with formula (1):

[0109]

[0110] Where sx is the scaling factor in the horizontal direction; sy is the scaling factor in the vertical direction; tx is the offset in the horizontal direction; ty is the offset in the vertical direction; and [xy 1] is the vertex coordinate of the UI component.

[0111] A vertex represents a position in two-dimensional space, represented by coordinates (x, y) in 2D graphics. Vertices are the foundation for constructing more complex shapes (lines, triangles, etc.). Scaling and transferring graphics can be achieved by performing transformations on the vertices. UV coordinates are a coordinate system in 3D computer graphics used to map 2D textures onto the surface of 3D models. In the UV coordinate system, U and V represent the horizontal and vertical directions of the texture image, respectively. These coordinates are equivalent to the "address" of the texture, specifying how a point on the 3D model should correspond to a pixel on the texture image. One UV coordinate corresponds to one pixel. Moving and transforming the UV coordinates of each pixel in the image transforms the entire image. Transformation operations are achieved by multiplying the pixel's UV coordinates by a transformation matrix.

[0112] Based on the scaling principle described above, when the pre-configured scaling ratio is 0.5, the transformation matrix corresponding to the pre-configured scaling ratio can be represented as follows: Transformation matrix Multiplying the first rendering transformation matrix by the initial rendering transformation matrix of the virtual scene yields the first rendering transformation matrix, which is represented by the following formula (2):

[0113]

[0114] The first rendering transformation matrix is ​​multiplied by each grid cell of the initial image to obtain the first scaled-down image. Each grid cell of the initial image can be represented by the vertex coordinates contained in the grid. During the scaling process using the first rendering transformation matrix, the width of the grid cells is reduced to half of their original width, while the length is maintained. Consequently, the width of the initial image is halved, forming the first scaled-down image.

[0115] In step 3012, under the first canvas coordinate system established based on the first drawing origin, a drawing operation is performed based on the data of the first scaled-down image to obtain the first left view.

[0116] For example, the drawing origin is a point on the canvas, typically with coordinates (0, 0), serving as a reference point for positioning all graphic elements. In most cases, the drawing origin is located at the top-left corner of the canvas. The first drawing origin can be an initial drawing origin, for example, with coordinates (0, 0). The coordinate system established based on the drawing origin defines the position of each point on the canvas relative to the origin. In this coordinate system, the horizontal direction is typically represented by the x-axis, and the vertical direction by the y-axis. The first canvas coordinate system is a coordinate system established with the first drawing origin as its origin, and it has both horizontal and vertical directions.

[0117] Drawing operations refer to the process of placing graphical elements (such as points, lines, rectangles, text, or more complex shapes and models) onto a rendering target (usually the screen or a framebuffer) using a graphics application programming interface (API) during a specific stage of graphics rendering.

[0118] The data of the first scaled-down image includes size and pixel values. The drawing operation based on the data of the first scaled-down image can be implemented in the following way: the drawing area is determined in the first canvas coordinate system established according to the first drawing origin based on the size of the first scaled-down image. In the drawing area, the first left view is formed by drawing according to the pixel values ​​of the pixels of the first scaled-down image.

[0119] In step 3013, the first drawing origin is translated according to the first pre-configured distance between the right eye view and the left eye view to obtain the second drawing origin.

[0120] For example, based on the principles above, translation can be achieved using a transformation matrix. The first pre-configured distance between the right-eye and left-eye views is the view width set for the left and right views. In practical applications, technicians set the view width according to the requirements of the display device's parameters. For example, if the initial view size displayed in the human-computer interaction interface of the display device is M, and the view width of the left or right view displayed by the display device is m, where m is less than or equal to M, then the first pre-configured distance is also m. The second drawing origin is the origin required for performing drawing operations on the first right view. The size of the first right view is the same as the first scaled-down image, and the first right view is drawn based on the data of the first scaled-down image. Therefore, the transformation matrix required to translate the first drawing origin is obtained by adjusting the first drawing transformation matrix corresponding to the first scaled-down image. Specifically, this adjustment can be achieved by multiplying the first drawing transformation matrix by the transformation matrix corresponding to the first pre-configured distance to form a new transformation matrix. Multiplying the new transformation matrix by the first drawing origin yields the second drawing origin.

[0121] In some embodiments, step 3013 can be implemented as follows: multiply the transformation matrix corresponding to the first pre-configured distance with the first drawing transformation matrix to obtain a second drawing transformation matrix, wherein the second drawing transformation matrix is ​​used to perform size transformation operation and displacement operation, and the first pre-configured distance is pre-configured for the left eye view and the right eye view; perform a translation operation on the first drawing origin based on the second drawing transformation matrix to obtain the second drawing origin.

[0122] Assuming the first pre-configured distance is represented by width, the first drawing transformation matrix is Representing the first pre-configured distance in matrix form yields the transformation matrix corresponding to the first pre-configured distance. The transformation matrix corresponding to the first pre-configured distance is then multiplied by the first drawing transformation matrix to obtain the second drawing transformation matrix, which can be represented by the following formula (3):

[0123]

[0124] Multiply the second drawing transformation matrix by the first drawing origin, thereby shifting the first drawing origin to obtain the second drawing origin.

[0125] In step 3014, under the second canvas coordinate system established based on the second drawing origin, a drawing operation is performed based on the data of the first scaled-down image to obtain the first right view.

[0126] In practical applications, the left and right views can be drawn simultaneously, or the left view can be drawn first and then the right view, or vice versa. The principle of drawing the right view is similar to that of drawing the left view. The drawing operation is performed based on the data of the first scaled-down image, which can be achieved as follows: the drawing area is determined in the second canvas coordinate system established according to the second drawing origin based on the size of the first scaled-down image. In the drawing area, the pixel values ​​of the pixels of the first scaled-down image are filled into the corresponding positions in the drawing area to form the first right view.

[0127] In this embodiment, by scaling down the initial image, a new view is drawn based on the data from the scaled-down image. During network transmission or data synchronization, scaling down the initial image reduces the amount of data that needs to be transmitted, thus lowering bandwidth requirements. The left and right views drawn based on the scaled-down image can better adapt to the screen size. Processing the scaled-down image simplifies the rendering process by avoiding complex image transformations and resampling operations.

[0128] In some embodiments, when the initial image of the virtual scene is an image of a UI component, the drawing operation can be implemented using the Flutter rendering engine. Flutter is an open-source UI software development kit used to build high-performance cross-platform applications. Flutter compiles to native machine code, providing a single codebase that supports six platforms: Android, iOS, Web, Windows, macOS, and Linux. When the application starts, Flutter traverses and creates all the widget data to form a widget tree. The widget tree is a configuration data structure used to describe the application's user interface. This data structure stores the rendered content, and widget tree creation is very lightweight, being rebuilt at any time during page refreshes.

[0129] Flutter creates each element object by calling the `createElement()` method on the Widget, forming an element tree. Finally, it calls the `createRenderObject()` method of the element to create each render object, forming a render tree. The render tree is used for the layout and drawing of the application interface, responsible for the actual rendering, and stores information such as element size and layout. The element acts as an intermediary layer separating the Widget Tree from the actual render object. The widget tree describes the corresponding element properties, holds both the Widget and RenderObject properties, stores context information, and is used to traverse the view tree, supporting the UI structure. For easier understanding, refer to [reference needed]. Figure 8A , Figure 8A This is a schematic diagram of the component tree structure provided in this application embodiment; a render object tree 802A is built based on the element tree 801A, and the render object tree 802A is painted to obtain a layer tree 803A. The layer tree is a drawing instruction tree, where each node is used to implement the corresponding drawing function. In the layer tree, the drawing performed by a node refers to the process of converting the content of each layer into pixels and finally displaying it on the screen. The layer tree is a tree structure composed of multiple layers, each layer representing a drawable entity, such as a text layer, an image layer, or a container layer. Each node (usually corresponding to a layer in the layer tree) collects the associated drawing commands. These commands describe how to draw the appearance of the node, including color, shape, text content, etc. Each node may be divided into multiple tiles, especially for large or complex layers. This tile division can optimize the rendering process because only the visible or redrawable parts are processed.

[0130] To facilitate understanding of the graphical layer tree in the embodiments of this application, the following description is provided in conjunction with the accompanying drawings. Figure 8B , Figure 8B This is a schematic diagram of the graphics layer tree structure provided in an embodiment of this application. The graphics layer tree 800B contains multiple nodes. Node 801B is the left and right view node, used to perform drawing operations for the first left view and the first right view. Node 801B corresponds to multiple child nodes, such as child nodes 8011 to 801N, where N is a positive even number greater than or equal to 2. Each child node is used to perform drawing of each sub-block (i.e., a part of the first left view or the first right view) in the UI component. After the drawing operation corresponding to each child node is completed, the parts drawn by each child node are composited together to form a final frame image. The graphics layer tree includes nodes at multiple levels; during compositing, lower-level nodes may be covered by pixels from higher-level nodes.

[0131] Based on the above graphics layer tree, step 3012 can be implemented in the following way: In the first canvas coordinate system established based on the first drawing origin, each pixel of the first left view is drawn according to the data of the first scaled-down image through the left view node of the graphics layer tree to form the first left view.

[0132] Based on the above graphics layer tree, step 3014 can be implemented in the following way: In the second canvas coordinate system established based on the second drawing origin, each pixel of the first right view is drawn according to the data of the first scaled-down image through the right view node of the graphics layer tree to form the first right view.

[0133] Drawing the graphics layer tree typically occurs in the later stages of the rendering pipeline. It's an efficient rendering optimization technique that redraws only the changed parts, rather than the entire page. Drawing is performed on the first left and first right views through the graphics layer tree. Each first left and first right view can be a separate node in the tree, allowing independent control over its rendering properties and interactive behaviors, such as scaling, scrolling, and animations, without affecting the other view. Treating the two views as separate nodes helps the rendering engine update or redraw only the changed parts, rather than the entire screen. This reduces computation and improves rendering efficiency. Managing the left and right views through nodes in the graphics layer tree allows developers to handle each view logically more clearly, reducing code complexity and maintenance difficulty. Furthermore, because the graphics layer tree performs drawing through nodes, adding nodes makes it easier to adapt to screens of different sizes and aspect ratios, and layout can be optimized by adjusting the size and position of each view.

[0134] In some embodiments, the graphics layer tree corresponds to a rendering transformation matrix, and this rendering transformation matrix is ​​reused in actual applications to perform corresponding transformation operations. Reusing the rendering transformation matrix can be achieved through data pointers, which point to different storage locations to access the transformation matrix stored therein. For example, the initial rendering transformation matrix, the first rendering transformation matrix corresponding to the left-eye view, and the second rendering transformation matrix corresponding to the right-eye view are stored in different storage locations. When the first rendering transformation matrix corresponding to the left-eye view is needed during the current processing, the data pointer is pointed to the storage location of the first rendering transformation matrix corresponding to the left-eye view, and the first rendering transformation matrix corresponding to the left-eye view is accessed.

[0135] Reusing rendering transformation matrices can be achieved by storing the reused transformation matrix in a fixed storage location. When needed, the initial value of the transformation matrix is ​​adjusted according to requirements such as position transformation operations and scaling operations, and then restored to its initial value after use. For example, if the initial rendering transformation matrix is ​​stored in the storage location, during the rendering of the first left view, the initial rendering transformation matrix is ​​converted into the first rendering transformation matrix corresponding to the left-eye view. Similarly, during the rendering of the first right view, the first rendering transformation matrix corresponding to the left-eye view is converted into the second rendering transformation matrix corresponding to the right-eye view.

[0136] In this embodiment, by reusing the transformation matrix, consistency of transformations can be ensured across a series of graphics operations. This helps avoid complex errors caused by the accumulation of continuous transformations. Reusing the matrix avoids recalculating for each transformation, thus saving computational resources. Reusing and restoring the transformation matrix simplifies code logic, making transformation operations easier to understand and maintain. After multiple transformations, if each transformation process does not restore the original matrix, transformation errors may accumulate, leading to inaccurate final rendering results. Restoring the original matrix can avoid this accumulation of errors.

[0137] Continue to refer to Figure 3A In step 302, the left view offset value under the left eye view and the right view offset value under the right eye view are obtained.

[0138] For example, binocular offset refers to the offset of the image relative to a reference baseline between the eyes when viewing an image. The left-view offset and right-view offset are opposites. The concepts of left-view offset and right-view offset are similar; taking the left-view offset as an example, the left-view offset of the left eye perspective is the lateral or longitudinal displacement of the left view relative to the initial image or reference position. The left-view offset of the left eye perspective can be obtained by determining the left-view offset based on the image parameters of the display device, the viewer's interpupillary distance, and the distance between the viewer's eyes and the display device. Since the left-view offset of the left eye perspective and the right-view offset of the right eye perspective are opposites, obtaining either the left-view offset of the left eye perspective or the right-view offset of the right eye perspective allows the determination of the other offset value.

[0139] In some embodiments, reference Figure 3C , Figure 3C This is a schematic diagram of the third process of the image processing method for virtual scenes provided in the embodiments of this application. Figure 3A Step 302 in the process can be achieved through Figure 3C Steps 3021 to 3024 are implemented, and the details are explained below.

[0140] In step 3021, the three-dimensional image parameters of the human-computer interaction interface displaying the virtual scene are obtained.

[0141] For example, 3D image parameters are used to create the 3D effect of a scene. These parameters can be calculated based on the image's depth value z, depth of field d, and distance from the plane o. Depth value (z), depth of field (d), and distance from the plane o are important parameters describing the position and visual effect of objects in 3D space. The depth value z, depth of field d, and distance from the plane o of a virtual scene can be set according to the user's needs.

[0142] To facilitate understanding, explanations are provided in conjunction with the accompanying diagrams. Figure 6 , Figure 6 This is a schematic diagram illustrating the principle of the left-right image parallax algorithm provided in this application embodiment. The depth of field d (pre-configured depth of field) is the distance from the far-distance plane to the near-distance plane (out-of-field plane). The far-distance plane to the near-distance plane is a plane artificially set to achieve a 3D effect. The near-distance plane is located outside the screen and is closer to the viewer than the far-distance plane. The far-distance plane is located inside the screen. The out-of-field distance o represents the distance between the screen and the near-distance plane, and is an initial value provided by the display device. Users can modify this globally through settings. The near-distance plane is the upper limit of the 3D image height. The interpupillary distance p refers to the distance between the user's pupils. The eye-to-screen distance s can be provided with an initial value based on the display device's specifications, and users can modify it globally to a value that suits their actual needs.

[0143] The depth value z (pre-configured image depth) is preset by developers based on the display requirements of the actual application scenario. It can be set for individual UI components or a group of UI elements. Figure 6 Taking the 3D UI component 601 as an example, UI component 601 is a UI component that allows users to perceive a 3D effect protruding above the screen height. With the interpupillary distance p between the left and right eyes, the visual position of point A formed in the user's brain is on the 3D UI component 601 protruding from the screen. The actual pixels of the 3D UI component 601 are displayed on the screen. The points that the user actually sees are A_Right and A_Left on the screen. A_Right and A_Left are observed by the user on the plane, and point A is formed in the brain according to the left and right parallax.

[0144] by Figure 6 Taking the 3D UI component 602 as an example, UI component 602 is a 3D UI component that allows users to feel that it protrudes above the height of the screen. When the interpupillary distance p is between the left and right eyes, the visual position of point B formed in the user's brain is on the 3D UI component 602 that sinks into the screen. The actual pixels of the 3D UI component 602 are displayed on the screen. The points that the user actually sees are B_Right and B_Left on the screen. B_Right and B_Left are observed by the user on the plane, and point B is formed in the brain according to the left and right parallax.

[0145] In some embodiments, step 3021 can be implemented by: obtaining the pre-configured image depth, pre-configured depth of field, and out-of-plane distance of the human-computer interaction interface, wherein the out-of-plane distance is the upper limit of the height of the three-dimensional image formed in the naked-eye three-dimensional mode; determining the second product between the pre-configured image depth and the pre-configured depth of field; and using the first difference between the second product and the out-of-plane distance as a three-dimensional image parameter.

[0146] Based on the above parameters, the second product between the pre-configured image depth z and the pre-configured depth of field d is z*d, and the three-dimensional image parameter is the first difference (z*do) between the second product z*d and the distance o from the out-of-plane.

[0147] In step 3022, a first product of the pre-configured pupil distance and the three-dimensional image parameters is determined, and a first summation between the three-dimensional image parameters and the third distance is determined.

[0148] For example, continuing with the example above, the first product of the pre-configured interpupillary distance *p* and the 3D image parameter (z*do) is *p*(z*do) / 2. Since in practical applications, the human left and right eyes are usually symmetrically developed, this first product is actually half the pre-configured interpupillary distance *p* multiplied by the 3D image parameter (z*do). The third distance *s* is the distance between the eye and the human-computer interaction interface. The first summation between the 3D image parameter and the third distance is (z*d - o + s).

[0149] In step 3023, the ratio of the first product to the first sum is used as the left view offset value.

[0150] For example, continuing from the example above, the left view offset value left_eye_transform is represented by the pre-configured formula (4.1): left_eye_transform=-p*(z*do) / (2*(z*d-o+s))(4.1)

[0151] In step 3024, the opposite of the left view offset value is used as the right view offset value.

[0152] For example, based on the left view offset value, the right view offset value, right_eye_transform, can be represented by formula (4.2): right_eye_transform=p*(z*do) / (2*(z*d-o+s))(4.2)

[0153] As can be seen from the above formulas (4.1) and (4.2), the offset value of the right view is the opposite of the offset value of the left view.

[0154] In this embodiment, by using 3D image parameters such as depth of field, depth value, and distance from the plane of the display device, depth of field can be used to simulate the visual effect of the human eye observing the real world, enhancing stereoscopic sense and spatial perception. The depth value can be used to accurately determine the position of each object in the scene, thereby correctly calculating the parallax of each object when generating left and right eye views. This improves the quality and realism of the stereoscopic image. Furthermore, determining the right view offset value using the left view offset value accelerates the calculation speed.

[0155] In some embodiments, prior to step 3021, the pre-configured interpupillary distance can be determined in the following manner, as detailed below:

[0156] Obtain a binocular image dataset, which includes multiple sample binocular images; call an image segmentation model to segment the multiple sample binocular images to obtain the first eye region image in each sample binocular image; normalize the size of each first eye region image, and determine the first pupillary distance based on the size-normalized first eye region image; use the mean of each first pupillary distance as the pre-configured pupillary distance.

[0157] For example, the sample binocular images in the binocular image dataset were taken with the user's authorization, and the sample binocular images show both eyes fully open to improve the accuracy of determining the pre-configured interpupillary distance.

[0158] Image segmentation models are models used to segment images into multiple parts or objects. Common image segmentation models include: thresholding models, edge detection models, deep learning segmentation models (e.g., Convolutional Neural Networks (CNNs) using deep learning techniques), Fully Convolutional Networks (FCNs), and U-Nets. Taking deep learning segmentation models as an example, the training dataset for training an image segmentation model includes sample images and the location labels of the target objects to be segmented for each sample image. For example, collecting a large amount of image data containing binocular regions. This data can be obtained from public datasets or collected from networks using image scraping tools. The collected images are then labeled, marking the precise boundaries of the binocular regions. This usually requires manual work, but semi-automated labeling tools can be used to assist (such as pre-trained neural network models for labeling). Simultaneously, to improve the generalization ability of the deep learning segmentation model, operations such as rotation, scaling, cropping, and flipping can be used to increase the diversity of sample images. During training, the sample data can also be divided into training, validation, and test sets. The training set is fed into the model, and the predicted values ​​are calculated through forward propagation. Then, the loss function is calculated, and the network weights are updated through backpropagation. The validation set is used to evaluate the model's performance, and the model is tuned according to performance metrics (such as IOU, Dice coefficient, etc.). The test set is used to evaluate the model's final performance to ensure that the model has good generalization ability.

[0159] Image segmentation models can use loss functions such as Cross-Entropy Loss and Dice Coefficient Loss. Cross-Entropy Loss measures the difference between two probability distributions and is typically used for classification problems. In object extraction, it measures the difference between the predicted and ground truth label distributions. Cross-Entropy Loss is suitable for pixel-level classification in object extraction, where each pixel is classified as either an object or background. Dice Loss is a loss function that measures the similarity between two sets and is commonly used for semantic segmentation in tasks such as image segmentation and object detection. Dice Loss is particularly suitable for handling class imbalance in object extraction because it focuses on the overlap between predicted and ground truth labels, rather than the error of a single pixel. It is robust to the size and shape of the predicted region and therefore, in some cases, reflects the segmentation quality better than Cross-Entropy Loss.

[0160] In this embodiment, the target segmentation object is both eyes, and the first eye region image is an image containing both eyes. The image segmentation model extracts features from the sample eye images to obtain sample image features. Based on these features, it performs classification processing to identify pixels in the sample eye images that belong to the target segmentation object. Each pixel is then combined according to its corresponding coordinates to form the first eye region image. Each first eye region image is scaled to a uniform size. The interpupillary distance is determined based on the ratio of the interpupillary distance between pixels in the first eye region image to the interpupillary distance between pixels of the calibrated object in the first eye region image. The average of all interpupillary distances is used as the pre-configured interpupillary distance. During application, the actual size of the calibrated object is known, and the following parameters remain the same: the ratio between the size of the calibrated object in the image and its actual size, and the ratio between the interpupillary distance between the eyes in the first eye region image and the interpupillary distance in the real world. Therefore, the interpupillary distance in the real world for both eyes in the first eye region image = the actual size of the calibrated object * the ratio of the interpupillary distance between pixels in the first eye region image to the interpupillary distance between pixels of the calibrated object in the first eye region image. The formula is A = B * (a / b), where B is the actual size of the calibrated object, A is the interpupillary distance between the two eyes in the first eye region image in the real world, a is the interpupillary distance between pixels in the first eye region image, and b is the pixel spacing of the calibrated object in the first eye region image.

[0161] In this embodiment, a large amount of sample data is statistically analyzed, and artificial intelligence technology is used to determine the pre-configured interpupillary distance (IPD). Statistical analysis can better generalize to different IPD distributions, improving the applicability and versatility of the pre-configured IPD in different populations, thereby enhancing the stereoscopic effect of the naked-eye 3D mode. In related technologies, manual measurement of interpupillary distance may introduce subjective judgment bias, and users inevitably experience eye movements during the measurement process. However, in this embodiment, the measurement is automated through an image segmentation model, and statistical analysis is performed using static images, which can reduce these biases and provide more objective and consistent results. It is applicable to complex scenarios. The image segmentation model can handle interpupillary distance measurement in complex scenarios such as occlusion, lighting changes, and different facial expressions, while manual measurement may encounter difficulties in such situations.

[0162] In some embodiments, prior to step 3021, the pre-configured interpupillary distance can be determined in the following manner, as detailed below:

[0163] The process involves acquiring reference binocular images and current binocular images, where the current binocular image is an image captured by the eyes currently viewing the human-computer interaction interface. An image segmentation model is then used to segment the current binocular image, resulting in a second eye region image. This second eye region image is then scaled to obtain a third eye region image, where the image size of the eyes in the third eye region image is the same as that in the reference binocular image. Finally, the pixel spacing between the eyes in the third eye region image is enlarged according to a preset ratio to obtain a pre-configured interpupillary distance, where the preset ratio is the ratio between the reference interpupillary distance corresponding to the reference binocular image and the pixel spacing between the eyes in the reference binocular image.

[0164] For example, the principle of the image segmentation model is the same as above, and will not be repeated here. The two eyes currently viewing the human-computer interaction interface are the eyes of the user currently viewing the screen. Based on the reference image, the eye region in the image taken for the current user is scaled to the same size as the eyes in the reference image. Using the eyes themselves as the calibration objects, the indirect distance formed by multiple pixels in the image is converted into the distance in the real scene to obtain the actual interpupillary distance of the current user.

[0165] For example, the process of capturing the current binocular image can be implemented as follows: The real-time image captured by the user is displayed on the human-computer interaction interface of the terminal device, along with a preset frame indicating the area where the eyes are positioned during image capture. After placing both eyes in the corresponding area according to the preset frame, the user can manually trigger the terminal device's capture operation or wait for the terminal device to perform the capture operation automatically. The preset frame prompts the user, allowing for a more accurate current binocular image, thereby improving the accuracy of measuring the actual interpupillary distance. In practical applications, multiple current binocular images can be captured. The user can select the most accurate image for confirmation, and the terminal device extracts the user's actual interpupillary distance based on the confirmed current binocular image.

[0166] For example, the actual reference interpupillary distance (IPD) corresponding to the reference binocular image is known. The preset ratio is the ratio between the reference IPD corresponding to the reference binocular image and the pixel pitch between the eyes in the reference binocular image. The reference binocular image is a human eye image conforming to the normal human eye size. The size of the eyes in the third eye region image is the same as the size of the eyes in the reference binocular image. When the eye sizes are the same, the following parameters are the same: the ratio between the reference IPD corresponding to the reference binocular image and the pixel pitch between the eyes in the reference binocular image; the ratio between the actual IPD of the eyes corresponding to the third eye region image (i.e., the IPD used as the preset IPD) and the pixel pitch between the eyes in the third eye region image. This is represented by the formula C / c = S / s, where C is the reference IPD corresponding to the reference binocular image, c is the pixel pitch between the eyes in the reference binocular image, S is the actual IPD corresponding to the eyes in the third eye region image, and s is the pixel pitch between the eyes in the third eye region image. Furthermore, the actual interpupillary distance of the two eyes corresponding to the third eye region image = the ratio between the reference interpupillary distance of the reference two-eye image and the pixel spacing between the two eyes in the reference two-eye image * the pixel spacing between the two eyes in the third eye region image.

[0167] The pre-configured interpupillary distance is obtained by magnifying the pixel spacing between the two eyes in the third eye region image according to a preset ratio. This can be achieved by multiplying the preset ratio by the pixel spacing between the two eyes in the third eye region image.

[0168] In this embodiment, determining the current user's interpupillary distance (IPD) using an image segmentation model improves the efficiency of determining the pre-configured IPD. This allows users to adjust the naked-eye 3D mode according to their preferences and visual acuity, providing a more personalized service and experience. Customizing the preset IPD for the current user adapts to different users, as different users may have varying sensitivities to 3D effects. Adjusting the IPD can meet the needs of different users, especially those sensitive to 3D effects. It also reduces eye strain during naked-eye 3D mode use. Inappropriate IPD settings may cause discomfort or fatigue when viewing naked-eye 3D content. Personalized IPD adjustment can reduce this discomfort. Based on the pre-configured IPD, the stereoscopic effect and realism of the naked-eye 3D mode can be improved in practical applications.

[0169] In some embodiments, prior to step 3022, the third distance between the eyes and the human-computer interaction interface can be determined in the following manner, as detailed below:

[0170] Acquire a first pair of binocular images and a second pair of binocular images, wherein the first pair of binocular images and the second pair of binocular images are images captured by the eyes viewing the human-computer interaction interface, and the first pair of binocular images and the second pair of binocular images are captured by different cameras; call an image recognition model to extract key points from the first pair of binocular images and the second pair of binocular images respectively, to obtain a first key point in the first pair of binocular images and a second key point in the second pair of binocular images, wherein the first key point and the corresponding second key point represent the same position in the real scene; determine a second difference between the following positions: the first key point at a first position in the first pair of binocular images and the second key point at a second position in the second pair of binocular images; take the ratio of the third product between the camera focal length and the camera distance to the second difference as the third distance, wherein the camera distance is the distance between the two cameras that captured the first pair of binocular images and the second pair of binocular images.

[0171] For example, the first and second binocular images can be captured in real time. Two cameras mimic human binocular vision to measure the distance to objects in a real-world scene; that is, the cameras mimic binocular ranging. The principle of this binocular ranging is based on the principle of stereoscopic vision in the human visual system. In human vision, when two eyes observe the same scene from slightly different angles, the brain uses the differences in the images seen by these two eyes to perceive depth and distance. Stereoscopic vision refers to using two cameras to capture images of the same scene from different positions, simulating the stereoscopic effect seen by the human eye. The two cameras are equivalent to two human eyes, and there is a certain distance between them, called the baseline. When two cameras capture the same scene, due to their different positions, the images they capture will differ in the horizontal direction; this difference is called disparity. Disparity is the difference in the horizontal position of the same object in the images from the two cameras, and it decreases as the distance to the object increases. Binocular ranging utilizes the principle of triangulation; that is, by measuring the disparity in the images of the same object captured by the two cameras, the distance from the object to the cameras can be calculated.

[0172] The first keypoint and its corresponding second keypoint represent the same location in a real-world scene, meaning they are the same location on the same object in the real scene. There can be multiple keypoints. The type of keypoint can be set according to the actual application scenario. For example, for eye images, the keypoint location can be the inner or outer corner of the eye. The image recognition model performing keypoint extraction can be a convolutional neural network (CNN) model, used to classify pixels in the image to determine which pixels belong to keypoints and extract them. A CNN model typically contains multiple convolutional and pooling layers to extract image features and fully connected layers to predict keypoint locations.

[0173] The training set of an image recognition model includes a large amount of image data containing keypoints to be detected. These keypoints can be set according to actual needs. Specifically, the image data consists of binocular images, with each sample image containing complete images of both eyes. During the training process, Mean Squared Error (MSE) loss or cross-entropy loss can be used as the loss function. Taking MSE loss as an example, we will explain this further. MSE loss measures the error between the model's predicted keypoint coordinates and the true coordinates; specifically, it measures the average squared difference between the predicted and actual values. MSE loss allows the image recognition model to focus on inaccurately predicted keypoints during training, thereby improving the overall keypoint detection accuracy. During training, based on the loss function, optimization algorithms such as backpropagation and gradient descent are used to adjust the parameters of the image recognition model to reduce prediction errors.

[0174] The image recognition model extracts key points from the first and second binocular images separately, which can be achieved as follows: Feature extraction is performed on the first and second binocular images respectively to obtain first image features and second image features. For the first image features, a first prediction probability is made for each pixel in the first binocular image that it may be a key point, and the pixel with the highest first prediction probability is selected as the first key point. Similarly, for the second image features, a second prediction probability is made for each pixel in the second binocular image that it may be a key point, and the pixel with the highest second prediction probability is selected as the second key point. The first position of the first key point in the first binocular image and the second position of the second key point in the second binocular image are represented by coordinate values, and the difference between the coordinate values ​​is converted into a straight-line distance to obtain the second difference.

[0175] The ratio of the third product of the camera's focal length and the camera distance to the second difference, as the third distance Z, can be represented as: Z = f * B / D, where D is the parallax, B is the distance between the centers of the two cameras (baseline distance), and f is the camera's focal length. The cameras are mounted on the display device corresponding to the terminal device, and the specific placement can be determined according to actual needs. For example, if the terminal device is a computer, the two cameras can be symmetrically positioned on either side of the center of the top of the monitor, and the baseline distance between the two cameras is set according to the size of the monitor.

[0176] In this embodiment, two cameras are used to mimic human binoculars to measure the distance to objects in a real-world scene. The advantages of this binocular ranging approach include low cost, high precision, strong environmental adaptability, and the ability to acquire 3D information in real-time. The principle of binocular ranging is similar to human vision, making it easier to understand and explain. It also facilitates the design of related applications and algorithms, thereby improving the response speed for measuring distances between the user and a plane, enhancing the smoothness of the image in naked-eye 3D mode, and reducing the probability of image stuttering.

[0177] In some embodiments, before step 3022, the third distance between the eyes and the human-computer interaction interface can be determined in the following manner, as specifically described below: the distance between the eyes and the human-computer interaction interface is obtained by measuring the distance of the user viewing the human-computer interaction interface through an infrared ranging camera set on the display device.

[0178] In some embodiments, the third distance between the eyes and the human-computer interface can be a fixed value, typically based on ergonomic research and user comfort considerations. This fixed value can be set by the display device at the factory according to ergonomic requirements. Users can reset the fixed value corresponding to the third distance between their eyes and the human-computer interface during actual use, and view the display device at their set distance to achieve a better stereoscopic effect. Users can adjust the viewing distance according to their own feelings and habits. By setting the third distance between the eyes and the human-computer interface, it is possible to adapt to the usage needs of different users, improve the stereoscopic effect of the naked-eye 3D mode during viewing, reduce visual fatigue, and enhance viewing comfort.

[0179] Continue to refer to Figure 3A In step 303, the first left view is offset according to the left view offset value to obtain the second left view, and the first right view is offset according to the right view offset value to obtain the second right view.

[0180] For example, the principle of offsetting the first left view and offsetting the first right view is similar. Taking the first left view as an example, the offset can be achieved in the following way: offset the coordinates of each pixel of the first left view based on the transformation matrix to form the second left view.

[0181] In some embodiments, step 303 can be implemented in the following ways:

[0182] Multiply the transformation matrix corresponding to the offset value of the left view with the first drawing transformation matrix corresponding to the first left view to obtain the third drawing transformation matrix; offset each pixel in the first left view based on the third drawing transformation matrix to obtain the second left view; multiply the transformation matrix corresponding to the offset value of the right view with the second drawing transformation matrix corresponding to the first right view to obtain the fourth drawing transformation matrix; offset each pixel in the first right view based on the fourth drawing transformation matrix to obtain the second right view.

[0183] For example, the transformation matrix corresponding to the left view offset value is represented as follows: The third transformation matrix can be represented by the following formula (5):

[0184]

[0185] The transformation matrix corresponding to the right view offset value is represented as follows: Similarly, the fourth transformation matrix is ​​represented by the following formula (6):

[0186]

[0187]

[0188] For example, the principles of offsetting the left view and offsetting the right view are similar. Taking the left view as an example, the offset can be achieved by multiplying the coordinates of each pixel in the first left view based on the third transformation matrix to achieve the offset processing and form the second left view.

[0189] To facilitate understanding of the relationship between the second left view, the second right view, and the initial image, the following explanation is provided in conjunction with the accompanying drawings. (Reference) Figure 7A , Figure 7A This is a first schematic diagram of the human-computer interaction interface provided in an embodiment of this application; the human-computer interaction interface 701 is a virtual scene in two-dimensional mode. (Reference) Figure 7B , Figure 7B This is a second schematic diagram of the human-computer interaction interface provided in this application embodiment; it shows the left eye view 702 and the right eye view 703 drawn for the human-computer interaction interface 701 in a two-dimensional manner. If the three-dimensional display function of the display device is invoked to display the left eye view 702 and the right eye view 703 in an interlaced manner, a three-dimensional effect can be formed.

[0190] In some embodiments, when the initial image of the virtual scene is an image of a UI component, the offset operation can be implemented using the Flutter rendering engine. Continuing with the example in step 301, refer to... Figure 8B , Figure 8BThis is a schematic diagram of the structure of the graphics layer tree provided in this application embodiment. The graphics layer tree 800B contains multiple nodes. Node 801B is a left / right view node, and node 801B corresponds to multiple child nodes, such as child nodes 8011 to 801N, where N is a positive even number greater than or equal to 2. Each child node is used to perform the drawing operation of each sub-block (i.e., a part of the first left view or the first right view) in the UI component. Taking child node 801N as an example, child node 801N is used to draw a part of the first right view. After drawing a part of the first right view, the first right view is offset according to the right view offset value to obtain the second right view. The graphics layer tree 800B also includes node 802B, which is a parallax node. Node 802B is used to offset the part corresponding to child node 801N in the first right view based on the corresponding right view offset value. Node 802B corresponds to multiple child nodes, such as child nodes 8021 to 802M, where M is a positive number greater than 1. Taking child node 802M as an example, child node 802M is used to call the corresponding instruction, offset the position of the pixel in the first right view according to the transformation matrix corresponding to the right view offset value, and draw the corresponding pixel value at the offset position to form a part of the offset second right view. After the drawing operation corresponding to each child node is completed, the part drawn by each child node will be composited together to form a second right view.

[0191] In some embodiments, based on the above-described graphics layer tree, step 303 can be implemented as follows: For each left-view disparity node in the graphics layer tree, multiply the transformation matrix corresponding to the left-view offset value with the first drawing transformation matrix corresponding to the first left view to obtain a third drawing transformation matrix. Based on the data of the first left view and the third drawing transformation matrix, offset the positions of the original pixels in the first left view in the first coordinate system, and draw new pixels at the offset positions. Combine each new pixel drawn by each left-view disparity node into a second left view. Store the third drawing transformation matrix and restore it to its initial state. For each right-view disparity node in the graphics layer tree, multiply the transformation matrix corresponding to the right-view offset value with the second drawing transformation matrix corresponding to the first right view to obtain a fourth drawing transformation matrix. Based on the data of the first right view and the fourth drawing transformation matrix, offset the positions of the original pixels in the first right view in the second coordinate system, and draw new pixels at the offset positions. Combine each new pixel drawn by each right-view disparity node into a second right view. After drawing the second right view, reset the fourth drawing transformation matrix to its initial state.

[0192] In this embodiment, the graphics layer tree allows for the selective creation of left and right views for at least a portion of the virtual scene. For example, parts of the virtual scene that are 3D models may not have their left and right views drawn; only 2D UI components may be drawn. This saves computational resources while achieving a glasses-free 3D display effect. It eliminates the need for a full-screen design for every screen size, saving computational resources required to achieve a glasses-free 3D effect. The graphics layer tree also allows developers more flexible control over various parts of the UI, enabling fine-grained management, such as applying different visual effects or animations to specific components.

[0193] In this embodiment, by reusing the transformation matrix, consistency of transformations can be ensured across a series of graphics operations. This helps avoid complex errors caused by the accumulation of continuous transformations. Reusing the matrix avoids recalculating for each transformation, thus saving computational resources. Reusing and restoring the transformation matrix simplifies code logic, making transformation operations easier to understand and maintain. After multiple transformations, if each transformation process does not restore the original matrix, transformation errors may accumulate, leading to inaccurate final rendering results. Restoring the original matrix can avoid this accumulation of errors.

[0194] In step 304, a virtual scene in naked-eye 3D mode is displayed based on the second left view and the second right view.

[0195] For example, a naked-eye 3D mode based on a second left view and a second right view can be achieved using a specific display device. The second left view and the second right view are projected onto their respective display areas, so that the viewer's left eye can only see the second left view and the right eye can only see the second right view. Then, the viewer's brain forms a corresponding 3D image based on the parallax observed by the left and right eyes, thus achieving a naked-eye 3D effect.

[0196] In some embodiments, step 304 can be implemented as follows: when the display device displaying the human-computer interaction interface is in naked-eye 3D mode, determine the left-eye view display area and the right-eye view display area of ​​the display device; display the second left view in the left-eye view display area and the second right view in the right-eye view display area to display the virtual scene in naked-eye 3D mode.

[0197] For example, the display areas for the left and right eyes are staggered, thus projecting the corresponding pixels for each eye onto the left and right eyes respectively, achieving image separation and creating a naked-eye 3D effect. Naked-eye 3D mode can be achieved using a slit-type liquid crystal grating or a lenticular lens. Slit-type liquid crystal grating technology adds a slit grating in front of the display screen. Based on this slit grating, when the image for the left eye is displayed on the LCD screen, opaque stripes block the right eye; similarly, when the image for the right eye is displayed on the LCD screen, opaque stripes block the left eye. By separating the visible images for the left and right eyes, the viewer sees a 3D image. Lenticular lens technology works by using the refraction principle of a lens to project the corresponding pixels for the left and right eyes respectively, achieving image separation.

[0198] For ease of understanding, the following explanation is provided in conjunction with the accompanying drawings. Figure 7C , Figure 7C This is a third schematic diagram of the human-computer interaction interface provided in this application embodiment. Assuming the screen of human-computer interaction interface 701 is the initial image, and human-computer interaction interface 704 is a three-dimensional effect screen formed from the screen of human-computer interaction interface 701. Specifically, left and right views are drawn for the UI components in human-computer interaction interface 701, so that... Figure 7A The two-dimensional UI component 706 of the human-computer interaction interface 701 is presented as a three-dimensional UI component 705 in the human-computer interaction interface 704. The left and right views of the three-dimensional model 707 in the human-computer interaction interface 701 are not drawn, as it is a three-dimensional graphic and does not require drawing. Instead, the two-dimensional UI component 706 is selectively drawn to form the screen of the human-computer interaction interface 704, thus saving the computational resources required to achieve a naked-eye 3D effect.

[0199] In this embodiment, two views, left and right, are created from an initial image. The visual illusion created by these two views is then combined to present a three-dimensional effect on a two-dimensional plane. Compared to related technologies that use two virtual cameras to create left and right parallax maps in a virtual scene for naked-eye 3D rendering, this embodiment does not require complex 3D modeling and rendering techniques. Drawing the left and right views typically requires less computational resources and storage space. The left and right views can be implemented in most existing 2D graphics software without requiring specialized 3D software or hardware support, thus improving the compatibility and versatility of the virtual scene image processing method in this embodiment.

[0200] In some embodiments, this application also provides an image processing method for a virtual scene, which will be described below with reference to the accompanying drawings. See also Figure 3D , Figure 3D This is a schematic diagram of the fourth process of the image processing method for virtual scenes provided in the embodiments of this application, which will be combined with Figure 3D The steps shown are explained. Figure 3D The entity responsible for executing the steps is Figure 1A Terminal device 400.

[0201] In step 305, the virtual scene is displayed in a two-dimensional plane mode on the human-computer interaction interface.

[0202] Example, reference Figure 7A , Figure 7A This is a first schematic diagram of the human-computer interaction interface provided in this application embodiment; the human-computer interaction interface 701 is a virtual scene in two-dimensional mode. Displaying the virtual scene in two-dimensional planar mode in the human-computer interaction interface can be presented as follows: Figure 7A Human-computer interaction interface 701.

[0203] In step 306, in response to the switching operation for the two-dimensional planar mode, the virtual scene is displayed in a naked-eye three-dimensional mode.

[0204] For example, the naked-eye 3D mode is implemented using the image processing method of the virtual scene in this application embodiment. The terminal device includes a screen, which can be controlled via an external device or a touchscreen. For example, the terminal device is a handheld game console, controlled via a game controller or touchscreen.

[0205] In some embodiments, the human-computer interaction interface is controlled by an electronic device (e.g., an external electronic device or the terminal device itself), which includes a first button (e.g., a button on an external device controlling the terminal device, or a button on the terminal device itself). Step 306 can be implemented as follows: in response to a click operation on the first button, the virtual scene displayed in a two-dimensional planar mode is switched to a virtual scene displayed in a naked-eye three-dimensional mode. The first button can be a gamepad button, a keyboard button, or a button or motion-sensing control on the terminal device. For example, if the terminal device is equipped with a gyroscope, and the terminal device tilts at a preset angle, the virtual scene is switched to naked-eye three-dimensional mode.

[0206] In some embodiments, the human-computer interaction interface includes a first control, which is a control for switching display modes; step 306 can be implemented in the following way: in response to a trigger operation on the first control, the virtual scene displayed in two-dimensional planar mode is switched to the virtual scene displayed in naked-eye three-dimensional mode.

[0207] For example, a control is a triggerable piece of information displayed in the form of a region, button, icon, link, text, checkbox, input box, and tab; wherein, the triggering method can be contact triggering, non-contact triggering, or instruction-based triggering, etc.; in addition, the various controls in the embodiments of this application can be a single control or a collective term for multiple controls.

[0208] Continuing with the example from the previous level, after the switching operation, refer to... Figure 7C , Figure 7C This is a third schematic diagram of the human-computer interaction interface provided in this application embodiment. The human-computer interaction interface 704 is a three-dimensional image formed from the screen of the human-computer interaction interface 701. Specifically, it involves drawing left and right views of the UI components in the human-computer interaction interface 701, so that... Figure 7A The two-dimensional UI component 706 of the human-computer interaction interface 701 is presented as a three-dimensional UI component 705 in the human-computer interaction interface 704. In practical applications, the image processing method for virtual scenes provided in this application can selectively create left and right views for the two-dimensional parts of the virtual scene, enabling the two-dimensional image to be presented as a three-dimensional effect. The three-dimensional parts of the virtual scene are not drawn with left and right views, thus saving computational resources while achieving a naked-eye three-dimensional display effect.

[0209] In this embodiment, a first left view from the left-eye perspective and a first right view from the right-eye perspective are generated based on the initial image of the virtual scene. Compared to related technologies that use a virtual camera to obtain a disparity map or estimate depth to create a naked-eye 3D effect, generating images based on the initial image saves the computational resources required to achieve a naked-eye 3D effect. The first left view is offset according to the left view offset value to obtain a second left view, and the first right view is offset according to the right view offset value to obtain a second right view. By creating left-eye and right-eye views and appropriately offsetting them, the stereoscopic effect seen by the human eye can be simulated, thereby achieving a naked-eye 3D effect. Processing the left and right eye views in stages allows for more precise control over the generation and offset process of each view, ensuring the accuracy of the final stereoscopic visual effect.

[0210] The following will describe an exemplary application of the image processing method for virtual scenes according to the embodiments of this application in a real-world application scenario.

[0211] Humans possess a degree of parallax, which is the difference in direction produced when observing the same target from two points at a certain distance. The angle between two points viewed from the target is called the parallax angle, and the line connecting the two points is called the baseline. Knowing the parallax angle and the baseline length allows us to calculate the distance between the target and the observer. Parallax causes the images seen by the left and right eyes to differ slightly, and the brain processes these differences to perceive depth and three-dimensional space. Naked-eye 3D (3D) technology generates left and right parallax maps, sends these maps to a special program and a 3D display device, and the display device outputs the left and right eye maps separately based on the observer's position, allowing each eye to see different content, thus creating a stereoscopic effect.

[0212] In related technologies, the methods for implementing naked-eye 3D left and right parallax maps can be divided into several categories: (1) Depth estimation method: Depth estimation is performed on the image to generate naked-eye 3D images; however, the depth estimation method consumes a lot of resources and the actual depth estimation is inaccurate, resulting in screen display stuttering, frame drops, and occasionally screen ghosting and blurring. (2) Hook method: Hook image depth information to generate naked-eye 3D images; (3) Engine plug-in method: Engine development is used, and the 3D mode uses dual virtual cameras to generate 3D images. That is, in this method, the left and right eye parallax is achieved by taking left and right eye images with two virtual cameras. Due to the increase in the number of virtual cameras, the computational resources are increased, which may cause the virtual scene screen to stutter. Both the Hook and engine plug-in methods require redevelopment through the 3D engine, which cannot utilize existing data, resulting in high development costs and long development time.

[0213] This application provides an image processing method for virtual scenes to address the aforementioned problems in related technologies. By converting the parallax of the left and right views using normalized depth values, the left and right views are offset separately, and a naked-eye 3D effect is achieved based on the offset left and right views. Furthermore, this solution can be implemented by adding components to the Flutter UI framework, improving development efficiency during the development of the naked-eye 3D mode UI; in the actual application of the naked-eye 3D mode, it saves the computing resources required for the terminal device to run the naked-eye 3D mode, and avoids screen stuttering and other issues in the naked-eye 3D mode.

[0214] refer to Figure 4 , Figure 4 This is a schematic diagram of the fifth process of the image processing method for virtual scenes provided in the embodiments of this application; Figure 1A The terminal device 400 in the process is the executing entity, and the following steps will be explained.

[0215] In step 401, a graph layer tree is generated based on the component tree data structure.

[0216] The image processing method for virtual scenes provided in this application embodiment can be implemented using the Flutter rendering engine. Flutter is an open-source UI software development kit used to build high-performance cross-platform applications; a single codebase can support six platforms: Android, iOS, Web, Windows, macOS, and Linux. Flutter compiles to native machine code, which improves the smoothness of the application and achieves beautiful animation effects.

[0217] A widget tree is a data structure used to describe an application's user interface. This data structure stores rendered content. It's a lightweight configuration data structure that is rebuilt during page refreshes. When the application starts, Flutter traverses and creates all widget data to form the widget tree. It creates each Element object by calling the `createElement()` method on the Widget, forming an Element Tree. Finally, it calls the `createRenderObject()` method of the Element to create each render object, forming a render tree. The render tree is used for the layout and drawing of the application interface, responsible for the actual rendering, and stores information such as element size and layout. The Element acts as an intermediary layer separating the widget tree from the actual render object. The widget tree describes the corresponding Element properties, holds both the widget and the render object, stores context information, and is used to traverse the view tree, supporting the UI structure.

[0218] For ease of understanding, please refer to Figure 8A , Figure 8A This is a schematic diagram of the component tree structure provided in an embodiment of this application; a rendering object tree 802A is built based on the element tree 801A, and the rendering object tree 802A is painted to obtain a graphics layer tree 803A. The graphics layer tree is a drawing instruction tree, and each node is used to implement the corresponding drawing function.

[0219] When the Flutter engine performs depth-first rendering on the generated graphics layer tree at the underlying level, it stores the offset and scaling data of UI components in a 3x3 transformation matrix. Scaling and shifting of UI component vertices can be achieved through rectangle multiplication. The parameters of the transformation matrix are explained below using formula (1):

[0220]

[0221] Where sx is the horizontal scaling factor; sy is the vertical scaling factor; tx is the horizontal offset; ty is the vertical offset; and [xy + 1] are the vertex coordinates of the UI component. A vertex represents a position in two-dimensional space, represented by coordinates (x, y) in two-dimensional graphics. Vertices are the foundation for constructing more complex graphics (lines, triangles, etc.). Scaling and shifting of graphics can be achieved by performing transformations on the vertices of the graphics.

[0222] In step 402, based on the nodes in the graphics layer tree, a drawing operation is performed on the data of the initial image of the virtual scene to form a first left view and a first right view.

[0223] For example, based on the aforementioned features of Flutter, nodes can be inserted into the graphics layer tree, and the transformation matrix can be modified to display naked-eye 3D UI components. A small UI component with a left and right parallax map is defined in the UI layer. This component inserts a left and right image node into the graphics layer tree and performs the drawing of the left and right images. (Reference) Figure 8B , Figure 8B This is a schematic diagram of the structure of the graphics layer tree provided in an embodiment of this application. The graphics layer tree 800B contains multiple nodes, with node 801B being the left and right image node, used to perform the drawing operations for the left and right images.

[0224] As mentioned above, during the drawing process, data related to image scaling and offset are stored in the transformation matrix, and the original drawing transformation matrix is ​​stored for the initial image. To facilitate understanding of the relationship between the initial image and the first left and first right views, the following explanation is provided in conjunction with the accompanying drawings.

[0225] refer to Figure 5 This is a schematic diagram illustrating the principle of the naked-eye 3D mode provided in this application embodiment. Taking the implementation of a naked-eye 3D mode for a UI component as an example, a left-eye view 502 and a right-eye view 503 are drawn based on the UI component diagram 501. The UI component diagram 504 with a naked-eye 3D effect is formed by interweaving the left-eye view 502 and the right-eye view 503. The process of creating the left and right eye disparity maps involves drawing operations and offset operations. The first left view and the first right view are images drawn for the left and right eyes before offset.

[0226] The width of the child node is scaled to half the size of the node by rectangular multiplication, which is represented by the following formula (2):

[0227]

[0228] The above diagram illustrates the transformation matrix. Based on this, start drawing the first left view and inform the child nodes that the left side of the screen is being drawn.

[0229] When the first left view is drawn, the current drawing transformation matrix is ​​saved. The drawing origin is moved to the first position above the middle of the node by rectangular multiplication. The moving distance is obtained according to the width value of the left and right view nodes. The width value is preset according to the application scenario. In the graphics layer tree, the width value of the node refers to the size of the graphics layer represented by the node in the horizontal direction. It is part of the layer tree node attributes and determines how much horizontal space of the screen or canvas the node can occupy during rendering. The movement of the drawing origin is represented by the following formula (3):

[0230]

[0231] Here, `width` is the width value. Specifically, it modifies the coordinates of the drawing origin to the coordinates above the center of the node. If you want to move the drawing origin to the center above the node, you need to calculate the upward offset distance. To achieve this movement, you can use a transformation matrix through rectangle multiplication to change the position of the drawing origin. This matrix will include a translation transformation, shifting the origin to the new position.

[0232] Based on the above matrix Start drawing the first right view and inform the child nodes that the right side of the screen is being drawn. The drawing principle of the first right view is similar to that of the first left view. Once the first right view is drawn, resume drawing the transformation matrix.

[0233] In step 403, the left view offset value under the left eye perspective and the right view offset value under the right eye perspective are determined.

[0234] For example, the left view offset value from the left-eye perspective and the right view offset value from the right-eye perspective represent the distance offset by the left and right eyes, respectively, with reference to the baseline. For ease of understanding, an explanation is provided in conjunction with the accompanying drawings. Figure 6 , Figure 6 This is a schematic diagram illustrating the principle of the left and right image disparity algorithm provided in the embodiments of this application.

[0235] Figure 6 The parameters involved include: depth of field d, which is the distance from the far plane to the near plane; out-of-plane distance o, which represents the distance between the screen and the near plane, and is an initial value provided by the display device. Users can modify it globally to enhance the 3D effect. The near plane is the upper limit of the height of the 3D image.

[0236] Interpupillary distance (p) refers to the distance between the pupils of a user's two eyes. The distance from the eye to the screen (s) can be provided with an initial value based on the display device. Users can then modify this value globally to suit their specific needs.

[0237] The depth value z is preset by the developers based on the display requirements of the actual application scenario. It can be set for individual UI components or a group of UIs.

[0238] The left view offset value `left_eye_transform` and the right view offset value `right_eye_transform` can be determined using pre-configured formulas (4.1) and (4.2):

[0239] left_eye_transform=-p*(z*do) / (2*(z*d-o+s))(4.1)

[0240] right_eye_transform=p*(z*do) / (2*(z*d-o+s))(4.2)

[0241] As can be seen from the above formulas (4.1) and (4.2), the offset value of the right view is the opposite of the offset value of the left view.

[0242] In step 404, the first left view is offset based on the left view offset value to obtain the second left view.

[0243] For example, after obtaining the left view offset value, further offset processing is performed through the graphics layer tree. A UI widget for left and right disparity maps is defined in the UI layer. This widget inserts a left-right disparity node into the graphics layer tree (based on the left and right images processed by the left-right image node, it performs further processing to create visual differences between the left and right images), implementing the disparity of its child left and right images. See further reference. Figure 8B Offset processing is performed based on node 802B. When drawing the left image, the drawing transformation matrix is ​​saved; the left view offset value `left_eye_transform` is updated to the transformation matrix using rectangular multiplication. In this context, it is represented by the following formula (5):

[0244]

[0245] Once the second left view is drawn, resume drawing the transformation matrix.

[0246] In step 405, the first right view is offset based on the right view offset value to obtain the second right view.

[0247] For example, when drawing the second right view, save the current drawing transformation matrix, and update the right view offset value to the transformation matrix through rectangle multiplication. In this context, it is represented by the following formula (6):

[0248]

[0249] Once the second right view is drawn, resume drawing the transformation matrix.

[0250] In step 406, a virtual scene in naked-eye 3D mode is displayed based on the second left view and the second right view.

[0251] For example, the second left view and the second right view are displayed interlaced to form a naked-eye 3D virtual scene. Naked-eye 3D can be achieved using a slit-type liquid crystal grating or a lenticular lens. Slit-type liquid crystal grating technology adds a slit grating in front of the display device's screen. Based on this slit grating, when the image for the left eye is displayed on the LCD screen, opaque stripes block the right eye; similarly, when the image for the right eye is displayed on the LCD screen, opaque stripes block the left eye. By separating the visible images for the left and right eyes, the viewer sees a 3D image. Lens technology works by using the refraction principle of lenses to project corresponding pixels for the left and right eyes separately, achieving image separation. The biggest advantage of slit-type grating technology is that the lens does not block light, thus greatly improving brightness.

[0252] The following description, in conjunction with the accompanying drawings, illustrates the effects of the image processing method for virtual scenes provided in this application. Figure 7A , Figure 7A This is a first schematic diagram of the human-computer interaction interface provided in this application embodiment; the human-computer interaction interface 701 is a virtual scene in two-dimensional mode. After launching the application, the user can automatically switch the application's display mode to 3D or 2D mode by querying the 3D / 2D switching switch of the 3D naked-eye mode device, or switch to 3D mode via a custom button or the buttons on the gamepad. (Reference) Figure 7B , Figure 7B This is a second schematic diagram of the human-computer interaction interface provided in this application embodiment; it displays the left-eye view 702 and right-eye view 703 drawn for the human-computer interaction interface 701 in a two-dimensional manner. By calling the device's three-dimensional display function and displaying the left-eye view 702 and right-eye view 703 in an interlaced manner, a three-dimensional effect can be formed. (Reference) Figure 7C , Figure 7C This is a third schematic diagram of the human-computer interaction interface provided in this application embodiment. The human-computer interaction interface 704 is a three-dimensional image formed from the screen of the human-computer interaction interface 701. Specifically, it involves drawing left and right views of the UI components in the human-computer interaction interface 701, so that... Figure 7A In the human-computer interaction interface 701, the two-dimensional UI component 706 is transformed into a three-dimensional UI component 705 in the human-computer interaction interface 704. In practical applications, the image processing method for virtual scenes provided in this application embodiment can selectively create left and right views for the two-dimensional parts of the virtual scene, enabling the two-dimensional image to be presented as a three-dimensional effect. The three-dimensional parts of the virtual scene are not drawn with left and right views, thus saving computational resources while achieving a naked-eye three-dimensional display effect.

[0253] Related technologies cannot utilize existing data or reuse existing Flutter UI framework projects, resulting in slower development speeds and lower performance compared to the Flutter engine. The image processing method for virtual scenes provided in this application allows existing Flutter projects to achieve naked-eye 3D left-right parallax maps with simple modifications, reducing the cost of repetitive development. Newly created Flutter projects can continue to use the original 2D development, only needing to integrate 3D components to achieve 3D left-right parallax maps. In practical applications, 3D left-right parallax maps can be enabled only for certain components, while images and videos that inherently possess 3D left-right parallax maps can have their 3D effects disabled, enabling browsing and control of 3D effects for videos and images.

[0254] The following description continues to illustrate the exemplary structure of the virtual scene image processing device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules in the image processing device 455 storing the virtual scene in the memory 450 may include: a rendering module 4551, used to generate a first left view from a left-eye perspective and a first right view from a right-eye perspective for an initial image of the virtual scene; a parameter acquisition module 4552, used to acquire the offset value of the left view from the left-eye perspective and the offset value of the right view from the right-eye perspective; the rendering module 4551 is further used to offset the first left view according to the left view offset value to obtain a second left view, and to offset the first right view according to the right view offset value to obtain a second right view; and a display module 4553, used to display the virtual scene in naked-eye 3D mode based on the second left view and the second right view.

[0255] In some embodiments, the rendering module 4551 is configured to reduce the size of the initial image of the virtual scene according to a pre-configured scaling ratio to obtain a first scaled-down image; perform a drawing operation based on the data of the first scaled-down image in a first canvas coordinate system established based on a first drawing origin to obtain a first left view; translate the first drawing origin according to a first pre-configured distance between the right eye view and the left eye view to obtain a second drawing origin; and perform a drawing operation based on the data of the first scaled-down image in a second canvas coordinate system established based on the second drawing origin to obtain a first right view.

[0256] In some embodiments, the rendering module 4551 is used to multiply the transformation matrix corresponding to the pre-configured scaling ratio with the initial drawing transformation matrix of the virtual scene to obtain a first drawing transformation matrix, wherein the first drawing transformation matrix is ​​used to perform a size transformation operation; and to perform a size transformation operation on each grid in the initial image based on the first drawing transformation matrix to obtain a first scaled-down image, wherein the length of the first scaled-down image is the same as that of the initial image, and the width of the first scaled-down image is smaller than that of the initial image.

[0257] In some embodiments, the rendering module 4551 is used to multiply the transformation matrix corresponding to the first pre-configured distance with the first drawing transformation matrix to obtain a second drawing transformation matrix, wherein the second drawing transformation matrix is ​​used to perform size transformation operation and displacement operation, and the first pre-configured distance is pre-configured for the left eye view and the right eye view; and to perform a translation operation on the first drawing origin based on the second drawing transformation matrix to obtain the second drawing origin.

[0258] In some embodiments, the rendering module 4551 is configured to draw each pixel of the first left view based on the data of the first scaled-down image through the left view node of the graphics layer tree in a first canvas coordinate system established based on the first drawing origin, thereby forming the first left view; and to draw each pixel of the first right view based on the data of the first scaled-down image through the right view node of the graphics layer tree in a second canvas coordinate system established based on the second drawing origin, thereby forming the first right view.

[0259] In some embodiments, the parameter acquisition module 4552 is configured to acquire three-dimensional image parameters of the screen displaying the human-computer interaction interface of the virtual scene; determine a first product of a pre-configured interpupillary distance and the three-dimensional image parameters, and determine a first sum between the three-dimensional image parameters and a third distance, wherein the third distance is the distance between the eyes and the human-computer interaction interface; use the ratio of the first product to the first sum as the left view offset value; and use the opposite of the left view offset value as the right view offset value.

[0260] In some embodiments, the parameter acquisition module 4552 is used to acquire the pre-configured image depth, pre-configured depth of field, and out-of-plane distance of the human-computer interaction interface, wherein the out-of-plane distance is the upper limit of the height of the three-dimensional image formed in the naked-eye three-dimensional mode; determine the second product between the pre-configured image depth and the pre-configured depth of field; and use the first difference between the second product and the out-of-plane distance as the three-dimensional image parameter.

[0261] In some embodiments, the parameter acquisition module 4552 is configured to acquire a binocular image dataset before determining the first product of the pre-configured interpupillary distance and the three-dimensional image parameters, wherein the binocular image dataset includes multiple sample binocular images; call an image segmentation model to segment the multiple sample binocular images to obtain a first eye region image in each sample binocular image; normalize the size of each first eye region image, and determine a first interpupillary distance based on the size-normalized first eye region image; and use the mean of each first interpupillary distance as the pre-configured interpupillary distance.

[0262] In some embodiments, the parameter acquisition module 4552 is configured to acquire a reference binocular image and a current binocular image before determining the first product of the pre-configured interpupillary distance and the three-dimensional image parameters, wherein the current binocular image is an image captured by the eyes currently viewing the human-computer interaction interface; call an image segmentation model to perform image segmentation processing on the current binocular image to obtain a second eye region image; scale the second eye region image to obtain a third eye region image, wherein the image size of the eyes in the third eye region image is the same as the image size of the eyes in the reference binocular image; and enlarge the pixel spacing between the eyes in the third eye region image according to a preset ratio to obtain the pre-configured interpupillary distance, wherein the preset ratio is the ratio between the reference interpupillary distance corresponding to the reference binocular image and the pixel spacing between the eyes in the reference binocular image.

[0263] In some embodiments, the parameter acquisition module 4552 is configured to acquire a first binocular image and a second binocular image before determining the first difference between the three-dimensional image parameters and the third distance, wherein the first binocular image and the second binocular image are images captured by the eyes viewing the human-computer interaction interface, and the first binocular image and the second binocular image are captured by different cameras; call an image recognition model to extract key points from the first binocular image and the second binocular image respectively, to obtain a first key point in the first binocular image and a second key point in the second binocular image, wherein the first key point and the corresponding second key point represent the same position in the real scene; determine a second difference between the following positions: the first key point is at a first position in the first binocular image, and the second key point is at a second position in the second binocular image; and take the ratio of the third product between the camera focal length and the camera distance to the second difference as the third distance, wherein the camera distance is the distance between the two cameras that captured the first binocular image and the second binocular image.

[0264] In some embodiments, the rendering module 4551 is configured to multiply the transformation matrix corresponding to the offset value of the left view with the first drawing transformation matrix corresponding to the first left view to obtain a third drawing transformation matrix; offset each pixel in the first left view based on the third drawing transformation matrix to obtain a second left view; multiply the transformation matrix corresponding to the offset value of the right view with the second drawing transformation matrix corresponding to the first right view to obtain a fourth drawing transformation matrix; and offset each pixel in the first right view based on the fourth drawing transformation matrix to obtain a second right view.

[0265] In some embodiments, the display module 4553 is configured to determine the left-eye view display area and the right-eye view display area of ​​the display device when the display device displaying the human-computer interaction interface is in the naked-eye 3D mode; display the second left view in the left-eye view display area and the second right view in the right-eye view display area to display the virtual scene in the naked-eye 3D mode.

[0266] In some embodiments, such as Figure 2 As shown, the software module in the image processing device 455 storing the virtual scene in the memory 450 may further include: a display module 4553, used to display the virtual scene in a two-dimensional plane mode in a human-computer interaction interface; the display module 4553 is used to display the virtual scene in a naked-eye three-dimensional mode in response to a switching operation of the two-dimensional plane mode, wherein the naked-eye three-dimensional mode is implemented by the image processing method of the virtual scene described in the embodiments of this application.

[0267] In some embodiments, the human-computer interaction interface is controlled by an electronic device, the electronic device including a first button; the display module 4553 is used to switch the virtual scene displayed in the two-dimensional planar mode to the virtual scene displayed in the naked-eye three-dimensional mode in response to a click operation on the first button.

[0268] In some embodiments, the human-computer interaction interface includes a first control; the display module 4553 is configured to switch the virtual scene displayed in the two-dimensional planar mode to the virtual scene displayed in the naked-eye three-dimensional mode in response to a trigger operation on the first control.

[0269] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer program or computer-executable instructions, causing the electronic device to perform the virtual scene image processing method described above in this application.

[0270] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the image processing method for a virtual scene provided in this application. For example, ... Figure 3A The image processing method for the virtual scene is shown.

[0271] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0272] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0273] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0274] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0275] In summary, this application's embodiments generate a first left view from the left-eye perspective and a first right view from the right-eye perspective based on the initial image of a virtual scene. Compared to related technologies that use virtual cameras to acquire disparity maps or estimate depth to create a naked-eye 3D effect, generating images based on the initial image saves the computational resources required to achieve a naked-eye 3D effect. The first left view is offset according to the left view offset value to obtain a second left view, and the first right view is offset according to the right view offset value to obtain a second right view. By creating left-eye and right-eye views and appropriately offsetting them, the stereoscopic effect seen by the human eye can be simulated, thereby achieving a naked-eye 3D effect. Processing the left and right eye views in stages allows for more precise control over the generation and offset process of each view, ensuring the accuracy of the final stereoscopic visual effect.

[0276] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. An image processing method for a virtual scene, characterized in that, The method includes: For the initial image of the virtual scene, generate a first left view from the left eye perspective and a first right view from the right eye perspective; Obtain the left view offset value from the left eye perspective and the right view offset value from the right eye perspective; The first left view is offset according to the left view offset value to obtain the second left view, and the first right view is offset according to the right view offset value to obtain the second right view; The virtual scene is displayed in naked-eye 3D mode based on the second left view and the second right view.

2. The method according to claim 1, characterized in that, The process of generating a first left view from a left-eye perspective and a first right view from a right-eye perspective based on the initial image of the virtual scene includes: The initial image of the virtual scene is reduced in size according to a pre-configured scaling ratio to obtain a first scaled-down image; In the first canvas coordinate system established based on the first drawing origin, a drawing operation is performed based on the data of the first scaled-down image to obtain the first left view; The first drawing origin is translated according to the first pre-configured distance between the right eye view and the left eye view to obtain the second drawing origin. In the second canvas coordinate system established based on the second drawing origin, a drawing operation is performed based on the data of the first scaled-down image to obtain the first right view.

3. The method according to claim 2, characterized in that, The step of reducing the size of the initial image of the virtual scene according to a pre-configured reduction ratio to obtain a first reduced image includes: Multiply the transformation matrix corresponding to the pre-configured scaling ratio with the initial rendering transformation matrix of the virtual scene to obtain the first rendering transformation matrix, wherein the first rendering transformation matrix is ​​used to perform the size transformation operation; Based on the first drawing transformation matrix, a size transformation operation is performed on each grid in the initial image to obtain a first reduced image, wherein the length of the first reduced image is the same as that of the initial image, and the width of the first reduced image is smaller than that of the initial image.

4. The method according to claim 3, characterized in that, The step of translating the first drawing origin based on a first pre-configured distance between the right eye view and the left eye view to obtain a second drawing origin includes: Multiply the transformation matrix corresponding to the first pre-configured distance with the first drawing transformation matrix to obtain the second drawing transformation matrix. The second drawing transformation matrix is ​​used to perform size transformation operation and displacement operation. The first pre-configured distance is pre-configured for the left eye view and the right eye view. The second drawing origin is obtained by performing a translation operation on the first drawing origin based on the second drawing transformation matrix.

5. The method according to any one of claims 2 to 4, characterized in that, The step of performing a drawing operation based on the data of the first scaled-down image in a first canvas coordinate system established based on a first drawing origin to obtain a first left view includes: In the first canvas coordinate system established based on the first drawing origin, each pixel of the first left view is drawn according to the data of the first scaled-down image through the left view node of the graphics layer tree to form the first left view; The step of performing a drawing operation based on the data of the first scaled-down image in a second canvas coordinate system established based on the second drawing origin to obtain a first right view includes: In the second canvas coordinate system established based on the second drawing origin, each pixel of the first right view is drawn according to the data of the first scaled-down image through the right view node of the graphics layer tree to form the first right view.

6. The method according to any one of claims 1 to 5, characterized in that, The process of obtaining the left view offset value from the left eye perspective and the right view offset value from the right eye perspective includes: Obtain the three-dimensional image parameters of the human-computer interaction interface displaying the virtual scene; Determine the first product of the pre-configured interpupillary distance and the three-dimensional image parameters, and determine the first sum between the three-dimensional image parameters and the third distance, wherein the third distance is the distance between the eye and the human-computer interaction interface; The ratio of the first product to the first sum is used as the left view offset value; The opposite of the left view offset value is used as the right view offset value.

7. The method according to claim 6, characterized in that, The step of acquiring the 3D image parameters of the human-computer interaction interface displaying the virtual scene includes: The pre-configured image depth, pre-configured depth of field, and out-of-plane distance of the human-computer interaction interface are obtained, wherein the out-of-plane distance is the upper limit of the height of the three-dimensional image formed in the naked-eye three-dimensional mode; Determine the second product between the pre-configured image depth and the pre-configured depth of field; The first difference between the second product and the distance to the outgoing plane is used as the three-dimensional image parameter.

8. The method according to claim 6, characterized in that, Before determining the first product of the pre-configured interpupillary distance and the three-dimensional image parameters, the method further includes: Obtain a binocular image dataset, wherein the binocular image dataset includes multiple sample binocular images; An image segmentation model is invoked to segment the multiple sample binocular images to obtain a first eye region image in each sample binocular image; The size of each first eye region image is normalized, and the first interpupillary distance is determined based on the size-normalized first eye region image; The average value of each of the first pupillary distances is used as the pre-configured pupillary distance.

9. The method according to claim 6, characterized in that, Before determining the first product of the pre-configured interpupillary distance and the three-dimensional image parameters, the method further includes: Acquire a reference binocular image and a current binocular image, wherein the current binocular image is an image captured by the eyes currently viewing the human-computer interaction interface; The image segmentation model is invoked to perform image segmentation processing on the current binocular image to obtain the second eye region image; The second eye region image is scaled to obtain a third eye region image, wherein the image size of both eyes in the third eye region image is the same as the image size of both eyes in the reference eye image; The pixel spacing between the two eyes in the third eye region image is magnified according to a preset ratio to obtain the pre-configured interpupillary distance, wherein the preset ratio is the ratio between the reference interpupillary distance corresponding to the reference eye image and the pixel spacing between the two eyes in the reference eye image.

10. The method according to any one of claims 6 to 9, characterized in that, Before determining the first difference between the three-dimensional image parameters and the third distance, the method further includes: Acquire a first pair of binocular images and a second pair of binocular images, wherein the first pair of binocular images and the second pair of binocular images are images captured by the eyes viewing the human-computer interaction interface, and the first pair of binocular images and the second pair of binocular images are captured by different cameras; The image recognition model is called to extract key points from the first binocular image and the second binocular image respectively, to obtain the first key point in the first binocular image and the second key point in the second binocular image, wherein the first key point and the corresponding second key point represent the same location in the real scene. Determine a second difference between the following locations: the first key point is at a first position in the first binocular image, and the second key point is at a second position in the second binocular image; The ratio of the third product between the camera focal length and the camera distance to the second difference is taken as the third distance, wherein the camera distance is the distance between the two cameras that captured the first binocular image and the second binocular image.

11. The method according to any one of claims 1 to 10, characterized in that, The step of offsetting the first left view according to the left view offset value to obtain the second left view, and offsetting the first right view according to the right view offset value to obtain the second right view, includes: Multiply the transformation matrix corresponding to the left view offset value with the first drawing transformation matrix corresponding to the first left view to obtain the third drawing transformation matrix; Based on the third rendering transformation matrix, each pixel in the first left view is offset to obtain the second left view; Multiply the transformation matrix corresponding to the right view offset value with the second drawing transformation matrix corresponding to the first right view to obtain the fourth drawing transformation matrix; The second right view is obtained by offsetting each pixel in the first right view based on the fourth drawing transformation matrix.

12. The method according to any one of claims 1 to 11, characterized in that, The virtual scene displaying naked-eye 3D mode based on the second left view and the second right view includes: When the display device displaying the human-computer interaction interface is in the naked-eye 3D mode, the left-eye view display area and the right-eye view display area of ​​the display device are determined; The second left view is displayed in the left-eye view display area, and the second right view is displayed in the right-eye view display area to display the virtual scene in the naked-eye 3D mode.

13. An image processing method for a virtual scene, characterized in that, The method includes: The virtual scene is displayed in a two-dimensional plane mode in the human-computer interaction interface; In response to a switching operation for the two-dimensional planar mode, the virtual scene is displayed in a naked-eye three-dimensional mode, wherein the naked-eye three-dimensional mode is implemented by the image processing method for the virtual scene according to any one of claims 1 to 12.

14. The method according to claim 13, characterized in that, The human-computer interaction interface is controlled by an electronic device, which includes a first button; The step of displaying the virtual scene in a naked-eye 3D mode in response to a switching operation for the two-dimensional planar mode includes: In response to a click operation on the first button, the virtual scene displayed in the two-dimensional planar mode is switched to the virtual scene displayed in the naked-eye three-dimensional mode.

15. The method according to claim 14, characterized in that, The human-computer interaction interface includes a first control; The step of displaying the virtual scene in a naked-eye 3D mode in response to a switching operation for the two-dimensional planar mode includes: In response to a trigger operation on the first control, the virtual scene displayed in the two-dimensional planar mode is switched to the virtual scene displayed in the naked-eye three-dimensional mode.

16. An image processing device for a virtual scene, characterized in that, The device includes: The rendering module is used to generate a first left view from the left eye perspective and a first right view from the right eye perspective based on the initial image of the virtual scene. The parameter acquisition module is used to acquire the left view offset value under the left eye view and the right view offset value under the right eye view. The rendering module is further configured to offset the first left view according to the left view offset value to obtain the second left view, and to offset the first right view according to the right view offset value to obtain the second right view; The display module is used to display the virtual scene in naked-eye 3D mode based on the second left view and the second right view.

17. An image processing device for a virtual scene, characterized in that, The device includes: The display module is used to display virtual scenes in a two-dimensional plane mode in the human-computer interaction interface; A display module is configured to display the virtual scene in a naked-eye 3D mode in response to a switching operation for the two-dimensional plane mode, wherein the naked-eye 3D mode is implemented by the image processing method of the virtual scene according to any one of claims 1 to 14.

18. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the image processing method for the virtual scene according to any one of claims 1 to 15.

19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the image processing method for the virtual scene according to any one of claims 1 to 15 is implemented.

20. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the image processing method for the virtual scene according to any one of claims 1 to 15 is implemented.