A method, apparatus, electronic device, and storage medium for image processing

By sending initial pose data to the server in the virtual reality terminal and receiving initial image data, the terminal adjusts the image data to match the target pose, solving the problem of untimely image refresh caused by network delay by virtual reality terminals, achieving efficient image display and reducing hardware requirements.

CN114078092BActive Publication Date: 2025-07-01ZTE CORP

Patent Information

Application Number
CN202010803372.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-11
Publication Date
2025-07-01
Estimated Expiration
2040-08-11

AI Technical Summary

Technical Problem

Due to network delay, virtual reality terminals cannot synchronously render images corresponding to the terminal poses during the refresh cycle, resulting in the problem of untimely image refresh.

Method used

By sending the initial pose data of the terminal to the server, the server returns the initial image data, the terminal adjusts the two-dimensional image data based on the initial pose data and the target pose data, and generates the target image data corresponding to the target pose data.

Benefits of technology

Ensure that the image data displayed by the terminal always corresponds to the current position, reduces hardware requirements, avoids the problem of untimely image refresh caused by network delay, and improves the effect and speed of image display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114078092B_ABST
    Figure CN114078092B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention relate to the technical field of computer graphics, and disclose a method, device, electronic device, and storage medium for image processing. The method for image processing in the present invention includes: sending the initial pose data of the terminal at the initial moment to the server, where the initial moment is a moment before the current moment; receiving the initial image data returned by the server according to the initial pose data, and the initial image data includes two-dimensional image data corresponding to the initial pose data; obtaining the target pose data of the terminal at the current moment; and adjusting the two-dimensional image data according to the initial pose data and the target pose data to obtain the target image data corresponding to the current moment. By adopting the embodiments of the present application, it is possible to avoid the problem that the virtual reality terminal cannot display an image corresponding to the terminal pose within the refresh period due to network delay, and at the same time reduce the hardware requirements of the virtual device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer graphics, and particularly to a method, device, electronic device and storage medium for image processing. Background Art

[0002] Currently, when using a virtual reality (VR) terminal, it often occurs that the terminal cannot synchronously render a picture corresponding to the terminal's actions. That is, a part of the picture displayed by the virtual reality terminal is the image displayed last time, and a part is the newly rendered image, resulting in problems such as lag and stutter in the continuously displayed picture seen by the user of the virtual reality terminal, and even the user may experience dizziness. Currently, asynchronous distortion techniques such as asynchronous rotation and displacement of the picture are used to solve the problem of picture refresh display caused by untimely local VR rendering. For example, Asynchronous Timewarp (ATW), Asynchronous Space warp (ASW), etc.

[0003] During the implementation of the asynchronous distortion technique, the rendering thread and the ATW thread in the GPU run asynchronously. Before each synchronization between the rendering thread and the ATW thread, the ATW thread generates a new image to be displayed based on the last frame image of the rendering thread. However, when it is detected that the rendering thread cannot render a new image to be displayed within the refresh period, the ATW thread needs to preempt the rendering thread; therefore, the GPU hardware needs to support a reasonable preemption granularity. For example, at 90 Hz, the interval between frames is approximately 11 ms (1 / 90 Hz). To ensure that the ATW thread can generate a new frame of image, the ATW thread needs to be able to preempt the rendering thread and run for less than 11 ms; at the same time, it is also required that the operating system and the driver support the GPU preemption. It can be seen that the asynchronous distortion technique has relatively high requirements for the hardware of the virtual reality terminal; if the server performs image rendering, due to network latency, it will also cause the terminal to be unable to synchronously render a picture corresponding to the terminal's actions. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a method for image processing, which can avoid the problem that the virtual reality terminal cannot display an image corresponding to the terminal's pose within the refresh period due to network latency, and at the same time reduce the hardware requirements of the virtual device.

[0005] To achieve the above object, an embodiment of the present application provides a method for image processing, including: sending the initial pose data of the terminal at the initial moment to the server, where the initial moment is a moment before the current moment; receiving the initial image data returned by the server according to the initial pose data, and the initial image data includes two-dimensional image data corresponding to the initial pose data; obtaining the target pose data of the terminal at the current moment; and adjusting the two-dimensional image data according to the initial pose data and the target pose data to obtain the target image data corresponding to the target pose data.

[0006] To achieve the above object, an embodiment of the present application further provides an apparatus for image processing, including: a sending module, a receiving module, an obtaining module, and an adjusting module; the sending module is configured to send the initial pose data of the terminal at the initial moment to the server, where the initial moment is a moment before the current moment; the receiving module is configured to receive the initial image data returned by the server according to the initial pose data, and the initial image data includes two-dimensional image data corresponding to the initial pose data; the obtaining module is configured to obtain the target pose data of the terminal at the current moment; and the adjusting module is configured to adjust the two-dimensional image data according to the initial pose data and the target pose data to obtain the target image data corresponding to the target pose data.

[0007] To achieve the above object, an embodiment of the present application further provides an electronic device, including: at least one processor; and,

[0008] a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above method for image processing.

[0009] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method for image processing is implemented.

[0010] The image processing method proposed in this application enables the server to return initial image data based on the initial pose data of the terminal. The initial image data includes two-dimensional image data corresponding to the initial pose data. The terminal adjusts the two-dimensional image data according to the target pose data and the initial pose data at the current moment to generate target image data corresponding to the target pose data. This ensures that the two-dimensional image data displayed on the terminal always corresponds to the current pose of the terminal. Since the two-dimensional image data has a low dimension and can be adjusted quickly, even if there is network latency, it can ensure that the adjusted two-dimensional image data is refreshed in a timely manner within the refresh cycle, thus avoiding the problem of untimely image refresh within the refresh cycle due to network latency and improving the image display effect on the terminal. In addition, the server sends the initial image data to the terminal, enabling the terminal to avoid complex rendering operations, improving the image display speed, and reducing the hardware requirements of the terminal device. Moreover, during the process of the terminal generating the target image, it does not require the support of the terminal's GPU hardware for reasonable preemption granularity, nor does it require the support of the terminal's operating system and driver program to enable GPU preemption, further reducing the hardware requirements of the terminal and being more conducive to the popularization of virtual reality terminals. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a flowchart of the image processing method according to the first embodiment of the present invention;

[0012] Figure 2 is a flowchart of the image processing method according to the second embodiment of the present invention;

[0013] Figure 3 is a flowchart of the image processing method according to the third embodiment of the present invention;

[0014] Figure 4 is a structural block diagram of the image processing device according to the fourth embodiment of the present invention;

[0015] Figure 5 is a structural block diagram of the electronic device according to the fifth embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will elaborate on each embodiment of this application with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of this application, many technical details are presented to help readers better understand this application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in this application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined and cross-referenced with each other on the premise of not being contradictory.

[0017] Cloud-based virtual reality technology renders the images to be displayed on virtual reality terminals through a server, which can greatly reduce the device power consumption caused by the rendering of local virtual reality terminals and lower the hardware requirements for virtual reality terminals. Moreover, with the continuous development of 5G technology, the network transmission speed has been improved, further enhancing the application of cloud-based virtual reality technology. However, network transmission usually has network latency, and the VR terminal cannot synchronously render the images corresponding to the terminal actions within the refresh cycle, resulting in problems such as stuttering and jitter in the continuous images seen by users, seriously affecting the user experience.

[0018] The first embodiment of the present invention relates to a method for image processing, which is applied to a virtual reality terminal. The VR terminal can be a head-mounted VR glasses, helmet, etc. The process is as follows Figure 1 shown:

[0019] Step 101: Send the initial pose data of the terminal at the initial moment to the server, where the initial moment is a moment before the current moment.

[0020] Step 102: Receive the initial image data returned by the server according to the initial pose data. The initial image data includes two-dimensional image data corresponding to the initial pose data.

[0021] Step 103: Obtain the target pose data of the terminal at the current moment.

[0022] Step 104: Adjust the two-dimensional image data according to the initial pose data and the target pose data to obtain the target image data corresponding to the target pose data.

[0023] The image processing method proposed in this application enables the server to return initial image data based on the initial pose data of the terminal. The initial image data includes two-dimensional image data corresponding to the initial pose data. The terminal adjusts the two-dimensional image data according to the target pose data and the initial pose data at the current moment to generate target image data corresponding to the target pose data. This ensures that the two-dimensional image data displayed on the terminal always corresponds to the current pose of the terminal. Since the two-dimensional image data has a low dimension and can be adjusted quickly, even if network latency occurs, it can ensure that the adjusted two-dimensional image data is refreshed in time within the refresh cycle, thus avoiding the problem of untimely image refresh within the refresh cycle due to network latency and improving the image display effect on the terminal. In addition, the server sends the initial image data to the terminal, enabling the terminal to avoid complex rendering operations, improving the image display speed, and reducing the hardware requirements of the terminal device. Moreover, during the process of the terminal generating the target image, it does not require the GPU hardware of the terminal to support a reasonable preemption granularity, nor does it require the operating system and driver of the terminal to support GPU preemption, further reducing the hardware requirements of the terminal and being more conducive to the popularization of virtual reality terminals.

[0024] The second embodiment of the present invention relates to an image processing method. The second embodiment specifically describes steps 101-103 in the first embodiment, and its process is specifically as Figure 2 shown:

[0025] Step 201: Send the initial pose data of the terminal at the initial moment to the server. The initial moment is a moment before the current moment.

[0026] Specifically, the virtual reality terminal can obtain its own pose data in real time. The pose data includes the attitude data and position data of the terminal. The initial moment is a moment before the current moment. For example, the current moment is t1; the initial moment t0 < t1; the interval between the initial moment and the current moment can be less than a preset interval. For example, the preset interval can be 5s, 1ms, 0.1ms; the preset interval can also be the same as the time interval for collecting the pose data of the terminal.

[0027] It should be noted that the initial moment is not a fixed moment but a moment that changes dynamically according to the current moment. For example, if the current moment is the 80th second, the initial moment can be less than the 80th second, such as the 79th second, the 78th second, etc.

[0028] The form of the initial pose data can be various. In this example, the type of the initial pose data is not restricted. For example, in this example, the expression form of the initial pose data can be a matrix. The terminal sends the collected initial pose data to the server.

[0029] The terminal can obtain its own pose data in real time through its own sensors, and the sensors can be gyroscopes, laser positioning devices, etc.

[0030] Step 202: Receive the initial image data returned by the server according to the initial pose data. The initial image data includes two-dimensional image data corresponding to the initial pose data.

[0031] Specifically, the server can pre-store the three-dimensional image data of the space where the virtual reality terminal is located; for example, pre-store the three-dimensional image data of multiple game scenes, such as game scenes in places like forests, secret rooms, cruise ships, etc. After the server receives the initial pose data, according to the initial pose data, the server can render the stereoscopic image data observed by the terminal at the pose at the initial moment into two-dimensional image data corresponding to the initial pose data. This two-dimensional image data corresponds to the initial moment and also corresponds to the initial pose data.

[0032] It is worth mentioning that rendering an image is to convert high-dimensional image data into low-dimensional image data. After the terminal receives the two-dimensional image data, it enables the terminal to quickly draw the low-dimensional two-dimensional image data, which can reduce the resource usage of the terminal and improve the speed of displaying the image.

[0033] It should be noted that the initial image data includes the two-dimensional image data, the depth data corresponding to each pixel point in the two-dimensional image data, and the color information corresponding to each pixel point.

[0034] In addition, the server can also return the two-dimensional image data and the corresponding initial pose data to the terminal at the same time, so as to ensure that the initial pose data and the two-dimensional image data of the terminal correspond one by one.

[0035] Step 203: Obtain the target pose data of the terminal at the current moment.

[0036] Specifically, after receiving the initial image data, the target pose data of the terminal at the current moment can be obtained. The obtaining method is roughly the same as the process of Step 201. Similarly, the target pose data can include: the attitude data and position data of the terminal at the current moment. The attitude data can be, for example, the angle of rotation of the terminal, and the position data can be the space coordinates where the terminal is located at the current moment.

[0037] Step 204: Obtain the initial coordinate data of the two-dimensional image data in the screen coordinate system and each depth data.

[0038] Specifically, the initial image data includes the coordinates of the two-dimensional image data in the screen coordinate system. Take this coordinate as the initial coordinate data of the two-dimensional image data. For example, the initial coordinate data of the two-dimensional image in the screen coordinate system can be expressed as (Xcurrent, Ycurrent)T The initial image data further includes depth data and color information corresponding to each pixel point in the two-dimensional image data. The color information can specifically be the color values of the RGB primary colors; the depth data can be expressed as: DepthBuffer(Xcurrent, Ycurrent).

[0039] Step 205: According to the initial coordinate data, each depth data, and a preset coordinate system conversion relationship, convert the initial coordinate data into first conversion coordinate data based on the world coordinate system.

[0040] In an example, according to the first conversion relationship and each depth data, convert the initial coordinate data into first standard device coordinates based on the standard device coordinate system; obtain a pose matrix and a projection matrix from the initial pose data; according to the pose matrix, the projection matrix, the first standard device coordinates, and the second conversion relationship, obtain the first conversion coordinate data.

[0041] Specifically, information such as the width and height of the terminal screen can be obtained. For example, the screen width ScreenWidth and the screen height ScreenHeight, then 0 ≤ Xcurrent ≤ ScreenWidth, 0 ≤ Ycurrent ≤ ScreenHeight. The color information corresponding to each pixel point in the two-dimensional image data is expressed as ColorBuffer, and the depth data is expressed as DepthBuffer. The depth corresponding to each pixel point in the screen space coordinate system is DepthBuffer(Xcurrent, Ycurrent).

[0042] The preset coordinate conversion relationship includes: a first conversion relationship for converting between the screen coordinate system and the standard device coordinate system, and a second conversion relationship for converting between the standard device coordinate system and the world coordinate system.

[0043] In this example, the first conversion relationship is:

[0044] x = 2 * Xcurrent / ScreenWidth - 1; y = 2 * Ycurrent / ScreenHeight - 1;

[0045] z = DepthBuffer(Xcurrent, Ycurrent); w = 1.0 Formula (1);

[0046] Among them, the initial coordinate data can be expressed as (Xcurrent, Ycurrent) T, the screen width ScreenWidth and the screen height ScreenHeight, and 0 ≤ Xcurrent ≤ ScreenWidth, 0 ≤ Ycurrent ≤ ScreenHeight; w is a preset value used to characterize the feature that the image data shows objects larger when closer and smaller when farther away; where z represents the vertical axis in the standard device coordinate system.

[0047] According to formula (1), the initial coordinate data (Xcurrent, Ycurrent) T can be converted into the first standard device coordinates in the standard device coordinate system, that is, the obtained first standard device coordinates are expressed as (x, y, z, w) T .

[0048] Specifically, the second conversion relationship between the standard device coordinate system and the world coordinate system can be expressed as:

[0049] (x_world, y_world, z_world, w_world) T = Rcurrent -1 * P -1 *(x, y, z, w) T Formula (2);

[0050] where, (x_world, y_world, z_world, w_world) T represents the coordinates in the world coordinate system, which is expressed as the first conversion coordinate data in this example, and (x, y, z, w) T represents the first standard device coordinates, P represents the projection matrix for converting the three-dimensional image data into two-dimensional image data; Rcurrent represents the initial pose data of the terminal at the initial moment, and this initial pose data can be represented in the form of a matrix.

[0051] According to the second conversion relationship represented by formula (2), the first standardized device coordinates (x, y, z, w)T can be converted into the first conversion coordinate data (x_world, y_world, z_world, w_world) in the world coordinate system T .

[0052] Step 206: Use the first conversion coordinate data as the second conversion coordinate data corresponding to the target pose data in the world coordinate system.

[0053] In one example, use this first conversion coordinate as the second conversion coordinate data corresponding to the target pose data in the world coordinate system, that is:

[0054] (x_world_n, y_world_n, z_world_n, w_world_n)T =(x_world, y_world, z_world, w_world) T

[0055] where (x_world_n, y_world_n, z_world_n, w_world_n) T represents the coordinates corresponding to the target pose data in the world coordinate system; (x_world, y_world, z_world, w_world) T represents the first transformed coordinate data.

[0056] Step 207: Generate target image data according to the second transformed coordinate data and the coordinate transformation relationship.

[0057] In one example, according to the second transformed coordinate data, the target pose data, the projection matrix, and the second transformation relationship, obtain the second standard device coordinate data corresponding to the target pose data in the standard device coordinate system; according to the first transformation relationship, convert the second standard device coordinate data into target coordinate data based on the screen coordinate system; map the color information of each pixel point in the two-dimensional image data to the corresponding pixel point in the target coordinate data to form the target image data.

[0058] Specifically, according to the second transformation relationship, as can be seen from formula (2);

[0059] (x_world_n, y_world_n, z_world_n, w_world_n) T = Rnext -1 * P -1 * (x’, y’, z’, w’) T

[0060] Formula (3);

[0061] where (x’, y’, z’, w’) T represents the second standard device coordinate data corresponding to the current moment. By simply transforming formula (3), we can obtain:

[0062] (x’, y’, z’, w’) T = P * Rnext * (x_world_n, y_world_n, z_world_n, w_world_n) T

[0063] Since

[0064] (x_world_n, y_world_n, z_world_n, w_world_n) T=(x_world, y_world, z_world, w_world) T

[0065] Then, (x’, y’, z’, w’) T = P * Rnext * (x_world, y_world, z_world, w_world) T Formula (4);

[0066] Wherein, Rnext represents the target pose data, which is represented in matrix form in this example; P is the projection matrix. (x’, y’, z’, w’) T represents the second standard device coordinate data corresponding to the target pose data.

[0067] According to the first conversion relationship, the second standard device coordinate data can be converted into the coordinates in the screen coordinate system at the current moment;

[0068] Xnext = ScreenWidth * (x’ + 1) / 2; Ynext = ScreenHeight * (y’ + 1) / 2; Formula (5);

[0069] The screen coordinates at the current moment are represented as (Xnext, Ynext) T .

[0070] After obtaining the target coordinate data of the image data at the current moment based on the screen coordinate system, the color information of each pixel point in the two-dimensional image data is mapped to the corresponding pixel point in the target coordinate data to form the target image data.

[0071] For example, the color1 corresponding to the point with coordinates (0, 0) in the two-dimensional image data. The coordinates (0, 0) are transformed to the position (3, 0). Then, the pixel point at the position (3, 0) is filled with color1.

[0072] It should be noted that steps 204 to 207 are specific descriptions of step 103.

[0073] The third embodiment of the present invention relates to a method for image processing. This embodiment is a further improvement of the second embodiment. The main improvement lies in that: in this embodiment, the terminal searches for the initial pose data corresponding to the two-dimensional image data according to the time identification information. The process of this image processing method is as Figure 3 shown.

[0074] Step 301: Send the initial pose data of the terminal at the initial moment to the server.

[0075] This step 301 is substantially the same as step 201 and will not be elaborated here.

[0076] Step 302: Send the time identification information used to represent the initial moment to the server.

[0077] Specifically, the time identification information can be set according to the actual application. For example, the time identification information can be the time stamp of the initial moment. Send the time stamp of the initial moment to the server. This time stamp can be sent to the server simultaneously with the initial pose data. Or, after the initial pose data is sent, upload this time stamp to the server.

[0078] Step 303: Use the time identification information as an index to store the corresponding initial pose data in the pose buffer queue.

[0079] Specifically, the time identification information can be used as an index to cache the initial pose data in the terminal, so that the terminal can quickly find the initial pose data according to the time stamp in the future. For example, a pose matrix buffer queue can be set in the terminal, and the initial pose data is stored in this pose matrix buffer queue, and the time stamp is used as the index of the initial pose data.

[0080] Step 304: Receive the initial image data returned by the server according to the initial pose data.

[0081] Specifically, the initial image data further includes: time identification information. The server can send the initial image data and this time identification information to the terminal together. A receiving thread and a decoding thread can be set in the terminal, and the receiving thread and the decoding thread can run asynchronously to improve the speed of the terminal to display images. The receiving thread is used to receive the two-dimensional image data and the corresponding time identification information. The decoding thread is used to decode the two-dimensional image data.

[0082] Step 305: Store the two-dimensional image data in the first buffer queue and store the time identification information in the second buffer queue.

[0083] Specifically, the terminal can set a first buffer queue and a second buffer queue. The first buffer queue is used to cache the received two-dimensional image data, and the second buffer queue is used to store the time identification information corresponding to the initial moment. The decoding thread is used to obtain the two-dimensional image data from the first buffer queue for decoding.

[0084] Step 306: If it is detected that the two-dimensional image data has been decoded, then read the time identification information from the second buffer queue.

[0085] Specifically, since the receiving thread and the decoding thread run asynchronously, if it is detected that the two-dimensional image data has been decoded, the time identification information is read from the second buffer queue, so as to ensure that the initial pose data corresponding to the two-dimensional image data can be obtained. For example, if each frame of two-dimensional image data is decoded, the time identification information corresponding to the two-dimensional image data is read from the second buffer queue.

[0086] Step 307: Read the initial pose data from the pose buffer queue according to the time identification information.

[0087] Specifically, since the initial pose data is stored with the time identification information, the initial pose data corresponding to the time identification information can be read through the time identification information. For example, the initial time is time T1, and the corresponding initial pose data is Rcurret. According to this T1, the Rcurrent can be read from the pose matrix buffer queue.

[0088] Step 308: Obtain the target pose data of the terminal at the current moment.

[0089] It is substantially the same as step 203 in the second embodiment, and will not be elaborated here.

[0090] Step 309: Adjust the two-dimensional image data according to the initial pose data and the target pose data to obtain the target image data corresponding to the current moment.

[0091] It is substantially the same as steps 204 to 207 in the second embodiment, and will not be elaborated here.

[0092] In the image processing method provided in this embodiment, the initial image data includes the time identification information corresponding to the initial time, so that the server does not need to return the initial pose data. The terminal can read the initial pose data according to the time identification information, which improves the speed of obtaining the initial pose data. At the same time, the first buffer queue and the second buffer queue are set, so that decoding can be performed asynchronously, further improving the speed of displaying image data.

[0093] The fourth embodiment of the present invention relates to an image processing device, and its structural block diagram is as Figure 4 shown, including: a sending module 401, a receiving module 402, an obtaining module 403, and an adjusting module 404.

[0094] The sending module 401 is used to send the initial pose data of the terminal at the initial moment to the server, where the initial moment is the moment before the current moment; the receiving module 402 is used to receive the initial image data returned by the server according to the initial pose data, and the initial image data includes two-dimensional image data corresponding to the initial pose data; the obtaining module 403 is used to obtain the target pose data of the terminal at the current moment; the adjusting module 404 is used to adjust the two-dimensional image data according to the initial pose data and the target pose data to obtain the target image data corresponding to the target pose data.

[0095] It is not difficult to find that this embodiment is a device embodiment corresponding to the first embodiment, and this embodiment can be implemented in cooperation with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the first embodiment.

[0096] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of the present invention, units that are not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0097] The fifth embodiment of the present invention relates to an electronic device, and its structural block diagram is as Figure 5 shown. The electronic device includes: at least one processor 501; and a memory 502 communicatively connected to the at least one processor 501; wherein, the memory 502 stores instructions executable by the at least one processor 501, and the instructions are executed by the at least one processor 501 to enable the at least one processor 501 to execute the above-mentioned image processing method.

[0098] Among them, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus links various circuits of one or more processors and memories together. The bus can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0099] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory can be used to store data used by the processor during operation.

[0100] The sixth embodiment of the present invention relates to a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned image processing method.

[0101] Those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0102] Those of ordinary skill in the art can understand that the above embodiments are specific examples for implementing the present invention, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present invention.

Claims

1. A method for image processing, characterized in that, Including: Sending the initial pose data of the terminal at the initial moment to the server, where the initial moment is a moment before the current moment; Receiving the initial image data generated by the server in real time according to the initial pose data, where the initial image data includes two-dimensional image data corresponding to the initial pose data; Obtaining the target pose data of the terminal at the current moment; Adjusting the two-dimensional image data according to the initial pose data and the target pose data to obtain target image data corresponding to the target pose data; Wherein, after the server generates the initial image data in real time, the initial image data is directly pushed to the terminal, and the adjustment of the two-dimensional image data by the terminal is completed within a single refresh cycle.

2. The method for image processing according to claim 1, wherein The initial image data further includes: depth data corresponding to each pixel point in the two-dimensional image data; The adjusting the two-dimensional image data according to the initial pose data and the target pose data to obtain target image data corresponding to the target pose data includes: Obtaining the initial coordinate data of the two-dimensional image data in the screen coordinate system and each of the depth data; Converting the initial coordinate data into first conversion coordinate data based on the world coordinate system according to the initial coordinate data, each of the depth data, and a preset coordinate conversion relationship; Taking the first conversion coordinate data as second conversion coordinate data corresponding to the target pose data based on the world coordinate system; Generating the target image data according to the second conversion coordinate data and the coordinate conversion relationship.

3. The method for image processing according to claim 1 or 2, characterized in that, Before receiving the initial image data returned by the server according to the initial pose data, the method further includes: Sending time identification information for representing the initial moment to the server; Storing the corresponding initial pose data in a pose buffer queue with the time identification information as an index.

4. The method for image processing according to claim 3, wherein The initial image data further includes: the time identification information; Before adjusting the two-dimensional image data according to the initial pose data and the target pose data to obtain target image data corresponding to the target pose data, the method further includes: Reading the initial pose data from the pose buffer queue according to the time identification information.

5. The method for image processing according to claim 4, wherein Before reading the initial pose data from the pose buffer queue according to the time identification information, the method further includes: Storing the two-dimensional image data in a first buffer queue and storing the time identification information in a second buffer queue; If it is detected that the two-dimensional image data is decoded, then reading the time identification information from the second buffer queue.

6. The method for image processing according to claim 2, characterized in that, The coordinate conversion relationship includes: a first conversion relationship for converting between the screen coordinate system and the standard device coordinate system, and a second conversion relationship for converting between the standard device coordinate system and the world coordinate system; The converting the initial coordinate data into first conversion coordinate data based on the world coordinate system according to the initial coordinate data, each of the depth data, and a preset coordinate conversion relationship includes: Convert the initial coordinate data into first standard device coordinate data based on the standard device coordinate system according to the first conversion relationship and the respective depth data; Obtain a pose matrix and a projection matrix from the initial pose data; Obtain the first converted coordinate data according to the pose matrix, the projection matrix, the first standard device coordinate data, and the second conversion relationship.

7. The method for image processing according to claim 6, wherein The generating the target image data according to the second converted coordinate data and the coordinate conversion relationship includes: Obtain second standard device coordinate data corresponding to the target pose data in the standard device coordinate system according to the second converted coordinate data, the target pose data, the projection matrix, and the second conversion relationship; Convert the second standard device coordinate data into target coordinate data based on the screen coordinate system according to the first conversion relationship; Map the color information of each pixel point in the two-dimensional image data to the corresponding pixel point in the target coordinate data to form the target image data.

8. An apparatus for image processing, characterized in that, including: A sending module, a receiving module, an obtaining module, and an adjusting module; The sending module is configured to send the initial pose data of the terminal at an initial moment to the server, where the initial moment is a moment before the current moment; The receiving module is configured to receive the initial image data generated by the server in real time according to the initial pose data, where the initial image data includes two-dimensional image data corresponding to the initial pose data; The obtaining module is configured to obtain the target pose data of the terminal at the current moment; The adjusting module is configured to adjust the two-dimensional image data according to the initial pose data and the target pose data to obtain target image data corresponding to the target pose data; wherein, after the server generates the initial image data in real time, the server directly pushes the initial image data to the terminal, and the adjustment of the two-dimensional image data by the terminal is completed within a single refresh cycle.

9. An electronic device, characterized in that, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the image processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Virtual reality interactive system and method based on machine vision

    CN104699247A

  • Method and systems used for presenting AR (Augmented Reality) information

    CN107977082A

Cited By

  • Image processing method and apparatus, and electronic device and storage medium

    WO2022033389A1