Image processing method and system applied to display area of virtual exhibition hall
Through eye tracking and convolutional neural network processing image data, combined with Fourier transform and filtering technology, the problems of image delay and shadowing in VR are solved, adaptive clarity adjustment and content preloading are realized, and the user experience of the virtual exhibition hall is improved.
Patent Information
- Application Number
- CN202510384080.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
AI Technical Summary
In the existing VR technology, the low GPU processing efficiency leads to the image rendering speed that cannot keep up with the user's head movement and field of view movement, resulting in image delay and smear, especially in serious dizziness.
The eye tracking sensor collects viewing distance and viewing angle offset angle data, uses a convolutional neural network to process image pixel density and focus area coordinates, generates adaptive resolution images, and performs fast Fourier transform and high-pass filtering processing in the focus area, and preloads the network delay content based on user behavior analysis.
It realizes automatic adjustment of image clarity based on user location and perspective, reducing latency and scheduling, providing a smooth virtual exhibition hall experience, and reducing waiting time when network delays.
Smart Images

Figure CN120234485A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to an image processing method and system applied to a display area of a virtual exhibition hall. Background Art
[0002] VR (Virtual Reality), also known as spiritual realm technology, is an advanced computer human-computer interface with immersion, interactivity and imagination as its basic characteristics. It makes comprehensive use of computer graphics, simulation technology, multimedia technology, artificial intelligence technology, computer network technology, parallel processing technology and multi-sensor technology to simulate the functions of human sense organs such as vision, hearing and touch, so that people can immerse themselves in the virtual realm generated by the computer and interact with it in real time through language, gestures, mouse and keyboard, etc., creating a human-friendly multi-dimensional information space. It is a new technology in development with far-reaching potential application directions.
[0003] With the development of VR, the processing efficiency of GPU (Graphic Processing Unit) is increasingly required. When the GPU processing efficiency is low, the image rendering speed cannot keep up with the speed of the user's head movement and field of view, causing image delay. Studies have shown that the delay in head movement and field of view cannot exceed 20ms, otherwise there will be a very obvious sense of ghosting, and long-term ghosting can easily cause dizziness. Summary of the invention
[0004] In order to reduce the delay and smearing of images caused by movement, the present invention provides an image processing method applied to the display area of a virtual exhibition hall, and the method mainly comprises the steps of: S101. Obtain viewing distance data and viewing angle offset data collected by the eye tracking sensor; S102. Determine the head position state according to the viewing distance data and the viewing angle offset data, and generate a dynamic perception result; S103. Processing the original image pixel density and the focus area coordinates corresponding to the dynamic perception result using a convolutional neural network to determine a resolution threshold range; if the resolution threshold range exceeds a preset standard value, performing bicubic interpolation resampling on the original image pixel density to generate an adaptive resolution image; S104. When it is detected that the focal area coordinates are displaced, a fast Fourier transform is performed on the adaptive resolution image to obtain frequency domain data of a specified window size, and a high-pass filtering process is performed on the frequency domain data according to a preset frequency domain smoothing coefficient to obtain filtered frequency domain data; S105. Performing inverse Fourier transform precision control processing on the filtered frequency domain data to restore and generate an optimized focus detail image; S106. Input the focus detail image and the pre-established low-frequency background template into a real-time rendering engine to generate final display frame data, and output the final display frame data to the display area of the virtual exhibition hall; S107. Input the final display frame data into the user behavior analysis and prediction module to obtain the display content preloaded during network delay.
[0005] Preferably, the step 102 includes: Collecting the viewing distance data and the viewing angle deviation data through a sensor, and preprocessing the collected data; Matching the preprocessed data with a predefined head position state to determine the user's real-time head state; Generate a dynamic perception result according to the real-time head state.
[0006] Preferably, the step S103 includes: Acquiring pixel density information of an original image and determining focal region coordinates in the original image by using saliency analysis based on a dynamic perception result; Inputting the pixel density information and the focal area coordinates into a convolutional neural network, and outputting a resolution threshold range, wherein a convolutional layer is used to extract local features of the focal area, a pooling layer is used to reduce the dimension of the local features, and a fully connected layer is used to output the resolution threshold range; The resolution threshold range is compared with a preset standard value. If the resolution threshold range is within the preset standard value, the original image pixel density is directly used for processing. If the resolution threshold range exceeds the preset standard value, the original image is resampled by bicubic interpolation to generate an adaptive resolution image.
[0007] Preferably, the step S104 includes: Compare the current focus area coordinates with the focus area coordinates of the previous frame to determine whether displacement occurs; If it is determined that displacement occurs, a fast Fourier transform is performed on the adaptive resolution image to convert the adaptive resolution image from the spatial domain to the frequency domain, and in the frequency domain, a window of a predetermined size is selected to extract frequency domain data near the coordinates of the current focal area to obtain frequency domain data of a predetermined window size; A high-pass filter is constructed by using a preset frequency domain smoothing coefficient to filter the frequency domain data to obtain filtered frequency domain data, and the filtered frequency domain data is converted back to the spatial domain to obtain a frequency domain focal image.
[0008] Preferably, the step S107 includes: Clean, denoise, and extract features from the final display frame data to obtain the user's current behavior trajectory; Identify the user's typical behavior patterns by analyzing the user's historical behavior data; Extract the areas of interest to the user based on the user's historical operation frequency and historical stay time. Based on the user's current behavior trajectory, combined with the user's typical behavior patterns and the areas of interest to the user, obtain the pre-loaded content, and dynamically adjust the pre-loaded content in combination with the current network status.
[0009] Preferably, the dynamically adjusting the pre-loaded content in combination with the current network status includes: Construct a delay prediction model based on historical network delay data to predict the possible delay situation under the current network conditions; Rank the content to be displayed according to the user behavior prediction results; dynamically adjust the pre-loaded content in combination with the predicted network delay.
[0010] In a second aspect, the present invention provides an image processing system applied to a virtual exhibition hall display area, the system includes: A data acquisition module that acquires the viewing distance data and the viewing angle offset data collected by the eye tracking sensor; A dynamic perception result generation module that determines the head position state according to the viewing distance data and the viewing angle offset data, and generates a dynamic perception result; An adaptive resolution image generation module that processes the original image pixel density and the focus area coordinates corresponding to the dynamic perception result by using a convolutional neural network to determine the resolution threshold range; if the resolution threshold range exceeds the preset standard value, perform bicubic interpolation resampling on the original image pixel density to generate an adaptive resolution image; A frequency domain focus image generation module that, when detecting a displacement of the focus area coordinates, performs a fast Fourier transform on the adaptive resolution image to obtain frequency domain data of a specified window size, and performs high-pass filtering processing on the frequency domain data according to a preset frequency domain smoothing coefficient to obtain a frequency domain focus image; A focus detail image generation module that performs inverse Fourier transform precision control processing on the frequency domain focus image to restore and generate an optimized focus detail image; A display module that inputs the focus detail image and a pre-established low-frequency background template into a real-time rendering engine to generate final display frame data, and outputs the final display frame data to the display area of the virtual exhibition hall; A display content prediction module that inputs the final display frame data into a user behavior analysis and prediction module to obtain the display content pre-loaded during network delay.
[0011] In a third aspect, the present invention provides an electronic device, including a memory, an image processor, and a computer program stored on the memory and executable on the image processor. When the image processor executes the program, the image processing method applied to the display area of the virtual exhibition hall is implemented.
[0012] In a fourth aspect, the present aspect provides a non-transitory computer-readable storage medium, on which a computer program is stored. The computer program, when executed by an image processor, implements the image processing method applied to the display area of the virtual exhibition hall.
[0013] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: By developing an adaptive clarity enhancement algorithm, the clarity of the display object is automatically adjusted according to the user's viewing distance and angle. When the user approaches a certain exhibit, the system can automatically increase the resolution and details of this area; when the user views the exhibit from different angles, the system can adjust the clarity and focus of the image in real time according to the perspective change, ensuring that the user can obtain the best viewing experience at any position. It can also predict the exhibits that the user may view next according to the user's browsing trajectory and interest preferences, and cache the relevant content in advance, reducing the waiting time even in case of network latency. At the same time, through the progressive image loading technology, the user can obtain a basic visual experience while waiting for the high-definition image. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent: Figure 1 The hardware structure block diagram of a mobile terminal of an image processing method applied to the display area of the virtual exhibition hall according to an embodiment is shown; Figure 2 The step flow chart of an image processing method applied to the display area of the virtual exhibition hall according to an embodiment is shown; Figure 3 The step flow chart of step S104 according to an embodiment is shown; The realization, functional characteristics, and advantages of the object of the present invention will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of this application, and are not used to limit the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by this application.
[0016] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of this application, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0017] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of this application, the descriptions of "first", "second", etc. are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0018] The method embodiments provided in the embodiments of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for an image processing method applied to a virtual exhibition hall display area in an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than those shown in Figure 1 the figure, or have a different configuration from that shown in Figure 1 the figure.
[0019] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to a data information security protection method in an embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0020] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0021] A schematic flowchart of an image processing method applied to a virtual exhibition hall display area in an embodiment of the present application, as shown in the appendix Figure 2 As shown, an embodiment of the present application provides an image processing method applied to a virtual exhibition hall display area, and the method includes the following steps: S101. Obtain the viewing distance data and the viewing angle offset data collected by the eye tracking sensor; In this embodiment, first, determine the eye tracking device to be used. Different devices may have different APIs and data output formats. Installing and configuring this software is the first step to obtain data. Taking the Tobii Eye Tracker as an example, download and install the Tobii Pro SDK from the Tobii official website; ensure that your development environment supports the Tobii Pro SDK, such as Python, C#, C++; initialize the eye tracking device in the code to start data collection, and parse out the viewing distance and the viewing angle offset according to the data format provided by the device.
[0022] S102. Determine the head position state according to the viewing distance data and the viewing angle offset data, and generate a dynamic perception result; In this embodiment, the viewing distance data and the viewing angle offset data are collected by sensors, and the collected data is preprocessed; the preprocessed data is matched with predefined head position states to determine the user's real-time head state; and a dynamic perception result is generated according to the real-time head state.
[0023] Specifically, the collected data is preprocessed to ensure the accuracy and consistency of the data. This includes: using a low-pass filter or a Kalman filter to remove noise, ensuring the zero calibration of the sensor, and avoiding systematic errors. Different head position states are defined, for example: straight ahead: the head is in front of the screen directly, and the viewing angle offset is close to 0 degrees. Left deviation: the head deviates to the left, and the viewing angle offset is negative. Right deviation: the head deviates to the right, and the viewing angle offset is positive. Up deviation: the head deviates upward, and the viewing angle offset is negative (in the vertical direction). Down deviation: the head deviates downward, and the viewing angle offset is positive (in the vertical direction). Far away: the head is far from the screen. Close: the head is close to the screen. According to the collected viewing distance and viewing angle offset data and the predefined head position states, the real-time head position state of the user is judged; The preprocessed data is compared with a predefined head position state model. By learning different head pose features, a pattern recognition model that can accurately identify various head position states is established. When the actually collected data is input into this model, the system can automatically judge which state the current user's head is in.
[0024] A dynamic perception result is generated according to the head position state. For example, when the user's head is far from the screen, it is recommended to get closer for a better viewing experience; when the user's head is close to the screen, it is recommended to move away appropriately to protect eyesight; when the user's head deviates, it is recommended to adjust the head position for a better viewing experience; when the user's head deviates up and down, it is recommended to adjust the head position for a better viewing experience.
[0025] S103. Use a convolutional neural network to process the original image pixel density and the focus area coordinates corresponding to the dynamic perception result, and determine the resolution threshold range; if the resolution threshold range exceeds a preset standard value, perform bicubic interpolation resampling on the original image pixel density to generate an adaptive resolution image; In this embodiment, the pixel density information of the original image is obtained, and the coordinates of the focus area in the original image are determined by using saliency analysis based on the dynamic perception result; the pixel density information and the focus area coordinates are input into a convolutional neural network, and a resolution threshold range is output. Among them, the convolutional layer is used to extract the local features of the focus area, the pooling layer is used to reduce the dimension of the local features, and the fully connected layer is used to output the resolution threshold range; the resolution threshold range is compared with a preset standard value. If the resolution threshold range is within the preset standard value, the original image pixel density is directly used for processing. If the resolution threshold range exceeds the preset standard value, the original image is resampled by bicubic interpolation to generate an adaptive resolution image.
[0026] Specifically, the pixel density information is obtained through image metadata (such as EXIF information) or an image processing library (such as OpenCV, PIL, etc.). The pixel density usually refers to the number of pixels per inch (PPI, Pixels Per Inch), which is used to measure the clarity of an image. For a digital image, the pixel density can be calculated by dividing the width and height of the image (in pixels) by the physical size (in inches); The coordinates of the focus area are determined through a saliency map. The saliency map is a grayscale image of the same size as the original image, and the value of each pixel represents the saliency degree of that pixel in the image. The coordinates of the focus area are determined by threshold processing or selecting the area with the highest saliency value.
[0027] The pixel density information and the focus area coordinates are used as input features. These features can be encoded into a vector or a part of the image and then input into the convolutional neural network, and the convolutional neural network outputs the resolution threshold range; among them, the convolutional neural network (CNN): is a deep learning model, especially suitable for image processing tasks. It extracts the features of the image and performs classification or regression through convolutional layers, pooling layers, and fully connected layers. Convolutional layer: used to extract the local features of the image. The convolutional layer applies filters to the image through a sliding window operation to generate a feature map. Pooling layer: used to reduce the dimension of the feature map, reduce the amount of calculation, and extract higher-level features. Common pooling operations include max pooling and average pooling. Fully connected layer: used to map the extracted features to the output space. The fully connected layer is usually used for classification or regression tasks and outputs the final result.
[0028] The resolution threshold range output by the convolutional neural network; wherein, the resolution threshold range is used to determine whether the resolution of the image needs to be adjusted. If the output resolution threshold range is within the preset standard value, it is considered that the resolution of the image is appropriate and the pixel density of the original image can be directly used for processing, where the preset standard value is a predefined threshold for determining whether the resolution of the image meets the requirements; otherwise, the resolution of the image needs to be adjusted. For example, when the pixel density of the original image exceeds the preset resolution threshold, resampling processing needs to be performed on it to reduce its resolution. The bicubic interpolation image scaling algorithm is adopted to determine the color value of each new pixel point by calculating the weighted average of 16 nearest neighbor pixel points (4x4 grid) around it. This method can better maintain the quality and details of the image and avoid obvious jagged or blurred phenomena. After the above processing, the finally obtained image is an "adaptive resolution" version, which means that its resolution has been adjusted to a level more suitable for the current application scenario.
[0029] Figure 3 The step flowchart of step S103 of an embodiment is shown. S104. When it is detected that the coordinates of the focus area are displaced, perform a fast Fourier transform on the adaptive resolution image to obtain frequency domain data with a specified window size, and perform high-pass filtering on the frequency domain data according to a preset frequency domain smoothing coefficient to obtain the filtered frequency domain data; In this embodiment, compare the current focus area coordinates with the focus area coordinates of the previous frame to determine whether a displacement occurs; if it is determined that a displacement occurs, perform a fast Fourier transform on the adaptive resolution image to convert the adaptive resolution image from the spatial domain to the frequency domain. In the frequency domain, select a window of a predetermined size to extract the frequency domain data near the current focus area coordinates to obtain the frequency domain data with a predetermined window size; construct a high-pass filter through a preset frequency domain smoothing coefficient, filter the frequency domain data to obtain the filtered frequency domain data, and convert the filtered frequency domain data back to the spatial domain to obtain the frequency domain focus image.
[0030] Specifically, compare the coordinates of the focus area in the current frame with those in the previous frame, and calculate the offset between the two (such as Euclidean distance or Manhattan distance). If the offset exceeds a preset threshold, it is determined that the focus area has shifted. If it is determined that the focus area has shifted, perform a fast Fourier transform (FFT) on the adaptive resolution image to convert the image from the spatial domain to the frequency domain, obtaining frequency domain data. The frequency domain data contains the frequency components of the image and can reflect the details and structural information of the image. In the frequency domain, with the current focus area coordinates as the center, select a window of a predetermined size and extract the frequency domain data within the window as the input for subsequent processing. According to a preset frequency domain smoothing coefficient, construct a high-pass filter. Filter the extracted frequency domain data to retain the high-frequency components (such as edges and details) while filtering out the low-frequency components (such as smooth areas), obtaining the filtered frequency domain data.
[0031] S105. Perform an inverse Fourier transform precision control process on the filtered frequency domain data to restore and generate an optimized focus detail image; In this embodiment, perform an inverse fast Fourier transform (IFFT) on the filtered frequency domain data to convert it back from the frequency domain to the spatial domain. Perform an inverse Fourier transform on the filtered frequency domain data to convert it back from the frequency domain to the spatial domain. Obtain a preliminary restored image that contains the detailed information processed by the high-pass filter. To further improve the quality of the restored image, the following precision control measures can be taken: truncate or quantize the data after the inverse Fourier transform to remove noise or irrelevant low-frequency components. Set a reasonable truncation threshold or quantization step size to retain important detailed information. Perform edge enhancement processing on the restored image to further highlight the details of the focus area. This can be achieved using edge detection algorithms (such as Canny or Sobel) or adaptive sharpening filters. While retaining the details, perform local smoothing processing on the image to remove high-frequency noise. Smoothing algorithms such as Gaussian filtering or bilateral filtering can be used to balance detail retention and noise suppression. Adjust the dynamic range of the pixel values of the image (such as histogram equalization or contrast stretching) to enhance the visual effect of the image. Ensure that the details of the focus area are more clearly visible. After the above precision control processing, generate the final optimized focus detail image. This image focuses on the detailed information of the focus area and has high clarity and visual effect. By performing an inverse Fourier transform on the filtered frequency domain data and combining precision control processing, the system can generate high-quality optimized focus detail images.
[0032] S106. Input the focus detail image and a pre-established low-frequency background template into a real-time rendering engine to generate final display frame data, and output the final display frame data to the display area of the virtual exhibition hall; In this embodiment, the focus detail image is fused with the low-frequency background template using an image overlay or blending algorithm (such as weighted average or Alpha blending) to ensure seamless connection between the focus detail and the background. The focus detail image usually serves as the foreground, and the low-frequency background template serves as the background. After fusion, a complete scene image is generated. The fused image is input into the real-time rendering engine. The rendering engine further optimizes the image, such as adjusting lighting, color correction, anti-aliasing, etc., to enhance the visual effect. According to the scene requirements of the virtual exhibition hall, the rendering engine may also add dynamic effects (such as light and shadow changes or interactive elements). The rendering engine outputs the final display frame data, which contains the complete scene information and highlights the details of the focus area. The resolution and format of the final display frame data need to be adapted to the display requirements of the output device. The final display frame data is transmitted to the display area of the virtual exhibition hall. The display area can be a display screen, a VR headset, or other display devices, used to present a high-quality virtual exhibition hall scene to users.
[0033] Through combining the focus detail image with the low-frequency background template and using the real-time rendering engine to generate the final display frame data in this embodiment, the system can provide users with a high-definition visual experience in the virtual exhibition hall. It not only retains the details of the focus area but also ensures the smoothness and consistency of the overall scene, and is applicable to application scenarios that require dynamic focus adjustment and high-quality rendering.
[0034] S107. Input the final display frame data into the user behavior analysis and prediction module to obtain the display content pre-loaded during network latency.
[0035] In this implementation, while transmitting the final display frame data to the display area of the virtual exhibition hall, the system can also input the final display frame data into the user behavior analysis and prediction module to realize content preloading during network delay. The specific process is as follows: First, the system needs to collect information such as the user's browsing behavior, interaction mode, etc. This includes but is not limited to the user's movement path, residence time, clicked or touched exhibits in the virtual exhibition hall. Then, the collected data is processed and analyzed by machine learning algorithms or other data analysis technologies to identify the user's points of interest and possible behavior patterns. For example, if most users will first view a specific exhibit after entering a certain exhibition hall, then this exhibit can be regarded as a high-probability point of interest. Based on the above analysis results, the system can predict which content or exhibits the user may visit next. This prediction can help the system load related content in advance, thereby reducing the waiting time caused by network delay. When the system predicts that the user may access certain content, it can download these contents to the local cache in advance. In this way, when the user actually accesses these contents, the system can read them directly from the cache without requesting data from the server again, thereby improving the user experience. To further optimize the effect of preloading, the system can dynamically adjust the preloaded content based on the user's real-time behavior. For example, if the user suddenly changes direction or point of interest, the system can quickly adjust the preloading strategy to ensure that the most relevant content is always provided to the user. Finally, the system can continuously optimize its prediction model and preloading strategy by continuously collecting user feedback and behavior data, forming a closed-loop optimization process. In this way, the virtual exhibition hall can provide a smooth user experience even when network delays occur, respond to users' personalized needs more intelligently, and improve the overall visit effect.
[0036] In the embodiment of the present application, an adaptive clarity enhancement algorithm is developed to automatically adjust the clarity of the displayed object according to the user's viewing distance and angle. When the user approaches an exhibit, the system can automatically improve the resolution and details of the area; when the user views the exhibit from different angles, the system can adjust the clarity and focus of the image in real time according to the change in viewing angle, ensuring that the user can get the best viewing experience at any position. It is also possible to predict the exhibits that the user may view next based on the user's browsing trajectory and interest preferences, and cache related content in advance, which can reduce waiting time even when there is a network delay. At the same time, through progressive image loading technology, users can get a basic visual experience while waiting for high-definition images.
[0037] The embodiment of the present application also provides an image processing system applied to a display area of a virtual exhibition hall, characterized in that the system includes: A data acquisition module, which acquires viewing distance data and viewing angle deviation data collected by an eye tracking sensor; A dynamic perception result generation module determines the head position state based on the viewing distance data and the viewing angle offset data, and generates a dynamic perception result; An adaptive resolution image generation module processes the original image pixel density and the focus area coordinates corresponding to the dynamic perception result using a convolutional neural network to determine a resolution threshold range; if the resolution threshold range exceeds a preset standard value, perform bicubic interpolation resampling on the original image pixel density to generate an adaptive resolution image; A frequency domain focus image generation module, when detecting a displacement of the focus area coordinates, performs a fast Fourier transform on the adaptive resolution image to obtain frequency domain data with a specified window size, and performs high-pass filtering on the frequency domain data according to a preset frequency domain smoothing coefficient to obtain a frequency domain focus image; A focus detail image generation module performs inverse Fourier transform precision control processing on the frequency domain focus image to restore and generate an optimized focus detail image; A display module inputs the focus detail image and a pre-established low-frequency background template into a real-time rendering engine to generate final display frame data, and outputs the final display frame data to the display area of the output virtual exhibition hall; A display content prediction module inputs the final display frame data into a user behavior analysis and prediction module to obtain the display content pre-loaded during network latency.
[0038] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0039] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0040] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more of the blocks and / or processes. Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0041] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0042] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0043] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0044] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0045] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0046] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An image processing method applied to a display area of a virtual exhibition hall, characterized in that: The method comprises the steps of: S101. Obtain viewing distance data and viewing angle deviation data collected by the eye tracking sensor; S102. Determine the head position state according to the viewing distance data and the viewing angle offset data, and generate a dynamic perception result; S103. Using a convolutional neural network to process the original image pixel density and focus area coordinates corresponding to the dynamic perception result to determine a resolution threshold range; If the resolution threshold range exceeds a preset standard value, performing bicubic interpolation resampling on the original image pixel density to generate an adaptive resolution image; S104. When it is detected that the focal area coordinates are displaced, a fast Fourier transform is performed on the adaptive resolution image to obtain frequency domain data of a specified window size, and a high-pass filtering is performed on the frequency domain data according to a preset frequency domain smoothing coefficient to obtain filtered frequency domain data; S105. Performing inverse Fourier transform precision control processing on the filtered frequency domain data to restore and generate an optimized focus detail image; S106. Inputting the focus detail image and the pre-established low-frequency background template into a real-time rendering engine to generate final display frame data, and outputting the final display frame data to the display area of the virtual exhibition hall; S107. Input the final display frame data into the user behavior analysis and prediction module to obtain the display content preloaded during network delay.
2. The method according to claim 1, characterized in that The step 102 includes: Collecting the viewing distance data and the viewing angle deviation data through a sensor, and preprocessing the collected data; Matching the preprocessed data with a predefined head position state to determine the user's real-time head state; Generate a dynamic perception result according to the real-time head state.
3. The method according to claim 1, characterized in that The step S103 includes: Acquiring pixel density information of an original image and determining focal region coordinates in the original image by using saliency analysis based on a dynamic perception result; Inputting the pixel density information and the focal area coordinates into a convolutional neural network, and outputting a resolution threshold range, wherein a convolutional layer is used to extract local features of the focal area, a pooling layer is used to reduce the dimension of the local features, and a fully connected layer is used to output the resolution threshold range; The resolution threshold range is compared with a preset standard value. If the resolution threshold range is within the preset standard value, the original image pixel density is directly used for processing. If the resolution threshold range exceeds the preset standard value, the original image is resampled by bicubic interpolation to generate an adaptive resolution image.
4. The method according to claim 1, characterized in that: The step S104 includes: Compare the current focus area coordinates with the focus area coordinates of the previous frame to determine whether displacement occurs; If it is determined that displacement occurs, a fast Fourier transform is performed on the adaptive resolution image to convert the adaptive resolution image from the spatial domain to the frequency domain, and in the frequency domain, a window of a predetermined size is selected to extract frequency domain data near the coordinates of the current focal area to obtain frequency domain data of a predetermined window size; A high-pass filter is constructed by using a preset frequency domain smoothing coefficient to filter the frequency domain data to obtain filtered frequency domain data, and the filtered frequency domain data is converted back to the spatial domain to obtain a frequency domain focal image.
5. The method according to claim 1, characterized in that The step S107 includes: Cleaning, denoising, and extracting features from the final display frame data to obtain the user's current behavior trajectory; Identify the user's typical behavior patterns by analyzing the user's historical behavior data; According to the user's historical operation frequency and historical stay time, the area of interest to the user is extracted, and the preloaded content is obtained based on the user's current behavior trajectory combined with the user's typical behavior pattern and the area of interest to the user. In combination with the current network status, the preloaded content is dynamically adjusted.
6. The method according to claim 5, characterized in that The dynamically adjusting the preloaded content in combination with the current network status includes: Based on historical network delay data, a delay prediction model is built to predict possible delays under current network conditions. Based on the predicted user behavior results, the content to be displayed is prioritized; combined with the predicted network latency, the preloaded content is dynamically adjusted.
7. An image processing system applied to a virtual exhibition hall display area, characterized in that: The system comprises: A data acquisition module, which acquires viewing distance data and viewing angle deviation data collected by an eye tracking sensor; A dynamic perception result generating module, which determines the head position state according to the viewing distance data and the viewing angle offset angle data, and generates a dynamic perception result; An adaptive resolution image generation module, which uses a convolutional neural network to process the original image pixel density and the focus area coordinates corresponding to the dynamic perception result to determine a resolution threshold range; if the resolution threshold range exceeds a preset standard value, bicubic interpolation resampling is performed on the original image pixel density to generate an adaptive resolution image; A frequency domain focus image generation module, when detecting that the focus area coordinates are displaced, performs fast Fourier transform on the adaptive resolution image, obtains frequency domain data of a specified window size, and performs high-pass filtering on the frequency domain data according to a preset frequency domain smoothing coefficient to obtain a frequency domain focus image; A focus detail image generation module, which performs inverse Fourier transform precision control processing on the frequency domain focus image to restore and generate an optimized focus detail image; A display module inputs the focus detail image and a pre-established low-frequency background template into a real-time rendering engine to generate final display frame data, and outputs the final display frame data to a display area of the virtual exhibition hall; The display content prediction module inputs the final display frame data into the user behavior analysis and prediction module to obtain the display content preloaded during network delay.
8. An electronic device comprising a memory, an image processor, and a computer program stored in the memory and executable on the image processor, wherein the image processor implements the image processing method for a virtual exhibition hall display area as described in any one of claims 1 to 6 when executing the program.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by an image processor, the image processing method applied to a virtual exhibition hall display area according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Virtual video optimization method and system, electronic equipment and storage medium
CN121000833A