Method, device, equipment and readable storage medium for displaying media data
By obtaining the gaze-affected area in the XR device and performing adaptive filtering, the problem of insufficient clarity of the virtual display screen is solved, the display effect of media data is improved, and resources are saved.
Patent Information
- Application Number
- CN202411310952.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-09-19
AI Technical Summary
Because the pixels of the virtual display screen are smaller than the resolution supported by the device, downsampling is used in XR devices. Although the media data can be fully presented, the clarity cannot be guaranteed, affecting the viewing experience.
By obtaining the viewing object's gaze-affected area on the virtual display, the media content is divided into the gazed and non-gaze media content, and different filtering methods are used to process them, thereby improving the clarity of the gazed area and reducing resource consumption in the non-gaze area.
While maintaining the overall clarity of media data, it optimizes the presentation of media data, reduces computing resource consumption, and improves the viewing experience.
Smart Images

Figure CN119211624B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, device, and readable storage medium for displaying media data. Background Art
[0002] With the development of science and technology, extended-range (XR) technology has become more mature, and XR devices have also emerged. XR devices are widely used in various scenarios; for example: gaming scenes, cinema scenes, etc. Among them, in XR devices, the display screen used to carry the image can be understood as the "display" in the XR device. The "display" can be called the virtual display screen or virtual display of the XR device. The virtual environment in which the virtual display screen is located can be a virtual cinema, gaming environment, etc.
[0003] With the continuous improvement of XR technology, the hardware specifications of XR devices have also been greatly improved. The ultra-high-resolution display capability (such as 4K) makes it the best viewing device for media data (such as video data). However, since the pixels of the virtual display screen in the virtual environment are much smaller than the resolution supported by the device (such as 4K), when using the device to view media data, a downsampling processing method will be used to enable the virtual display screen to display the media data. Although this downsampling processing method can fully present the media data, it cannot guarantee the clarity of the media data presented on the screen. Summary of the Invention
[0004] Embodiments of the present application provide a method, apparatus, device, and readable storage medium for displaying media data, which can perform adaptive filtering on media data to improve the clarity of the media data when displaying media data using an extended reality device.
[0005] In one aspect, an embodiment of the present application provides a method for displaying media data. The method is applied to an extended reality device, wherein the extended reality device includes a virtual display, and the virtual display is used to display the media data. The method includes:
[0006] In response to a display request for a target media frame in the media data, obtaining a gaze influence area of a viewing object of the media data on a virtual display;
[0007] Dividing the media content of the target media frame according to the gaze influence area to obtain gaze media content and non-gaze media content; the gaze media content is the media content that needs to be displayed in the gaze influence area, and the non-gaze media content is the media content that needs to be displayed outside the gaze influence area;
[0008] Different filtering methods are used to filter the attention media content and the non-attention media content, and the filtered target media frames are displayed on a virtual display.
[0009] On the one hand, an embodiment of the present application provides a device for displaying media data, the device comprising:
[0010] a region acquisition module configured to, in response to a display request for a target media frame in the media data, acquire a gaze-affected region of a virtual display of the media data by a viewing subject; the virtual display being configured to display the media data; and the virtual display being provided on the extended reality device.
[0011] a content segmentation module for segmenting the media content of the target media frame according to the gaze influence area to obtain gaze media content and non-gaze media content; the gaze media content is the media content to be displayed in the gaze influence area, and the non-gaze media content is the media content to be displayed outside the gaze influence area;
[0012] The filtering module is used to filter the attention media content and the non-attention media content using different filtering methods, and display the filtered target media frames on the virtual display.
[0013] In one embodiment, a specific implementation of the region acquisition module acquiring the gaze influence region of the viewing subject of the media data on the virtual display in response to a display request for a target media frame in the media data includes:
[0014] obtaining a positional relationship between the virtual display and a viewing object of the media data;
[0015] Get the display parameters of the target media frame;
[0016] Based on the positional relationship and the display parameters of the target media frame, a gaze influence area of the viewing object on the virtual display is determined.
[0017] In one embodiment, a specific implementation of the region acquisition module acquiring the positional relationship between the virtual display and the viewing object of the media data includes:
[0018] Acquiring a virtual screen displayed by an extended reality device to a viewing object of the media data; the virtual display is in the virtual screen;
[0019] Acquire a first area of the virtual display and a second area of the virtual screen;
[0020] determining a target area ratio between the first area and the second area;
[0021] The positional relationship between the virtual display and the viewing object is determined according to the target area ratio.
[0022] In one embodiment, a specific implementation method of the region acquisition module determining the positional relationship between the virtual display and the viewing object according to the target area ratio includes:
[0023] Obtaining a standard area ratio between the virtual display and the virtual screen, and a standard virtual distance corresponding to the standard area ratio;
[0024] determining a ratio between the standard area ratio and the target area ratio;
[0025] The ratio between the standard area ratio and the target area ratio is calculated and processed with the standard virtual distance to obtain the target virtual distance between the virtual display and the viewing object;
[0026] The positional relationship between the virtual display and the viewing object is determined according to a target virtual distance between the virtual display and the viewing object.
[0027] In one embodiment, the positional relationship includes a target virtual distance between the virtual display and the viewing object;
[0028] The specific implementation method of the region acquisition module for determining the gaze influence region of the viewing object on the virtual display based on the position relationship and the display parameters of the target media frame includes:
[0029] Constructing a range mapping table based on the display parameters of the target media frame; the range mapping table includes a mapping relationship between a configuration distance set and a configuration focus influence parameter set, and a mapping relationship exists between a configuration distance in the configuration distance set and a configuration focus influence parameter in the configuration focus influence parameter set;
[0030] In the range mapping table, a target focus influence parameter that has a mapping relationship with the target virtual distance is obtained;
[0031] The gaze influence area of the viewing object on the virtual display is determined according to the target focus influence parameter.
[0032] In one embodiment, a specific implementation of the region acquisition module constructing a range mapping table based on display parameters of the target media frame includes:
[0033] According to the display parameters of the target media frame, a near focus influence parameter and a far focus influence parameter are set; the near focus influence parameter refers to a parameter of the viewing object's gaze influence on the virtual display when the area of the virtual display is equal to the area of the virtual screen and the virtual distance between the viewing object and the virtual display is a first distance threshold, and the virtual screen refers to the screen displayed to the viewing object by the extended reality device; the far focus influence parameter refers to a parameter of the viewing object's gaze influence on the virtual display when the area of the virtual display is equal to the area of the virtual screen and the virtual distance between the viewing object and the virtual display is a second distance threshold;
[0034] A range mapping table is constructed according to the first distance threshold, the second distance threshold, the near focus influence parameter, and the far focus influence parameter.
[0035] In one embodiment, a specific implementation of the region acquisition module constructing a range mapping table according to the first distance threshold, the second distance threshold, the near focus influence parameter, and the far focus influence parameter includes:
[0036] selecting one or more candidate distances among distances between a first distance threshold and a second distance threshold;
[0037] Among the influencing parameters between the near focus influencing parameters and the far focus influencing parameters, a corresponding influencing parameter is configured for each candidate distance to obtain a focus influencing parameter corresponding to each candidate distance;
[0038] Determine the first distance threshold, the second distance threshold, and each candidate distance as a configuration distance, and determine the near focus influence parameter, the far focus influence parameter, and the focus influence parameter corresponding to each candidate distance as a configuration focus influence parameter;
[0039] A mapping relationship is established between each configuration distance and the configuration focus influence parameter corresponding to each configuration distance to obtain a range mapping table.
[0040] In one embodiment, the target focus impact parameters include a width parameter and a height parameter;
[0041] The specific implementation method of the region acquisition module determining the gaze influence region of the viewing object on the virtual display according to the target focus influence parameter includes:
[0042] Acquire a mask image of the virtual display; the mask image completely overlaps with the virtual display in the virtual screen;
[0043] Obtaining the gaze focus of the viewing object on the virtual display, and determining the focal position coordinates of the gaze focus in the mask image;
[0044] A rectangular area is determined in the mask image with the position point indicated by the focus position coordinates as the center point, the width parameter in the target focus influence parameter as the rectangle width, and the height parameter in the target focus influence parameter as the rectangle height;
[0045] An area completely covered by the rectangular area in the virtual display is determined as a gaze influence area of the viewing object on the virtual display.
[0046] In one embodiment, the specific implementation of the filtering module using different filtering methods to filter the attention media content and the non-attention media content includes:
[0047] Using a first filtering method to filter the attention media content;
[0048] The non-attention media content is filtered using the second filtering method; the filtering effect indicated by the first filtering method is better than the filtering effect indicated by the second filtering method.
[0049] In one embodiment, a specific implementation of the filtering module using the first filtering method to filter the attention media content includes:
[0050] Calling the image enhancement model according to the first filtering method;
[0051] Perform content enhancement on gaze-based media content through image enhancement models.
[0052] In one embodiment, a specific implementation of the filtering module using the second filtering method to filter the non-attention media content includes:
[0053] Calling the smoothing filter model according to the second filtering mode;
[0054] The non-attention media content is smoothed and filtered using a smoothing filter model.
[0055] In one aspect, an embodiment of the present application provides a computer device, including: a processor and a memory;
[0056] The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method in the embodiment of the present application.
[0057] On one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the method in the embodiment of the present application is executed.
[0058] In one aspect of the present application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the method provided in one aspect of the embodiments of the present application.
[0059] In an embodiment of the present application, a display processing scheme for media data is provided, which can dynamically filter the target media frame in the media data based on the gaze influence area of the viewing object on the virtual display in the extended reality device, so as to optimize the effect of the media content displayed in different areas of the virtual display and improve the clarity of the media frame. Specifically, for a certain media frame in the media data (which can be called the target media frame), after receiving its display request, the extended reality device can obtain the gaze influence area of the viewing object on the virtual display in the device; then, according to the gaze influence area, the media content of the target media frame can be divided to obtain gaze media content and non-gaze media content, wherein the gaze media content is the media content that needs to be displayed in the gaze influence area, and the non-gaze media content is the media content that needs to be displayed outside the gaze influence area. Furthermore, different filtering methods can be used to filter the gaze media content and the non-gaze media content, and the filtered target media frame is displayed in the virtual display. It should be understood that for any media frame, the area of gaze influence of the viewing object on the virtual display can be determined. This area can be understood as the area where the viewing object has a higher degree of visual attention. Then, the media content of the target media frame can be divided according to the gaze influence area, and the media content with a higher degree of visual attention of the viewing object can be obtained. These contents are the gaze media content, and the media content with a lower degree of visual attention of the viewing object can also be obtained. These contents are the non-gaze media content. For gaze media content and non-gaze media content, the present application can select different filtering methods according to business needs to perform adaptive filtering processing on them. For example, the gaze media content can be filtered according to a filtering method that can improve the quality to ensure its clarity. Compared with a one-size-fits-all downsampling processing method, the present application can perform adaptive filtering processing based on the gaze media content and the non-gaze media content, which can not only fully present the media data, but also well maintain the overall clarity of the media data. It can be seen that the present application can perform adaptive filtering processing on media data in the business of displaying media data using an extended reality device to improve the clarity of the media data. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0061] Figure 1 This is a schematic diagram of the architecture of a solution system provided by an exemplary embodiment of the present application;
[0062] Figure 2 This is a schematic diagram of a scenario for displaying media data provided by an embodiment of the present application;
[0063] Figure 3 is a flowchart of a method for displaying media data provided by an exemplary embodiment of the present application;
[0064] Figure 4 This is a schematic diagram of a change in the gaze influence range provided by an embodiment of the present application;
[0065] Figure 5 This is a schematic diagram of a process for determining a gaze influence area provided by an embodiment of the present application;
[0066] Figure 6 is a schematic diagram of determining a gaze influence area provided by an embodiment of the present application;
[0067] Figure 7 This is a schematic diagram of a system logic architecture provided by an embodiment of the present application;
[0068] Figure 8 This is a schematic structural diagram of a media data display device provided in an embodiment of the present application;
[0069] Figure 9 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0070] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0071] The embodiments of the present application involve technologies such as extended reality. For ease of understanding, some of the terms involved in the embodiments of the present application will be explained below.
[0072] 1. Augmented Reality (AR)
[0073] Augmented reality is a computer technology that applies virtual information to the real world, superimposing the real environment and virtual objects on the same screen or space, allowing them to coexist in real time. Things are real, but they have been enhanced. Essentially, augmented reality is a new interface technology that integrates positioning, presentation, and interaction software and hardware technologies. Its purpose is to allow users (e.g., users) to sense the spatiotemporal connection and integration of virtual and real space, thereby enhancing their perception and cognition of the real environment.
[0074] 2. Virtual Reality (VR)
[0075] Virtual reality uses computers and other devices to simulate a realistic three-dimensional world. It provides users with 3D vision, hearing, touch, and other sensory experiences, creating an immersive experience in which users can observe objects in the three-dimensional space without any limitations. In other words, everything they see is virtual and non-existent. It encompasses computer, electronic information, and simulation technologies. Virtual reality is primarily implemented using computer technology, leveraging and integrating the latest advances in high-tech technologies such as 3D graphics, multimedia, simulation, display, and servo technology.
[0076] 3. Mixed Reality (MR)
[0077] Mixed reality is a new visual environment created by merging the real and virtual worlds. It is a further development of virtual reality technology. By presenting virtual scene information in real scenes, this technology establishes an interactive feedback loop between the real world, the virtual world, and the user, thereby enhancing the user experience. In this new visual environment, physical and digital objects coexist and interact in real time. It is a combination of real world + virtual world + digital information.
[0078] 4. Extended-Range (XR)
[0079] Extended reality refers to the use of computers to combine the real and virtual, creating a digital environment that combines the real and virtual, as well as a new mode of human-computer interaction. This can provide users with an immersive experience of seamless transitions between the virtual and real worlds. Unlike traditional 2D interfaces such as computer screens and mobile phones, extended reality allows users to enter a different world and create a new reality by fusing the physical and digital worlds. This is achieved through the use of various devices (such as handheld devices) and sensory feedback systems. Extended reality includes the aforementioned virtual reality, augmented reality, and mixed reality. It is a general term for virtual reality, augmented reality, and mixed reality. In general, these technologies change the user's perception of reality, allowing them to interact with digital objects in a natural way.
[0080] The following is a brief introduction to the design concept of the embodiments of this application.
[0081] In actual applications, XR devices are widely used in various scenarios, such as gaming and cinema scenarios. In XR devices, the display screen used to carry images can be understood as the "display" in the XR device. The "display" can be referred to as the virtual display screen or virtual display of the XR device (referred to as a virtual display in this application). The virtual environment in which the virtual display screen is located can be a virtual cinema, gaming environment, etc.
[0082] With the continuous improvement of XR technology, the hardware specifications of XR devices have also been greatly improved. The ultra-high-resolution (such as 4K) display capability makes it the best viewing device for media data (such as video data). However, since the pixels of the virtual display screen in the virtual environment are much smaller than the resolution supported by the device (such as 4K), when using the device to view media data, a downsampling processing method will be used to enable the virtual display screen to fully display the media data. Although this downsampling processing method can fully present the media data, it will cause the resolution of the media content outside the viewpoint area of the viewing object to be reduced. The lower resolution will cause the appearance of jagged edges, which will greatly affect the overall presentation effect of the media data and also affect the viewing experience of the viewing object.
[0083] Based on this, in order to optimize the effect of media data displayed by the virtual display in the XR device and improve the clarity of the media data, the present application provides a display scheme for media data. Specifically, the display scheme for media data provided by the present application can be applied to any extended reality device (for example, AR device, VR device, etc.), and the extended reality device includes a virtual display for displaying media data, and the virtual display can be used to display any media data (for example, video data). For any media data, before displaying a certain media frame of the media data (for example, a video frame. The media frame can be referred to as a target media frame), the viewing object of the media data (that is, the object viewing the media data, for example, a user) can first obtain the gaze influence area of the virtual display (this area can be understood as the area on the virtual display where the viewing object's line of sight has a higher degree of attention); then, the media content of the target media frame can be divided according to the gaze influence area. For example, the media content that needs to be displayed in the gaze influence area can be divided into gaze media content, which the viewing object will focus on. The media content that needs to be displayed in the area outside the gaze influence area can be divided into non-gaze media content, which is not within the viewing object's line of sight.
[0084] Furthermore, for the attention media content and the non-attention media content, the present application can use different filtering methods to filter them. For example, a higher-quality filtering method (the high-quality filtering method here can refer to a filtering method with filtering performance higher than a performance threshold) can be used to filter the attention media content. In this way, the clarity of the attention media content can be improved. At the same time, a lower-quality filtering method (the low-quality filtering method here can refer to a filtering method with filtering performance lower than a performance threshold) can be used to filter the non-attention media content. In this way, while maintaining a certain clarity, filtering resources can be reduced. Therefore, the media data display method that adaptively and dynamically adjusts the filtering method based on the attention media content and the non-attention media content can use fewer filtering resources, so that the attention media content has a higher clarity, while allowing the non-attention media content to maintain a certain clarity, thereby optimizing the overall presentation effect of the media data. Finally, the filtered target media frame can be displayed on the virtual display.
[0085] It should be understood that, for a target media frame of media data to be displayed, the present application can divide the media content of the target media frame according to the gaze influence area after obtaining the viewing object's gaze influence area on the virtual display, so as to obtain the gaze media content and the non-gaze media content. In this way, the content with higher and lower visual attention of the viewing object on the target media frame can be accurately found; thereafter, different filtering methods are used to perform different filtering processing on it, so that the clarity of the gaze media content and the non-gaze media content can be adaptively adjusted based on actual business needs. For example, a higher-quality filtering method can be used to filter the gaze media content, so that the clarity of the gaze media content can be improved. At the same time, a lower-quality filtering method can be used to filter the non-gaze media content, so that while maintaining a certain clarity, filtering resources can be reduced, thereby reducing resource consumption. In general, the solution provided by the present application can improve the clarity of media data and optimize the overall presentation effect of media data.
[0086] The solution provided in the embodiments of the present application can be applied to any application scenario in which an extended reality device is used to view media data, including but not limited to cinema scenarios, game scenarios, etc.
[0087] The theater scene may refer to a scene in which a virtual environment such as a virtual theater is presented to the user, a virtual display is displayed in the virtual theater, and a movie is played on the virtual display.
[0088] A game scene may refer to a scene in which a user uses an extended reality device to play an immersive game.
[0089] To sum up, the solution provided in the embodiment of the present application can improve the clarity of media data and optimize the overall presentation effect of media data in various application scenarios using extended reality devices. At the same time, it can reduce the resource consumption caused by computing, and to a certain extent, it can effectively improve business coverage (such as expanding applicable scenarios).
[0090] It should be noted that the several application scenarios given above are only examples and do not limit the application scenarios to which the solutions provided in the embodiments of the present application are applicable.
[0091] To facilitate understanding of the solution provided in the embodiment of this application, Figure 1 The system architecture diagram shown in the figure introduces the application scenarios involved in the embodiments of the present application; Figure 1 This is a schematic diagram of the architecture of a solution system provided by an exemplary embodiment of the present application. Figure 1 As shown, the system includes a server 1000 and an extended reality device cluster. The extended reality device cluster may include one or more extended reality devices. The number of extended reality devices is not limited here. Figure 1 As shown, the plurality of extended reality devices may include an extended reality device 100a, an extended reality device 100b, an extended reality device 100c, ..., an extended reality device 100n; Figure 1 As shown, the extended reality device 100a, the extended reality device 100b, the extended reality device 100c, ..., the extended reality device 100n can each be connected to the server 1000 through a network connection, so that each extended reality device can exchange data with the server 1000 through the network connection (for example, the extended reality device 100c can exchange data with the server 1000 through the network connection). To facilitate understanding of the functions implemented by the extended reality device or server, the extended reality device or server will be described below:
[0092] 1) Extended reality device: a device that combines the real and the virtual to create a virtual environment for human-computer interaction. Users can experience the "immersive feeling" of seamless transition between the virtual world and the real world by wearing an extended reality device. The extended reality device may include a virtual display, which can be used to carry media images, and the environment in which the virtual display is located can be understood as a virtual environment, which can be a virtual cinema, a game scene, etc. Depending on the application scenarios and fields applied by this solution, the extended reality devices provided in the embodiments of the present application are different. Extended reality devices may include but are not limited to: head-mounted display devices (for example: smart glasses), smart watches and other smart devices with multimedia data processing functions (for example, video data playback function, music data playback function, text data playback function), but are not limited to this. Extended reality devices can also be other smart devices. The embodiments of the present application do not limit the type of extended reality devices, which are explained here.
[0093] 2) The server can be a background server corresponding to the extended reality device, which is used to interact with the extended reality device for data. In this way, the data interaction server can provide computing and application service support for the extended reality device. Specifically, the server is a background server corresponding to the application deployed in the extended reality device, which is used to interact with the extended reality device to provide computing and application servers for the application. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0094] The extended reality device and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in this application. In addition, the embodiments of this application do not limit the number of extended reality devices and servers.
[0095] Based on the solution and system architecture described above, the following points need to be explained:
[0096] ① The above-mentioned embodiments of this application Figure 1The system shown is for the purpose of more clearly illustrating the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. It is known to those skilled in the art that with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems. For example, in an application scenario, the execution subject of the embodiments of the present application may include an extended reality device and a server, that is, the extended reality device and the server jointly execute the solution provided by the embodiments of the present application; it should be understood that in actual applications, the execution subject may also be an extended reality device or server, that is, support for the extended reality device or server to execute the solution provided by the embodiments of the present application alone. In general, the execution subject in the present application may be a computer device, and the computer device may refer to an extended reality device and a server, and the execution subject may also be the extended reality device or server.
[0097] ② The collection and processing of relevant data in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations. The acquisition of personal information must be subject to the knowledge or consent of the individual subject (or the legal basis for obtaining the information), and the subsequent use and processing of data shall be carried out within the scope of authorization of laws and regulations and the subject of personal information. For example, when the embodiments of this application are applied to specific products or technologies, such as obtaining relevant information of users (such as
[0098] : user's current location information, historical location information, etc.) or data, the user's permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of the relevant region.
[0099] The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. For ease of understanding, please refer to Figure 2 , Figure 2 1 is a schematic diagram of a scenario for displaying media data provided by an embodiment of the present application, wherein the scenario is described by taking the media data as video data as an example.
[0100] like Figure 2As shown, the viewing subject can wear an extended reality device 100 to view media data, and the extended reality device 100 can display the media data through a virtual display (that is, the media data viewed by the viewing subject through the extended reality device 100 is presented on the virtual display). For any media frame of the media data, after receiving its display request, the extended reality device can obtain the viewing subject's gaze influence area on the virtual display of the media data. The gaze influence area is the area affected by the viewing subject's line of sight, which can be understood as the area that the viewing subject will focus on (or the area where the viewing subject's line of sight has a high degree of attention). This gaze influence area can be determined based on the positional relationship between the viewing subject and the virtual display. The process of determining the gaze influence area can be described in the subsequent related embodiments.
[0101] like Figure 2 As shown, after determining the gaze-affected area, the media content that will be displayed within the gaze-affected area when the media frame is displayed can be determined. This content can be classified as gaze-affected media content, which the viewer will focus on. Media content that will not be displayed within the gaze-affected area can be classified as non-gaze-affected media content. Furthermore, after determining the gaze-affected and non-gaze-affected media content, different filtering methods can be used to filter the gaze-affected and non-gaze-affected media content. This allows for dynamic allocation of filtering resources between the gaze-affected and non-gaze-affected media content, achieving adaptive adjustment and optimization of clarity.
[0102] It should be understood that for any media frame, the area of gaze influence of the viewing object on the virtual display can be determined. This area can be understood as the area where the viewing object's line of sight has a higher degree of attention (that is, the area where the line of sight is focused / concentrated). Then, the media content of the target media frame is divided according to the gaze influence area, and the media content where the viewing object has a higher degree of attention can be obtained. These contents are the attention media content, and the media content where the viewing object has a lower degree of attention can also be obtained. These contents are the non-attention media content. For the attention media content and the non-attention media content, the present application can select different filtering methods according to business needs to perform adaptive filtering processing on them. For example, the attention media content can be filtered according to a high-quality filtering method to ensure its clarity, and the non-attention media content can be filtered according to a smoothing filtering method to reduce its aliasing problem. Compared with a one-size-fits-all downsampling processing method, the present application uses the adaptive filtering processing method based on the attention media content and the non-attention media content, which can not only fully present the media data, but also well maintain the overall clarity of the media data and reduce the aliasing problem. Moreover, the present application determines the gaze influence area of the viewing object based on the real-time positional relationship between the current virtual display and the viewing object when displaying each media frame. That is, the gaze influence area is determined based on the real-time gaze position of the viewing object when it looks at the virtual display, and is not fixed. In this way, different gaze influence areas can be determined efficiently and accurately as the viewing object moves.
[0103] Based on the above-described solutions and application scenarios, the embodiments of the present application propose a more detailed method for displaying media data. The method for displaying media data proposed in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0104] See Figure 3 , Figure 3 : This is a flow chart of a method for displaying media data provided by an exemplary embodiment of the present application. The flow may refer to the flow of a display processing solution for media data provided by an embodiment of the present application. The display method may be executed by a computer device in the aforementioned system, such as an extended reality device and / or a server. Here, the method is described as being executed by an extended reality device. In the case where the extended reality device includes a virtual display (for displaying media data), the display method may include at least the following steps S301-S303:
[0105] Step S301 : in response to a display request for a target media frame in media data, obtaining a gaze influence area of a viewing object of the media data on a virtual display.
[0106] In this application, media data may refer to media content in any media presentation form, wherein the media presentation form may include but is not limited to image form, video form, text form, audio form, etc., and the media presentation form may also refer to a combination of any two or more of the above forms. In other words, the media data in this application may refer to video data, image data, text data, and audio data, etc. Since video data may contain text data, image data, audio data, etc., this application may preferably use video data as the media data, and will subsequently explain the media data as video data as an example. It should be understood that video data is composed of multiple continuous video frames, so the media frame here may refer to a frame that constitutes the media data, and the target media frame may refer to any media frame.
[0107] A viewing object can refer to any object viewing media data, and an object can be a user, an intelligent robot, or the like. In other words, the viewing object can refer to any user, any intelligent robot, or the like. When a viewing object desires to view media data, they can wear an extended reality device and initiate a playback request for the media data on the extended reality device. The extended reality device then plays and displays each media frame of the media data sequentially on a virtual display, following the order in which the media frames are arranged.
[0108] Among them, it should be noted that the extended reality device can refer to a wearable device that integrates different components. The components included in the extended reality device include: head-mounted display, camera, sensor, computer processor, etc. The camera and sensor can capture and collect the surrounding environment information of the viewed object, and the computer processor can calculate and process the collected surrounding environment information to generate virtual reality content and display it on the head-mounted display.
[0109] It should be understood that the present application may refer to the display included in the extended reality device as a virtual display, and the virtual display is mainly used to display media data. That is, the extended reality device needs to display each media frame of the media data on the virtual display. The virtual display can be placed in the virtual environment presented by the extended reality device. The present application may determine the viewing influence area of the viewing object on the virtual display based on the placement position of the virtual display and the position where the viewing object looks at the virtual display. This area is the area affected by the viewing object's line of sight, that is, the area where the viewing object's line of sight will focus. This gaze influence area is a partial area in the virtual display, for example, Figure 2 As shown, the gaze influence area is a part of the area in the virtual display, and the gaze influence area changes as the viewing object's line of sight changes. For the specific method of determining the gaze influence area, please refer to the subsequent Figure 5 Related description in the corresponding embodiment.
[0110] Step S302 : dividing the media content of the target media frame according to the gaze influence area to obtain gaze media content and non-gaze media content; the gaze media content is the media content that needs to be displayed in the gaze influence area, and the non-gaze media content is the media content that needs to be displayed outside the gaze influence area.
[0111] In the present application, after determining the gaze influence area of the viewing object on the virtual display, the media content of the target media frame can be divided according to the gaze influence area. For example, the media content displayed in the gaze influence area can be divided into gaze media content. Since the viewing object will focus on the content presented in the gaze influence area, this part of the media content is the content that the viewing object will focus on, so it can be called gaze media content; at the same time, the media content displayed in the area outside the gaze influence area (that is, the area in the virtual display except the gaze influence area, this part of the area can be called the non-gaze influence area) can be divided into non-gaze media content. Since the viewing object will focus on the content presented in the gaze influence area, and the attention to the content presented in the area outside the gaze influence area is not high, it can be called non-gaze media content.
[0112] Step S303 : filtering the attention media content and the non-attention media content using different filtering methods, and displaying the filtered target media frames on the virtual display.
[0113] In the present application, after determining the attention media content and non-attention media content of the target media frame, different filtering methods can be used to filter the attention media content and the non-attention media content, and the filtered target media frame can be displayed on a virtual display. Among them, since the viewing object's line of sight is concentrated on the attention-affected area and the attention to the attention media content is high, the clarity of the attention media content should be higher and the presentation effect should be better. Based on this, the present application can use a high-quality filtering method (a filtering method with filtering performance higher than a performance threshold, which consumes more filtering resources) to filter it; and since the viewing object's line of sight is not concentrated on the non-attention-affected area and the attention to the non-attention media content is low, the clarity of the non-attention media content does not need to be very high. Based on this, the present application can use a lower-quality filtering method (a filtering method with filtering performance lower than a performance threshold, which consumes less filtering resources) to filter it. In this way, while maintaining the clarity of the media content with higher attention, the filtering resource consumption of the media content in the non-attention-affected area can be reduced, thereby saving filtering costs.
[0114] In other words, the specific implementation process of filtering the attended media content and the non-attended media content using different filtering methods may include at least, but is not limited to: using a first filtering method to filter the attended media content; and using a second filtering method to filter the non-attended media content; wherein the filtering effect indicated by the first filtering method is superior to the filtering effect indicated by the second filtering method. The first filtering method may be a high-quality filtering method that requires more filtering resources to obtain a higher-definition display, while the second filtering method may be a lower-quality filtering method that does not require as many filtering resources.
[0115] Among them, it is worth noting that, since the gaze media content in the gaze influence area needs to allocate more filtering resources to obtain a higher-definition display, the present application can adopt the following strategies for the filtering strategy of the gaze influence area: 1. High-performance strategy; 2. High-quality strategy. Among them, the high-performance strategy can refer to a strategy of not performing any filtering processing to obtain a faster filtering speed; and the high-quality strategy can refer to a strategy of enhancing the media frame to improve the details of its various media contents. The present application may preferably use the high-quality strategy as the filtering strategy for the gaze influence area in the present application. When the filtering strategy is the high-quality strategy, the specific implementation process of filtering the gaze media content using the first filtering method may at least include but is not limited to: first, the image enhancement model can be called according to the first filtering method, and then the gaze media content can be enhanced by the image enhancement model. Among them, the image enhancement model here can refer to an edge enhancement model, and the edge enhancement model can include an edge enhancement operator. The edge enhancement operator can be used to perform edge enhancement processing on the gaze media content to enhance the content details and obtain higher content quality; the image enhancement model can also be an enhancement model based on deep learning.
[0116] It is also worth noting that, based on the above, conventional technologies downsample to accommodate the pixels of a virtual display. However, this can lead to aliasing, significantly impacting the visual experience of the media data. To reduce the aliasing problem in the non-attention-affected area, the present application can filter the non-attention media content using a smoothing filter, which can reduce high-frequency information and thus reduce aliasing. That is, the specific implementation process of filtering the non-attention media content using the second filtering method can include, but is not limited to, at least: first, invoking a smoothing filter model according to the second filtering method; then, smoothing filtering the non-attention media content using the smoothing filter model. Smoothing filtering can include filtering methods such as mean filtering, median filtering, and Gaussian filtering. When the smoothing filter model performs smoothing filtering, a smoothing filter method of corresponding size can be selected based on actual business needs. For example, a convolution kernel of a certain size can be used to perform mean filtering on the non-attention media content. The size of the convolution kernel is, for example, [N*N], where its value is 1 / (N*N). For the filtering processing of the attention media content and the non-attention media content, the present application can use a GPU to perform the filtering in parallel.
[0117] In an embodiment of the present application, a display processing scheme for media data is provided, which can dynamically and adaptively filter the target media frame in the media data based on the gaze influence area of the viewing object on the virtual display in the extended reality device, so as to optimize the effect of the media content displayed in different areas of the virtual display and improve the clarity of the media frame. Specifically, for a certain media frame in the media data (which can be referred to as the target media frame), after receiving its display request, the extended reality device can obtain the gaze influence area of the viewing object on the virtual display in the device; then, according to the gaze influence area, the media content of the target media frame can be divided to obtain gaze media content and non-gaze media content, wherein the gaze media content is the media content that needs to be displayed in the gaze influence area, and the non-gaze media content is the media content that needs to be displayed outside the gaze influence area. Furthermore, different filtering methods can be used to filter the gaze media content and the non-gaze media content, and the filtered target media frame is displayed in the virtual display. It should be understood that for any media frame, the area of gaze influence of the viewing object on the virtual display can be determined. This area can be understood as the area where the viewing object's line of sight has a higher degree of attention (that is, the area where the line of sight is focused / concentrated). Then, the media content of the target media frame is divided according to the gaze influence area, and the media content where the viewing object has a higher degree of attention can be obtained. These contents are the attention media content, and the media content where the viewing object has a lower degree of attention can also be obtained. These contents are the non-attention media content. For the attention media content and the non-attention media content, the present application can select different filtering methods according to business needs to perform adaptive filtering processing on them. For example, the attention media content can be filtered according to a high-quality filtering method to ensure its clarity, and the non-attention media content can be filtered according to a smoothing filtering method to reduce its aliasing problem. Compared with a one-size-fits-all downsampling processing method, the present application uses the adaptive filtering processing method based on the attention media content and the non-attention media content, which can not only fully present the media data, but also well maintain the overall clarity of the media data and reduce the aliasing problem. It can be seen that the present application can perform adaptive filtering on media data in the business of displaying media data using an extended reality device, so as to improve the clarity of the media data and optimize the overall presentation effect of the media.
[0118] Based on the above, the filtering selection of the media frame provided in this application is dynamically selected based on the viewing object's gaze-affected area on the virtual display. For the gaze media content that needs to be displayed in the gaze-affected area, a filtering method with better performance can be selected to filter it, allocating more filtering resources to it. For the non-gaze media content that needs to be displayed in the non-gaze-affected area, a smoothing filtering method can be selected to filter it, without allocating more filtering resources, and at the same time reducing the aliasing problem. It can be seen that the viewing object's gaze-affected area on the virtual display is crucial, and see Figure 4 A schematic diagram of the change of the gaze influence range is provided, such as Figure 4 As shown, assuming that virtual displays of the same size are placed at the near plane and the far plane respectively (the distance between the near plane and the viewing object will be smaller than the distance between the far plane and the viewing object), when the virtual display is located at the near plane, the viewing object's gaze influence range is gaze influence range 1, and when the virtual display is located at the far plane, the viewing object's gaze influence range is gaze influence range 2, and the gaze influence range will be much larger than the gaze influence range. It can be seen that as the virtual display is placed farther away, the viewing object's gaze influence range will become larger and larger, and the gaze influence range is also the viewing object's gaze influence area on the virtual display, that is, the viewing object's gaze influence area will change with the placement position of the virtual display and the viewing object's gaze position on the virtual display. Based on this, the present application provides a method for determining the viewing object's gaze influence area on the virtual display based on the positional relationship between the virtual display and the viewing object. The following describes the process of determining the viewing object's gaze influence area on the virtual display in conjunction with the accompanying drawings. Please refer to them. Figure 5 , Figure 5 This is a schematic diagram of a process for determining a gaze influence area provided by an embodiment of the present application, wherein the process may correspond to the above Figure 3 The corresponding embodiment is about the implementation process of obtaining the gaze influence area of the virtual display by the viewing object of the media data. Figure 5 As shown, the process may at least include the following steps S501 to S503:
[0119] Step S501: Acquire a positional relationship between a virtual display and a viewing object of media data.
[0120] Specifically, while viewing media data using an extended reality device, the viewing object may move, and the placement of the virtual display may also change. Therefore, for each media frame to be displayed, after receiving its display request, the positional relationship between the current virtual display and the viewing object can be obtained. Similarly, for the target media frame, after receiving its display request, the positional relationship between the virtual display and the viewing object can be obtained.
[0121] Specific implementations for obtaining the positional relationship between the virtual display and the viewing object of the media data may include, but are not limited to, the following: first, a virtual screen displayed by the extended reality device to the viewing object of the media data may be obtained. This virtual screen may refer to the complete screen seen by the viewing object through the extended reality device, which can be understood as the virtual environment screen presented by the extended reality device, and the virtual display is located in the virtual screen. For example, a virtual theater can be understood as a complete virtual screen, and the virtual screen used to play and display movies in the virtual theater is the virtual display; then, the area of the current virtual display (which may be referred to as the first area) and the area of the virtual screen (which may be referred to as the second area) may be obtained; and an area ratio between the first area and the second area (which may be referred to as a target area ratio) may be determined. For example, if the first area is 50 and the second area is 100, then the target area ratio may be 50 / 100=1 / 2. That is, the area ratio between the first area and the second area may be obtained using the first area as the numerator and the second area as the denominator. The positional relationship between the virtual display and the viewing object may be determined based on the target area ratio.
[0122] In a specific implementation, the implementation process of determining the positional relationship between the virtual display and the viewing object based on the target area ratio may at least include but is not limited to: first, the standard area ratio between the virtual display and the virtual screen, and the standard virtual distance corresponding to the standard area ratio can be obtained. Specifically, the standard area ratio here can refer to a ratio of a preset virtual display occupying the virtual screen. This standard area ratio can be set accordingly based on actual business needs. For example, it can be set to 1 / 4. In actual applications, taking the standard area ratio of 1 / 4 as an example, the position of the setting object (i.e., the object used to assist in parameter setting) can be kept unchanged, and then the position of the virtual display can be continuously adjusted so that the virtual display occupies 1 / 4 of the entire virtual screen. When the area of the virtual display occupies 1 / 4 of the area of the entire virtual screen, the virtual distance between the position of the setting object and the virtual display at this time can be determined, and this virtual distance can be used as the standard virtual distance corresponding to the standard area ratio of 1 / 4. That is to say, the standard area ratio refers to a value in which the area of the preset virtual display occupies the area of the virtual screen, and the standard virtual distance refers to a virtual distance between the virtual display and the set object when the ratio of the area of the virtual display to the area of the virtual screen is the standard area ratio.
[0123] After obtaining the standard area ratio, the ratio between the standard area ratio and the target area ratio can be determined. For example, assuming the standard area ratio is 1 / 4 and the target area ratio is 1 / 3, the ratio between the standard area ratio and the target area ratio can be It is worth noting that the ratio between the standard area ratio and the above-mentioned target area ratio can be obtained with the standard area ratio as the numerator and the target area ratio as the denominator, or with the target area ratio as the numerator and the standard area ratio as the denominator; then, the ratio between the standard area ratio and the target area ratio is calculated with the standard virtual distance to obtain the target virtual distance between the virtual display and the viewing object, wherein, when the ratio between the standard area ratio and the above-mentioned target area ratio is obtained with the target area ratio as the numerator, the standard virtual distance can be multiplied by the ratio between the standard area ratio and the above-mentioned target area ratio to obtain the target virtual distance between the virtual display and the viewing object; and when the ratio between the standard area ratio and the above-mentioned target area ratio is obtained with the standard area ratio as the numerator, the target virtual distance can be obtained with the standard virtual distance as the numerator and the ratio between the standard area ratio and the above-mentioned target area ratio as the denominator. After determining the target virtual distance, the positional relationship between the virtual display and the viewing object can be determined based on the target virtual distance between the virtual display and the viewing object. Specifically, the present application can preset different gear position relationships, each gear position relationship can correspond to a different distance gear. The larger the target virtual distance, the higher the corresponding distance gear will be, and the gear position relationship will belong to the gear position relationship corresponding to the higher distance gear. In other words, the positional relationship between the virtual display and the viewing object can include the target virtual distance. In a feasible embodiment, the target virtual distance can be directly determined as the positional relationship between the virtual display and the viewing object.
[0124] Step S502: Obtain display parameters of the target media frame.
[0125] In this application, the display parameters of the target media frame may refer to the display parameters of the media data, which may refer to the resolution, pixels and other parameters of the media data. Taking resolution as an example, the display parameters of the media data may include but are not limited to: 4K, 1080p, 720p, etc. In the process of determining the gaze influence area of the viewing object on the virtual display, the display parameters of the media data can be referred to for determination.
[0126] Step S503 : determining the gaze influence area of the viewing object on the virtual display based on the positional relationship and the display parameters of the target media frame.
[0127] In the present application, after obtaining the current target virtual distance between the virtual display and the viewing object, and the display parameters of the target media frame, the viewing object's gaze influence area on the virtual display can be determined based on the positional relationship and the display parameters of the target media frame. The specific implementation process may at least include but is not limited to: constructing a range mapping table based on the display parameters of the target media frame, wherein the range mapping table includes a mapping relationship between a configuration distance set and a configuration focus influence parameter set, and a mapping relationship exists between a configuration distance in the configuration distance set and a configuration focus influence parameter in the configuration focus influence parameter set. It should be understood that the present application can pre-set the focus influence ranges corresponding to different distances based on the display parameters of the target media frame to obtain a range mapping table.
[0128] Specifically, the specific implementation process of constructing the range mapping table may at least include but is not limited to: first, according to the display parameters of the target media frame, the near focus influence parameter and the far focus influence parameter may be set; wherein the near focus influence parameter refers to the gaze influence parameter of the viewing object on the virtual display when the area of the virtual display is equal to the area of the virtual screen and the virtual distance between the viewing object and the virtual display is a first distance threshold, the virtual screen refers to the screen displayed by the extended reality device to the viewing object, and the far focus influence parameter refers to the gaze influence parameter of the viewing object on the virtual display when the area of the virtual display is equal to the area of the virtual screen and the virtual distance between the viewing object and the virtual display is a second distance threshold.
[0129] Furthermore, a range mapping table can be constructed according to the first distance threshold, the second distance threshold, the near focus influence parameter and the far focus influence parameter. Specifically, one or more candidate distances can be selected from the distance between the first distance threshold and the second distance threshold; then, corresponding influence parameters can be configured for each candidate distance in the influence parameters between the near focus influence parameter and the far focus influence parameter, thereby obtaining the focus influence parameter corresponding to each candidate distance; then, the first distance threshold, the second distance threshold and each candidate distance can be determined as the configuration distance, and the near focus influence parameter, the far focus influence parameter and the focus influence parameter corresponding to each candidate distance can be determined as the configuration focus influence parameter; a mapping relationship is constructed between each configuration distance and the configuration focus influence parameter corresponding to each configuration distance, so as to obtain a range mapping table containing different configuration distances and their corresponding configuration focus influence parameters.
[0130] It should be understood that when the area of the virtual display is completely equal to the area of the virtual screen (that is, when the virtual display completely occupies the virtual screen), a near focus influence parameter can be set (that is, assuming that the virtual distance between the viewing object and the virtual display is the first distance threshold, the viewing object's gaze influence parameter / range on the current virtual display). The near focus influence parameter can include a width parameter and a height parameter. The near focus influence parameter can be expressed as: F near (w1, h1); Similarly, when the area of the virtual display is completely equal to the area of the virtual screen (that is, when the virtual display completely occupies the virtual screen), a telefocus influence parameter can be set (that is, assuming that the virtual distance between the viewing object and the virtual display is the second distance threshold, the viewing object's gaze influence parameter / range on the current virtual display). The telefocus influence parameter can include a width parameter and a height parameter. The telefocus influence parameter can be expressed as: F far (w2,h2). The second distance threshold should be much larger than the first distance threshold, and the near focus influence parameter and the far focus influence parameter will increase with the increase of the display parameter. That is to say, the larger the display parameter is, the larger the near focus influence parameter and the far focus influence parameter will be set accordingly. Here, taking the display parameter of 4K as an example, when the display parameter is 4K, the near focus influence parameter can be set to F near (w1,h1)=[30*30], that is, the width parameter is 30, the height parameter is also 30; the telephoto effect parameter can be set to F far (w2, h2) = [500*500], that is, the width parameter is 500 and the height parameter is also 500. In actual applications, the near focus influence parameter and the far focus influence parameter can be specifically set based on actual business requirements and display parameters.
[0131] After determining the near focus influence parameter and the far focus influence parameter, the present application can set corresponding focus influence parameters for each distance between the first distance threshold and the second distance threshold. Specifically, the present application can select multiple different candidate distances in the first distance threshold and the second distance threshold, and then set corresponding focus influence parameters for each candidate distance in the near focus influence parameter and the far focus influence parameter, wherein the larger the candidate distance, the larger the corresponding focus influence parameter. When selecting the candidate distance, the candidate distance can be selected according to a certain rule (for example, an increasing rule), and the corresponding focus influence parameter will also increase according to the corresponding rule. For example, taking the first distance threshold as 10 and the second distance threshold as 100 as an example, assuming that the near focus influence parameter corresponding to the first distance threshold is [30*30] and the far focus influence parameter corresponding to the second distance threshold is [200*200], and assuming that the candidate distances are selected according to the tolerance of 20, the selected candidate distances may be: 30 (the difference from the first distance threshold is 20), 50 (the difference from the previous candidate distance 30 is 20), 70 (the difference from the previous candidate distance 50 is 20), and 90 (the difference from the previous candidate distance 70 is 20). It can be seen that from the first distance threshold to the last candidate distance 90, it increases based on the tolerance of 20. Next, the corresponding focus influence parameter can be assigned to each candidate distance from the focus influence parameters in [30*30]-[200*200]. For example, the focus influence parameter corresponding to candidate distance 30 can be set to [60*60], the focus influence parameter corresponding to candidate distance 50 can be set to [90*90], the focus influence parameter corresponding to candidate distance 70 can be set to [120*120], and the focus influence parameter corresponding to candidate distance 90 can be set to [150*150]. In this way, it can be seen that the focus influence parameters corresponding to each candidate distance are increased based on the tolerance 30 (the width parameter and the height parameter are both summed with the tolerance 30). In this way, the focus influence parameters corresponding to each candidate distance can also show an increasing pattern.
[0132] Furthermore, after determining the focus influence parameters corresponding to each candidate distance, the first distance threshold, the second distance threshold and each candidate distance can be determined as configuration distances, and the various configuration distances can be aggregated to obtain a set (this set can be called a configuration distance set); the near focus influence parameters, the far focus influence parameters and the focus influence parameters corresponding to each candidate distance can also be determined as configuration focus influence parameters, and the various configuration focus influence parameters can be aggregated to obtain a set (this set can be called a configuration focus influence parameter set). Then, a mapping relationship can be constructed between each configuration distance and the corresponding configuration focus influence parameter (the configuration focus influence parameter corresponding to the first distance threshold is the near focus influence parameter; the configuration focus influence parameter corresponding to the second distance threshold is the far focus influence parameter), thereby obtaining a mapping relationship between different configuration distances and their corresponding configuration focus influence parameters, and these mapping relationships can be added to the range mapping table.
[0133] It should be understood that after determining the range mapping table, a configuration focus influence parameter that is mapped to the target virtual distance can be obtained from the range mapping table. The configuration focus influence parameter can be used as the target focus influence parameter corresponding to the target virtual distance. Based on the target focus influence parameter, the viewing object's gaze influence area on the virtual display can be determined. It is worth noting that if the target virtual distance is not included in the configuration distance set included in the range mapping table, then each configuration distance smaller than the target virtual distance (referred to as a first configuration distance) can be obtained from the range mapping table, and each configuration distance greater than the target virtual distance (referred to as a second configuration distance) can also be obtained from the range mapping table. Then, a maximum first configuration distance can be obtained from each first configuration distance, and a minimum second configuration distance can be obtained from each second configuration distance. Furthermore, a configuration focus influence parameter corresponding to the maximum first configuration distance (referred to as a first reference parameter) and a configuration focus influence parameter corresponding to the minimum second configuration distance (referred to as a second reference parameter) can be obtained. From the focus influence parameters between the first reference parameter and the second reference parameter, an appropriate focus influence parameter can be selected as the target focus influence parameter corresponding to the target virtual distance.
[0134] After obtaining the target focus influence parameter corresponding to the target virtual distance, the gaze influence area of the viewing object on the virtual display can be determined according to the target focus influence parameter. The specific implementation process can be as follows: first, a mask image of the virtual display can be obtained; wherein, the mask image completely overlaps with the virtual display in the virtual screen. It can be understood that the mask image in this application can refer to the mask image of the virtual display (a binary image in which pixels are marked as black or white to represent the shape, position or attribute of certain areas or objects). For the virtual display, a mask image of the same size and display parameters can be placed at the placement position of the virtual display. The mask image can completely overlap with the virtual display, and the scaling factor of the mask image is the same as that of the virtual display. It can be enlarged as the virtual display is enlarged, and can also be reduced as the virtual display is reduced. When a display request for each media frame is received, the initial value of the mask image (that is, the initial pixel value) can be set to 1. That is to say, when filtering each media frame, the gaze influence area of the viewing object on the current virtual display can be determined based on the mask image with an initial pixel value of 1.
[0135] After acquiring the mask image, the viewing subject's gaze focus on the virtual display can be obtained, and the focal position coordinates of the gaze focus in the mask image can be determined. The gaze focus can be understood as the point where the viewing subject's line of sight is focused. Generally, the viewing subject's eye focus refers to focusing attention on the object, scene, or thing being observed in order to perceive and understand it more clearly. This focus not only refers to focusing attention on a specific area through the eyeballs, but also includes concentrating the mind and senses on this area and relaxing attention in other areas. The gaze focus is the center point of the area where the line of sight is focused. After obtaining the line of sight focus, it can be converted into coordinates in the mask image. Since the mask image and the virtual display are completely overlapped, the point where the line of sight focus is located in the mask image can be obtained, and the coordinates corresponding to the point (which can be called the focus position coordinates) can be determined. Then, a rectangular area can be determined in the mask image with the position point indicated by the focus position coordinates as the center point, and the width parameter in the target focus influence parameter as the rectangle width, and the height parameter in the target focus influence parameter as the rectangle height; the area completely covered by the rectangular area in the virtual display can be determined as the gaze influence area of the viewing object on the virtual display.
[0136] To understand the process of determining the gaze influence area, please also refer to Figure 6 , Figure 6 This is a schematic diagram of determining the gaze influence area provided by an embodiment of the present application. Figure 6As shown, after the viewing object wears the extended reality device 100, it can view media data through the virtual display. For a certain media frame, after receiving its display request, the target virtual distance between the current virtual display and the viewing object can be obtained. Then, the target focus influence parameter / range corresponding to the target virtual distance can be obtained through the range mapping table. The target focus influence parameter will include a width parameter and a height parameter. Furthermore, the current viewing object's gaze focus in the virtual display can be determined when gazing at the virtual display. The gaze focus can be converted into a mask image, and the focus position coordinates of the gaze focus in the mask image can be obtained, for example, as shown in FIG. Figure 6 As shown, in the mask image, the focus is located at the center of the mask image, and each pixel value in the mask image is set to 1.
[0137] Furthermore, in the mask image, a rectangular area can be determined with the point indicated in the gaze position table as the center point, the width parameter in the target virtual distance as the rectangle width, and the height parameter in the target virtual distance as the rectangle height. The width of the rectangular area is the width parameter, and the height of the rectangular area is the height parameter. After the mask image determines the rectangular area, since the virtual display completely overlaps with the mask image, the same area can be determined in the virtual display. This area, that is, the area that completely overlaps with the rectangular area, can serve as the gaze influence area in the virtual display.
[0138] It is worth noting that the gaze influence area can be a rectangular area. Of course, it can also be an area in other shapes. For example, it can be a circular area. When the gaze influence area is a circular area, a circular area can be determined in the mask image with the point indicated by the focus position coordinates as the center point and the width parameter or height parameter in the target virtual distance as the radius. The area that completely overlaps with the circular area in the virtual display can be used as the gaze influence area. The specific presentation form of the gaze influence area can be specifically set based on actual business needs. This application does not limit it, but it is specifically explained here.
[0139] In an embodiment of the present application, a display processing scheme for media data is provided, which can dynamically and adaptively filter the target media frame in the media data based on the gaze influence area of the viewing object on the virtual display in the extended reality device, so as to optimize the effect of the media content displayed in different areas of the virtual display and improve the clarity of the media frame. Specifically, for a certain media frame in the media data (which can be called the target media frame), after receiving its display request, the extended reality device can obtain the gaze influence area of the viewing object on the virtual display in the device; then, according to the gaze influence area, the media content of the target media frame can be divided to obtain gaze media content and non-gaze media content, wherein the gaze media content is the media content that needs to be displayed in the gaze influence area, and the non-gaze media content is the media content that needs to be displayed outside the gaze influence area. Furthermore, different filtering methods can be used to filter the gaze media content and the non-gaze media content, and the filtered target media frame is displayed in the virtual display. It should be understood that for any media frame, the area of gaze influence of the viewing object on the virtual display can be determined. This area can be understood as the area where the viewing object's line of sight has a higher degree of attention (that is, the area where the line of sight is focused / concentrated). Then, the media content of the target media frame is divided according to the gaze influence area, and the media content where the viewing object has a higher degree of attention can be obtained. These contents are the attention media content, and the media content where the viewing object has a lower degree of attention can also be obtained. These contents are the non-attention media content. For the attention media content and the non-attention media content, the present application can select different filtering methods according to business needs to perform adaptive filtering processing on them. For example, the attention media content can be filtered according to a high-quality filtering method to ensure its clarity, and the non-attention media content can be filtered according to a smoothing filtering method to reduce its aliasing problem. Compared with a one-size-fits-all downsampling processing method, the present application uses the adaptive filtering processing method based on the attention media content and the non-attention media content, which can not only fully present the media data, but also well maintain the overall clarity of the media data and reduce the aliasing problem. It can be seen that the present application can perform adaptive filtering on media data in the business of displaying media data using an extended reality device to improve the clarity of the media data and optimize the overall presentation effect of the media. Moreover, the present application determines the gaze influence area of the viewing object based on the real-time positional relationship between the current virtual display and the viewing object when displaying each media frame. In other words, the gaze influence area is determined based on the real-time gaze position of the viewing object looking at the virtual display, and is not fixed. In this way, different gaze influence areas can be determined efficiently and accurately as the viewing object moves.In summary, the present application provides a method for filtering media data based on the real-time gaze position of the viewing object. It can dynamically allocate filtering resources to the media content of each part of the media frame by real-time perception of the layout in the virtual environment and the gaze position of the viewing object. This can maintain clarity while reducing the probability of aliasing problems. In addition, the media content of each part can be filtered in parallel, which is suitable for GPU computing and is specifically real-time, which can effectively improve the user's viewing experience.
[0140] Further, to facilitate understanding of the overall logical structure of the method provided in the embodiment of the present application, please refer to Figure 7 , Figure 7 This is a schematic diagram of a system logic architecture provided by an embodiment of the present application. Figure 7 As shown, the system architecture may include at least the following components: an image receiving component, a gaze position determination component, a display position determination component, a mask image generation component, a filter selection component, a processor component, and an image output component. For ease of understanding, the functions implemented by each component are briefly described below:
[0141] Image receiving component: The image receiving component can be used to receive a media frame to be rendered / filtered (for example, a target media frame). After receiving a display request for a media frame, the media frame to be displayed can be obtained.
[0142] Gaze position determination component: The gaze position determination component can be used to obtain the gaze position where the viewing object is currently looking at the virtual display.
[0143] Display determination component: The display determination component can be used to obtain the current position of the virtual display.
[0144] Mask Image Generation Component: This component initializes a mask image whose scaling, size, and display parameters are identical to those of the virtual display. The component sets the initial pixel values of the mask image to 1. Once the gaze-affected region (i.e., the aforementioned rectangular region) is determined, the pixel values within that region are set to 0, allowing the gaze-affected region to be identified through different pixel values.
[0145] Filter selection component: The filter selection component can be used to divide the media content in the media frame based on the gaze influence area to obtain gaze media content and non-gaze media content. The gaze media content can be filtered using a first filtering method, while the non-gaze media content can be filtered using a second filtering method.
[0146] Processor component: The processor component may include a GPU for performing calculations to implement different filtering processes on the attention media content and the non-attention media content according to different filtering methods.
[0147] Image output component: The image output component is used to output the media frame after display filtering processing.
[0148] For the specific implementation of each component, please refer to the relevant description in the previous embodiments, and will not be elaborated here.
[0149] In an embodiment of the present application, a display processing scheme for media data is provided, which can dynamically and adaptively filter the target media frame in the media data based on the gaze influence area of the viewing object on the virtual display in the extended reality device, so as to optimize the effect of the media content displayed in different areas of the virtual display and improve the clarity of the media frame. Specifically, for a certain media frame in the media data (which can be referred to as the target media frame), after receiving its display request, the extended reality device can obtain the gaze influence area of the viewing object on the virtual display in the device; then, according to the gaze influence area, the media content of the target media frame can be divided to obtain gaze media content and non-gaze media content, wherein the gaze media content is the media content that needs to be displayed in the gaze influence area, and the non-gaze media content is the media content that needs to be displayed outside the gaze influence area. Furthermore, different filtering methods can be used to filter the gaze media content and the non-gaze media content, and the filtered target media frame is displayed in the virtual display. It should be understood that for any media frame, the area of gaze influence of the viewing object on the virtual display can be determined. This area can be understood as the area where the viewing object's line of sight has a higher degree of attention (that is, the area where the line of sight is focused / concentrated). Then, the media content of the target media frame is divided according to the gaze influence area, and the media content where the viewing object has a higher degree of attention can be obtained. These contents are the attention media content, and the media content where the viewing object has a lower degree of attention can also be obtained. These contents are the non-attention media content. For the attention media content and the non-attention media content, the present application can select different filtering methods according to business needs to perform adaptive filtering processing on them. For example, the attention media content can be filtered according to a high-quality filtering method to ensure its clarity, and the non-attention media content can be filtered according to a smoothing filtering method to reduce its aliasing problem. Compared with a one-size-fits-all downsampling processing method, the present application uses the adaptive filtering processing method based on the attention media content and the non-attention media content, which can not only fully present the media data, but also well maintain the overall clarity of the media data and reduce the aliasing problem. It can be seen that the present application can perform adaptive filtering on media data in the business of displaying media data using an extended reality device to improve the clarity of the media data and optimize the overall presentation effect of the media. Moreover, the present application determines the gaze influence area of the viewing object based on the real-time positional relationship between the current virtual display and the viewing object when displaying each media frame. In other words, the gaze influence area is determined based on the real-time gaze position of the viewing object looking at the virtual display, and is not fixed. In this way, different gaze influence areas can be determined efficiently and accurately as the viewing object moves.In summary, the present application provides a method for filtering media data based on the real-time gaze position of the viewing object. It can dynamically allocate filtering resources to the media content of each part of the media frame by real-time perception of the layout in the virtual environment and the gaze position of the viewing object. This can maintain clarity while reducing the probability of aliasing problems. In addition, the media content of each part can be filtered in parallel, which is suitable for GPU computing and is specifically real-time, which can effectively improve the user's viewing experience.
[0150] Further, see Figure 8 , Figure 8 This is a schematic diagram of the structure of a media data display device provided in an embodiment of the present application. The media data display device may be a computer program running on a computer device, for example, the media data display device is an application software; the media data display device may be used to execute Figure 3 As shown in the method. Figure 8 As shown, the media data display device 1 may include: a region acquisition module 11 , a content division module 12 and a filtering module 13 .
[0151] The region acquisition module 11 is configured to acquire a gaze influence region of a viewing object of the media data on the virtual display in response to a display request for a target media frame in the media data;
[0152] A content segmentation module 12 is configured to segment the media content of the target media frame according to the gaze influence area to obtain gaze media content and non-gaze media content; the gaze media content is the media content to be displayed in the gaze influence area, and the non-gaze media content is the media content to be displayed outside the gaze influence area;
[0153] The filtering module 13 is configured to filter the attention media content and the non-attention media content using different filtering methods, and display the filtered target media frames on the virtual display.
[0154] The specific implementation of the region acquisition module 11, the content division module 12 and the filtering module 13 can be found in the above Figure 3 The description of steps S301 to S303 in the corresponding embodiment will not be repeated here.
[0155] In one embodiment, the region acquisition module 11 obtains the gaze influence region of the virtual display of the viewing object of the media data in response to a display request for a target media frame in the media data, and specifically implements the following:
[0156] obtaining a positional relationship between the virtual display and a viewing object of the media data;
[0157] Get the display parameters of the target media frame;
[0158] Based on the positional relationship and the display parameters of the target media frame, a gaze influence area of the viewing object on the virtual display is determined.
[0159] In one embodiment, the specific implementation of the region acquisition module 11 acquiring the positional relationship between the virtual display and the viewing object of the media data includes:
[0160] Acquiring a virtual screen displayed by an extended reality device to a viewing object of the media data; the virtual display is in the virtual screen;
[0161] Acquire a first area of the virtual display and a second area of the virtual screen;
[0162] determining a target area ratio between the first area and the second area;
[0163] The positional relationship between the virtual display and the viewing object is determined according to the target area ratio.
[0164] In one embodiment, the specific implementation of the region acquisition module 11 determining the positional relationship between the virtual display and the viewing object according to the target area ratio includes:
[0165] Obtaining a standard area ratio between the virtual display and the virtual screen, and a standard virtual distance corresponding to the standard area ratio;
[0166] determining a ratio between the standard area ratio and the target area ratio;
[0167] The ratio between the standard area ratio and the target area ratio is calculated and processed with the standard virtual distance to obtain the target virtual distance between the virtual display and the viewing object;
[0168] The positional relationship between the virtual display and the viewing object is determined according to a target virtual distance between the virtual display and the viewing object.
[0169] In one embodiment, the positional relationship includes a target virtual distance between the virtual display and the viewing object;
[0170] The specific implementation method of the region acquisition module 11 for determining the gaze influence region of the viewing object on the virtual display based on the position relationship and the display parameters of the target media frame includes:
[0171] Constructing a range mapping table based on the display parameters of the target media frame; the range mapping table includes a mapping relationship between a configuration distance set and a configuration focus influence parameter set, and a mapping relationship exists between a configuration distance in the configuration distance set and a configuration focus influence parameter in the configuration focus influence parameter set;
[0172] In the range mapping table, a target focus influence parameter that has a mapping relationship with the target virtual distance is obtained;
[0173] The gaze influence area of the viewing object on the virtual display is determined according to the target focus influence parameter.
[0174] In one embodiment, the specific implementation of the region acquisition module 11 constructing the range mapping table based on the display parameters of the target media frame includes:
[0175] According to the display parameters of the target media frame, a near focus influence parameter and a far focus influence parameter are set; the near focus influence parameter refers to a parameter of the viewing object's gaze influence on the virtual display when the area of the virtual display is equal to the area of the virtual screen and the virtual distance between the viewing object and the virtual display is a first distance threshold, and the virtual screen refers to the screen displayed to the viewing object by the extended reality device; the far focus influence parameter refers to a parameter of the viewing object's gaze influence on the virtual display when the area of the virtual display is equal to the area of the virtual screen and the virtual distance between the viewing object and the virtual display is a second distance threshold;
[0176] A range mapping table is constructed according to the first distance threshold, the second distance threshold, the near focus influence parameter, and the far focus influence parameter.
[0177] In one embodiment, the specific implementation of the region acquisition module 11 constructing the range mapping table according to the first distance threshold, the second distance threshold, the near focus influence parameter, and the far focus influence parameter includes:
[0178] selecting one or more candidate distances among distances between a first distance threshold and a second distance threshold;
[0179] Among the influencing parameters in the near focus influencing parameters and the far focus influencing parameters, a corresponding influencing parameter is configured for each candidate distance to obtain a focus influencing parameter corresponding to each candidate distance;
[0180] Determine the first distance threshold, the second distance threshold, and each candidate distance as a configuration distance, and determine the near focus influence parameter, the far focus influence parameter, and the focus influence parameter corresponding to each candidate distance as a configuration focus influence parameter;
[0181] A mapping relationship is established between each configuration distance and the configuration focus influence parameter corresponding to each configuration distance to obtain a range mapping table.
[0182] In one embodiment, the target focus impact parameters include a width parameter and a height parameter;
[0183] The specific implementation method of the region acquisition module 11 determining the gaze influence region of the viewing object on the virtual display according to the target focus influence parameter includes:
[0184] Acquire a mask image of the virtual display; the mask image completely overlaps with the virtual display in the virtual screen;
[0185] Obtaining the gaze focus of the viewing object on the virtual display, and determining the focal position coordinates of the gaze focus in the mask image;
[0186] A rectangular area is determined in the mask image with the position point indicated by the focus position coordinates as the center point, the width parameter in the target focus influence parameter as the rectangle width, and the height parameter in the target focus influence parameter as the rectangle height;
[0187] An area completely covered by the rectangular area in the virtual display is determined as a gaze influence area of the viewing object on the virtual display.
[0188] In one embodiment, the specific implementation of the filtering module 13 using different filtering methods to filter the attention media content and the non-attention media content includes:
[0189] Using a first filtering method to filter the attention media content;
[0190] The non-attention media content is filtered using the second filtering method; the filtering effect indicated by the first filtering method is better than the filtering effect indicated by the second filtering method.
[0191] In one embodiment, the first filtering mode is a high-quality filtering mode;
[0192] The specific implementation of the filtering module 13 using the first filtering method to filter the attention media content includes:
[0193] Call the image enhancement model according to the high-quality filtering method;
[0194] Perform content enhancement on gaze-based media content through image enhancement models.
[0195] In one embodiment, the specific implementation of the filtering module 13 using the second filtering method to filter the non-attention media content includes:
[0196] Calling the smoothing filter model according to the second filtering mode;
[0197] The non-attention media content is smoothed and filtered using a smoothing filter model.
[0198] In an embodiment of the present application, a display processing scheme for media data is provided, which can dynamically and adaptively filter the target media frame in the media data based on the gaze influence area of the viewing object on the virtual display in the extended reality device, so as to optimize the effect of the media content displayed in different areas of the virtual display and improve the clarity of the media frame. Specifically, for a certain media frame in the media data (which can be called the target media frame), after receiving its display request, the extended reality device can obtain the gaze influence area of the viewing object on the virtual display in the device; then, according to the gaze influence area, the media content of the target media frame can be divided to obtain gaze media content and non-gaze media content, wherein the gaze media content is the media content that needs to be displayed in the gaze influence area, and the non-gaze media content is the media content that needs to be displayed outside the gaze influence area. Furthermore, different filtering methods can be used to filter the gaze media content and the non-gaze media content, and the filtered target media frame is displayed in the virtual display. It should be understood that for any media frame, the area of gaze influence of the viewing object on the virtual display can be determined. This area can be understood as the area where the viewing object's line of sight has a higher degree of attention (that is, the area where the line of sight is focused / concentrated). Then, the media content of the target media frame is divided according to the gaze influence area, and the media content where the viewing object has a higher degree of attention can be obtained. These contents are the attention media content, and the media content where the viewing object has a lower degree of attention can also be obtained. These contents are the non-attention media content. For the attention media content and the non-attention media content, the present application can select different filtering methods according to business needs to perform adaptive filtering processing on them. For example, the attention media content can be filtered according to a high-quality filtering method to ensure its clarity, and the non-attention media content can be filtered according to a smoothing filtering method to reduce its aliasing problem. Compared with a one-size-fits-all downsampling processing method, the present application uses the adaptive filtering processing method based on the attention media content and the non-attention media content, which can not only fully present the media data, but also well maintain the overall clarity of the media data and reduce the aliasing problem. It can be seen that the present application can perform adaptive filtering on media data in the business of displaying media data using an extended reality device to improve the clarity of the media data and optimize the overall presentation effect of the media. Moreover, the present application determines the gaze influence area of the viewing object based on the real-time positional relationship between the current virtual display and the viewing object when displaying each media frame. In other words, the gaze influence area is determined based on the real-time gaze position of the viewing object looking at the virtual display, and is not fixed. In this way, different gaze influence areas can be determined efficiently and accurately as the viewing object moves.In summary, the present application provides a method for filtering media data based on the real-time gaze position of the viewing object. It can dynamically allocate filtering resources to the media content of each part of the media frame by real-time perception of the layout in the virtual environment and the gaze position of the viewing object. This can maintain clarity while reducing the probability of aliasing problems. In addition, the media content of each part can be filtered in parallel, which is suitable for GPU computing and is specifically real-time, which can effectively improve the user's viewing experience.
[0199] Further, see Figure 9 , Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 9 As shown, the above-mentioned computer device 8000 may include: a processor 8001, a network interface 8004 and a memory 8005. In addition, the above-mentioned computer device 8000 also includes: a user interface 8003, and at least one communication bus 8002. The communication bus 8002 is used to realize the connection and communication between these components. The user interface 8003 may include a display screen (Display), a keyboard (Keyboard), and the user interface 8003 may optionally include a standard wired interface and a wireless interface. The network interface 8004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 8005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 8005 may optionally also be at least one storage device located away from the aforementioned processor 8001. As Figure 9 As shown, the memory 8005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a computer program.
[0200] exist Figure 9 In the computer device 8000 shown, the network interface 8004 can provide network communication functions; the user interface 8003 is mainly used to provide an input interface for the user; and the processor 8001 can be used to call the computer program stored in the memory 8005 to implement the steps in the various method embodiments of the present application.
[0201] It should be understood that the computer device 8000 described in the embodiment of the present application can execute the above Figures 3 to 5 The description of the method for displaying the media data in the corresponding embodiment can also be performed as described above. Figure 7 The description of the media data display device 1 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.
[0202] In addition, it should be noted that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the computer device 8000 for data processing mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, the computer program can execute the above-mentioned data processing. Figures 3 to 5 The description of the method for displaying the above-mentioned media data in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0203] The computer-readable storage medium may be a display device for media data provided in any of the aforementioned embodiments or an internal storage unit of the computer device, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may include both an internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0204] In one aspect of the present application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the method provided in one aspect of the embodiments of the present application.
[0205] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0206] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0207] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0208] The methods and related devices provided by the embodiments of the present application are described with reference to the method flow charts and / or structural diagrams provided by the embodiments of the present application. Specifically, each process and / or block in the method flow charts and / or structural diagrams, as well as the combination of processes and / or blocks in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 The flow or flows and / or structures illustrate the steps of the functions specified in one block or multiple blocks.
[0209] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A method for displaying media data, characterized in that: The method is applied to an extended reality device, wherein the extended reality device includes a virtual display, and the virtual display is used to display the media data; the method includes: In response to a display request for a target media frame in the media data, a virtual screen displayed by the extended reality device to a viewing object of the media data is obtained; the virtual display is located in the virtual screen; a first area of the virtual display and a second area of the virtual screen are obtained; a target area ratio between the first area and the second area is determined; a positional relationship between the virtual display and the viewing object is determined based on the target area ratio; display parameters of the target media frame are obtained; and a gaze influence area of the viewing object on the virtual display is determined based on the positional relationship and the display parameters of the target media frame. Dividing the media content of the target media frame according to the gaze influence area to obtain gaze media content and non-gaze media content; the gaze media content is media content to be displayed in the gaze influence area, and the non-gaze media content is media content to be displayed outside the gaze influence area; The attention media content and the non-attention media content are filtered using different filtering methods, and the filtered target media frames are displayed on the virtual display.
2. The method according to claim 1, characterized in that The determining the positional relationship between the virtual display and the viewing object according to the target area ratio includes: Acquire a standard area ratio between the virtual display and the virtual screen, and a standard virtual distance corresponding to the standard area ratio; determining a ratio between the standard area ratio and the target area ratio; performing a calculation on the ratio between the standard area ratio and the target area ratio and the standard virtual distance to obtain a target virtual distance between the virtual display and the viewing object; A positional relationship between the virtual display and the viewing object is determined according to a target virtual distance between the virtual display and the viewing object.
3. The method according to claim 1, characterized in that The positional relationship includes a target virtual distance between the virtual display and the viewing object; The determining, based on the positional relationship and the display parameters of the target media frame, a gaze-affected area of the viewing object on the virtual display includes: Constructing a range mapping table based on the display parameters of the target media frame; wherein the range mapping table includes a mapping relationship between a set of configured distances and a set of configured focus influence parameters, and a mapping relationship exists between a configured distance in the set of configured distances and a configured focus influence parameter in the set of configured focus influence parameters; In the range mapping table, obtaining a target focus influence parameter that has a mapping relationship with the target virtual distance; A gaze influence area of the viewing object on the virtual display is determined according to the target focus influence parameter.
4. The method according to claim 3, characterized in that The constructing a range mapping table based on the display parameters of the target media frame includes: A near focus influence parameter and a far focus influence parameter are set according to the display parameters of the target media frame; the near focus influence parameter refers to a parameter of the viewing object's gaze influence on the virtual display when the area of the virtual display is equal to the area of the virtual screen and the virtual distance between the viewing object and the virtual display is a first distance threshold, and the virtual screen refers to a screen displayed by the extended reality device to the viewing object; the far focus influence parameter refers to a parameter of the viewing object's gaze influence on the virtual display when the area of the virtual display is equal to the area of the virtual screen and the virtual distance between the viewing object and the virtual display is a second distance threshold; A range mapping table is constructed according to the first distance threshold, the second distance threshold, the near focus influencing parameter, and the far focus influencing parameter.
5. The method according to claim 4, characterized in that The constructing a range mapping table according to the first distance threshold, the second distance threshold, the near focus influencing parameter, and the far focus influencing parameter includes: selecting one or more candidate distances among distances between the first distance threshold and the second distance threshold; Among the influencing parameters between the near focus influencing parameters and the far focus influencing parameters, configuring a corresponding influencing parameter for each candidate distance, to obtain a focus influencing parameter corresponding to each candidate distance; Determine the first distance threshold, the second distance threshold, and each candidate distance as a configuration distance, and determine the near focus influence parameter, the far focus influence parameter, and the focus influence parameter corresponding to each candidate distance as a configuration focus influence parameter; A mapping relationship is established between each of the configuration distances and the configuration focus influence parameter corresponding to each of the configuration distances to obtain a range mapping table.
6. The method according to claim 3, characterized in that The target focus influencing parameters include a width parameter and a height parameter; The determining, according to the target focus influence parameter, a gaze influence area of the viewing object on the virtual display includes: Acquiring a mask image of the virtual display; wherein the mask image completely overlaps with the virtual display in the virtual screen; Acquiring a gaze focus of the viewing object on the virtual display, and determining a focal position coordinate of the gaze focus in the mask image; Determine a rectangular area in the mask image with the position point indicated by the focus position coordinates as the center point, with the width parameter in the target focus influence parameter as the rectangle width, and with the height parameter in the target focus influence parameter as the rectangle height; An area in the virtual display that is completely covered by the rectangular area is determined as a gaze influence area of the viewing object on the virtual display.
7. The method according to claim 1, characterized in that The filtering process of the attention media content and the non-attention media content by using different filtering methods includes: Performing filtering processing on the watched media content using a first filtering method; The non-attention media content is filtered using a second filtering method; the filtering effect indicated by the first filtering method is better than the filtering effect indicated by the second filtering method.
8. The method according to claim 7, characterized in that The filtering process of the attention media content by using the first filtering method includes: Calling the image enhancement model according to the first filtering method; Content enhancement processing is performed on the gaze media content using the image enhancement model.
9. The method according to claim 7, characterized in that The filtering process of the non-attention media content by using the second filtering method includes: Calling the smoothing filter model according to the second filtering mode; The non-attention media content is subjected to smoothing filtering processing by using the smoothing filtering model.
10. A device for displaying media data, characterized in that: The device comprises: An area acquisition module is configured to, in response to a display request for a target media frame in the media data, acquire a virtual screen displayed by an extended reality device to a viewing object of the media data; the extended reality device includes a virtual display, the virtual display is configured to display the media data, and the virtual display is located in the virtual screen; acquire a first area of the virtual display and a second area of the virtual screen; determine a target area ratio between the first area and the second area; determine a positional relationship between the virtual display and the viewing object based on the target area ratio; acquire display parameters of the target media frame; and determine a gaze influence area of the viewing object on the virtual display based on the positional relationship and the display parameters of the target media frame. a content division module, configured to divide the media content of the target media frame according to the gaze influence area to obtain gaze media content and non-gaze media content; the gaze media content is media content to be displayed in the gaze influence area, and the non-gaze media content is media content to be displayed outside the gaze influence area; The filtering module is configured to filter the attention media content and the non-attention media content using different filtering methods, and display the filtered target media frames on the virtual display.
11. A computer device, characterized in that: include: processor, memory, and network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide a network communication function, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product comprises a computer program stored in a computer-readable storage medium. The computer program is suitable for being read and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Real-time VR image filtering method and system based on fixation point information and storage medium
CN111757090A
Information processing device, information processing method, and recording medium
US20230132045A1