Immersive Video Quality Evaluation Method and Device Guided by Six-Degree-of-Freedom Information
By constructing and training an immersive video quality evaluation model containing multimodal encoder and decoder, the accuracy problem of immersive video quality evaluation is solved, and video quality evaluation is achieved consistent with human visual perception.
Patent Information
- Application Number
- CN202510346077.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-24
AI Technical Summary
During the acquisition, transmission and display of immersive videos, distortion problems such as motion blur and Gaussian noise are easily encountered, which seriously affects the video quality and user experience. It is difficult for the existing technology to accurately evaluate and conform to the quality of immersive videos that are in line with human visual perception.
Using an immersive video quality evaluation method based on six-degree of freedom information guidance, an evaluation model of a speech decoder including visual information encoding module, spatiotemporal mapping module, language encoder and large language model is constructed and trained, texture videos and depth videos of multiple viewpoints are obtained, video blocks and keyframes are extracted, feature extraction and spatiotemporal mapping are performed, and immersive video quality scores are generated by combining language encoding and speech decoding.
The perception ability of different visual characteristics is improved, the cross-modal understanding ability of large language models is enhanced, and by integrating visual information processing and text instruction guidance, visual marking and instruction marking are effectively learned and processed, and the accurate evaluation of immersive video quality is completed. The results are highly consistent with the human visual perception characteristics.
Smart Images

Figure CN119863744B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and video evaluation, and particularly relates to an immersive video quality evaluation method and device guided by six-degree-of-freedom information. Background Art
[0002] Today, with the rapid development of digital multimedia technology, the concept of the metaverse has taken root in people's hearts and given rise to a series of new applications that integrate virtual and real technologies. Among them, immersive videos are particularly eye-catching. Immersive video technology has shown its broad application potential in many fields such as distance education and medical treatment. Compared with traditional videos, immersive videos not only contain richer texture details but also incorporate complex parallax information, bringing users an unprecedented visual experience.
[0003] However, during the acquisition, transmission, and display of immersive videos, distortion problems such as motion blur and Gaussian noise will inevitably occur. These problems seriously affect the video quality and user experience. Therefore, developing an algorithm that can accurately evaluate the quality of immersive videos and conform to human visual perception not only has profound research significance but also indicates broad application prospects, which will provide a solid foundation for the further development and optimization of immersive video technology. Summary of the Invention
[0004] The purpose of the present invention is to provide an immersive video quality evaluation method and device guided by six-degree-of-freedom information, which can accurately evaluate the quality of immersive videos and conform to human visual perception.
[0005] The present invention adopts the following technical solutions:
[0006] In a first aspect, an immersive video quality evaluation method guided by six-degree-of-freedom information includes:
[0007] Constructing and training an immersive video quality evaluation model guided by six-degree-of-freedom information to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a visual information encoding module, a spatio-temporal mapping module, a language encoder, and a speech decoder of a large language model;
[0008] Obtaining an immersive video containing texture videos and depth videos with multiple viewpoints, extracting several texture video blocks and texture key frames from the texture videos with multiple viewpoints, and extracting several depth key frames from the depth videos with multiple viewpoints;
[0009] Input a number of texture video blocks, texture key frames, and depth key frames into the trained immersive video quality evaluation model. The visual information encoding module extracts features from the texture video blocks, texture key frames, and depth key frames of multiple viewpoints respectively to obtain corresponding visual features. Input the visual features into the spatio-temporal mapping module to obtain temporal visual tags and spatial visual tags. Encode the instruction information and six-degree-of-freedom viewpoint position information through the language encoder to obtain text instruction tags and viewpoint position tags. Combine the temporal visual tags, spatial visual tags, text instruction tags, and viewpoint position tags to obtain combined tags, and input the combined tags into the speech decoder to obtain the immersive video quality score.
[0010] Preferably, extract a number of texture video blocks and texture key frames from the texture videos of multiple viewpoints, and extract a number of depth key frames from the depth videos of multiple viewpoints, specifically as follows:
[0011] For the texture video of each viewpoint and the depth video Perform block division to obtain K consecutive texture video blocks and K consecutive depth video blocks; where, represents the j-th texture video block, represents the j-th depth video block, N represents the number of frames of the texture video or depth video of each viewpoint, the number of frames of the texture video or depth video of each viewpoint is equal, n represents the n-th frame of each viewpoint, and τ represents the number of video frames included in each texture video block or depth video block, Take the first frame of each texture video block and depth video block respectively as the key frame of the block, and obtain the texture key frame f tj = x τj and the depth key frame f dj = y τj .
[0012] Preferably, the visual information encoding module extracts features from the texture video blocks, texture key frames, and depth key frames of multiple viewpoints to obtain different visual features, specifically including:
[0013] Through the depth visual encoder and spatial visual encoder of the visual information encoding module, perform feature extraction on the depth key frame f dj = y τj and the texture key frame f tj = x τj respectively to obtain the depth key frame feature and texture key frame feature of each viewpoint. Through the temporal visual encoder of the visual information encoding module, perform feature extraction on the texture video blocks to obtain texture video block features of multiple viewpoints as shown in the following formula:
[0014]
[0015] F t i = VE spatial (Concat(f t1 , f t2 ,..., f tK ));
[0016]
[0017] where Concat(·) represents feature concatenation; VE deep (·), VE spatial (·), and VE temporal (·) represent the depth visual encoder, the spatial visual encoder, and the temporal visual encoder respectively; i represents the i-th viewpoint, and H, W, and C represent the height, width, and number of channels of the feature, represents the set of real numbers.
[0018] Preferably, inputting the visual feature into the spatio-temporal mapping module to obtain the temporal visual token and the spatial visual token, specifically including:
[0019] Inputting the depth key-frame feature and the texture key-frame feature into the spatial mapping unit of the spatio-temporal mapping module to obtain the spatial visual token, and inputting the texture video block feature into the temporal mapping unit to obtain the temporal visual token, as shown in the following formula:
[0020]
[0021] where SP(·) represents the spatial mapping unit; TP(·) represents the temporal mapping unit; represents the spatial visual token; represents the temporal visual token.
[0022] Preferably, encoding the instruction information and the six-degree-of-freedom viewpoint position information through a language encoder to obtain the text instruction token and the viewpoint position token, specifically including:
[0023] Encoding the instruction information through the language encoder; the instruction information includes a system instruction, a guiding instruction, and a response limit, the system instruction is used to inform the large language model of its function, the guiding instruction is used to inform the large language model of its specific task, and the response limit is used to inform the large language model of the specific requirements for answering questions;
[0024] The text instruction token obtained after encoding the instruction information is as shown in the following formula:
[0025] F text = LE(I text );
[0026] Among them, I text represents instruction information; LE(·) represents a language encoder; F text represents a text instruction marker;
[0027] The six-degree-of-freedom viewpoint position information is encoded by the language encoder to obtain a viewpoint position marker, as shown in the following formula:
[0028]
[0029] Among them, represents the position information of viewpoint i; represents the position marker of viewpoint i.
[0030] Preferably, the obtained immersive video quality score is as follows:
[0031]
[0032] Among them, LD(·) represents the speech decoder of the large language model; Q represents the obtained immersive video quality score; represents the synthetic feature of each viewpoint.
[0033] In a second aspect, an immersive video quality evaluation device guided by six-degree-of-freedom information includes:
[0034] A model construction and training module, configured to construct and train an immersive video quality evaluation model guided by six-degree-of-freedom information to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a visual information encoding module, a spatio-temporal mapping module, a language encoder, and a speech decoder of a large language model;
[0035] A data extraction module, configured to obtain an immersive video including texture videos and depth videos of multiple viewpoints, extract a plurality of texture video blocks and texture key frames from the texture videos of multiple viewpoints, and extract a plurality of depth key frames from the depth videos of multiple viewpoints;
[0036] The video quality evaluation module is configured to input a plurality of texture video blocks, texture key frames, and depth key frames into a trained immersive video quality evaluation model. The visual information encoding module extracts features from the texture video blocks, texture key frames, and depth key frames of multiple viewpoints respectively to obtain corresponding visual features. The visual features are input into the spatio-temporal mapping module to obtain temporal visual tags and spatial visual tags. The instruction information and six-degree-of-freedom viewpoint position information are encoded by the language encoder to obtain text instruction tags and viewpoint position tags. The temporal visual tags, spatial visual tags, text instruction tags, and viewpoint position tags are combined to obtain combined tags, and the combined tags are input into the speech decoder to obtain the immersive video quality score.
[0037] In a third aspect, an electronic device includes:
[0038] One or more processors;
[0039] A storage device for storing one or more programs;
[0040] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the immersive video quality evaluation methods guided by six-degree-of-freedom information.
[0041] In a fourth aspect, a computer-readable storage medium stores a computer program, which when executed by a processor implements any of the immersive video quality evaluation methods guided by six-degree-of-freedom information.
[0042] In a fifth aspect, a computer program product includes a computer program, which when executed by a processor implements any of the immersive video quality evaluation methods guided by six-degree-of-freedom information.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] (1) The spatio-temporal mapping module in the immersive video quality evaluation model constructed by the present invention maps spatial visual features and temporal visual features into the text space respectively, which can improve the perception ability of different visual features;
[0045] (2) The language encoder in the immersive video quality evaluation model constructed by the present invention guides visual features by encoding six-degree-of-freedom viewpoint position information and text guidance instructions, thereby improving the cross-modal understanding ability of the large language model;
[0046] (3) The present invention integrates the quality perception marking of immersive videos from two parts: visual information processing and text instruction guidance. Through the speech decoder of the large language model, it can effectively learn and process visual marks and instruction marks, thereby completing the quality evaluation of immersive videos. The evaluation results are highly consistent with the characteristics of human visual perception, providing a scientific and effective solution for the objective evaluation of immersive video quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flowchart of the immersive video quality evaluation method based on six-degree-of-freedom information guidance provided by an embodiment of the present invention;
[0048] Figure 2 is a schematic structural diagram of the immersive video quality evaluation model provided by an embodiment of the present invention;
[0049] Figure 3 is a block diagram of the structure of the immersive video quality evaluation device based on six-degree-of-freedom information guidance provided by an embodiment of the present invention;
[0050] Figure 4 is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0052] Refer to Figure 1 As shown, an immersive video quality evaluation method based on six-degree-of-freedom information guidance in this embodiment includes the following steps.
[0053] S101. Construct an immersive video quality evaluation model based on six-degree-of-freedom information guidance and train it to obtain a trained immersive video quality evaluation model.
[0054] Specifically, refer to Figure 2 As shown, it is a schematic structural diagram of the immersive video quality evaluation model provided by an embodiment of the present invention. The immersive video quality evaluation model includes a visual information encoding module, a spatio-temporal mapping module, a language encoder, and a speech decoder of the large language model. The visual information encoding module includes a depth visual encoder, a spatial visual encoder, and a temporal visual encoder. The spatio-temporal mapping module includes a spatial mapping unit and a temporal mapping unit.
[0055] The proposed immersive video quality evaluation model based on six-degree-of-freedom information guidance is built using Pytorch and experiments are conducted using an NVIDIA RTX A6000 GPU. The experimental dataset is IMVD, 80% of which is used for training. In the training phase, the input resolution is adjusted to 448×448 videos for visual encoding, and the corresponding text guidance information is fused. The model uses a pre-trained Vision Transformer as the spatial visual encoder, Resnet as the depth visual encoder, and Slowfast as the temporal visual encoder, while the language codec uses the language codec of the large language model Llama. Specifically, when implementing, the parameters of the language codec and the visual encoder are fixed, and only the spatio-temporal mapping module is trained and adjusted to achieve the modality alignment of visual text tokens. S102, Obtain an immersive video containing texture videos and depth videos with multiple viewpoints, extract several texture video blocks and texture key frames from the texture videos of multiple viewpoints, and extract several depth key frames from the depth videos of multiple viewpoints.
[0056] Specifically, the texture video and depth video of each viewpoint are segmented. The length of each video block is the same. Each viewpoint obtains K consecutive texture video blocks and K consecutive depth video blocks. The first frame in each block is taken as the key frame of the block, and thus several video blocks and key frames are obtained.
[0057] In this embodiment, for the texture video and depth video of each viewpoint, they are segmented to obtain K consecutive texture video blocks and K consecutive depth video blocks; where, represents the j-th texture video block, represents the j-th depth video block, N represents the number of frames of the texture video or depth video of each viewpoint, the number of frames of the texture video or depth video of each viewpoint is equal, n represents the n-th frame of each viewpoint, and τ represents the number of video frames included in each texture video block or depth video block. For each texture video block and depth video block, the first frame is respectively taken as the key frame of the block, and the texture key frame f tj = x τj and the depth key frame f dj = y τj .
[0058] S103. Input a number of texture video blocks, texture key frames, and depth key frames into the trained immersive video quality evaluation model. The visual information encoding module extracts features from the texture video blocks, texture key frames, and depth key frames of multiple viewpoints respectively to obtain corresponding visual features. Input the visual features into the spatio-temporal mapping module to obtain temporal visual tags and spatial visual tags. Encode the instruction information and six-degree-of-freedom viewpoint position information through the language encoder to obtain text instruction tags and viewpoint position tags. Combine the temporal visual tags, spatial visual tags, text instruction tags, and viewpoint position tags to obtain combined tags, and input the combined tags into the speech decoder to obtain the immersive video quality score.
[0059] In this embodiment, the visual information encoding module extracts features from the texture video blocks, texture key frames, and depth key frames of multiple viewpoints to obtain different visual features, specifically including:
[0060] The depth visual encoder and spatial visual encoder of the visual information encoding module extract features from the depth key frame f dj = y τj and the texture key frame f tj = x τj respectively to obtain the depth key frame features and texture key frame features of each viewpoint. The temporal visual encoder of the visual information encoding module extracts features from the texture video blocks to obtain the texture video block features of multiple viewpoints as shown in the following formula:
[0061]
[0062] F t i = VE spatial (Concat(f t1 , f t2 ,..., f tK ));
[0063]
[0064] where Concat(·) represents feature concatenation; VE deep (·), VE spatial (·), and VE temporal (·) represent the depth visual encoder, spatial visual encoder, and temporal visual encoder respectively; i represents the i-th viewpoint, and H, W, and C represent the height, width, and number of channels of the feature respectively, representing the set of real numbers.
[0065] Input the visual features into the spatio-temporal mapping module to obtain temporal visual tags and spatial visual tags, specifically including:
[0066] Input the depth key-frame features and texture key-frame features into the spatial mapping unit of the spatio-temporal mapping module to obtain spatial visual tags, and input the texture video block features into the temporal mapping unit to obtain temporal visual tags, as shown in the following formula:
[0067]
[0068] where, SP(·) represents the spatial mapping unit; TP(·) represents the temporal mapping unit; represents the spatial visual tag; represents the temporal visual tag.
[0069] Encode the instruction information and six-degree-of-freedom viewpoint position information through the language encoder to obtain text instruction tags and viewpoint position tags, specifically including:
[0070] Encode the instruction information through the language encoder; the instruction information includes system instructions, guiding instructions, and response limitations. The system instructions are used to inform the large language model of its functions, the guiding instructions are used to inform the large language model of its specific tasks, and the response limitations are used to inform the large language model of the specific requirements for answering questions;
[0071] The text instruction tags obtained after encoding the instruction information are as shown in the following formula:
[0072] F text =LE(I text );
[0073] where, I text represents the instruction information; LE(·) represents the language encoder; F text represents the text instruction tag;
[0074] Encode the six-degree-of-freedom viewpoint position information through the language encoder to obtain the viewpoint position tag, as shown in the following formula:
[0075]
[0076] where, represents the position information of viewpoint i; represents the position tag of viewpoint i.
[0077] Combine the temporal visual tags, spatial visual tags, text instruction tags, and viewpoint position tags to obtain combined tags, and input the combined tags into the speech decoder to obtain the immersive video quality score, as shown below:
[0078]
[0079] Among them, LD(·) represents the speech decoder of the large language model; Q represents the obtained immersive video quality score; represents the synthesized feature of each viewpoint.
[0080] See Figure 3 As shown, the present invention also discloses an immersive video quality evaluation device guided by six-degree-of-freedom information, including:
[0081] A model construction and training module 301, configured to construct and train an immersive video quality evaluation model guided by six-degree-of-freedom information to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a visual information encoding module, a spatio-temporal mapping module, a language encoder, and a speech decoder of the large language model;
[0082] A data extraction module 302, configured to obtain an immersive video including texture videos and depth videos of multiple viewpoints, extract several texture video blocks and texture key frames from the texture videos of multiple viewpoints, and extract several depth key frames from the depth videos of multiple viewpoints;
[0083] A video quality evaluation module 303, configured to input several texture video blocks, texture key frames, and depth key frames into the trained immersive video quality evaluation model, respectively extract features of the texture video blocks, texture key frames, and depth key frames of multiple viewpoints through the visual information encoding module to obtain corresponding visual features; input the visual features into the spatio-temporal mapping module to obtain temporal visual markers and spatial visual markers; encode the instruction information and six-degree-of-freedom viewpoint position information through the language encoder to obtain text instruction markers and viewpoint position markers; combine the temporal visual markers, spatial visual markers, text instruction markers, and viewpoint position markers to obtain combined markers, and input the combined markers into the speech decoder to obtain the immersive video quality score.
[0084] The specific implementation of each module of an immersive video quality evaluation device guided by six-degree-of-freedom information is the same as that of an immersive video quality evaluation method guided by six-degree-of-freedom information, and will not be repeated in this embodiment.
[0085] See Figure 4 Shown is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. As Figure 4 shown, the electronic device of this embodiment includes: a processor 401 and a memory 402; wherein the memory 402 is used to store computer execution instructions; the processor 401 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. For specific details, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0086] Optionally, the memory 402 can be either independent or integrated with the processor 401.
[0087] When the memory 402 is independently provided, the electronic device further includes a bus 403 for connecting the memory 402 and the processor 401.
[0088] An embodiment of the present invention further provides a computer storage medium, in which computer-executable instructions are stored. When the processor 401 executes the computer-executable instructions, the above method is implemented.
[0089] An embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 401, the above method is implemented.
[0090] In the embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical or other forms.
[0091] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0092] In addition, in each embodiment of the present invention, the various functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The units formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0093] The above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or the processor 401 to execute some steps of the methods in various embodiments of the present application.
[0094] It should be understood that the above-mentioned processor 401 may be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor 401 may also be any conventional processor 401, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being completed by the hardware processor 401, or by a combination of the hardware and software modules in the processor 401.
[0095] The memory 402 may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk, or an optical disc, etc.
[0096] The bus 403 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus 403 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the bus 403 in the drawings of this application is not limited to only one bus 403 or one type of bus 403.
[0097] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0098] An exemplary storage medium is coupled to the processor 401, enabling the processor 401 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 401. The processor 401 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 401 and the storage medium can also exist as discrete components in an electronic device or a master device.
[0099] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An immersive video quality evaluation method based on six-degree-of-freedom information guidance, characterized in that: include: An immersive video quality assessment model guided by six degrees of freedom information is constructed and trained to obtain a trained immersive video quality assessment model; the immersive video quality assessment model includes a visual information encoding module, a spatiotemporal mapping module, a language encoder and a speech decoder of a large language model; Acquire an immersive video including texture videos and depth videos from multiple viewpoints, extract a number of texture video blocks and texture key frames from the texture videos from multiple viewpoints, and extract a number of depth key frames from the depth videos from multiple viewpoints; Several texture video blocks, texture key frames and depth key frames are input into the trained immersive video quality evaluation model, and the texture video blocks, texture key frames and depth key frames of multiple viewpoints are respectively extracted with the visual information encoding module to obtain the corresponding visual features; the visual features are input into the spatiotemporal mapping module to obtain the temporal visual mark and the spatial visual mark; the instruction information and the six-degree-of-freedom viewpoint position information are encoded with the language encoder to obtain the text instruction mark and the viewpoint position mark; The temporal visual marker, the spatial visual marker, the text instruction marker and the viewpoint position marker are combined to obtain a combined marker, and the combined marker is input into a speech decoder to obtain an immersive video quality score; The visual features are input into the spatiotemporal mapping module to obtain temporal visual tags and spatial visual tags, including: The depth key frame features and texture key frame features are input into the space mapping unit of the spatiotemporal mapping module to obtain spatial visual tags, and the texture video block features are input into the time mapping unit to obtain temporal visual tags, as shown in the following formula: Wherein, SP(·) represents the spatial mapping unit; TP(·) represents the temporal mapping unit; Represents spatial visual markers; Indicates a visual marker of time; F t i Represents texture keyframe features; Represents the deep keyframe feature; Concat represents feature concatenation; Representing texture video block features; The command information and the six-degree-of-freedom viewpoint position information are encoded by a language encoder to obtain text command tags and viewpoint position tags, which specifically include: The instruction information is encoded by the language encoder; the instruction information includes system instructions, guide instructions and response restrictions, the system instructions are used to inform the large language model of its functions, the guide instructions are used to inform the large language model of its specific tasks, and the response restrictions are used to inform the large language model of the specific requirements for answering questions; The text instruction tag obtained after encoding the instruction information is as shown in the following formula: F text =LE(I text ); Among them, I text Indicates command information; LE(·) indicates language encoder; F text Represents a text instruction tag; The six-degree-of-freedom viewpoint position information is encoded by the language encoder to obtain a viewpoint position mark, as shown in the following formula: in, Represents the position information of viewpoint i; Represents the position tag of viewpoint i.
2. The immersive video quality assessment method based on six-degree-of-freedom information guidance according to claim 1, characterized in that: A plurality of texture video blocks and texture key frames are extracted from texture videos of multiple viewpoints, and a plurality of depth key frames are extracted from depth videos of multiple viewpoints, as follows: Texture video for each viewpoint and Deep Video Divide the video into blocks to obtain K consecutive texture video blocks and K consecutive depth video blocks; wherein, represents the j-th texture video block, represents the jth depth video block, N represents the number of frames of texture video or depth video for each viewpoint, the number of frames of texture video or depth video for each viewpoint is equal, n represents the nth frame of each viewpoint, τ represents the number of video frames contained in each texture video block or depth video block, For each texture video block and depth video block, take the first frame as the key frame of the block to get the texture key frame f tj =x τj and the depth keyframe f dj =y τj .
3. The immersive video quality assessment method based on six-degree-of-freedom information guidance according to claim 2, characterized in that: The visual information encoding module extracts features of texture video blocks, texture key frames and depth key frames from multiple viewpoints to obtain corresponding visual features, including: The deep visual encoder and spatial visual encoder of the visual information encoding module are used to encode the deep key frame f dj =y τj and texture keyframe f tj =x τj Perform feature extraction separately to obtain the deep key frame features of each viewpoint and texture keyframe features The temporal visual encoder of the visual information encoding module extracts features of the texture video block and obtains the texture video block features of multiple viewpoints. As shown below: F t i =VE spatial (Concat(f t1 ,f t2 ,...,f tK )); Among them, VE deep (·), VE spatial (·) and VE temporal (·) represents the deep visual encoder, spatial visual encoder, and temporal visual encoder, respectively; i represents the i-th viewpoint, H, W, and C represent the height, width, and number of channels of the feature, respectively. Represents the set of real numbers.
4. The immersive video quality assessment method based on six-degree-of-freedom information guidance according to claim 3 is characterized in that: The resulting immersive video quality score is as follows: Where LD(·) represents the speech decoder of the large language model; Q represents the obtained immersive video quality score; Represents the synthetic features of each viewpoint.
5. An immersive video quality assessment device based on six-degree-of-freedom information guidance, characterized in that: include: A model building and training module is configured to build and train an immersive video quality assessment model guided by six degrees of freedom information to obtain a trained immersive video quality assessment model; the immersive video quality assessment model includes a visual information encoding module, a spatiotemporal mapping module, a language encoder and a speech decoder of a large language model; A data extraction module is configured to obtain an immersive video including texture videos and depth videos of multiple viewpoints, extract a number of texture video blocks and texture key frames from the texture videos of the multiple viewpoints, and extract a number of depth key frames from the depth videos of the multiple viewpoints; The video quality assessment module is configured to input a plurality of texture video blocks, texture key frames and depth key frames into a trained immersive video quality assessment model, extract features of the texture video blocks, texture key frames and depth key frames of multiple viewpoints respectively through a visual information encoding module to obtain corresponding visual features; input the visual features into a spatiotemporal mapping module to obtain temporal visual tags and spatial visual tags; encode the instruction information and the six-degree-of-freedom viewpoint position information through a language encoder to obtain text instruction tags and viewpoint position tags; The temporal visual marker, the spatial visual marker, the text instruction marker and the viewpoint position marker are combined to obtain a combined marker, and the combined marker is input into a speech decoder to obtain an immersive video quality score; The visual features are input into the spatiotemporal mapping module to obtain temporal visual tags and spatial visual tags, including: The depth key frame features and texture key frame features are input into the space mapping unit of the spatiotemporal mapping module to obtain spatial visual tags, and the texture video block features are input into the time mapping unit to obtain temporal visual tags, as shown in the following formula: Wherein, SP(·) represents the spatial mapping unit; TP(·) represents the temporal mapping unit; Represents spatial visual markers; Indicates a visual marker of time; F t i Represents texture keyframe features; Represents the deep keyframe feature; Concat represents feature concatenation; Representing texture video block features; The command information and the six-degree-of-freedom viewpoint position information are encoded by a language encoder to obtain text command tags and viewpoint position tags, which specifically include: The instruction information is encoded by the language encoder; the instruction information includes system instructions, guide instructions and response restrictions, the system instructions are used to inform the large language model of its functions, the guide instructions are used to inform the large language model of its specific tasks, and the response restrictions are used to inform the large language model of the specific requirements for answering questions; The text instruction tag obtained after encoding the instruction information is as shown in the following formula: F text =LE(I text ); Among them, I text Indicates command information; LE(·) indicates language encoder; F text Represents a text instruction tag; The six-degree-of-freedom viewpoint position information is encoded by the language encoder to obtain a viewpoint position mark, as shown in the following formula: in, Represents the position information of viewpoint i; Represents the position tag of viewpoint i.
6. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Objective quality evaluation method for window six-degree-of-freedom synthetic video
CN114745546A
Immersive video quality evaluation method and device based on multi-feature network
CN118506168A