Three-dimensional streaming media video quality evaluation method and device, equipment and storage medium

By extracting the two-dimensional video and audio information of three-dimensional streaming videos, and using multi-modal large models for quality evaluation and regression analysis, the problem of low accuracy in the quality evaluation of three-dimensional streaming videos is solved, achieving more efficient evaluation results.

CN120355647APending Publication Date: 2025-07-22PENG CHENG LAB
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510275100.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art has low accuracy and poor efficiency in three-dimensional streaming video quality evaluation, and it is impossible to effectively utilize information in depth dimensions.

Method used

By obtaining three-dimensional streaming videos, extracting two-dimensional video and audio information, using multi-modal large models for quality evaluation, integrating video features, and comparing audio and video synchronization features and audio features for quality regression analysis.

Benefits of technology

It improves the accuracy and efficiency of 3D streaming video quality evaluation, can analyze video quality in multiple dimensions, and improves the accuracy of evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355647A_ABST
    Figure CN120355647A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a three-dimensional streaming media video quality evaluation method and device, equipment and a storage medium, and belongs to the technical field of image processing. The method comprises the following steps: acquiring an initial three-dimensional streaming media video to be evaluated; extracting two-dimensional video information of the target object from the initial three-dimensional streaming media video, and extracting audio information from the initial three-dimensional streaming media video; based on a plurality of different types of multi-modal large models, performing quality evaluation processing on the two-dimensional video information, and fusing a plurality of evaluation features obtained after the quality evaluation processing to obtain video features; comparing the consistency of the two-dimensional video information and the audio information to obtain audio and picture synchronization features, and performing short-time analysis processing on the audio information to obtain audio features; and performing quality regression analysis on the initial three-dimensional streaming media video based on the video features, the audio and video synchronization features and the audio features to obtain a target quality score. According to the invention, the accuracy and efficiency of three-dimensional streaming media video quality evaluation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and particularly to a method, device, equipment and storage medium for evaluating the quality of three-dimensional streaming media video. Background Art

[0002] Three-dimensional streaming media is a technology that sends three-dimensional data (such as 3D models, scenes or animations) from the server side to the client side in a streaming transmission manner. With the continuous development of Internet technology and the continuous improvement of user interaction requirements, in practical applications, videos generated by three-dimensional streaming media are often used to achieve immersive video communication.

[0003] The evaluation of video communication quality involved in the related technology focuses on two-dimensional videos, which only contain information in two dimensions, horizontal and vertical. However, three-dimensional streaming media videos also involve information in the depth dimension. The large amount and richness of data and the strong continuity between data of different dimension information make the related technology have problems of low accuracy and poor efficiency in evaluating the quality of three-dimensional streaming media videos. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a method, device, equipment and storage medium for evaluating the quality of three-dimensional streaming media video, aiming to improve the accuracy and efficiency of evaluating the quality of three-dimensional streaming media video.

[0005] To achieve the above object, the first aspect of the embodiments of the present application proposes a method for evaluating the quality of three-dimensional streaming media video, including:

[0006] Obtain an initial three-dimensional streaming media video to be evaluated, where the initial three-dimensional streaming media video includes at least one target object in a three-dimensional space;

[0007] Extract two-dimensional video information of the target object from the initial three-dimensional streaming media video, and extract audio information from the initial three-dimensional streaming media video;

[0008] Based on multiple different types of multimodal large models, respectively perform quality evaluation processing on the two-dimensional video information, and fuse the multiple evaluation features obtained after the quality evaluation processing to obtain video features;

[0009] Compare the consistency between the two-dimensional video information and the audio information to obtain an audio-visual synchronization feature, and perform short-time analysis processing on the audio information to obtain an audio feature;

[0010] Based on the video features, audio-visual synchronization features and audio features, perform quality regression analysis on the initial three-dimensional streaming media video to obtain a target quality score.

[0011] In some embodiments, extracting two-dimensional video information of the target object from the initial three-dimensional streaming media video includes:

[0012] Locate at least one target object from the initial three-dimensional streaming media video;

[0013] Determine the geometric center of the target object and establish a three-dimensional coordinate system based on the geometric center;

[0014] Project the initial three-dimensional streaming media video onto a two-dimensional plane based on the three-dimensional coordinate system to obtain the corresponding two-dimensional video information of the target object.

[0015] In some embodiments, based on multiple different types of multimodal large models, perform quality assessment processing on the two-dimensional video information respectively, and fuse multiple evaluation features obtained after the quality assessment processing to obtain video features, including:

[0016] Construct a unified prompt for multiple different types of multimodal large models;

[0017] Based on a preset sampling frequency, perform frame sampling processing on the two-dimensional video information to obtain a target set, where the target set contains target image frames at multiple different sampling times;

[0018] Input the target set into multiple different types of multimodal large models respectively, and according to the prompt, guide each multimodal large model to perform quality assessment processing on the target set to obtain multiple evaluation features corresponding to the multiple target image frames;

[0019] Fuse multiple evaluation features output by different types of multimodal large models respectively to obtain the video features of the two-dimensional video information.

[0020] In some embodiments, fuse multiple evaluation features output by different types of multimodal large models respectively to obtain the video features of the two-dimensional video information, including:

[0021] Based on the Ebbinghaus forgetting curve, fuse multiple evaluation features output by different multimodal large models respectively to obtain the video features corresponding to the two-dimensional video information;

[0022] Among them, the Ebbinghaus forgetting curve includes fusion weights at multiple different sampling times, and the fusion weight at the current sampling time is greater than the fusion weight at the historical sampling time.

[0023] In some embodiments, compare the consistency between the two-dimensional video information and the audio information to obtain the audio-visual synchronization feature, including:

[0024] When the target object indicated by the two-dimensional video information is a speaking face, based on a preset synchronization network, compare the lip movement of the speaking face with the audio information to obtain the audio-lip synchronization error distance value and the audio-lip synchronization error confidence level;

[0025] Based on the audio-lip synchronization error distance value and the audio-lip synchronization error confidence level, an audio-visual synchronization feature is obtained.

[0026] In some embodiments, short-time analysis processing is performed on the audio information to obtain audio features, including:

[0027] Determine the speech waveform time-domain signal corresponding to the target image frame from the audio information;

[0028] Determine the short-time energy of the audio information according to the speech waveform time-domain signal, and determine the short-time zero-crossing rate of the audio information according to the speech waveform time-domain signal;

[0029] Compare the differences between the short-time energy and the short-time zero-crossing rate and a preset threshold to obtain audio features.

[0030] In some embodiments, based on the video features, the audio-visual synchronization features, and the audio features, quality regression analysis is performed on the initial three-dimensional streaming media video to obtain a target quality score, including:

[0031] In the first half of the initial three-dimensional streaming media video, quality regression analysis is performed on the video features, the audio-visual synchronization features, and the audio features according to a preset first weight set to obtain a first quality score;

[0032] In the second half of the initial three-dimensional streaming media video, quality regression analysis is performed on the video features, the audio-visual synchronization features, and the audio features according to a preset second weight set to obtain a second quality score;

[0033] Based on the first quality score and the second quality score, a target quality score is obtained, where the first weight set includes a first audio-visual synchronization weight corresponding to the audio-visual synchronization feature, the second weight set includes a second audio-visual synchronization weight corresponding to the audio-visual synchronization feature, and the first audio-visual synchronization weight is greater than the second audio-visual synchronization weight.

[0034] To achieve the above object, a second aspect of the embodiments of the present application proposes a three-dimensional streaming media video quality evaluation device, including:

[0035] An acquisition module, configured to acquire an initial three-dimensional streaming media video to be evaluated, where the initial three-dimensional streaming media video includes at least one target object in a three-dimensional space;

[0036] An extraction module, configured to extract two-dimensional video information of the target object from the initial three-dimensional streaming media video, and extract audio information from the initial three-dimensional streaming media video;

[0037] A first feature determination module, configured to respectively perform quality evaluation processing on the two-dimensional video information based on multiple different types of multimodal large models, and fuse multiple evaluation features obtained after the quality evaluation processing to obtain video features;

[0038] A second feature determination module, configured to compare the consistency between two-dimensional video information and audio information to obtain an audio-visual synchronization feature, and perform short-time analysis and processing on the audio information to obtain an audio feature;

[0039] A target quality scoring module, configured to perform quality regression analysis on the initial three-dimensional streaming media video based on the video feature, the audio-visual synchronization feature, and the audio feature to obtain a target quality score.

[0040] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the three-dimensional streaming media video quality evaluation method in the first aspect above is implemented.

[0041] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the three-dimensional streaming media video quality evaluation method in the first aspect above is implemented.

[0042] The three-dimensional streaming media video quality evaluation method, device, equipment, and storage medium provided by the present application obtain an initial three-dimensional streaming media video to be evaluated, where the initial three-dimensional streaming media video includes at least one target object in a three-dimensional space; extract two-dimensional video information of the target object from the initial three-dimensional streaming media video, and extract audio information from the initial three-dimensional streaming media video; based on multiple different types of multimodal large models, perform quality evaluation processing on the two-dimensional video information respectively, and fuse the multiple evaluation features obtained after the quality evaluation processing to obtain a video feature; in the case where the initially obtained initial three-dimensional streaming media video contains complex and a large amount of three-dimensional feature data, convert the three-dimensional data into two-dimensional data and utilize the efficient data processing ability of the multimodal large model itself to perform quality evaluation processing on the two-dimensional video information, which can reduce the computational complexity and improve the quality evaluation efficiency; then, compare the consistency between the two-dimensional video information and the audio information to obtain an audio-visual synchronization feature, and perform short-time analysis and processing on the audio information to obtain an audio feature; based on the video feature, the audio-visual synchronization feature, and the audio feature, perform quality regression analysis on the initial three-dimensional streaming media video to obtain a target quality score. In this way, by performing quality regression analysis on the initial three-dimensional streaming media video in a multi-dimensional and comprehensive manner, the accuracy of the finally obtained target quality score can be improved. Description of the Drawings

[0043] Figure 1 is a schematic diagram of an application scenario of the three-dimensional streaming media video quality evaluation device provided by the embodiments of the present application;

[0044] Figure 2 is an optional flowchart of the three-dimensional streaming media video quality evaluation method provided by the embodiments of the present application;

[0045] Figure 3 It is an optional quality evaluation schematic diagram of the three-dimensional streaming media video quality evaluation method provided by the embodiments of the present application;

[0046] Figure 4 It is an optional Ebbinghaus forgetting curve schematic diagram of the three-dimensional streaming media video quality evaluation method provided by the embodiments of the present application;

[0047] Figure 5 It is another optional quality evaluation schematic diagram of the three-dimensional streaming media video quality evaluation method provided by the embodiments of the present application;

[0048] Figure 6 It is an optional schematic diagram of the three-dimensional streaming media video quality evaluation device provided by the embodiments of the present application;

[0049] Figure 7 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0051] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0053] First, several nouns involved in the present application are analyzed:

[0054] Three-dimensional data (3D Data) refers to data that describes an object, scene or phenomenon in three-dimensional space, which contains information in three dimensions (usually length, width, and height) and is used to represent the spatial position, shape, structure or other related attributes of an object.

[0055] Immersive video communication is a new type of communication that combines three-dimensional vision, spatial audio, and real-time interaction technologies. It aims to provide users with an immersive experience as if they were in the same physical space as remote participants. It goes beyond traditional two-dimensional video calls and significantly enhances the realism and naturalness of remote communication through richer sensory information and interaction means.

[0056] Three-dimensional mesh body, three-dimensional mesh flow: Three-dimensional mesh flow is a three-dimensional geometric data stream transmitted in a time series form. It consists of a series of continuous three-dimensional mesh bodies. Each three-dimensional mesh body represents the geometric structure of a three-dimensional object or scene at a certain moment and contains information such as vertices, edges, and faces. Three-dimensional mesh flow is usually used to transmit dynamic three-dimensional content in real-time or near real-time.

[0057] Three-dimensional streaming media is a technology that sends three-dimensional data from the server side to the client side in a streaming transmission manner. With the continuous development of Internet technology and the continuous improvement of user interaction requirements, in practical applications, videos generated by three-dimensional streaming media are often used to achieve immersive video communication.

[0058] The video communication quality assessment involved in related technologies focuses on two-dimensional videos, which only contain information in the horizontal and vertical dimensions. However, three-dimensional streaming media videos also involve information in the depth dimension. The large amount and richness of data, as well as the strong continuity between data in different dimensions, lead to problems such as low accuracy and poor efficiency in the application of related technologies to three-dimensional streaming media videos.

[0059] Based on this, the embodiments of the present application provide a method, device, equipment, and storage medium for three-dimensional streaming media video quality assessment, aiming to improve the accuracy and efficiency of three-dimensional streaming media video quality assessment.

[0060] Exemplarily, as Figure 1 shown, Figure 1FIG. 0 is a schematic diagram of an application scenario of a three-dimensional streaming media video quality evaluation device provided by an embodiment of the present application. In an optional application scenario, a client 11 is communicatively connected to a server 12, and the three-dimensional streaming media video quality evaluation device proposed by the embodiment of the present application is deployed in the server 12. Among them, the three-dimensional streaming media video quality evaluation device obtains an initial three-dimensional streaming media video to be evaluated, and the initial three-dimensional streaming media video includes at least one target object in a three-dimensional space; then, two-dimensional video information of the target object is extracted from the initial three-dimensional streaming media video, and audio information is extracted from the initial three-dimensional streaming media video; then, based on multiple different types of multimodal large models, the two-dimensional video information is respectively subjected to quality evaluation processing, and multiple evaluation features obtained after the quality evaluation processing are fused to obtain video features; after that, the consistency between the two-dimensional video information and the audio information is compared to obtain a lip-sync feature, and the audio information is subjected to short-time analysis processing to obtain audio features; finally, based on the video features, the lip-sync feature, and the audio features, a quality regression analysis is performed on the initial three-dimensional streaming media video to obtain a target quality score. The server 12 may send the target quality score to the client 11 so that the client 11 can further optimize the three-dimensional streaming media video based on the target quality score in the future.

[0061] It should be noted that in the embodiment of the present application, when it comes to obtaining user permission or consent in advance when it is necessary to obtain information related to user characteristics such as user basic information or user identity, and moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain sensitive personal information of users, it will first obtain the user's separate permission or separate consent. After clearly obtaining the user's separate permission or separate consent, the necessary data for the normal operation of the embodiment of the present application will be obtained. For example, before obtaining the three-dimensional streaming media video to be evaluated, the embodiment of the present application will first obtain the consent of the relevant personnel involved in the three-dimensional streaming media video and the relevant regulatory personnel, otherwise the three-dimensional streaming media video that cannot be applied to the embodiment of the present application will be obtained. In addition, other relevant data obtained by the three-dimensional streaming media video quality evaluation device in the embodiment of the present application are all authorized data, which will not be elaborated here one by one.

[0062] In the embodiment of the present application, the description will be made from the dimension of a three-dimensional streaming media video quality evaluation device (for the convenience of description, hereinafter may also be simply referred to as the "quality evaluation device"), and this quality evaluation device may be integrated in a computer device, such as a server. As Figure 2 shown, Figure 2 FIG. 10 is an optional flowchart of a three-dimensional streaming media video quality evaluation method provided by an embodiment of the present application. Figure 2The method in Figure 2 does not specifically limit the order of steps 101 to 105 in

[0063] Step 101: Obtain the initial 3D streaming video to be evaluated, where the initial 3D streaming video includes at least one target object in a three-dimensional space.

[0064] The following provides a detailed description of Step 101.

[0065] Among them, 3D Streaming Video is a type of dynamic content presented in three dimensions, which is transmitted to users in real-time or near real-time in a streaming manner over the network. Compared with traditional 2D videos, 3D streaming videos not only contain information about length and width but also depth information (i.e., the third dimension), enabling a more realistic restoration of scenes and objects in the real world. In the embodiments of this application, a 3D streaming video that has not undergone quality evaluation or other processing is referred to as an initial 3D streaming video.

[0066] Among them, Streaming refers to a data transmission technology that allows data to be processed or played before it has been fully downloaded, and is commonly used for real-time or near real-time transmission of audio, video, or other continuous media content, enabling users to start experiencing the content without waiting for the entire file to be downloaded.

[0067] Among them, the target object refers to an entity or element existing in a three-dimensional space. The target object can be a person, an object, a part of a scene, or anything that can be three-dimensionally modeled and presented. In the initial 3D streaming video, the target object refers to the main body with three-dimensional characteristics in the video content.

[0068] Furthermore, the initial 3D streaming video can be obtained in real-time by the user through a special device, or can be pre-recorded based on a special device. The specific method for obtaining the initial 3D streaming video can be set according to the actual situation, and this application does not limit this. Among them, the special device can be a virtual reality device (Virtual Reality Devices, VR), an augmented reality device (Augmented Reality Devices, AR), or other devices that provide perceptual interaction for users to experience the real environment. Of course, the special device for obtaining the initial 3D streaming video can also be set according to the actual situation, and this application also does not limit this.

[0069] Exemplarily, a multinational company needs to hold an important remote meeting. The participants (target objects) are distributed all over the world. The company uses three-dimensional streaming media video to achieve immersive video communication, improving communication efficiency and realism. Each participant wears a special device to enter the virtual meeting room. In the virtual meeting, participants can see the three-dimensional images of others and feel the spatial positions of their actions, expressions, and voices. In this way, multiple participants in different geographical locations seem to be in the same physical space, reducing the "screen barrier" of traditional video conferences.

[0070] Step 102: Extract the two-dimensional video information of the target object from the initial three-dimensional streaming media video, and extract the audio information from the initial three-dimensional streaming media video.

[0071] The following provides a detailed description of step 102.

[0072] Among them, the two-dimensional video information refers to the visual content of the target object extracted from the three-dimensional streaming media video and presented in a planar form. Essentially, it is the result of a dimensionality reduction process. The audio information refers to the sounds emitted by the target object and the background sound effects in the environment.

[0073] In some embodiments, extracting the two-dimensional video information of the target object from the initial three-dimensional streaming media video includes the following steps:

[0074] (102.a.1) Locate at least one target object in the initial three-dimensional streaming media video;

[0075] (102.a.2) Determine the geometric center of the target object and establish a three-dimensional coordinate system based on the geometric center;

[0076] (102.a.3) Project the initial three-dimensional streaming media video onto a two-dimensional plane based on the three-dimensional coordinate system to obtain the corresponding two-dimensional video information of the target object.

[0077] The following provides a detailed description of steps (102.a.1) to (102.a.3).

[0078] In some embodiments, at least one target object can be located by the depth difference between the foreground and the background. Exemplarily, in a virtual meeting scenario, a depth camera is used to capture the three-dimensional data of the participants, and the three-dimensional models of each person are extracted from the background by analyzing the depth map to achieve the positioning of the target participants.

[0079] In some embodiments, as Figure 3 shown Figure 3It is an optional quality assessment schematic diagram of the 3D streaming media video quality assessment method provided by the embodiments of the present application. Since the face-to-face speaking face is the most common scenario in immersive video communication, the embodiments of the present application will take the target object in the initial 3D streaming media video as the speaking face as an example and combine Figure 3 to explain (including the relevant content from "the first step" to "the seventh step") so that readers can better understand the specific implementation manner of this case and the corresponding beneficial effects.

[0080] First step, given the 3D mesh body and 3D mesh flow of the speaking face, so as to establish a 3D coordinate system by calculating the geometric center of the speaking face later:

[0081] M = {P, V, E}

[0082] S = {M i , i = 1, 2,... L}

[0083] Among them, P represents the set containing all vertices of the texture network, V represents the set containing all normal vectors of the texture network, and E represents the set containing all edges of the texture network. Specifically, a single 3D mesh body M composed of P, V, and E; M is arranged in chronological order and can further form a 3D mesh flow S, and M i represents the i-th 3D mesh body in S, and the total number of 3D mesh bodies is L.

[0084] Next, calculate the geometric center for each speaking face M in the 3D mesh flow S:

[0085]

[0086] Among them, N represents the total number of mesh bodies, and C F represents the geometric center of the speaking face, and pi represents the i-th vertex of the texture network; after C F is determined, a 3D coordinate system is established with the positive half-axis of the x-axis and the positive half-axis of the y-axis on the front of the speaking face, and the top direction as the positive half-axis of the z-axis.

[0087] Furthermore, the front projection of each 3D mesh body can be realized by using toolkits (such as the Open3D package) in programming tools (such as Python). Among them, Python is a high-level programming language with concise and easy-to-read syntax and a powerful ecosystem, and has rich third-party libraries and can support a variety of application scenarios; the Open3D package is an open-source library in Python that focuses on 3D data processing and visualization.

[0088] Second step, based on the established 3D coordinate system, collect the front projection of each 3D mesh body:

[0089] Ii = F(M i )

[0090] where F(·) represents the acquisition process of the front projection, and M i represents the i-th three-dimensional grid body in S, and I i represents the two-dimensional video information obtained after the front projection.

[0091] Furthermore, the two-dimensional video information can be further processed. For example, the two-dimensional video information can be randomly cropped in regions to obtain updated two-dimensional video information, so as to reduce the computational complexity and consumption of computing resources. The random region cropping includes but is not limited to fixed-size cropping, random ratio cropping, random center cropping, cropping based on the target object, multi-region cropping and merging, and adaptive cropping, etc. The specific random region cropping method can be set according to the actual situation.

[0092] Step 103: Based on multiple different types of multi-modal large models, perform quality assessment processing on the two-dimensional video information respectively, and fuse the multiple evaluation features obtained after the quality assessment processing to obtain video features.

[0093] The following provides a detailed description of Step 103.

[0094] Among them, the multi-modal large model refers to a large artificial intelligence model with a huge number of parameters, rich training data, powerful functions, and capable of simultaneously processing and understanding multiple types of data (such as text, images, audio, video, etc.). The multi-modal large model can better simulate the comprehensive perception and processing method of humans for complex information, so as to efficiently and accurately process the complex tasks input.

[0095] Furthermore, in the case where the initially acquired initial three-dimensional streaming media video contains complex and a large amount of three-dimensional feature data, the embodiment of the present application first converts the three-dimensional data into two-dimensional data, and then uses the efficient data processing ability of the multi-modal large model itself to perform quality assessment processing on the two-dimensional video information, so as to reduce the computational complexity.

[0096] Furthermore, multiple different types of multi-modal large models can respectively perform collaborative processing on the quality assessment of the two-dimensional video information from different dimensions, effectively avoiding the evaluation limitations caused by using only a single large model for quality assessment, giving full play to the advantages of each multi-modal large model, and improving the accuracy of the quality assessment of the initial three-dimensional streaming media video while reducing the computational complexity and improving the quality assessment efficiency.

[0097] In some other embodiments, in addition to inputting two-dimensional video information into multiple multi-modal large models respectively, initial three-dimensional streaming media video or supplementary data related to the two-dimensional video information can also be input simultaneously. Since some depth information or spatial structure information is often lost when three-dimensional data is converted into two-dimensional data, the supplementary data can make up for the information loss caused in this process, thereby helping multiple multi-modal large models to more comprehensively understand the two-dimensional video information and capture more detailed features when evaluating the quality of the two-dimensional video information, so as to improve the accuracy and reliability of the finally obtained target quality score.

[0098] Among them, the supplementary data can be the depth map, semantic label, text description information, etc. of the initial three-dimensional streaming media video or two-dimensional video information. The specific form of the supplementary data can be set according to the actual situation, and the embodiments of the present application do not limit this.

[0099] In some embodiments, the specific types of the multi-modal large models can be:

[0100] (1) Generative Pre-trained Transformer (GPT), which is developed by OpenAI. The GPT series of large models are known for their powerful text generation capabilities and wide applicability, and can complete various complex reasoning tasks such as writing, programming, and image feature extraction; Exemplarily, GPT-4o can be selected as a multi-modal large model in the embodiments of the present application.

[0101] (2) Gemini Models (Gemini), which are famous for their powerful cross-modal capabilities and can simultaneously process various types of data such as text, images, and audio, and are suitable for complex multi-task scenarios. Exemplarily, Gemini-1.5-Pro can be selected as a multi-modal large model in the embodiments of the present application.

[0102] (3) Large Language and Vision Alignment Model (LLaVA), which is an advanced artificial intelligence model that combines natural language processing and computer vision technologies. LLaVA can simultaneously understand and generate various types of data, such as text, images, audio, etc., so as to achieve cross-modal task processing; Exemplarily, LLaVA-Next can be selected as a multi-modal large model in the embodiments of the present application.

[0103] Furthermore, the embodiments of the present application do not limit the number of multi-modal large models. In a feasible implementation, the number of multi-modal large models under different types is 30. Additionally, it should be noted that in practical applications, other types of multi-modal large models can also be selected as one of the multiple different types of multi-modal large models in the embodiments of the present application. The above are only some types of multi-modal large models listed for the convenience of readers' understanding, and do not represent limitations in the embodiments of the present application.

[0104] In some embodiments, based on multiple different types of multi-modal large models, quality assessment processing is respectively performed on two-dimensional video information, and multiple evaluation features obtained after the quality assessment processing are fused to obtain video features, including the following steps:

[0105] (103.a.1) Construct a unified prompt for multiple different types of multi-modal large models;

[0106] (103.a.2) Based on a preset sampling frequency, perform frame sampling processing on the two-dimensional video information to obtain a target set, where the target set contains target image frames at multiple different sampling times;

[0107] (103.a.3) Input the target set into multiple different types of multi-modal large models respectively, and according to the prompt, guide each multi-modal large model to perform quality assessment processing on the target set to obtain multiple evaluation features corresponding to the multiple target image frames;

[0108] (103.a.4) Fuse multiple evaluation features respectively output by different types of multi-modal large models to obtain the video features of the two-dimensional video information.

[0109] The following describes steps (103.a.1) to (103.a.4) in detail.

[0110] Among them, the prompt is used to provide guiding information for multiple different types of multi-modal large models to guide the model to complete the quality assessment task. As Figure 3 shown, the prompt can be "How to evaluate the quality of this picture", and by clarifying the task requirements through the prompt, the efficiency of quality assessment can be improved. Moreover, the number of prompts can be one or more, and the specific prompt content and the number of prompts can be set according to the actual situation, and the embodiments of the present application do not limit this.

[0111] Among them, the sampling frequency refers to the frequency of extracting target image frames from a continuous signal per unit time to perform frame sampling processing on the two-dimensional video information based on a set time interval; the entire multiple sampled target image frames are obtained to form the target set.

[0112] Among them, the evaluation features refer to the specific features extracted after each multimodal large model performs quality evaluation processing on the target set (i.e., multiple target image frames) according to the prompt words, and these features are used to describe the quality-related attributes of a single target image frame. The specific manifestation form of the evaluation features can be a quality score. In some alternative embodiments, the quality score can be further subdivided into a video clarity score, a video coherence score, a video realism score, etc. Further, by fusing the evaluation features of different multimodal large models, video features that comprehensively reflect the overall quality of two-dimensional video information are obtained. Among them, the process of fusing evaluation features is also called group decision-making.

[0113] Next, following the "second step", continue to combine Figure 3 for example:

[0114] The third step is to Figure 3 as shown, perform frame sampling on the two-dimensional video information, conduct quality evaluation by guiding N multimodal large models based on a unified prompt word, and then perform group decision-making:

[0115] f j = Sampler(V)

[0116]

[0117] Among them, f j represents the j-th target image frame sampled, Sampler(·) represents the frame sampling processing function, and the sampling frequency in the embodiments of the present application is one frame per second; LMk(·) represents the process of the k-th multimodal large model performing quality evaluation, represents the process of group decision-making by N multimodal large models. prompt represents the prompt word used. Finally, the multimodal large model group makes a decision on the quality score q j of f j .

[0118] In some embodiments, to obtain the video features of two-dimensional video information by fusing multiple evaluation features corresponding to the outputs of different types of multimodal large models, the following steps are included:

[0119] (A.1) Based on the Ebbinghaus forgetting curve, fuse multiple evaluation features corresponding to the outputs of different multimodal large models to obtain the corresponding video features of two-dimensional video information; among them, the Ebbinghaus forgetting curve includes fusion weights at multiple different sampling times, and the fusion weight at the current sampling time is greater than the fusion weight at the historical sampling time.

[0120] The following provides a detailed description of step (A.1).

[0121] Among them, the Ebbinghaus forgetting curve describes the law of forgetting of new things by the human brain, reveals the process of memory gradually declining over time. In the embodiments of this application, the law of human forgetting is simulated to perform quality evaluation processing on target image frames, so as to improve the efficiency and accuracy of evaluation. As Figure 4 shown, Figure 4 is an optional schematic diagram of the Ebbinghaus forgetting curve for the three-dimensional streaming media video quality evaluation method provided by the embodiments of this application. Since the target set includes multiple target image frames at multiple sampling times, multiple multimodal large models will assign corresponding different fusion weights based on the Ebbinghaus forgetting curve for different sampling times when making a group decision. Exemplarily, there are two sampling times, such as sampling time A and sampling time B. The sampling time far from the origin (point 0) is called the historical sampling time, and the sampling time close to point 0 is called the current sampling time, and the fusion weight at the current sampling time is greater than the fusion weight at the historical sampling time.

[0122] Next, following "the third step", continue to combine Figure 3 for illustrative description:

[0123] Step 4, according to the group decision feedback of the multimodal large model, use the fusion weight w j discretely sampled by the Ebbinghaus forgetting curve to comprehensively weight the group decision quality feedback of each frame returned by the multimodal large model, so that the just-played streaming media part has a higher weight, thereby obtaining the video feature Q v after fusion processing:

[0124]

[0125] Among them, J represents the total number of target image frames, and w j conforms to the exponential decay of the Ebbinghaus forgetting curve and satisfies

[0126] It can be understood that the initially obtained initial three-dimensional streaming media video has the characteristic of dynamic transformation, and its content changes every moment over time. By assigning a higher fusion weight to the current sampling time, the quality evaluation device can pay more attention to the recent dynamic changes of the initial three-dimensional streaming media video, taking the time dimension into account, and improving the accuracy of quality evaluation.

[0127] Step 104, compare the consistency of the two-dimensional video information and the audio information to obtain the audio-visual synchronization feature, and perform short-time analysis processing on the audio information to obtain the audio feature.

[0128] The following will describe step 104 in detail.

[0129] Among them, the audio-visual synchronization feature refers to measuring the matching degree of visual content and auditory content in the video in time by comparing the consistency of two-dimensional video information and audio information. Short-time analysis and processing is a technology for segmenting and analyzing audio information to extract the features of audio within a short-time window and capture the local characteristics of audio information.

[0130] In some embodiments, obtaining the audio-visual synchronization feature by comparing the consistency of two-dimensional video information and audio information includes the following steps:

[0131] (104.a.1) When the target object indicated by the two-dimensional video information is a speaking face, based on a preset synchronization network, compare the lip movement of the speaking face with the consistency of the audio information to obtain a lip-audio synchronization error distance value and a lip-audio synchronization error confidence level;

[0132] (104.a.2) Based on the lip-audio synchronization error distance value and the lip-audio synchronization error confidence level, obtain the audio-visual synchronization feature.

[0133] The following will describe steps (104.a.1) to (104.a.2) in detail.

[0134] Among them, the lip-audio synchronization error distance value is used to measure the time deviation amount between the lip movement of the speaking face and the audio information, quantifying the synchronization degree between the lip movement (visual signal) and the speech (auditory signal). The lip-audio synchronization error confidence level is an evaluation index for the reliability of the lip-audio synchronization error distance value, which reflects whether the calculated error distance value is credible.

[0135] The following continues with the "fourth step" and continues to combine Figure 3 for illustrative examples:

[0136] Step 5: Use the SyncNet network to perform lip-audio consistency detection on the two-dimensional video information and the audio information to obtain the audio-visual synchronization feature:

[0137] Q s = Sync(V)

[0138] Among them, Sync(·) represents the process of performing lip-audio consistency detection on the speaking face video using the SyncNet network, and Q s is the extracted synchronization feature.

[0139] Specifically, first extract the Mel-frequency cepstral coefficient features from the audio information, and intercept the grayscale image rendering the mouth area of the speaking face. Subsequently, input the corresponding audio-visual information into the SyncNet network to output the lip-audio synchronization error distance value (LSE-D) and the lip-audio synchronization error confidence level (LSE-C), and then obtain the audio-visual synchronization feature.

[0140] Among them, the Mel Frequency Cepstra Coefficient (MFCC) is a voice feature, which is a cepstrum parameter extracted in the Mel scale frequency domain. The Mel scale describes the non-linear characteristics of the human ear's frequency. SyncNet is a deep learning network used to solve the audio-visual synchronization problem. It analyzes the temporal alignment relationship between audio and video signals to determine whether they are synchronized or calculate the offset between them.

[0141] In some embodiments, short-time analysis and processing of audio information to obtain audio features includes the following steps:

[0142] (104.b.1) Determine the speech waveform time-domain signal corresponding to the target image frame from the audio information;

[0143] (104.b.2) Determine the short-time energy of the audio information according to the speech waveform time-domain signal, and determine the short-time zero-crossing rate of the audio information according to the speech waveform time-domain signal;

[0144] (104.b.3) Compare the differences between the short-time energy and the short-time zero-crossing rate and a preset threshold to obtain audio features.

[0145] The following details steps (104.b.1) to (104.b.3).

[0146] Among them, the speech waveform time-domain signal refers to the original representation form of the audio signal on the time axis, which describes the change of sound over time; short-time analysis and processing is a method of dividing a long-time signal into several short-time segments for analysis. Short-time analysis simplifies the analysis process by decomposing the audio information into "quasi-stationary" parts within short-time windows; short-time energy describes the total energy of the audio information within a small period of time; short-time zero-crossing rate describes the number of times the signal crosses zero from positive to negative or from negative to positive within a small period of time.

[0147] The following continues with the "fifth step" and continues to combine Figure 3 for illustrative examples:

[0148] Sixth step, perform short-time analysis and processing on the audio to extract audio features:

[0149] Q c =ST(A)

[0150] Among them, Q cRepresents the extracted continuous features; ST(·) represents the statistical short-time energy E(i) and short-time zero-crossing rate Z(i) from the audio information, and compares the two features with preset thresholds. When both the short-time energy and the short-time zero-crossing rate are lower than the respectively set thresholds, it is considered that the speaker's face has buffer or freeze:

[0151]

[0152] Among them, E(i) represents the audio short-time energy of the i-th target image frame, Z(i) represents the audio short-time zero-crossing rate of the i-th target image frame, y i represents the speech waveform time-domain signal of the i-th target image frame, L represents the frame length; sgn is the abbreviation of the Sign Function, which is used to return the sign information of the input value without considering its specific numerical size.

[0153] Step 105, based on the video features, audio-visual synchronization features, and audio features, perform quality regression analysis on the initial 3D streaming media video to obtain the target quality score.

[0154] The following is a detailed description of Step 105.

[0155] In some embodiments, as Figure 5 shown, Figure 5 is another optional quality assessment schematic diagram of the 3D streaming media video quality assessment method provided by the embodiments of the present application. In the case of no true label for the speaker face quality assessment, the quality assessment device uses video features, audio-visual synchronization features, and audio features to perform quality regression analysis on the initial 3D streaming media video in a multi-dimensional and comprehensive manner, thereby improving the accuracy of the finally obtained target quality score.

[0156] In some embodiments, based on the video features, audio-visual synchronization features, and audio features, performing quality regression analysis on the initial 3D streaming media video to obtain the target quality score includes the following steps:

[0157] (105.a.1) In the first half of the initial 3D streaming media video, according to the preset first weight group, perform quality regression analysis on the video features, audio-visual synchronization features, and audio features to obtain the first quality score;

[0158] (105.a.2) In the second half of the initial 3D streaming media video, according to the preset second weight group, perform quality regression analysis on the video features, audio-visual synchronization features, and audio features to obtain the second quality score;

[0159] (105.a.3) Based on the first quality score and the second quality score, obtain the target quality score, where the first weight group includes the first audio-visual synchronization weight corresponding to the audio-visual synchronization feature, the second weight group includes the second audio-visual synchronization weight corresponding to the audio-visual synchronization feature, and the first audio-visual synchronization weight is greater than the second audio-visual synchronization weight.

[0160] The following provides a detailed description of steps (105.a.1) to (105.a.3).

[0161] Among them, the initial three-dimensional streaming media video can be divided into two half-segments in the order of occurrence along the time axis, including the first half-segment and the second half-segment. Moreover, quality regression analysis is respectively performed on the first half-segment and the second half-segment using different first weight groups and second weight groups, and then the final target quality score is obtained by superimposing the quality scores of the two parts.

[0162] Among them, the first weight group includes the first video feature, the first audio-visual synchronization feature, and the first audio feature; the second weight group includes the second video feature, the second audio-visual synchronization feature, and the second audio feature.

[0163] It can be understood that in the first half-segment of the initial three-dimensional streaming media video, the audio-visual synchronization feature often determines the audio-visual consistency of the entire video. Therefore, the attention to the audio-visual synchronization feature can be increased in the first half-segment; while in the second half-segment, the importance of the audio-visual synchronization feature decreases, and the overall content of the video dominates. Thus, the attention to the video feature can be increased and the attention to the audio-visual synchronization feature can be reduced.

[0164] The following continues with the "Sixth Step" and continues to combine Figure 3 for an example illustration:

[0165] Seventh step, perform quality regression analysis on the extracted features:

[0166] Q = SVR(Q1, Q2, Q3)

[0167] Among them, SVR(·) represents the support vector regression processing process, Q represents the predicted target quality score; Q1 is the video feature, Q2 is the audio-visual synchronization feature, and Q3 is the audio feature.

[0168] Among them, support vector regression is a method for quality regression analysis based on the support vector machine (SVM). It finds an optimal hyperplane in a high-dimensional space to maximize the interval between data points and the hyperplane, thereby realizing the prediction of continuous variables.

[0169] As Figure 6 shown, Figure 6FIG. 0 is an alternative schematic diagram of a three-dimensional streaming media video quality evaluation device provided by an embodiment of the present application. The three-dimensional streaming media video quality evaluation device includes the following modules 201 to 205:

[0170] An acquisition module 201, configured to acquire an initial three-dimensional streaming media video to be evaluated, where the initial three-dimensional streaming media video includes at least one target object in a three-dimensional space;

[0171] An extraction module 202, configured to extract two-dimensional video information of the target object from the initial three-dimensional streaming media video, and extract audio information from the initial three-dimensional streaming media video;

[0172] A first feature determination module 203, configured to perform quality evaluation processing on the two-dimensional video information respectively based on multiple different types of multimodal large models, and fuse multiple evaluation features obtained after the quality evaluation processing to obtain video features;

[0173] A second feature determination module 204, configured to compare the consistency between the two-dimensional video information and the audio information to obtain a lip-sync feature, and perform short-time analysis processing on the audio information to obtain audio features;

[0174] A target quality scoring module 205, configured to perform quality regression analysis on the initial three-dimensional streaming media video based on the video features, the lip-sync feature, and the audio features to obtain a target quality score.

[0175] The three-dimensional streaming media video quality evaluation method, device, equipment, and storage medium proposed in the present application acquire an initial three-dimensional streaming media video to be evaluated, where the initial three-dimensional streaming media video includes at least one target object in a three-dimensional space; extract two-dimensional video information of the target object from the initial three-dimensional streaming media video, and extract audio information from the initial three-dimensional streaming media video; perform quality evaluation processing on the two-dimensional video information respectively based on multiple different types of multimodal large models, and fuse multiple evaluation features obtained after the quality evaluation processing to obtain video features; in the case where the initially acquired initial three-dimensional streaming media video contains complex and a large amount of three-dimensional feature data, convert the three-dimensional data into two-dimensional data and utilize the efficient data processing ability of the multimodal large model itself to perform quality evaluation processing on the two-dimensional video information, which can reduce the computational complexity and improve the quality evaluation efficiency; then, compare the consistency between the two-dimensional video information and the audio information to obtain a lip-sync feature, and perform short-time analysis processing on the audio information to obtain audio features; perform quality regression analysis on the initial three-dimensional streaming media video based on the video features, the lip-sync feature, and the audio features to obtain a target quality score. In this way, by performing quality regression analysis on the initial three-dimensional streaming media video in a multi-dimensional and comprehensive manner, the accuracy of the finally obtained target quality score can be improved.

[0176] The specific implementation of the three-dimensional streaming media video quality evaluation device is basically the same as the specific embodiments of the above three-dimensional streaming media video quality evaluation method, and will not be elaborated here.

[0177] An embodiment of this application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above three-dimensional streaming media video quality evaluation method. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0178] As Figure 7 shown, Figure 7 is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of this application. The electronic device includes:

[0179] A processor 301, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this application;

[0180] A memory 302, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 302 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 302, and the processor 301 is called to execute the three-dimensional streaming media video quality evaluation method of the embodiments of this application;

[0181] An input / output interface 303, which is used to implement information input and output;

[0182] A communication interface 304, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);

[0183] A bus 305, which transmits information between various components of the device (such as the processor 301, the memory 302, the input / output interface 303, and the communication interface 304);

[0184] Among them, the processor 301, the memory 302, the input / output interface 303, and the communication interface 304 are communicatively connected to each other inside the device through the bus 305.

[0185] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned three-dimensional streaming media video quality evaluation method is implemented.

[0186] As a non-transitory computer-readable storage medium, a memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0187] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0188] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0189] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0190] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0191] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0192] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0193] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0194] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0195] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0196] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0197] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.

Claims

1. A three-dimensional streaming media video quality assessment method, characterized in that Including: Obtain an initial three-dimensional streaming media video to be evaluated, where the initial three-dimensional streaming media video includes at least one target object in a three-dimensional space; Extract two-dimensional video information of the target object from the initial three-dimensional streaming media video, and extract audio information from the initial three-dimensional streaming media video; Based on multiple different types of multimodal large models, perform quality evaluation processing on the two-dimensional video information respectively, and fuse multiple evaluation features obtained after the quality evaluation processing to obtain video features; Compare the consistency between the two-dimensional video information and the audio information to obtain an audio-visual synchronization feature, and perform short-time analysis processing on the audio information to obtain an audio feature; Based on the video features, the audio-visual synchronization feature, and the audio feature, perform quality regression analysis on the initial three-dimensional streaming media video to obtain a target quality score.

2. The three-dimensional streaming media video quality evaluation method according to claim 1, wherein The extracting two-dimensional video information of the target object from the initial three-dimensional streaming media video includes: Locate at least one of the target objects in the initial three-dimensional streaming media video; Determine the geometric center of the target object, and establish a three-dimensional coordinate system according to the geometric center; Project the initial three-dimensional streaming media video onto a two-dimensional plane based on the three-dimensional coordinate system to obtain the corresponding two-dimensional video information of the target object.

3. The three-dimensional streaming media video quality evaluation method according to claim 1, wherein The performing quality evaluation processing on the two-dimensional video information respectively based on multiple different types of multimodal large models, and fusing multiple evaluation features obtained after the quality evaluation processing to obtain video features includes: Construct a unified prompt for multiple different types of the multimodal large models; Perform frame sampling processing on the two-dimensional video information based on a preset sampling frequency to obtain a target set, where the target set includes target image frames at multiple different sampling times; Input the target set into multiple different types of the multimodal large models respectively, and guide each multimodal large model to perform quality evaluation processing on the target set according to the prompt to obtain multiple evaluation features corresponding to the multiple target image frames; Fuse multiple evaluation features output by different types of the multimodal large models to obtain the video features of the two-dimensional video information.

4. The three-dimensional streaming media video quality evaluation method according to claim 3, wherein The fusing multiple evaluation features output by different types of the multimodal large models to obtain the video features of the two-dimensional video information includes: Based on the Ebbinghaus forgetting curve, fuse multiple evaluation features output by different multimodal large models to obtain the corresponding video features of the two-dimensional video information; Among them, the Ebbinghaus forgetting curve includes fusion weights at multiple different sampling times, and the fusion weight at the current sampling time is greater than the fusion weight at the historical sampling time.

5. The three-dimensional streaming media video quality evaluation method according to claim 1, characterized in that The comparing the consistency between the two-dimensional video information and the audio information to obtain an audio-visual synchronization feature includes: When the target object indicated by the two-dimensional video information is a speaking face, compare the lip movement of the speaking face with the consistency of the audio information based on a preset synchronization network to obtain a lip-sync error distance value and a lip-sync error confidence level; Based on the audio-lip synchronization error distance value and the audio-lip synchronization error confidence level, the audio-visual synchronization feature is obtained.

6. The three-dimensional streaming media video quality evaluation method according to claim 1, characterized in that The short-time analysis and processing of the audio information to obtain audio features includes: Determining a speech waveform time-domain signal corresponding to the target image frame from the audio information; Determining the short-time energy of the audio information according to the speech waveform time-domain signal, and determining the short-time zero-crossing rate of the audio information according to the speech waveform time-domain signal; Comparing the differences between the short-time energy and the short-time zero-crossing rate and a preset threshold to obtain the audio features.

7. The three-dimensional streaming media video quality evaluation method according to claim 1, characterized in that The quality regression analysis of the initial three-dimensional streaming media video based on the video features, the audio-visual synchronization features, and the audio features to obtain a target quality score includes: In the first half of the initial three-dimensional streaming media video, performing quality regression analysis on the video features, the audio-visual synchronization features, and the audio features according to a preset first weight set to obtain a first quality score; In the second half of the initial three-dimensional streaming media video, performing quality regression analysis on the video features, the audio-visual synchronization features, and the audio features according to a preset second weight set to obtain a second quality score; Based on the first quality score and the second quality score, the target quality score is obtained, where the first weight set includes a first audio-visual synchronization weight corresponding to the audio-visual synchronization feature, the second weight set includes a second audio-visual synchronization weight corresponding to the audio-visual synchronization feature, and the first audio-visual synchronization weight is greater than the second audio-visual synchronization weight.

8. A three-dimensional streaming media video quality evaluation device, characterized in that, Including: An acquisition module for acquiring an initial three-dimensional streaming media video to be evaluated, where the initial three-dimensional streaming media video includes at least one target object in a three-dimensional space; An extraction module for extracting two-dimensional video information of the target object from the initial three-dimensional streaming media video and extracting audio information from the initial three-dimensional streaming media video; A first feature determination module for respectively performing quality evaluation processing on the two-dimensional video information based on multiple different types of multimodal large models and fusing multiple evaluation features obtained after the quality evaluation processing to obtain video features; A second feature determination module for comparing the consistency between the two-dimensional video information and the audio information to obtain audio-visual synchronization features and performing short-time analysis and processing on the audio information to obtain audio features; A target quality score module for performing quality regression analysis on the initial three-dimensional streaming media video based on the video features, the audio-visual synchronization features, and the audio features to obtain a target quality score.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the three-dimensional streaming media video quality evaluation method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the three-dimensional streaming media video quality evaluation method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Video data processing method and device and electronic equipment

    CN121725402A