Live broadcast picture quality evaluation method and device, equipment, medium and product

By using acquisition parameters and scoring models based on live streaming scene types, combined with scoring methods based on image and text features, the problem of inaccurate live streaming image quality assessment in existing technologies has been solved, achieving more accurate image quality assessment and automated scoring.

CN121037580APending Publication Date: 2025-11-28GUANGZHOU KUGOU COMP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511139314.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies that infer the quality of live stream footage from network transmission parameters cannot reflect the visual experience of end users, leading to inaccurate assessments.

Method used

Based on the scene type of the live broadcast room, the acquisition parameters are obtained, and the live broadcast screen is scored using a pre-trained scoring model. Combining image and feedback text features, the scoring level is automatically evaluated through an image encoder, a text encoder, a generative language model, and a scoring adaptation layer.

Benefits of technology

It improves the accuracy of live stream quality assessment, reduces labor costs, and better matches the subjective feelings of end users, thus achieving automated scoring of picture quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121037580A_ABST
    Figure CN121037580A_ABST
Patent Text Reader

Abstract

The invention discloses a live broadcast picture quality evaluation method and device, equipment, a medium and a product, and relates to the technical field of live broadcast. The method comprises the following steps: acquiring acquisition parameters of a first live broadcast room based on a scene type of the first live broadcast room; acquiring at least one live broadcast picture of the first live broadcast room based on the acquisition parameters of the first live broadcast room; inputting the at least one live broadcast image into a scoring model, and obtaining a first scoring level corresponding to the at least one live broadcast image; the first scoring grade is used for indicating the image quality of the live broadcast picture; wherein the scoring model is obtained by training based on a sample image and a sample scoring grade of the sample image. According to the method, the accuracy of image quality evaluation of the scoring model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the live broadcast technical field, and particularly relates to a live broadcast picture quality evaluation method and device, equipment, medium and product. BACKGROUND

[0002] The live broadcast picture quality is a key factor affecting the user's experience of watching live broadcast, and how to measure the live broadcast picture quality of a live broadcast room is crucial to a live broadcast platform.

[0003] In the related art, the picture quality of a live broadcast room can be inferred by analyzing key network indicators (such as code rate, frame rate, packet loss rate, etc.) of a live broadcast stream in the transmission process.

[0004] However, the method of inferring the picture quality of a live broadcast room through network transmission parameters cannot reflect the visual perception of terminal users. SUMMARY

[0005] The present application provides a live broadcast picture quality evaluation method, device, equipment, medium and product, which can improve the accuracy of the picture quality evaluation of a scoring model.

[0006] According to an aspect of the present application, a live broadcast picture quality evaluation method is provided, which comprises:

[0007] Based on the scene type of a first live broadcast room, the acquisition parameters of the first live broadcast room are obtained;

[0008] Based on the acquisition parameters of the first live broadcast room, at least one live broadcast picture of the first live broadcast room is obtained;

[0009] The at least one live broadcast picture is input into a scoring model to obtain a first score level corresponding to each of the at least one live broadcast picture; the first score level is used to indicate the image quality of the live broadcast picture;

[0010] The scoring model is obtained based on sample images and sample score levels of the sample images.

[0011] According to an aspect of the present application, a live broadcast picture quality evaluation device is provided, which comprises:

[0012] The parameter acquisition module is configured to obtain the acquisition parameters of the first live broadcast room based on the scene type of the first live broadcast room;

[0013] The image acquisition module is configured to obtain at least one live broadcast picture of the first live broadcast room based on the acquisition parameters of the first live broadcast room;

[0014] The score obtaining module is configured to input the at least one live picture into a score model to obtain a first score level corresponding to each of the at least one live picture; and the first score level is used to indicate the image quality of the live picture.

[0015] The score model is trained based on sample images and sample score levels of the sample images.

[0016] In some embodiments, the score obtaining module is configured to input the at least one live picture into a first score model to obtain the first score level corresponding to each of the at least one live picture; and the first score model is the score model corresponding to the scene type of the first live room.

[0017] In some embodiments, the live picture quality evaluation apparatus further includes a text obtaining module configured to obtain feedback text of the first live room; and the feedback text is used to indicate the picture feedback of a terminal account to the first live room.

[0018] The score obtaining module is configured to input the at least one live picture and the feedback text into a score model to obtain a first score level corresponding to each of the at least one live picture.

[0019] In some embodiments, the score model includes an image encoder, a first text encoder, a generative language model, and a score adaptation layer; the score obtaining module is configured to input the at least one live picture into the image encoder to obtain image features corresponding to the at least one live picture; input the feedback text into the first text encoder to obtain text features corresponding to the feedback text; input the image features corresponding to the at least one live picture and the text features corresponding to the feedback text into the generative language model to obtain a high-dimensional feature vector of the at least one live picture output by a feature extraction part of the generative language model; the high-dimensional feature vector is used to indicate the score information of the image; and input the high-dimensional feature vector of the at least one live picture into the score adaptation layer to obtain the first score level corresponding to each of the at least one live picture.

[0020] In some embodiments, the live picture quality evaluation device further comprises a first training module configured to obtain a first sample image, auxiliary annotation information of the first sample image, and a sample score level of the first sample image; input the first sample image into the image encoder to obtain an image feature of the first sample image; input the auxiliary annotation information of the first sample image into the first text encoder to obtain a text feature corresponding to the auxiliary annotation information of the first sample image; input the image feature of the first sample image and the text feature corresponding to the auxiliary annotation information of the first sample image into the generative language model to obtain a high-dimensional feature vector of the first sample image output by a feature extraction part of the generative language model; input the high-dimensional feature vector of the first sample image into the score adaptation layer to obtain a score level of the first sample image; obtain a loss function value based on the score level of the first sample image and the sample score level of the first sample image; and update parameters of the first text encoder, the generative language model, and the score adaptation layer based on the loss function value.

[0021] In some embodiments, the live picture quality evaluation device further comprises a second training module configured to obtain at least two second sample images and a sample score level of each of the at least two second sample images; input the at least two second sample images into the image encoder to obtain at least two reference image features; each reference image feature corresponds to an image feature of each second sample image; input the sample score level of each of the at least two second sample images into a second text encoder to obtain at least two reference text features; each reference text feature corresponds to a text feature corresponding to the sample score level of each second sample image; obtain a first similarity and a second similarity based on the at least two reference image features and the at least two reference text features; the first similarity is used to indicate a similarity between a first reference image feature and a first reference text feature, and the second similarity is used to indicate a similarity between the first reference image feature and at least one second reference text feature; the second sample image corresponding to the first reference image feature is the same as the second sample image corresponding to the first reference text feature, and the second sample image corresponding to the first reference image feature is different from the second sample image corresponding to the second reference text feature; and update parameters of the image encoder to maximize the first similarity and minimize the second similarity.

[0022] In some embodiments, the live picture quality evaluation apparatus further comprises a sample acquisition module configured to acquire a first candidate sample image and a second candidate sample image; perform distortion processing on the first candidate sample image and the second candidate sample image to obtain a distorted first candidate sample image and a distorted second candidate sample image; and determine the first candidate sample image and the distorted first candidate sample image as the first sample image, and determine the second candidate sample image and the distorted second candidate sample image as the second sample image.

[0023] In some embodiments, the first sample image and the second sample image comprise artificial intelligence generated content (AIGC) images.

[0024] In some embodiments, the acquisition parameters of the first live room comprise at least one of the following: frame rate, resolution, and exposure.

[0025] In some embodiments, the live picture quality evaluation apparatus further comprises an application module configured to acquire a ranking score of the first live room based on the first score level corresponding to each of the at least one live picture; and update a recommendation index of the first live room based on the ranking score of the first live room, wherein the recommendation index is used to indicate a recommendation priority of the live room.

[0026] According to another aspect of the present application, a computer device is provided, which comprises a processor and a memory, and the memory stores at least one computer instruction, which is loaded and executed by the processor to implement the live picture quality evaluation method according to the above aspect.

[0027] According to another aspect of the present application, a computer readable storage medium is provided, which stores at least one computer instruction, which is loaded and executed by a processor to implement the live picture quality evaluation method according to the above aspect.

[0028] According to another aspect of the present application, a computer program product is provided, which comprises computer instructions stored in a computer readable storage medium, and a processor reads and executes the computer instructions from the computer readable storage medium to implement the live picture quality evaluation method according to the above aspect.

[0029] The technical scheme provided by the embodiments of the present application can bring the following beneficial effects:

[0030] The computer device can determine the collection parameter of the first live room according to the scene type of the first live room, then collect one or more live pictures of the first live room according to the collection parameter of the first live room, and then score the one or more live pictures by the pre-trained scoring model to obtain the score level corresponding to the one or more live pictures. In the scheme, different scene types of live rooms correspond to different picture collection strategies, which can improve the rationality of picture collection of various types of live rooms, and further improve the accuracy of subsequent picture quality evaluation by the scoring model.

[0031] In addition, by using the pre-trained scoring model, automatic scoring of live room pictures can be realized, which not only reduces the labor cost, but also improves the accuracy of live room picture quality evaluation. BRIEF DESCRIPTION OF DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0033] Figure 1 is a schematic diagram of the scheme implementation environment provided by an exemplary embodiment of the present application;

[0034] Figure 2 is a flowchart of a live picture quality evaluation method provided by an exemplary embodiment of the present application;

[0035] Figure 3 is a schematic diagram of the display interface of the audience client provided by an exemplary embodiment of the present application;

[0036] Figure 4 is a training and inference flowchart of the scoring model provided by an exemplary embodiment of the present application;

[0037] Figure 5 is a schematic diagram of reference image features and reference text features provided by an exemplary embodiment of the present application;

[0038] Figure 6 is a flowchart of a live picture quality evaluation method provided by an exemplary embodiment of the present application;

[0039] Figure 7 is a schematic diagram of a live picture quality evaluation method provided by an exemplary embodiment of the present application;

[0040] Figure 8 is a block diagram of a live picture quality evaluation device provided by an exemplary embodiment of the present application;

[0041] Figure 9 is a structural block diagram of a computer device provided by an embodiment of the present application.

[0042] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the present application. DETAILED DESCRIPTION

[0043] For the purposes of the present application, the technical solutions and advantages thereof will be more clearly apparent from the following further detailed description of embodiments of the present application, which will be described in conjunction with the accompanying drawings.

[0044] The exemplary embodiments will be described in detail herein below with reference to the accompanying drawings. In the following description, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.

[0045] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used in the present disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0046] In the embodiments of the present application, the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the feedback text and other user behaviors involved in the present application are obtained under sufficient authorization.

[0047] It should be understood that although the terms first, second, etc. can be employed in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to differentiate one piece of information from another piece of information of the same type. For example, a first sample image can also be referred to as a second sample image, and similarly, a second sample image can also be referred to as a first sample image, without departing from the scope of the present disclosure. Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining".

[0048] First, the terms involved in the present application are introduced and explained:

[0049] Live room: In a network live broadcast, a live room refers to a virtual space where a host creates real-time content. In a live room, the host and the audience interact in real time through video, audio, and text.

[0050] Please refer to Figure 1 , which shows a schematic diagram of a scheme implementation environment provided by an example embodiment of the present application. The scheme implementation environment can include a host terminal device 110, an audience terminal device 120, and a server 130. The host terminal device 110 and the audience terminal device 120 can communicate with the server 130 through a network.

[0051] The host terminal device 110 and the audience terminal device 120 can include audio and video capture devices and audio and video playback devices. Optionally, the number of host terminal devices 110 and audience terminal devices 120 can be one or more. For example, the host terminal device 110 and the audience terminal device 120 can be electronic devices such as a mobile phone, a tablet computer, a personal computer (PC), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a vehicle-mounted device, etc. The present application does not limit the host terminal device 110 and the audience terminal device 120. The client running the target application (such as a live broadcast application) in the host terminal device 110 and the audience terminal device 120. Optionally, the target application is an application that needs to be downloaded and installed, or a web-based application, or a mini-program, or other forms of applications, which are not limited by the embodiments of the present application. In the embodiments of the present application, the target application can include but is not limited to a live broadcast application, a social application, a shopping application, a video playback application, a remote conference application, a music playback application, a game application, an education application, an instant messaging application, or other types of applications, which are not limited by the embodiments of the present application.

[0052] The server 130 is used to provide background services for the client of the target application running in the host terminal device 110 and the audience terminal device 120. The server 130 can be a server, a server cluster composed of multiple servers, or a cloud computing service center. The server 130 can be a background server of the target application, used to provide background services for the client of the target application.

[0053] In the embodiments of the present application, the client of the target application program running in the anchor terminal device 110 is the anchor client, and the user corresponding to the anchor client is the anchor; the client of the target application program running in the audience terminal device 120 is the audience client, and the user corresponding to the audience client is the audience.

[0054] When the anchor performs live streaming through the anchor client, the audience can enter the live streaming room of the anchor through the network and watch the content of the live streaming room of the anchor through the live streaming interface displayed in the audience client. For example, the audience can interact with the anchor. For example, the anchor can also view the content of the live streaming room of the anchor through the live streaming interface displayed in the anchor client.

[0055] In some embodiments, the scheme implementation environment of the embodiments of the present application can be applied to a live streaming scene involving video live streaming. The anchor client and the audience client can be two different versions of clients of the target application program, and the two different versions of clients are respectively for anchor users and audience users, that is, the version for anchor users has the functions of the above-mentioned anchor client, and the version for audience users has the functions of the above-mentioned audience client; or, they can also be the same version of clients of the target application program, and the version of clients has both the functions of the above-mentioned anchor client and the functions of the above-mentioned audience client. For example, the audience client can not only watch video live streaming, but also perform video live streaming. For example, the anchor client can not only perform video live streaming, but also watch video live streaming of other anchors. The present application does not limit this.

[0056] For example, in the scenario where the anchor in the first live streaming room performs video live streaming, the server 130 can determine the collection parameters of the first live streaming room according to the scene type of the first live streaming room; then, the server 130 can collect a live streaming picture of the first live streaming room according to the collection parameters of the first live streaming room; then, the server 130 can input the live streaming picture into a scoring model to obtain a first score level corresponding to the live streaming picture, and the first score level is used to indicate the image quality of the live streaming picture; wherein the scoring model is trained based on sample images and sample score levels of the sample images. Correspondingly, the server 130 can score the first live streaming room according to the first score levels corresponding to the live streaming pictures collected multiple times to obtain a ranking score of the first live streaming room; then, the server 130 can update the recommendation index of the first live streaming room according to the ranking score of the first live streaming room, and the recommendation index is used to indicate the recommendation priority of the live streaming room. In the case where the recommendation index of the first live streaming room meets the specified requirements, the server 130 preferentially recommends the first live streaming room in the audience client.

[0057] Please refer to Figure 2Fig. 1 shows a flowchart of a live picture quality evaluation method according to an example embodiment of the present application. The subject performing each step of the method can be a computer device, which can be any electronic device with data storage and processing capabilities. Alternatively, the computer device is a terminal device 110 in the system shown in Fig. 2; or the computer device is a server 120 in the system shown in Fig. 2; or the computer device is the terminal device 110 and the server 120 in the system shown in Fig. 2. As shown in Fig. 1, the method can include steps 210-230. Figure 1 Figure 1 Figure 1 Figure 2

[0058] Step 210: Obtain the collection parameters of the first live room based on the scene type of the first live room.

[0059] In some embodiments, the scene type refers to the scene type of the picture display content of the live room. For example, the scene type includes at least one of the following: indoor, outdoor, game, and other scene types.

[0060] In some embodiments, the collection parameters refer to the setting parameters when collecting the picture of the first live room. For example, the collection parameters include at least one of the following: frame rate, resolution, and exposure, and other collection parameters.

[0061] In some embodiments, the computer device can create a database table or file structure to store the correspondence between the scene type of the live room and the collection parameters of the live room. Accordingly, the computer device can query the first collection parameters of the first live room from the correspondence according to the scene type of the first live room, so that the computer device can quickly obtain the collection parameters of the first live room.

[0062] For example, if the scene type of the first live room is an outdoor scene type (such as a sports event), the collection parameters of the first live room can include a high frame rate (such as 60 fps). For another example, if the scene type of the first live room is an indoor scene type (such as a good thing sharing), the collection parameters of the first live room can include a high local resolution.

[0063] Step 220: Obtain at least one live picture of the first live room based on the collection parameters of the first live room.

[0064] That is, the computer device collects the live picture of the first live room according to the collection parameters corresponding to the scene type of the first live room.

[0065] ​​​​In some embodiments, different live broadcast rooms correspond to different collection frequencies. For example, a live broadcast room with a higher requirement for live broadcast picture quality corresponds to a higher collection frequency (e.g., one live broadcast picture is collected every 1 second); and a live broadcast room with a lower requirement for live broadcast picture quality corresponds to a lower collection frequency (e.g., one live broadcast picture is collected every 3 seconds). Accordingly, the computer device acquires at least one live broadcast picture of the first live broadcast room according to the collection parameter and the collection frequency of the first live broadcast room.

[0066] Step 230: inputting the at least one live broadcast picture into a scoring model to acquire a first score level corresponding to each of the at least one live broadcast picture; the first score level is used to indicate the image quality of the live broadcast picture.

[0067] The scoring model is trained based on sample pictures and sample score levels of the sample pictures.

[0068] That is, the computer device can pre-train a scoring model according to sample pictures and sample score levels of the sample pictures, so as to acquire the score level of the live broadcast picture according to the live broadcast picture of the first live broadcast room.

[0069] In some embodiments, the first score level is one of at least two score levels, and different score levels represent different qualities of the live broadcast picture, such as definition, color performance, etc. For example, the score levels include five score levels: excellent, good, general, poor, and very poor.

[0070] In some embodiments, the scoring model also supports text instruction input, and the text instruction is used to instruct the scoring model to score a specified dimension of the input live broadcast picture. For example, the specified dimension can be to evaluate the definition, illumination, color, etc. of the live broadcast picture. For example, the text instruction can be “evaluate the illumination of the picture” or “analyze the performance of the color dimension”; for another example, the text instruction can be to output the picture quality level or the aesthetic level, etc. That is, the first score level can also be used to indicate the image aesthetics of the live broadcast picture.

[0071] In summary, the technical scheme provided by the embodiments of the present application can determine the collection parameter of the first live broadcast room according to the scene type of the first live broadcast room, then collect one or more live broadcast pictures of the first live broadcast room according to the collection parameter of the first live broadcast room, and then score the one or more live broadcast pictures by the pre-trained scoring model to obtain the score level corresponding to the one or more live broadcast pictures. In this scheme, different scene types of live broadcast rooms correspond to different picture collection strategies, which can improve the rationality of picture collection of various types of live broadcast rooms, and further improve the accuracy of subsequent picture quality evaluation by the scoring model, and align the picture quality evaluation method with human perception.

[0072] In addition, through the pre-trained scoring model, automatic scoring of the live streaming room picture can be realized, which can not only reduce the labor cost, but also improve the accuracy of the live streaming room picture quality evaluation.

[0073] Based on the above Figure 2 In a possible implementation, the step 230 can be implemented as step 230a in the solution of the embodiment shown.

[0074] Step 230a: input at least one live streaming picture into the first scoring model to obtain a first scoring grade corresponding to each of the at least one live streaming picture; the first scoring model is a scoring model corresponding to the scene type of the first live streaming room.

[0075] That is, different scene types of the live streaming room correspond to different scoring models. For example, the computer device can create a database table or file structure for storing the correspondence between the scene type of the live streaming room and the scoring model; accordingly, the computer device can query the scoring model of the first live streaming room from the above correspondence according to the scene type of the first live streaming room, so as to quickly obtain the first scoring model.

[0076] In some embodiments, the parameters in different scoring models are set according to different scene types of the live streaming room, such as different dimensions of scoring weights. For example, the above dimensions include at least one of the following: motion blur, light-related dimensions, and face-related dimensions. For example, outdoor live streaming scenes usually have the characteristics of large light changes and complex scenes. In terms of light, problems such as direct sunlight and shadow blocking may be encountered, so the light-related dimension scoring weight will be increased. For example, indoor live streaming scenes are relatively stable, but have high requirements for the display of characters and the arrangement of backgrounds. In terms of characters, the face-related dimension weight such as face clarity and skin color restoration will be higher.

[0077] For example, in the training process of the scoring model, the computer device can classify sample images to obtain sample images of at least two categories, and each category of sample images corresponds to a scene type of the live streaming room; then, the computer device trains the scoring model according to the sample images of at least two categories and the sample scoring grades corresponding to the sample images of at least two categories, to obtain the scoring model corresponding to different scene types of the live streaming room.

[0078] In the embodiment of the present application, the computer device can select the scoring model corresponding to the first live streaming room according to the scene type of the first live streaming room, so that the scoring model can match the corresponding parameters and scoring weights when processing the live streaming picture, thereby ensuring that the scoring grade of the live streaming picture is close to the subjective feeling of the terminal account, and improving the accuracy of the scoring model in picture quality evaluation.

[0079] Based on the scheme in each of the above embodiments of the present application, in a possible implementation, the live picture quality evaluation method further includes step 222, and step 230 can be implemented as step 230b.

[0080] Step 222: Obtain the feedback text of the first live room; the feedback text is used to indicate the picture feedback of the terminal account to the first live room.

[0081] In some embodiments, the audience user can interact with the live room while watching the live interface of the live room through interactive operations. For example, the interactive operations include at least one of the following: comments, sending a barrage, following, liking, forwarding, connecting a microphone, connecting a video, and sending virtual items. Optionally, the interactive operations can also include other manners, which are not limited in the present application. Please refer to Figure 3 which shows a schematic diagram of a display interface of an audience client provided by an example embodiment of the present application. As shown in Figure 3 The audience terminal device can display the live interface 310 of the live room and the interactive area 320. The audience user can perform the interactive operation of sending a barrage through the sending barrage control 320a of the interactive area 320.

[0082] In some embodiments, the feedback text can be the real-time feedback of the audience user to the first live room in the interactive operation of the audience client, including positive feedback and negative feedback. For example, the feedback text can be a text containing “stuttering”, “blur”, “clear”, etc.

[0083] In some embodiments, the feedback text is the feedback text in a specified time period before at least one live picture of the first live room is obtained. For example, the specified time period can be 1 second, 3 seconds, or other time periods, which are not limited in the present application.

[0084] Step 230b: input the at least one live picture and the feedback text into the scoring model to obtain the first score level corresponding to each of the at least one live picture.

[0085] That is, the computer device synchronously inputs the live picture of the first live room and the feedback text into the scoring model, and outputs the respective score level of the live picture by the scoring model.

[0086] In some embodiments, as shown in part (b) of Figure 4 The computer device can construct at least one image-text pair according to the at least one live picture and the corresponding feedback text, and then input the at least one image-text pair into the scoring model to obtain the first score level corresponding to each of the at least one live picture.

[0087] In the embodiment of the present application, the computer device can acquire real-time feedback (i.e., the feedback text) of the audience user of the first live room to the live picture while collecting the live picture of the first live room, and then synchronously input the feedback text and the live picture into the scoring model to acquire the score level output by the scoring model. The feedback text in the present solution can indicate real-time feedback of the audience user of the first live room to the live effect, and can improve the accuracy of the score level output by the scoring model.

[0088] Based on the solutions in the above-mentioned embodiments of the present application, in a possible implementation manner, the scoring model comprises an image encoder, a first text encoder, a generative language model, and a score adaptation layer; and the step 230b can be implemented as steps 230b1 to 230b4.

[0089] Step 230b1: input at least one live picture into the image encoder to acquire image features corresponding to the at least one live picture.

[0090] In some embodiments, the image encoder is pre-trained to extract global features of the live picture, including color, texture, motion blur, and other key information. That is, the image features can indicate the color, texture, motion blur, and other contents of the live picture.

[0091] Step 230b2: input the feedback text into the first text encoder to acquire text features corresponding to the feedback text.

[0092] In some embodiments, the first text encoder is used to extract semantic features representing picture description in the feedback text.

[0093] In some embodiments, the image encoder and the first text encoder can be calculated in parallel to adapt to the real-time requirements of live picture quality evaluation.

[0094] Step 230b3: input the image features corresponding to the at least one live picture and the text features corresponding to the feedback text into the generative language model to obtain a high-dimensional feature vector of the at least one live picture output by the feature extraction part of the generative language model; the high-dimensional feature vector is used to indicate the score information of the image.

[0095] In some embodiments, the computer device can process the input image features and text features through the feature extraction part of the generative language model after removing the encoding part and the output part to obtain the high-dimensional feature vector.

[0096] In some embodiments, the generative language model is used to fuse the image features corresponding to the live picture and the text features corresponding to the feedback text, and generate a high-dimensional feature vector indicating the score information of the live picture. For example, the generative language model can map the input image features and text features to the same semantic space.

[0097] Step 230b4: input the high-dimensional feature vector of at least one live picture into the score adaptation layer to obtain the first score grade corresponding to each live picture.

[0098] In some embodiments, different score grades correspond to different scores. For example, if the score grades include excellent, good, average, poor, and very poor, the excellent, good, average, poor, and very poor correspond to five scores of 5, 4, 3, 2, and 1, respectively.

[0099] In some embodiments, the score adaptation layer can map the high-dimensional feature vector of each live picture to the score grade space, generate the probability distribution of the high-dimensional feature vector of each live picture belonging to each score, determine the final score by calculating the mathematical expectation, and then obtain the first score grade corresponding to each live picture.

[0100] In the embodiments of the present application, the computer device can extract features from the live picture and the feedback text through the image encoder and the first text encoder, respectively, to obtain image features and text features, and then analyze and fuse the image features and the text features through the generative language model to map the image features and the text features to a unified semantic space to obtain a high-dimensional feature vector. Then, the high-dimensional feature vector is mapped to the score grade space through the score adaptation layer to obtain the first score grade corresponding to the live picture. Through the step-by-step reasoning of the image encoder, the first text encoder, the generative language model, and the score adaptation layer, the present scheme realizes automatic scoring, which not only reduces the labor cost, but also improves the accuracy of live room picture quality evaluation.

[0101] Based on the schemes in the above various embodiments of the present application, in one possible implementation, as shown in part (a) of Figure 4 The live picture quality evaluation method further includes a training process of the first text encoder, the generative language model, and the score adaptation layer.

[0102] Step A1: obtain a first sample image, auxiliary annotation information of the first sample image, and a sample score grade of the first sample image.

[0103] In some embodiments, the auxiliary annotation information is fine-grained annotation added to the first sample image. For example, the computer device can generate an image-text data pair with auxiliary annotation information by using a high-quality open source model to annotate the first sample image, so as to accelerate the generation speed of the auxiliary annotation information (benefit); for another example, the computer device can obtain the annotation of the first sample image made by a human being; for yet another example, the computer device can first pre-annotate the first sample image by using a high-quality open source model, and then filter and modify the data with errors in the annotation of the open source model by a human being.

[0104] For example, the auxiliary annotation information includes image blur in the upper left corner, image background blur, and the like.

[0105] In some embodiments, the developer can pre-collect a plurality of first sample images and auxiliary annotation information of the first sample images to obtain a plurality of first image-text pairs; then, the first image-text pairs are respectively input into an image encoder and a first text encoder to extract corresponding feature vectors.

[0106] Step A2: inputting the first sample image into the image encoder to obtain image features of the first sample image.

[0107] In some embodiments, the image encoder is pre-trained. In the pre-training stage of the image encoder, the computer device can inject live scene priori, so that the image encoder can preferentially focus on the key regions of the first sample image. For example, the biological feature region: face (eyes, mouth), human body posture key point; for another example, the action focus region: sports object (football, athlete) in sports live broadcast, character action region in game live broadcast.

[0108] Step A3: inputting the auxiliary annotation information of the first sample image into the first text encoder to obtain text features corresponding to the auxiliary annotation information of the first sample image.

[0109] In some embodiments, the first text encoder is used to extract semantic features in the auxiliary annotation information of the first sample image.

[0110] Step A4: inputting the image features of the first sample image and the text features corresponding to the auxiliary annotation information of the first sample image into a generative language model to obtain a high-dimensional feature vector of the first sample image output by a feature extraction part of the generative language model.

[0111] In some embodiments, the generative language model is used to fuse the image features of the first sample image and the text features of the auxiliary annotation information to generate a high-dimensional feature vector that implicitly indicates the score information of the first sample image. For example, the generative language model can map the input image features and text features to the same semantic space.

[0112] Step A5: input the high-dimensional feature vector of the first sample image into the score adaptation layer to obtain a score level of the first sample image.

[0113] In some embodiments, the score adaptation layer can map the high-dimensional feature vector of each first sample image to a score level space to generate a probability distribution of each first sample image belonging to each score level, determine the final score by taking the mathematical expectation, and then obtain the corresponding score level of each first sample image.

[0114] Step A6: based on the score level of the first sample image and the sample score level of the first sample image, obtain a loss function value.

[0115] The loss function, also known as the cost function, is a function used to evaluate the difference between the predicted value of the neural network model and the true value. The smaller the value of the loss function, the better the performance of the neural network model. The training process of the model is to minimize the value of the loss function by adjusting the parameters of the model. Common loss functions include 0-1 loss function, absolute value loss function, logarithmic loss function, exponential loss function, perception loss function, cross-entropy loss function, KL divergence loss function, and triplet loss function.

[0116] In some embodiments, when calculating the loss function value, the computer device can calculate the difference between the score level of the first sample image and the sample score level of the first sample image by using a pre-set loss function to obtain the loss function value.

[0117] Step A7: based on the loss function value, update the parameters of the first text encoder, the generative language model, and the score adaptation layer, respectively.

[0118] In the initial stage, the parameters of the first text encoder, the generative language model, and the score adaptation layer are default values or pre-set values. During the training process, the computer device can update the parameters of the first text encoder, the generative language model, and the score adaptation layer according to the loss function value to improve the accuracy of the first text encoder, the generative language model, and the score adaptation layer.

[0119] In some embodiments, the computer device repeats steps A1 to A7 to perform iterative updating and training until a pre-set condition is reached. For example, the developer can set the pre-set condition to a certain number of training iterations or a certain convergence condition.

[0120] In the embodiments of the present application, the first text encoder, the generative language model and the score adaptation layer obtained through iterative training can accurately output the score level of the live picture in the picture quality evaluation process of the first live room.

[0121] Based on the schemes in the above various embodiments of the present application, in a possible implementation manner, as shown in part (a) of Figure 4 The live picture quality evaluation method further includes a training process of the image encoder, as shown in part (a) of

[0122] Step B1: Obtain at least two second sample images and sample score levels corresponding to the at least two second sample images respectively.

[0123] In some embodiments, the developer can pre-collect multiple second sample images and manually label the sample score levels corresponding to each second sample image to obtain multiple second image-text pairs; then, the image encoder and the second text encoder are used to extract the corresponding feature vectors from each second image-text pair.

[0124] In some embodiments, the at least two second sample images include images corresponding to different score levels, so as to improve the comprehensiveness and integrity of the sample data.

[0125] Step B2: Input the at least two second sample images into the image encoder to obtain at least two reference image features; each reference image feature corresponds to the image feature of each second sample image.

[0126] Since the traditional visual encoder, such as the convolutional neural network (CNN), generates a large number of tokens (for example, 196 tokens may be generated after processing a 224x224 pixel image) when processing high-resolution images, the calculation amount increases in a square level with the resolution.

[0127] In some embodiments, the image encoder can adopt a lightweight visual encoder to reduce the calculation complexity and resource consumption of image encoding.

[0128] In some embodiments, the computer device can optimize through a visual abstractor, compress the number of image feature Tokens, reduce the computational load, and at the same time, through the attention mechanism, let the model focus on the key areas of the live room, such as faces, bodies, etc. For example, the Perceiver Resampler structure is adopted. Perceiver Resampler is a feature resampling technology based on attention mechanism, and its core design goal is to compress high-dimensional image features into low-dimensional feature sequences while maintaining the key semantic information of the image. Perceiver Resampler performs cross-attention calculation with the original image features through a learnable query vector (Query Tokens), selectively aggregates key information from global features, and compresses the number of Tokens from hundreds to tens (such as 32 or 64 Tokens), reducing the computational complexity by about 80%. The computational load is balanced. Taking a 1080P live picture as an example, the traditional model needs about 200 GFLOPS of calculation for a single frame, and after adopting Perceiver Resampler, it can be reduced to less than 40 GFLOPS, which adapts to the real-time processing needs of edge devices (such as mobile phones and live streaming servers).

[0129] For example, the above-mentioned image encoder can be a Transformer-based image encoder, such as Vision Transformer (ViT), which divides the second sample image into small patches and extracts features using a Transformer-based encoder part to convert the second sample image into a fixed-dimensional vector representation.

[0130] For example, the above-mentioned image encoder can be a Swin Transformer, which reduces the computational complexity through a local window attention mechanism and adapts to the long sequence processing needs of high-resolution live pictures. Among them, the global calculation is decomposed into the calculation within the local window through the hierarchical window attention mechanism, and each window only processes the self-attention of a 7x7 pixel area, and then realizes the global information interaction through cross-window connection. This design reduces the computational complexity from O(N 2 ) to O(N), which adapts to the processing of live pictures with a resolution of 1080P and above.

[0131] Exemplarily, the image encoder described above can combine the local feature extraction capability of convolution with the global modeling advantage of ViT to improve the modeling accuracy of complex scenes such as low light and motion blur. For example, using a ConvNeXt-ViT hybrid architecture, the bottom layer uses the deep convolution (such as a 7x7 convolution kernel) of ConvNeXt to extract local texture features (such as fabric texture and skin texture in a live broadcast picture); the upper layer uses the global self-attention modeling of ViT to model the cross-region dependence (such as the spatial relationship between the host and the background).

[0132] Exemplarily, the image encoder described above can use a non-Transformer architecture (such as a dynamic convolution network) to realize cross-modal alignment of images and text through a parameter sharing mechanism to reduce the number of model parameters. For example, using a dynamic convolution network (Dynamic Convolution), the input adaptive convolution kernel generation mechanism (such as dynamically adjusting the convolution kernel weight according to the picture content) is used to replace the self-attention of the Transformer; the parameter sharing mechanism: the image encoder and the text encoder share part of the parameters to reduce the number of model parameters.

[0133] Step B3: input the sample score level of each of the at least two second sample images into the second text encoder to obtain at least two reference text features; each reference text feature corresponds to a text feature corresponding to the sample score level of each second sample image.

[0134] In some embodiments, the second text encoder described above is used to extract semantic features in the sample score level of the second sample image.

[0135] In some embodiments, the image encoder and the second text encoder described above can be calculated in parallel to adapt to the real-time requirements of live picture quality evaluation. Exemplarily, the computer device uses a CLIP pre-training model to extract global features of the second sample image, retaining key information such as color, texture, and motion blur, and cross-modal text alignment: through contrastive learning of the second sample image-sample score level, the visual features and the language semantics are in the same feature space. Among them, the contrastive language-image pre-training (CLIP pre-training model) is composed of a visual encoder and a text encoder, the visual encoder is responsible for processing the input image and converting the image into a fixed-dimensional vector representation to capture the key features of the image, and to prepare for subsequent calculation of similarity scores with the text vector output by the text encoder in the shared embedding space.

[0136] Step B4: Obtain a first similarity and a second similarity based on the at least two reference image features and the at least two reference text features; the first similarity is used to indicate the similarity between the first reference image feature and the first reference text feature, and the second similarity is used to indicate the similarity between the first reference image feature and the at least one second reference text feature; the second sample image corresponding to the first reference image feature is the same as the second sample image corresponding to the first reference text feature, and the second sample image corresponding to the first reference image feature is different from the second sample image corresponding to the second reference text feature.

[0137] That is, the computer device can construct positive and negative sample pairs, the positive sample pair being the first reference image feature and the first reference text feature, and the negative sample pair being the first reference image feature and the at least one second reference text feature.

[0138] As shown in Figure 5 , input the N second sample images into the image encoder to obtain N reference image features I1, I2, I3,..., IN. N , input the sample score levels of the N second sample images into the second text encoder to obtain N reference text features T1, T2, T3,..., TN. N , the positive sample pair includes I1*T1, I2*T2, I3*T3,..., IN*TN. N , the negative sample pair includes I1*T2, I1*T3, I1*T4,..., I1*TN. N , the negative sample pair includes I1*T2, I1*T3, I1*T4,..., I1*TN. N , the negative sample pair includes I1*T2, I1*T3, I1*T4,..., I1*TN. N .

[0139] Step B5: Update the parameters of the image encoder to maximize the first similarity and minimize the second similarity.

[0140] In some embodiments, maximizing the first similarity means that the similarity between the reference image feature of each second sample image and the reference text feature corresponding to the sample score level thereof is high, and the difference is small, which can improve the contrast capability of the image encoder.

[0141] In some embodiments, minimizing the second similarity means that the similarity between the reference image feature of each second sample image and the reference text features corresponding to the sample score levels of other second sample images is low, and the difference is large, which can improve the distinguishing ability of the meta-learner.

[0142] For example, the computer device repeats steps B1 to B5 to perform iterative updating and training until a preset condition of the image encoder is reached. For example, the developer can set the preset condition to a certain number of training iterations or a certain convergence condition. The updated image encoder is used to generate the image feature corresponding to the first sample image by inputting the first sample image.

[0143] In some embodiments, the first similarity can be used to indicate the similarity between different second sample images under the same sample score level, and the second similarity can be used to indicate the similarity between different second sample images under different sample score levels, so as to reduce the complexity of the contrast learning. That is, the positive sample is different second sample images under the same sample score level, and the negative sample is different second sample images under different sample score levels. For example, the positive sample can be selected as a "good" picture under different scenes, and the negative sample can be selected as a "poor" game live picture with serious motion blur.

[0144] The training process of the image encoder is shown in the embodiments of the present application. By maximizing the feature similarity between the matched second sample images and the score level and minimizing the similarity with other unpaired data, the image encoder can be strengthened to distinguish the image features of different score levels, thereby improving the accuracy of the image features output by the image encoder. In addition, by pre-training the image encoder, the training difficulty in the subsequent training processes of the first text encoder, the generative language model, and the score adaptation layer can be reduced.

[0145] Based on the schemes in the above various embodiments of the present application, in one possible implementation, the live picture quality evaluation method further includes:

[0146] obtaining a first candidate sample image and a second candidate sample image;

[0147] performing distortion processing on the first candidate sample image and the second candidate sample image to obtain a distorted first candidate sample image and a distorted second candidate sample image;

[0148] determining the first candidate sample image and the distorted first candidate sample image as the first sample image;

[0149] determining the second candidate sample image and the distorted second candidate sample image as the second sample image.

[0150] In some embodiments, the distortion processing can be injecting artifacts, compression, etc. on the first candidate sample image and the second candidate sample image, for simulating live loss. For example, the computer device can simulate the packet loss rate in network transmission by randomly deleting image blocks (such as 16x16 pixel blocks) or adding salt and pepper noise. For another example, the computer device can use an x264 encoder to re-encode the image with different quantization parameter (QP) values to generate compression distortion such as blocking effect and blur. For another example, the computer device can calculate the object motion trajectory based on the optical flow field, and apply a Gaussian blur kernel (kernel size 5-15 pixels) to the fast motion area (such as the football in sports live), to simulate the hand-shake blur when shooting with a mobile phone.

[0151] The embodiments of the present application show a data augmentation scheme for training samples in the training process, which can realize data enhancement in the training process, improve the diversity of the first sample image and the second sample image, and further improve the accuracy when training the scoring model.

[0152] Based on the scheme in the above various embodiments of the present application, in a possible implementation, the first sample image and the second sample image include an Artificial Intelligence Generated Content (AIGC) image.

[0153] That is, the computer device can collect the AIGC image as the first sample image and the second sample image, so as to ensure that the scoring model can accurately evaluate both real images and generated images.

[0154] The embodiments of the present application show another data augmentation scheme for training samples in the training process, which can realize data enhancement in the training process, improve the diversity of the first sample image and the second sample image, and further improve the accuracy when training the scoring model.

[0155] Based on the scheme in the above various embodiments of the present application, in a possible implementation, the collection parameters of the first live room include at least one of the following: frame rate, resolution, exposure.

[0156] The above-mentioned frame rate refers to the number of image frames collected and output per unit time, with the unit being frames / second. The frame rate can indicate the smoothness of a dynamic scene, for example, the higher the frame rate, the more coherent the dynamic picture; the lower the frame rate, the more likely to appear stuttering or trailing.

[0157] The above-mentioned resolution refers to the number of pixels of an image, which can be expressed as width x height (such as 1920 x 1080). The resolution can indicate the richness of the details of the image, and the higher the resolution, the clearer the image.

[0158] The above-mentioned exposure refers to the duration and intensity of light received by the sensor. Exposure directly affects image brightness.

[0159] The embodiments of the present application show various implementable ways of the collection parameters of the first live room. The collection parameters can be one or more of the frame rate, the resolution, and the exposure. Different parameter combinations can affect the clarity, brightness, dynamic range, etc. of the collected picture. The computer device can determine the corresponding collection parameters according to the scene type of the first live room, thereby improving the applicability and diversity of the present scheme.

[0160] Please refer to Figure 6Fig. 1 shows a flowchart of a live picture quality evaluation method according to an example embodiment of the present application. As shown in Fig. 1, the live picture quality evaluation method comprises steps 100, 200 and 300. Figure 6 As shown in Fig. 2, in a possible implementation, the live picture quality evaluation method further comprises steps 240 and 250.

[0161] Step 240: obtaining a ranking score of the first live room based on the first score level corresponding to each of the at least one live picture.

[0162] In some embodiments, different score levels can correspond to different ranking scores. For example, the higher the image quality of the live picture indicated by the first score level, the higher the ranking score.

[0163] Step 250: updating a recommendation index of the first live room based on the ranking score of the first live room; the recommendation index is used to indicate a recommendation priority of the live room.

[0164] In some embodiments, different ranking scores can correspond to different recommendation indexes. For example, the higher the ranking score of the first live room, the higher the recommendation index of the first live room, the higher the recommendation priority of the first live room, and the easier the first live room is to be seen by the terminal account.

[0165] The example embodiments of the present application show that the dynamic recommendation ranking of the live room is achieved through the score level, which can effectively improve the exposure rate of the live room with high image quality.

[0166] Based on the schemes shown in the above various embodiments, in a possible implementation, please refer to Figure 7 Fig. 3 shows a schematic diagram of a live picture quality evaluation method according to an example embodiment of the present application. As shown in Fig. 3, the live picture quality evaluation method comprises steps 300 and 400. Figure 7 For example, the terminal device is a smart phone 10, and the application program is a mobile application program. The user can download the mobile application program from a server 20 through an application store of the smart phone 10. After the installation of the mobile application program is successful, the mobile application program is displayed on the display screen of the terminal device in the form of an icon. Then, the user can start the mobile application program by clicking the icon of the mobile application program.

[0167] The display screen of the smart phone 10 displays a user interface of the mobile application program, which includes entrances of different live rooms and other contents. The user can click the entrance of the live room N to enter the corresponding live room.

[0168] The live picture quality evaluation method comprises the following steps:

[0169] Step 701: obtaining the collection parameters of the live room N based on the scene type of the live room N;

[0170] At step 702, at least one live picture of the live room N is obtained based on the collection parameters of the live room N, and feedback text of the live room N is obtained.

[0171] The feedback text is used to indicate the picture feedback of the terminal account to the live room N.

[0172] At step 703, the at least one live picture and the feedback text are input into a scoring model to obtain a first score level corresponding to each of the at least one live picture:

[0173] The first score level is used to indicate the image quality of the live picture, and the scoring model is trained based on sample images and sample score levels of the sample images.

[0174] At step 704, a ranking score of the live room N is obtained based on the first score level corresponding to each of the at least one live picture.

[0175] At step 705, the recommendation index of the live room N is updated based on the ranking score of the live room N:

[0176] The recommendation index is used to indicate the recommendation priority of the live room N.

[0177] Correspondingly, the order of the entrances of different live rooms displayed in the user interface of the smartphone 10 is updated, and the exposure rate of the live room with higher image quality is improved.

[0178] It should be noted that the execution mode of each step is the same as the execution mode of each embodiment of the present application, and will not be repeated here.

[0179] Based on the scheme shown in each of the above embodiments, the present application proposes a live picture quality intelligent evaluation algorithm and system based on a multi-modal generative language model. The multi-modal generative language model (Multi-Moodal Large Language Model, MLLM) is a generative language model that can process multiple modalities (such as images, speech, text, etc.) information at the same time. The present method realizes accurate and subjective picture quality evaluation of terminal users through discrete text label grading and training of self-developed scoring model, and preferentially recommends live rooms with high-definition picture quality on the user side.

[0180] For example, the live picture quality intelligent evaluation system based on the multi-modal generative language model includes:

[0181] 1) Data acquisition module

[0182] Image acquisition: The picture of the live room is intercepted at a certain time interval according to the level, supporting dynamic switching of multiple resolutions and multiple code rates.

[0183] Among them, the above level can be a live streaming definition level. For example, the requirements for live streaming quality in high-definition special zones and general zones are different. Higher frequency monitoring quality is required for high-definition special zones, and substandard quality is immediately removed to avoid affecting the audience's experience.

[0184] Multi-modal data collection: Collect user comments, reviews and other text feedback (such as lag, blur) and other auxiliary information to build image-text input data.

[0185] Dynamic scene capture: For fast motion, low light and other complex scenes in live streaming, an adaptive strategy is designed to improve data coverage. Fast motion (such as sports events, etc.), low light (such as night, indoor weak light, etc.) can cause motion blur, noise surge, detail loss and other problems in the picture. For example, optical flow or motion detection is used to automatically increase the frame rate when high-speed motion is detected; for example, the content of a certain part of the picture is more concerned, and local high-resolution sampling is performed, and other areas remain normal.

[0186] 2) Visual encoder and abstractor

[0187] Lightweight visual coding: Use the CLIP pre-training model to extract the global features of the picture, and retain key information such as color, texture, and motion blur. Global features refer to abstract information extracted from the entire image that can represent the overall semantics and structure of the picture, which is different from local features (such as corner points and small area textures).

[0188] Visual abstractor optimization: Use the Perceiver Resampler structure to compress the number of image feature Tokens and reduce the computational load, while using attention mechanisms to focus the model on key areas in the live streaming room, such as faces and bodies.

[0189] 3) Visual score adaptation layer

[0190] Unified modeling of multi-modal tasks: Based on the modularity of mPLUG-Owl2, a visual score adaptation layer is added to support unified evaluation of image quality, image aesthetics, and video quality evaluation; that is, this scheme can be reused in multiple application scenarios such as live streaming aesthetics evaluation, live streaming cover quality and aesthetics evaluation. The mPLUG-Owl2 architecture integrates visual and language encoders, and the visual score adaptation layer is expanded based on this, integrating image quality, image aesthetics, and video quality evaluation tasks to build a unified evaluation system. The visual score adaptation layer uses deep learning methods to effectively interface visual features with text instructions, enabling the model to understand and handle different types of evaluation tasks and achieve deep fusion of multi-modal information.

[0191] Multi-task unified evaluation: Can handle multiple tasks such as image quality, image aesthetics, video quality evaluation, etc. For example, in a live streaming scenario, it can evaluate not only the clarity and color restoration of the picture, but also the composition and visual aesthetics of the picture, and the smoothness and stuttering of the video.

[0192] Discrete text label mapping: Combined with the Q-ALIGN method (image grading training strategy based on discrete text labels), the quality label is set to five rating levels (corresponding to five scores of 5 / 4 / 3 / 2 / 1 respectively). The model outputs the probability distribution of the current evaluation object belonging to each category, and determines the final score by calculating the mathematical expectation, realizing the semantic alignment of the quality label and human subjective perception.

[0193] Instruction-driven flexibility: Supports dynamic language instruction input. By default, it evaluates the overall quality of the image (blur / high definition), but can also be guided by text instructions to evaluate the score of a certain dimension of the image, such as "evaluate the lighting of the image" and "analyze the performance of the color dimension", etc., to meet the diverse evaluation needs.

[0194] Support for live streaming scenario recognition: indoor, outdoor, game, etc. live streaming scenario recognition.

[0195] 4) Score conversion module

[0196] Dynamic allocation of score conversion weights: According to different scenarios (game, outdoor) and other needs, dynamically adjust the score weights to ensure that the score results are highly consistent with human perception. In games and sports live streaming, the dynamic effects of the picture, skill special effects display, and motion blur processing are crucial to user experience. Therefore, the weight of the indicators related to sports will be relatively high, such as the weight of the "motion blur" indicator will be increased. Outdoor live streaming scenarios usually have characteristics such as large changes in light and complex and diverse scenes. In terms of light, there may be problems such as direct sunlight and shadow obstruction, so the score weight related to light will be increased. Indoor live streaming scenarios are relatively stable, but have high requirements for the display of characters and the arrangement of backgrounds. In terms of characters, the clarity of the face, the color restoration of the skin, and other indicators related to the face will have a higher weight.

[0197] In the case of collecting multiple live pictures, the long sequence modeling capability of mPLUG-DocOwl can support joint inference of multiple frames (such as continuous stall detection), replacing the limitations of single frame analysis. The sliding window attention (Sliding Window Attention) is introduced to support processing of continuous video sequences of more than 200 frames, capturing gradual quality degradation in live streaming (such as blurring caused by gradually reducing code rate); cross-frame feature fusion can also be achieved by aggregating the motion features of the previous 5 frames through the attention mechanism, improving the accuracy of stall detection.

[0198] In summary, the scheme realizes accurate evaluation of live picture quality that conforms to human subjective judgment, and through the end-to-end modeling capability of MLLM, the picture quality is matched with human perception, solving the problem of large deviation between traditional methods and subjective experience.

[0199] 1) Multi-modal quality evaluation paradigm

[0200] Discrete text label mapping technology: convert traditional numerical scores (such as 1-5) into learnable discrete text embeddings (excellent / good / average, etc.), and realize semantic alignment of quality labels and human subjective perception through Q-ALIGN method.

[0201] Cross-modal probability distribution modeling: use improved mPLUG-Owl2 to process image features and text instructions (such as "evaluate low-light noise") at the same time, generate multi-level probability distribution output, cover picture quality, aesthetics, and smoothness, etc. multi-dimensional evaluation.

[0202] 2) Dynamic evaluation mechanism

[0203] Scene adaptive weight allocation: dynamically adjust the weight of the score dimension according to the type of live streaming, for example, increase the weight proportion of the "motion blur" indicator in game live streaming.

[0204] Temporal continuity modeling: analyze the quality change trend of consecutive frames through sliding window, identify gradual blurring or sudden stall events.

[0205] The following is an embodiment of the device of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0206] Please refer to Figure 8 , which shows a block diagram of a live picture quality evaluation device provided by an example embodiment of the present application. The device has the functions of the examples described above, which can be implemented by hardware or by hardware executing corresponding software. As Figure 8 shown, the device can include a parameter acquisition module 801, an image acquisition module 802, and a score acquisition module 803.

[0207] The parameter obtaining module 801 is configured to obtain the collection parameter of the first live room based on the scene type of the first live room.

[0208] The image obtaining module 802 is configured to obtain at least one live picture of the first live room based on the collection parameter of the first live room.

[0209] The score obtaining module 803 is configured to input the at least one live picture into a scoring model to obtain a first score level corresponding to each of the at least one live picture, wherein the first score level is used to indicate the image quality of the live picture.

[0210] The scoring model is trained based on sample pictures and sample score levels of the sample pictures.

[0211] In some embodiments, the live picture quality evaluation apparatus further includes a text obtaining module configured to obtain a feedback text of the first live room, wherein the feedback text is used to indicate the picture feedback of the terminal account to the first live room.

[0212] The score obtaining module 803 is configured to input the at least one live picture and the feedback text into the scoring model to obtain a first score level corresponding to each of the at least one live picture.

[0213] In some embodiments, the score obtaining module 803 is configured to input the at least one live picture into a first scoring model to obtain a first score level corresponding to each of the at least one live picture, wherein the first scoring model is a scoring model corresponding to the scene type of the first live room.

[0214] In some embodiments, the scoring model includes an image encoder, a first text encoder, a generative language model, and a score adaptation layer. The score obtaining module 803 is configured to input the at least one live picture into the image encoder to obtain image features corresponding to the at least one live picture, input the feedback text into the first text encoder to obtain text features corresponding to the feedback text, input the image features corresponding to the at least one live picture and the text features corresponding to the feedback text into the generative language model to obtain a high-dimensional feature vector of the at least one live picture output by a feature extraction part of the generative language model, wherein the high-dimensional feature vector is used to indicate the score information of the image, and input the high-dimensional feature vector of the at least one live picture into the score adaptation layer to obtain a first score level corresponding to each of the at least one live picture.

[0215] In some embodiments, the live picture quality evaluation apparatus further comprises: a first training module configured to obtain a first sample image, auxiliary annotation information of the first sample image, and a sample score level of the first sample image; input the first sample image into an image encoder to obtain an image feature of the first sample image; input the auxiliary annotation information of the first sample image into a first text encoder to obtain a text feature corresponding to the auxiliary annotation information of the first sample image; input the image feature of the first sample image and the text feature corresponding to the auxiliary annotation information of the first sample image into a generative language model to obtain a high-dimensional feature vector of the first sample image output by a feature extraction part of the generative language model; input the high-dimensional feature vector of the first sample image into a score adaptation layer to obtain a score level of the first sample image; obtain a loss function value based on the score level of the first sample image and the sample score level of the first sample image; and update parameters of the first text encoder, the generative language model, and the score adaptation layer based on the loss function value.

[0216] In some embodiments, the live picture quality evaluation apparatus further comprises: a second training module configured to obtain at least two second sample images and a sample score level of each of the at least two second sample images; input the at least two second sample images into an image encoder to obtain at least two reference image features; each reference image feature corresponds to an image feature of each second sample image; input the sample score level of each of the at least two second sample images into a second text encoder to obtain at least two reference text features; each reference text feature corresponds to a text feature corresponding to the sample score level of each second sample image; obtain a first similarity and a second similarity based on the at least two reference image features and the at least two reference text features; the first similarity is used to indicate a similarity between a first reference image feature and a first reference text feature, and the second similarity is used to indicate a similarity between the first reference image feature and at least one second reference text feature; the second sample image corresponding to the first reference image feature is the same as the second sample image corresponding to the first reference text feature, and the second sample image corresponding to the first reference image feature is different from the second sample image corresponding to the second reference text feature; and update parameters of the image encoder to maximize the first similarity and minimize the second similarity.

[0217] In some embodiments, the live picture quality evaluation apparatus further comprises: a sample obtaining module configured to obtain a first candidate sample image and a second candidate sample image; perform distortion processing on the first candidate sample image and the second candidate sample image to obtain a distorted first candidate sample image and a distorted second candidate sample image; determine the first candidate sample image and the distorted first candidate sample image as the first sample image; and determine the second candidate sample image and the distorted second candidate sample image as the second sample image.

[0218] In some embodiments, the first sample image and the second sample image include artificial intelligence generated content (AIGC) images.

[0219] In some embodiments, the collection parameters of the first live room include at least one of the following: frame rate, resolution, exposure.

[0220] In some embodiments, the live picture quality evaluation device further includes an application module configured to: obtain a ranking score of the first live room based on the first score level corresponding to each of the at least one live picture; and update a recommendation index of the first live room based on the ranking score of the first live room, wherein the recommendation index is used to indicate a recommendation priority of the live room.

[0221] It should be noted that the device provided in the above embodiments is only used as an example to illustrate the division of the above functional modules in realizing its functions. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the content structure of the device is divided into different functional modules to complete all or part of the above-described functions. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be described here.

[0222] Please refer to Figure 9 which shows a structural block diagram of a computer device provided in an embodiment of the present application. The computer device 900 can be any electronic device with data computing, processing and storage capabilities. The computer device 900 can be used to implement the live picture quality evaluation method provided in the above embodiments.

[0223] Generally, the computer device 900 includes a processor 901 and a memory 902.

[0224] The processor 901 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 901 can be implemented in the form of at least one of a DSP (Digital Signal Processing), an FPGA (Field Programmable Gate Array), a PLA (Programmable Logic Array), and the like. The processor 901 can also include a main processor and a co-processor. The main processor is a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit). The co-processor is a low-power consumption processor for processing data in a standby state. In some embodiments, the processor 901 can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing of content to be displayed on a display screen. In some embodiments, the processor 901 can further include an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.

[0225] The memory 902 can include one or more computer-readable storage media that can be non-transitory. The memory 902 can also include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 902 is configured to store a computer program configured to be executed by one or more processors to implement the live picture quality evaluation method described above.

[0226] Those skilled in the art can understand that, Figure 9 The structure shown in the figure does not constitute a limitation on the computer device 900, and can include more or fewer components than those shown, or combine certain components, or adopt a different arrangement of components.

[0227] In an example embodiment, a computer readable storage medium is also provided, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the live picture quality evaluation method described above. Optionally, the computer readable storage medium can include a ROM (Read-Only Memory), a RAM (Random Access Memory), a SSD (Solid State Drives), an optical disc, or the like. Among them, the random access memory can include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0228] In an example embodiment, a computer program product is also provided, and the computer program product includes a computer program stored in a computer readable storage medium. A processor of a computer device reads the computer program from the computer readable storage medium, and the processor executes the computer program to enable the computer device to perform the live picture quality evaluation method described above.

[0229] It should be understood that "multiple" mentioned herein refers to two or more. The "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the associated objects before and after it are in an "or" relationship. In addition, the step numbers described herein only exemplarily show a possible execution order between steps, and in some other embodiments, the above steps can also be executed in a different order from the number order, such as executing two different numbered steps at the same time, or executing two different numbered steps in an order opposite to the illustration, and the embodiments of the present application do not limit this.

[0230] The above only describes example embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for live picture quality assessment, characterized in that, The method comprises: obtaining the acquisition parameters of the first live room based on the scene type of the first live room; obtaining at least one live picture of the first live room based on the acquisition parameters of the first live room; inputting the at least one live picture into a scoring model to obtain the first score level corresponding to the at least one live picture respectively; the first score level is used to indicate the image quality of the live picture; wherein the scoring model is obtained based on sample images and sample score levels of the sample images.

2. The method of claim 1, wherein, The method further comprises: obtaining the feedback text of the first live room; the feedback text is used to indicate the picture feedback of the terminal account to the first live room; 3. The method according to claim 1 or 2, characterized in that, inputting the at least one live picture and the feedback text into the scoring model to obtain the first score level corresponding to the at least one live picture respectively. The scoring model comprises an image encoder, a first text encoder, a generative language model, and a score adaptation layer; inputting the at least one live picture and the feedback text into the scoring model to obtain the first score level corresponding to the at least one live picture respectively comprises: inputting the at least one live picture into the image encoder to obtain the image features corresponding to the at least one live picture; inputting the feedback text into the first text encoder to obtain the text features corresponding to the feedback text; 4. The method of claim 3, wherein, inputting the image features corresponding to the at least one live picture and the text features corresponding to the feedback text into the generative language model to obtain the high-dimensional feature vector of the at least one live picture output by the feature extraction part of the generative language model; the high-dimensional feature vector is used to indicate the score information of the image; inputting the high-dimensional feature vector of the at least one live picture into the score adaptation layer to obtain the first score level corresponding to the at least one live picture respectively. The method further comprises: obtaining the first sample image, the auxiliary annotation information of the first sample image, and the sample score level of the first sample image; inputting the first sample image into the image encoder to obtain the image features of the first sample image; 5. The method of claim 4, wherein, inputting the auxiliary annotation information of the first sample image into the first text encoder to obtain the text features corresponding to the auxiliary annotation information of the first sample image; inputting the image features of the first sample image and the text features corresponding to the auxiliary annotation information of the first sample image into the generative language model to obtain the high-dimensional feature vector of the first sample image output by the feature extraction part of the generative language model; ​ ​ ​ input the high-dimensional feature vector of the first sample image into the score adaptation layer to obtain a score level of the first sample image; obtain a loss function value based on the score level of the first sample image and a sample score level of the first sample image; update parameters of the first text encoder, the generative language model, and the score adaptation layer based on the loss function value.

6. The method according to claim 4 or 5, characterized in that, The method further comprises: obtain at least two second sample images and respective sample score levels of the at least two second sample images; input the at least two second sample images into the image encoder to obtain at least two reference image features; each reference image feature corresponds to an image feature of each second sample image; input the respective sample score levels of the at least two second sample images into a second text encoder to obtain at least two reference text features; each reference text feature corresponds to a text feature corresponding to a sample score level of each second sample image; obtain a first similarity and a second similarity based on the at least two reference image features and the at least two reference text features; the first similarity indicates a similarity between a first reference image feature and a first reference text feature, and the second similarity indicates a similarity between the first reference image feature and at least one second reference text feature; the second sample image corresponding to the first reference image feature is the same as the second sample image corresponding to the first reference text feature, and the second sample image corresponding to the first reference image feature is different from the second sample image corresponding to the second reference text feature; update parameters of the image encoder to maximize the first similarity and minimize the second similarity.

7. The method according to claim 5 or 6, characterized in that, The method further comprises: obtain a first candidate sample image and a second candidate sample image; perform distortion processing on the first candidate sample image and the second candidate sample image to obtain a distorted first candidate sample image and a distorted second candidate sample image; determine the first candidate sample image and the distorted first candidate sample image as the first sample image; determine the second candidate sample image and the distorted second candidate sample image as the second sample image.

8. The method according to claim 5 or 6, characterized in that, The first sample image and the second sample image include artificial intelligence generated content (AIGC) images.

9. The method according to any one of claims 1 to 8, characterized in that, The acquisition parameters of the first live room include at least one of the following: frame rate, resolution, and exposure.

10. The method according to any one of claims 1 to 9, characterized in that, The method further comprises: obtain a ranking score of the first live room based on the respective first score levels of the at least one live picture; update a recommendation index of the first live room based on the ranking score of the first live room; the recommendation index indicates a recommendation priority of a live room.

11. A computer device, comprising: The computer device includes a processor and a memory, and the memory stores a computer program, which is loaded and executed by the processor to implement the live picture quality evaluation method according to any one of claims 1 to 10.

12. A computer-readable storage medium, characterized in that, The computer program is stored in the computer readable storage medium and is loaded and executed by the processor to implement the live picture quality evaluation method according to any one of claims 1 to 10.

13. A computer program product, characterised in that, The computer program product comprises a computer program stored in a computer readable storage medium, and the processor reads and executes the computer program from the computer readable storage medium to implement the live picture quality evaluation method according to any one of claims 1 to 10.