Evaluation method for explanation slice picture quality and electronic equipment
By evaluating the correlation of product information and the content of the images in the target keyframes of the narration slice, the problem of relying on manual scoring for the quality evaluation of the narration slice in the existing technology is solved, and standardized and accurate evaluation results are achieved.
Patent Information
- Application Number
- CN202511790542.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-03-06
Smart Images

Figure CN121617009A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and more particularly to a method and electronic device for evaluating the quality of sliced images. Background Technology
[0002] Live streaming involves setting up independent signal acquisition equipment at the filming location, importing the acquired audio and video signals into the production control unit, uploading them to a server via the network, and then publishing them to a website for viewers to watch. Due to its real-time nature, live streaming is widely used in various scenarios, such as e-commerce sales and education.
[0003] Live streaming is a crucial method for e-commerce platforms to sell products. Platforms typically allow merchants to record live streams, and the highlights of the product presentations extracted from the replays are called "presentation clips." These clips are generally used by merchants for advertising or placed on the product homepage for secondary promotion.
[0004] Explanatory segments need to be of high quality to achieve their promotional purpose. Since these segments are very short, visual quality is a crucial indicator of their overall quality. Currently, the visual quality of explanatory segments primarily relies on manual scoring, a method heavily influenced by personal preference and failing to provide a rapid and standardized assessment of segment quality. Summary of the Invention
[0005] The first aspect of this application provides a method for evaluating the quality of a video slice, including:
[0006] Receive the explanatory video clips to be evaluated, wherein the explanatory video clips are video segments that explain the target product;
[0007] Extract the target keyframe image from the explained slice;
[0008] Using a large model, product information correlation evaluation and image content evaluation are performed on the target keyframe image to obtain product information correlation evaluation results and image content evaluation results. The product information correlation represents the correlation between product information in the target keyframe image and the attribute information of the target product. The image content is the relevant information of the human figure in the target keyframe image.
[0009] Based on the evaluation results of the screen content and the evaluation results of the correlation between the product information, the evaluation results of the quality of the explanatory slice screen are obtained.
[0010] In one possible implementation, a large model is used to evaluate the image content of the target keyframe image, resulting in an image content evaluation result, including:
[0011] Obtain preset prompt words, which are used to indicate the analysis rules for the image quality of portraits in large model analysis;
[0012] The preset prompt words and the target keyframe image are sent to the large model so that the large model can determine the corresponding analysis rules based on the preset prompt words and analyze the target keyframe image using the analysis rules to obtain the image content evaluation result of the target keyframe image. The image content evaluation result indicates whether the human image in the target keyframe image passes the evaluation or fails the evaluation.
[0013] In one possible implementation, the analysis rules include: a human figure appears in the image, the human figure is located in a preset core area of the image, and the face of the human figure meets a preset clarity condition.
[0014] In one possible implementation, a large model is used to evaluate the relevance of product information to the target keyframe image, resulting in a product information relevance evaluation result, including:
[0015] Based on the slice identifiers of the explained slices, the attribute information of the target product is obtained from the database;
[0016] Using a large model, based on the attribute information, the correlation of product information is evaluated for the target keyframe image, and the product information correlation evaluation result is obtained.
[0017] In one possible implementation, a large model is used to evaluate the relevance of product information to the target keyframe image based on the attribute information, resulting in a product information relevance evaluation result, including:
[0018] Using a large model, product information is extracted from the target keyframe image to obtain the product information present in the image;
[0019] By using a large model, the attribute information and the product information are semantically compared to obtain matching information between the attribute information and the product information;
[0020] Based on the matching information, the correlation assessment results of product information are obtained.
[0021] In one possible implementation, comparing the attribute information and the product information using a large model to obtain matching information between the attribute information and the product information includes:
[0022] Using a large model, the product information is divided into at least one category based on at least one category of the attribute information;
[0023] Using a large model and based on categories, we analyze the matching information between the product information and the attribute information to obtain matching information for each category.
[0024] Based on the matching information for each category, the matching information between the attribute information and the product information is determined.
[0025] In one possible implementation, obtaining the evaluation result of the explanatory slice image quality based on the image content evaluation result and the product information relevance evaluation result includes:
[0026] Based on the evaluation result of the screen content, a first score is obtained. The evaluation result of the screen content indicates that the evaluation is passed, and the first score is a first value. If the evaluation result of the screen content indicates that the evaluation is not passed, the first score is a second value, and the first value is higher than the second value.
[0027] The second score is obtained based on the assessment results of the relevance of product information;
[0028] Based on the first integral and the second integral, the evaluation result of the quality of the explanatory slice image is obtained.
[0029] In one possible implementation, obtaining the target keyframe from the received explanatory slice includes:
[0030] Using a preset extraction tool, the first keyframe image is extracted after a preset time from the start of the explained slice;
[0031] Based on the extracted first keyframe image, the first keyframe image is used as the target keyframe image;
[0032] Since the first keyframe image was not extracted, the second keyframe image is extracted and used as the target keyframe image. The second keyframe image is a keyframe in the explanatory slice that is later than the first keyframe image. The second keyframe is a keyframe that is directly or indirectly adjacent to the first keyframe.
[0033] In one possible implementation, determining whether the first keyframe image has been extracted includes:
[0034] Based on the preset extraction tool, key frame images are extracted from the explanatory slices within a preset time period, and it is determined that the first key frame has been extracted.
[0035] If the preset extraction tool fails to extract keyframe images from the explanatory slices within a preset time period, it is determined that the first keyframe has not been extracted.
[0036] A second aspect of this application provides an electronic device, comprising:
[0037] Memory, used to store large models;
[0038] A processor is configured to receive a narration slice to be evaluated, wherein the narration slice is a video clip that narrates the target product;
[0039] The target keyframe image is extracted from the narration slice; using a large model, the product information correlation evaluation and image content evaluation are performed on the target keyframe image to obtain the product information correlation evaluation result and the image content evaluation result. The product information correlation represents the correlation between the product information in the target keyframe image and the attribute information of the target product, and the image content is the relevant information of the human figure in the target keyframe image; based on the image content evaluation result and the product information correlation evaluation result, the image quality evaluation result of the narration slice is obtained.
[0040] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the method for evaluating the quality of explanatory slice images described in the first aspect or any implementation thereof.
[0041] A fourth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to perform an evaluation method for the quality of a slice of video as described in the first aspect or any implementation thereof.
[0042] In summary, this application receives a video clip explaining a target product; extracts target keyframe images from the video clip; and uses a large model to evaluate the product information relevance and screen content of the target keyframe images, obtaining product information relevance evaluation results and screen content evaluation results. Product information relevance characterizes the correlation between product information in the target keyframe image and the attribute information of the target product, while screen content refers to the information related to the human figure in the target keyframe image. Based on the screen content evaluation results and product information relevance evaluation results, an evaluation result of the video quality of the video clip is obtained. Using a large model to evaluate the target keyframe images in the video clip from two dimensions—the human figure dimension and the product information relevance dimension—results with high accuracy. Moreover, because the large model's processing is stable, its evaluation of the input target keyframe images has a unified evaluation standard, is not influenced by personal bias, and can obtain standardized evaluation results. Attached Figure Description
[0043] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0044] Figure 1 This is a flowchart illustrating a method for evaluating the quality of sliced images provided in an embodiment of this application;
[0045] Figure 2 This is a schematic diagram of the image frame in the explanatory slice provided in the embodiments of this application;
[0046] Figure 3 This is a flowchart illustrating the process of evaluating the image content of a target keyframe image using a large model, as provided in an embodiment of this application, to obtain the image content evaluation result.
[0047] Figure 4 This is a flowchart illustrating the process of evaluating the relevance of product information to a target keyframe image using a large model, as provided in this embodiment of the application, to obtain the evaluation result.
[0048] Figure 5 This is a flowchart illustrating how, through a large model, product attribute information is used to evaluate the relevance of product information to a target keyframe image, thereby obtaining the product information relevance evaluation result.
[0049] Figure 6 This is a flowchart illustrating the process of obtaining the evaluation result of the explanatory slice image quality based on the evaluation result of the image content and the evaluation result of the correlation between the product information provided in this application embodiment;
[0050] Figure 7 This is a schematic diagram of the process of obtaining target keyframes from received explanatory slices provided in an embodiment of this application;
[0051] Figure 8 This is a flowchart illustrating an application scenario of a method for evaluating the quality of sliced images, as provided in an embodiment of this application.
[0052] Figure 9 This is a schematic diagram of the structure of an electronic device that illustrates a method for evaluating the quality of sliced images, as provided in an embodiment of this application.
[0053] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0054] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0055] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0056] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0057] Live streaming is a technology that involves setting up independent signal acquisition equipment at the filming location, importing the acquired audio and video signals into the broadcasting terminal, uploading them to a server via the network, and then publishing them to a website for people to watch.
[0058] The signal acquisition equipment set up at the shooting location can be the terminal devices used by the live broadcast personnel, such as mobile phones, tablets, or cameras and microphones that come with mobile terminals.
[0059] The terminal device runs live streaming software. The live streamer logs into the software and registers a live streaming room. The video captured by the live streamer's terminal device is transmitted to the server in real time as the video of the live streaming room. The server then sends the video to the terminal used by the audience via the Internet in real time.
[0060] A live streaming room is a virtual space where live streams are conducted, allowing viewers to watch the content and interact with the streamers.
[0061] Merchants extract engaging segments from live stream replays to create promotional clips. These clips typically include the live streamer explaining the product, as well as the product itself. These clips are then used by merchants for advertising or placed on the product's homepage for secondary promotion.
[0062] Because the quality of numerous presentation clips generated on live streaming platforms varies greatly, they may not only fail to promote products during distribution but also negatively impact user experience. Therefore, it is necessary to evaluate the video quality of presentation clips to select those with higher video quality for distribution, ensuring a positive user experience while still achieving the goal of product promotion.
[0063] To improve the consistency of image quality assessment in explanatory segments, a stable assessment method is needed.
[0064] Reference Figure 1 , Figure 1 This is a flowchart illustrating a method for evaluating the quality of sliced images provided in an embodiment of this application, as shown below. Figure 1 As shown in the embodiment of this application, a method for evaluating the quality of a video slice may include steps 101 to 104, which are described in detail below.
[0065] 101. Receive the explanatory video clips to be evaluated. The explanatory video clips are video segments that explain the target product.
[0066] The explanatory video clip can be extracted from the video generated during the live broadcast. During the live broadcast, the host explains the target product, and the explanatory video clip is a video segment explaining the target product.
[0067] In one possible implementation, the explanatory slices extracted from the live video can be stored in storage space. When evaluating the picture quality of the explanatory slices, the explanatory slices can be retrieved from storage space. Alternatively, after extracting the explanatory slices from the live video, the picture quality evaluation can be performed directly. This application does not restrict the source of the explanatory slices.
[0068] 102. Extract target keyframe images from the explanatory slices;
[0069] Since the image frames in a video include keyframes and prediction frames, the keyframe image contains more information and can be encoded independently, and the prediction frame is incremental information based on the keyframe image. The prediction frame, combined with the corresponding keyframe image, can be encoded to obtain the corresponding image.
[0070] Since the keyframe images contain more data than the prediction frames and contain important content from the video, keyframe images are extracted from the narration slice for evaluation instead of all image frames, in order to reduce the amount of data processing.
[0071] Since the keyframe images and predicted frame images in the video are spaced apart, they can be distinguished by marking or by setting a set number of frames. Therefore, the target keyframe image can be obtained from the narration slice by marking the keyframe image, or the current frame can be determined as a keyframe by counting the number of image frames from the beginning of the narration slice. If it is a keyframe, then the image is extracted; otherwise, it is not extracted.
[0072] This explanatory slice contains multiple image frames, which include keyframes and prediction frames.
[0073] The target keyframe image extracted from the explanatory slice can be either the first frame image in the explanatory slice or a non-first frame image in the explanatory slice.
[0074] Figure 2 This is a schematic diagram of an image frame in a illustrative slice provided in an embodiment of this application. The image frame in the illustrative slice includes a key frame image 201 and a prediction frame image 202. In this schematic diagram, one key frame image 201 corresponds to four prediction frame images 202.
[0075] It should be noted that... Figure 2 The text explains the arrangement of image frames in a slice as an example. Figure 2 The explanation states that the first frame of a slice is a keyframe image. However, in practice, the arrangement of image frames in a slice is not limited to this, and the first frame image can also be a predicted frame image from the original live video.
[0076] 103. Using a large model, evaluate the correlation of product information and the content of the image for the target keyframe image, and obtain the evaluation results of product information correlation and the evaluation results of the image content. Product information correlation represents the correlation between product information in the target keyframe image and the attribute information of the target product, and image content is the relevant information of the human image in the target keyframe image.
[0077] Since the explanatory video clip is a video segment explaining the target product, it should contain relevant information about the target product and images of the live streamer. Accordingly, this application evaluates the target keyframe image from two dimensions: the image dimension and the product information relevance dimension.
[0078] The product information relevance reflects the degree of correlation between the product information contained in the target keyframe image and the attribute information of the target product. Since the explanatory slice is for explaining the target product, this degree of correlation indicates whether the attribute information of the target product is described in detail during the explanation of the target product in the explanatory slice.
[0079] The product information in the target keyframe image can be information related to the target product extracted from the image content in the target keyframe image. The image content in the target keyframe image can include at least one of the following: product image, product text description, product illustration, etc.
[0080] As an example, when the product is a food item, the target keyframe image can include a product image, a text description of the product, a product illustration, or any two or all of the three. The product illustration can be a photo showing the product ingredients, an image illustrating the steps to claim a discount, etc.
[0081] As an example, when the product is a travel-related product, the target keyframe image can include product images, product text descriptions, or both. The product images could be photos of various attractions included in the travel itinerary, illustrations of how to claim discounts, etc.
[0082] This large model is a model with a massive number of parameters, which can include tens of thousands of parameters. Due to this massive number of parameters, the large model has strong capabilities in fields such as image processing and semantic understanding.
[0083] The image quality of the target keyframe image is analyzed using a large model. This large model can accurately identify the content in the image (including the correlation dimension between product information and the attribute information of the target product and the human image dimension), and obtain the product information correlation evaluation result and the image content evaluation result of the target keyframe image.
[0084] Because the large model has a stable processing procedure, it has a unified evaluation standard for evaluating the input target keyframe image, which can yield more objective evaluation results.
[0085] In one possible implementation, the product information relevance assessment results and the screen content assessment results may include passing or failing the assessment.
[0086] In one possible implementation, the product information relevance assessment result and the image content assessment result can be quantitatively represented by scores to indicate whether the product passes or fails, or whether it passes some of the assessment conditions.
[0087] The large model used in this application can be an AI (Artificial Intelligence) model, which has image understanding and reasoning capabilities. By utilizing the image understanding and reasoning capabilities of this large model, it is possible to evaluate the correlation between the product information and the attribute information of the target product in the target keyframe image and the human image dimension.
[0088] 104. Based on the evaluation results of the content of the screen and the evaluation results of the relevance of the product information, the evaluation results of the quality of the explanatory slice screen are obtained.
[0089] The large model evaluates the target keyframe image from two dimensions: the correlation between product information and the attribute information of the target product, and the human image dimension. After obtaining the product information correlation evaluation result and the image content evaluation result, the image content evaluation result and the product information correlation evaluation result are combined to obtain the image quality evaluation result of the explanatory slice.
[0090] Since the explanatory slices are generally short, only one or a limited number of keyframe images can be evaluated, and the evaluation result of the keyframe image can be used as the evaluation result of the picture quality of the explanatory slice.
[0091] In one possible implementation, for each narration slice, only one keyframe image is obtained, and the evaluation is based on this keyframe image. The evaluation result of this keyframe image is used as the evaluation result of the picture quality of the narration slice.
[0092] In one possible implementation, the number of keyframe images can be set according to the duration of the narration segment, with a larger number of keyframe images corresponding to a longer duration, in order to evaluate the overall picture quality of the narration segment.
[0093] As an example, if the duration of the narration segment is 30 seconds, one keyframe image can be used for evaluation; if the duration of the narration segment is 2 minutes, three keyframe images can be used for evaluation. By combining the evaluation results of the three keyframe images, the evaluation result of the picture quality of the narration segment is obtained. These three keyframe images can be evenly distributed in the narration segment to achieve the evaluation of the overall picture quality of the narration segment.
[0094] In this embodiment, a narration segment to be evaluated is received, which is a video clip explaining the target product. Target keyframe images are extracted from the narration segment. Using a large model, product information relevance and image content evaluations are performed on the target keyframe images to obtain product information relevance evaluation results and image content evaluation results. Product information relevance characterizes the correlation between product information in the target keyframe image and the attribute information of the target product. Image content refers to the information related to the human figure in the target keyframe image. Based on the image content evaluation results and product information relevance evaluation results, the image quality evaluation result of the narration segment is obtained. Using a large model to evaluate the target keyframe images in the narration segment from two dimensions—the human figure dimension and the product information relevance dimension—results in high accuracy. Moreover, because the large model's processing is stable, its evaluation of the input target keyframe images has a unified evaluation standard, is not influenced by personal bias, and can obtain standardized evaluation results.
[0095] Figure 3 This is a flowchart illustrating the process of evaluating the content of a target keyframe image using a large model to obtain the evaluation result, provided in an embodiment of this application. It may include steps 301 to 302, which are described in detail below.
[0096] 301. Obtain preset prompt words. Preset prompt words are used to indicate the analysis rules for the image quality of portraits in large model analysis.
[0097] The preset prompt words define the scene standard for the image quality of human figures in the large model analysis image. The large model can determine the analysis rules based on the preset prompt words, and then use the analysis rules to analyze the image quality of human figures in the received image.
[0098] In one possible implementation, the analysis rules include: a human figure appears in the image, the human figure is located in a preset core area of the image, and the human figure's face meets a preset clarity condition.
[0099] Correspondingly, the preset prompt can be a scenario definition for each item in the analysis rule that involves a name. This scenario definition defines the scope of the noun based on a general understanding. The scenario definition serves as a temporary knowledge base for the large model, enabling it to acquire and utilize the latest information or domain-specific knowledge without retraining.
[0100] As an example, the portrait is located in a preset core area of the image, and the definition of the core area is indicated by a preset prompt. For example, the core area is a rectangular area consisting of one-third of the length and one-third of the width of the image, with the center of the image as the center.
[0101] As an example, the portrait's face meets preset clarity criteria, with preset prompts indicating the definition of clarity. For instance, the eyes, nose, and mouth in the portrait occupy more than a set number of pixels (e.g., 10 pixels).
[0102] 302. Send the preset prompt words and the target keyframe image to the large model so that the large model can determine the corresponding analysis rules based on the preset prompt words, and use the analysis rules to analyze the target keyframe image to obtain the image content evaluation result of the target keyframe image. The image content evaluation result indicates whether the human image in the target keyframe image passes the evaluation or fails the evaluation.
[0103] The large model used in this application can be an AI model with image understanding and reasoning capabilities. The image understanding and reasoning capabilities of the large model are used to evaluate the image quality of the target keyframe image.
[0104] After extracting the target keyframe image, the target keyframe image and preset prompt words are sent to the large model. The large model can use the preset prompt words to determine the corresponding analysis rules, and then use the analysis rules to evaluate the human image in the target keyframe image to obtain the image content evaluation result.
[0105] Since the streamer is the most attractive element in a live broadcast, evaluating the video stream's quality requires considering their image. Therefore, we can first determine if a human figure appears in the frame. If a human figure is present, we can then perform two subsequent checks on the details of that figure. Otherwise, we don't need to perform these checks; the output will simply show that the evaluation failed. This reduces wasted model resources and prevents the large model from developing illusions over time.
[0106] In one possible implementation, the large model performs the following analysis process on the target keyframe image according to the analysis rules: First, it determines whether a human image appears in the image. If no human image appears, it is determined to be a non-live broadcast, the judgment ends, and the image content evaluation result is "failed evaluation". If a human image appears and is determined to be a live broadcast, it continues to determine whether the human image is located in the preset core area of the image. If the human image is not located in the preset core area, the judgment ends, and the image content evaluation result is "failed evaluation". If the human image is located in the preset core area, it further determines whether the human face meets the preset clarity conditions. If the human face does not meet the preset clarity conditions, the judgment ends, and the image content evaluation result is "failed evaluation". If the human face meets the preset clarity conditions, the judgment ends, and the image content evaluation result is "passed evaluation".
[0107] The three analysis rules in the above-mentioned image content evaluation process are because during the product explanation process, there may be situations where the product obscures the anchor's face or the anchor focuses on the product. Therefore, it is only necessary that there is a human figure in the target keyframe, and the human figure is located in the preset core area and the human figure's face is clear.
[0108] If the target keyframe image contains multiple portraits, each portrait must be evaluated if it is located in a preset core area of the image according to the analysis rules. Otherwise, the portrait in the target keyframe fails the evaluation.
[0109] During the detailed review of multiple portraits in the target keyframe image, the portraits can be judged sequentially according to their positions in the image, in a preset order. This preset order can be from the center of the image to the edge, or from left to right, from top to bottom, etc., and this application does not limit the order.
[0110] Because the analysis rules are distributed from broad to detailed, it can minimize the possibility of the model exhibiting illusions when dealing with images of lower resolution.
[0111] In one possible implementation, a calling thread for the large model can be set up. This calling thread is used to trigger the large model to perform an evaluation task on the human figure in the target keyframe image. After the target keyframe image is extracted, the calling thread is triggered. The calling thread calls the large model and sends the target keyframe image and preset prompts to the large model so that the large model can evaluate the image content of the target keyframe image.
[0112] In one possible implementation, the large model could be the Gemini-2.5-Flash model, which not only has excellent reasoning capabilities in image understanding but also maintains high-speed response and cost-effectiveness.
[0113] Since temporary files cannot be accessed by the large model, and the image hosting service of the media library in the electronic device can provide a link that can be accessed externally, after extracting the keyframe image, the target keyframe image can be stored in the storage area of the electronic device that performs the evaluation method for slice image quality described in this application, the target keyframe image can be uploaded to the media library, and the link of the target keyframe image can be sent from the media library to the large model so that the large model can obtain the target keyframe image from the media library according to the link.
[0114] In this embodiment, preset prompts are obtained, which are used to instruct the large model on the analysis rules for the image quality of the portrait. The preset prompts and the target keyframe image are sent to the large model, so that the large model determines the corresponding analysis rules based on the preset prompts and analyzes the target keyframe image using the analysis rules to obtain the image content evaluation result of the target keyframe image. The image content evaluation result indicates whether the portrait in the target keyframe image passes or fails the evaluation. Through the large model, the image content of the target keyframe image is evaluated from the perspective of portrait and image clarity to obtain the image content evaluation result of the target keyframe image. Since the analysis rules used in the large model utilize unified preset prompts, a unified image evaluation standard is achieved to evaluate the image content of the narration segment and obtain a standardized evaluation result.
[0115] Figure 4 This is a flowchart illustrating how a large model is used to evaluate the relevance of product information to a target keyframe image, and how product information relevance evaluation results are obtained. The flowchart may include steps 401 to 402, which are described in detail below.
[0116] 401. Based on the slice identifiers in the explanatory slices, retrieve the attribute information of the target product from the database;
[0117] During the process of extracting and elaborating segments from live video, a segment identifier is added to each segment. This segment identifier is unique, and each segment has a different identifier.
[0118] As an example, the slice identifier can be a slice ID (Identity document, unique code).
[0119] The live streaming platform assigns a unique product identifier to each product it sells. Furthermore, the platform stores all generated segments in a database. This database contains a segment list, where each data entry is a segment record. The segment list also records the segment identifier for each segment and the product identifier that each segment is explaining.
[0120] As an example, the product identifier could be a product ID.
[0121] In one possible implementation, the slice identifier is a parameter carried in the slice link of the tutorial slice. After receiving the tutorial slice, the link of the tutorial slice is analyzed to obtain its slice identifier. Then, the product identifier corresponding to the slice identifier is queried from the slice list. This product identifier is the identifier of the target product being taught in the tutorial slice. Finally, the attribute information of the target product is retrieved from the database.
[0122] In one possible implementation, the narration slice carries a slice identifier assigned to it by the live streaming platform. Therefore, by analyzing the narration slice, its slice identifier can be obtained from the information it carries.
[0123] It should be noted that since multiple hosts may explain the same product during the live stream, there may be multiple slice IDs corresponding to the same product ID.
[0124] In one possible implementation, the slice identifier is used to query the slice list in the database to obtain the product identifier corresponding to the slice identifier, and then the product identifier is used to query the attribute information of the product in the database.
[0125] The database stores product attribute information, which can include different attribute information for different types of products.
[0126] For example, for tourism products, attribute information may include product name, destination, departure point, product attributes (tourism products), etc.
[0127] For example, for food products, attribute information may include product name, food category, manufacturer, etc.
[0128] It should be noted that the examples of attribute information for tourism products and food products mentioned above are for illustrative purposes only and do not limit the specific content of the product attribute information.
[0129] In one possible implementation, if no corresponding attribute information is found in the database based on the product identifier, the evaluation can be terminated, and the product information relevance evaluation result can be determined as failing the evaluation.
[0130] 402. Using a large model and based on attribute information, evaluate the correlation of product information in the target keyframe image to obtain the product information correlation evaluation results.
[0131] The target keyframe image and the attribute information of the target product are input into the large model. The large model uses the attribute information to evaluate the product information in the target keyframe image and obtains the product information correlation evaluation result for the target keyframe image.
[0132] In one possible implementation, the attribute information of the target product can be used as a cue word, which is input into the large model along with the target keyframe image. The cue word serves as data in the temporary knowledge base of the large model, and the large model evaluates the correlation between the product information and attribute information in the target keyframe image.
[0133] In one possible implementation, the large model has image recognition capabilities. The large model can compare the attribute information of the product with the product information present in the target keyframe image to determine the degree of matching between the two, and then obtain the correlation evaluation result of the product information and attribute information.
[0134] The large model can establish a systematic evaluation standard for product presentation segments, thereby quantifying the image quality of the segments. This evaluation standard takes into account the image quality of the presentation segment and the correlation between the product being presented and its attribute information. Therefore, by using this large model to analyze the presentation segments, an evaluation result for the image quality of the presentation segments can be obtained. This evaluation result combines image quality, the product information being presented, and the correlation between the product attribute information and the product information.
[0135] In this embodiment, based on the slice identifier of the narration slice, the attribute information of the target product is obtained from the database. Using a large model, based on this attribute information, the product information relevance of the target keyframe image is evaluated, yielding a product information relevance evaluation result. Utilizing the existing attribute information of the target product in the database, the relevance of the product information contained in the target keyframe image to the target product's attribute information is evaluated, yielding the product information relevance evaluation result. Since the evaluation condition used in the large model is the product's attribute information in the database, a unified image evaluation standard is achieved to evaluate the image content of the narration slice, resulting in a standardized evaluation result.
[0136] Figure 5 This is a flowchart illustrating how a large model, based on attribute information, evaluates the relevance of product information to a target keyframe image to obtain the evaluation result. It may include steps 501 to 503, which are described in detail below.
[0137] 501. Using a large model, extract the product information from the target keyframe image to obtain the product information present in the image of the target keyframe.
[0138] This large model has image recognition capabilities. By using the large model to extract product information from the keyframe image of the target, the product information present in the image can be obtained.
[0139] Because the host may add promotional advertising language when introducing products during the live stream, this language may be displayed on the screen, or the product information to be explained may be placed on the screen. Therefore, the target keyframe image may contain a lot of information, including some real information about the target product, some advertising language, and even information about other products to be explained during the live stream. Therefore, the large model needs to distinguish between real product information and other information (advertising language, product information of other products) that appear in the target keyframe image.
[0140] In one possible implementation, product information prompts can be set according to the product information extracted from the target. These product information prompts are used to instruct the large model to analyze the target keyframe image to obtain the product information of the product. The product information is then communicated to the large model through these product information prompts.
[0141] For example, the product information prompt can specify the name of the product to be extracted, and this name can occupy the largest area in the product image.
[0142] The product information prompt can also be similar to the preset prompt in step 301 above, and can be sent to the large model as data in the large model's temporary knowledge base.
[0143] In one possible implementation, a request can be generated and sent to a large model. This request carries the product information prompt and the target keyframe image together with the product information prompt. The large model then extracts the product information from the target keyframe image based on the product information prompt to obtain the product information present in the image.
[0144] By defining clear product information prompts, background knowledge is injected into the large model during the extraction of product information from the target keyframe image. This allows the large model to learn the meaning of "product information," extract the product information needed for comparison, eliminate redundant information, and provide an information basis for subsequent comparison.
[0145] 502. Using a large model, semantic comparison is performed between attribute information and product information to obtain matching information between attribute information and product information;
[0146] After extracting the product information of the target product from the target keyframe image, the large model performs a semantic comparison between the attribute information of the target product obtained from the database and the product information to determine whether the two semantics match, thus obtaining matching information, which represents the degree of matching between the two.
[0147] The degree of matching corresponds to the semantic consistency. The more consistent the semantics of the attribute information and the product information are, the higher the degree of matching between them. Conversely, the degree of matching between them is lower.
[0148] Since the product's attribute information can be provided by the business party (the product provider), it serves as "template" information and can be sent to the large model as contextual information in the prompts. Correspondingly, the large model uses this product attribute information, which is used as "template" information, to compare it with the product information, thus obtaining matching information between the product's attributes and the product information.
[0149] In one possible implementation, the product information extracted from the target keyframe image can be converted into text format. The product attribute information obtained from the database is also in text format. The two text formats are compared to determine their matching information.
[0150] The matching information between the product's attribute information and the product information reflects the correlation between the product information extracted from the screen and the attribute information of the target product. The higher the degree of matching, the stronger the correlation between the two, and the more attribute information of the product that needs to be displayed in the screen, which meets the expectations of the business. Conversely, the lower the degree of matching, the weaker the correlation, and the less attribute information of the product that needs to be displayed in the screen, which does not meet the expectations of the business.
[0151] In one possible implementation, step 502 includes:
[0152] 5021. Using a large model, classify product information into at least one category based on at least one category of attribute information;
[0153] Since the essence of comparison is semantic understanding, and the product attribute information provided by the business side contains rich and sufficient content, once it exceeds the length range of the large model's input, the large model cannot accurately compare the product information and attribute information. Therefore, this attribute information can be pre-divided into multiple categories, and subsequent comparisons can be performed separately for each category.
[0154] In one possible implementation, the attribute information of the target product can be pre-classified according to a set category. The classified category is then sent to the large model as contextual information in the prompt words, serving as data in the large model's temporary knowledge base. This allows the large model to learn the different classifications of the product's attribute information, thereby classifying the extracted product information according to the category.
[0155] In one possible implementation, the database can store the attribute information of the target product according to the classification, and the attribute information obtained from the database is already classified; alternatively, the attribute information of the target product obtained from the database can be classified according to a preset classification category to obtain attribute information for each category; alternatively, after obtaining the attribute information of the target product from the database, the large model can classify the attribute information according to category prompts.
[0156] In one possible implementation, the large model can determine the categories into which attribute information can be divided, as well as the rules for categorization, based on the categories in the prompt words. The large model then categorizes the product information contained in the image according to these categories, resulting in at least one category.
[0157] As an example, this category includes product name, food category, and manufacturer. Accordingly, the large model classifies the extracted product information according to product name, food category, and manufacturer, resulting in a product name of "Crispy Potato", a food category of "Potato Chips", and a manufacturer of "A Food Company".
[0158] As an example, this category includes route name, number of attractions, attraction names, and number of days. Accordingly, the large model classifies the extracted product information according to route name, number of attractions, attraction names, and number of days, resulting in a product name of "Beijing One-Day Tour", 2 attractions, attraction names including the Summer Palace and Tiananmen Square, and 1 day.
[0159] Of course, the above categories are only for illustrative purposes. In actual implementation, the categories used by this large model to classify product information are not limited to the categories mentioned above.
[0160] Considering accuracy and cost factors, the extracted product information and the target product's attribute information can be input into the large model together, and the product information prompts corresponding to the category can also be sent to the large model. The large model then processes the product information prompts.
[0161] 5022. Using a large model and category as the basis, analyze product information and attribute information to obtain matching information for each category;
[0162] After classifying product information into categories, the large model performs a horizontal comparison between attribute information and product information based on the categories to obtain matching information between the two.
[0163] In one possible implementation, the large model performs semantic analysis and comparison on attribute information and product information within the same category. It can first expand any information (attribute information / product information) using synonyms or near-synonyms, then compares the expanded words. If the words for the other information (product information / attribute information) appear in the expanded word set, the two can be considered semantically similar. If the words used for attribute information and product information within the same category are exactly the same, the two can be considered semantically identical. Semantic similarity or identicalness can be considered semantically consistent. If the semantics of attribute information and product information within the same category are consistent, the two can be considered semantically matched; otherwise, they are not semantically matched.
[0164] During the comparison process, for sensitive information that requires complete textual consistency, an explanation should be added to the product information prompts to instruct such information to undergo text matching.
[0165] As an example, if the sensitive information that requires complete textual consistency is the product name, the product name can be converted to text format, and the product name read from the database can also be converted to text format. The two can then be matched. If the two texts are completely identical, the semantic match of the product names is determined; otherwise, they are not matched.
[0166] In one possible implementation, different score values can be assigned to each category based on the matching information of each category. For example, if the attribute information of a certain category semantically matches the product information, it can be assigned a higher score value; otherwise, it can be assigned a lower score value.
[0167] For the prompt words of the large model, a step-by-step instruction is adopted. For each step, the large model is instructed to execute by prompt words. The prompt words define the boundary method of the large model processing, and the attribute information categories are divided in detail. Moreover, for the judgment of the matching process, "consistency judgment" can be emphasized, which can effectively prevent instruction drift and improve the accuracy of large model processing.
[0168] 5023. Based on the matching information of each category, determine the matching information of the product's attribute information and product information.
[0169] Based on the category, the product attribute information and product information are compared horizontally within the same category. After obtaining the matching information for each category, the matching information of multiple categories is accumulated to obtain the matching information of the product attribute information and product information.
[0170] In one possible implementation, the large model can compare product information and corresponding target product attribute information in multiple explanatory slices at once, according to category. This allows for the classification and comparison of multiple sets of information in a single interaction. Since the large model understands relatively consistent classification standards, it can achieve evaluation based on a unified standard. This maintains stable evaluation quality while reducing the number of interactions with the large model. These multiple explanatory slices can be slices explaining the same target product or slices explaining different target products.
[0171] In one possible implementation, after obtaining the matching information for each category, the matching information for each category can be combined to obtain the matching information between the attribute information and the product information.
[0172] 503. Based on the matching information, the product information relevance assessment results are obtained.
[0173] This matching information represents the matching status between the target product's attribute information and the product information. Based on this matching information, the large model can obtain the product information correlation evaluation results of the target keyframe image.
[0174] The more semantic matches between the product information and attribute information in the matching information, the better the correlation between the product information in the target keyframe image and the attribute information of the target product. This can determine that the attribute information of the target product is presented in a better way during the live broadcast, thus achieving effective placement of key product information.
[0175] In one possible implementation, since different score values are assigned to each category according to different matching information, the score values assigned to each category can be accumulated during the process of determining the matching information between attribute information and product image information to obtain the score corresponding to the explanation slice.
[0176] As an example, this large model compares product image information and attribute information according to five categories, obtaining five matching information corresponding to each category. The matching information for categories 1 and 3-5 represents semantic matching within the corresponding category and can be assigned an integral score of 1. The matching information for category 2 represents semantic mismatch within the corresponding category and can be assigned an integral score of 0. The scores for the five categories are summed to obtain a total score of 4. Therefore, the large model's evaluation result for the relevance of product information in the target keyframe image is 4.
[0177] Since the process of evaluating the relevance of product information to the target keyframe image combines the content of product display and product key information patches, after obtaining the evaluation results of product information relevance, the subsequent evaluation results of the quality of the explanatory slice can be obtained. When outputting the evaluation results, suggestions can be given to live streamers or merchants for slice recording based on the evaluation results of product information relevance, especially for categories with semantic mismatch.
[0178] In this embodiment, firstly, a large model is used to extract product information from the target keyframe image, obtaining the product information present in the target keyframe image. Then, the large model compares the attribute information of the target product with the product information to obtain matching information between the product attributes and the product information. Finally, based on this matching information, a product information correlation assessment result is obtained. In this process, the large model extracts product information from the target keyframe image and compares it with the product attribute information to obtain matching information. Finally, the matching information is used to obtain a product information correlation assessment result for the target keyframe image. This achieves a standardized assessment of the correlation between product information and attribute information present in the explanatory slices through a unified evaluation standard, resulting in a standardized assessment result.
[0179] Figure 6 This is a flowchart illustrating the process of obtaining the quality assessment result of the explanatory slice based on the evaluation result of the screen content and the evaluation result of the correlation between the product information, as provided in this application embodiment. It may include steps 601 to 603, which are described in detail below.
[0180] 601. Based on the evaluation results of the screen content, the first score is obtained. The evaluation result of the screen content indicates that the evaluation is passed, and the first score is the first value. The evaluation result of the screen content indicates that the evaluation is not passed, and the first score is the second value. The first value is higher than the second value.
[0181] During the portrait evaluation of target keyframe images using a large model, evaluation rules are combined to obtain the corresponding image content evaluation result. This image content evaluation result includes whether the evaluation is passed or failed.
[0182] To quantify the evaluation results of the target keyframe images, an integral principle can be applied to the evaluation process. Different scoring rules can be applied to the evaluation results of the image content and the evaluation results of the relevance of product information. Satisfaction and dissatisfaction are sufficient to form different scoring records.
[0183] In one possible implementation, different scores are set for the evaluation results of the screen content, based on whether the evaluation passes or fails.
[0184] Generally, higher scores are given to those who pass the assessment, and lower scores are given to those who fail.
[0185] As an example, the first score is 1 and the second score is 0. This application does not impose any restrictions on the specific values of the first and second scores.
[0186] 602. Based on the assessment results of the relevance of product information, the second score is obtained;
[0187] In one possible implementation, the product information relevance assessment result can characterize whether each category in the product information passes or fails the assessment, and correspondingly, different scores can be set for different situations.
[0188] In one possible implementation, since the product information contains a lot of content, it can be assigned scores according to the categories in the evaluation process, and the sum of the scores of each category can be used as the score of the product information relevance evaluation result.
[0189] 603. Based on the first and second integrals, the evaluation results of the quality of the explanatory slice are obtained.
[0190] Since the evaluation of the quality of the explanatory video clip is conducted from two dimensions—the human figure dimension and the product information relevance dimension—and the two evaluations have different focuses, the two interactions with the large model can be isolated. The evaluation results of the two are then combined to obtain the overall evaluation result of the explanatory video clip's quality.
[0191] In one possible implementation, the first integral and the second integral are summed to obtain the score for evaluating the quality of the explanatory slice.
[0192] In one possible implementation, the presentation segment will have a cold start quality score when it is generated. This score is determined based on a variety of factors such as the streamer's historical data, number of likes, exposure, and product conversion rate.
[0193] The score of the quality assessment of the explanatory video segment is added to the cold start quality score to obtain the total quality score of the explanatory video segment. The priority of subsequent explanatory video segments on the live streaming platform is based on the quality score.
[0194] The live streaming platform sets up a delivery strategy based on the priority of the explanatory segments. The higher the priority, the more frequently the segments are delivered, in order to improve the overall quality of the segments on the live streaming platform and increase the coverage of products delivered by the segments.
[0195] In extreme cases, if two explanatory videos of the same product have the same quality score, they will be sorted according to the time when the videos were generated. The newer the video, the higher its priority. That is, the video with the highest quality score and the newest video will be in the first place.
[0196] In this embodiment, a first score is obtained based on the evaluation result of the screen content. The screen content evaluation result indicates that the evaluation is passed, and the first score is the first value. The screen content evaluation result indicates that the evaluation is failed, and the first score is the second value. The first value is higher than the second value. A second score is obtained based on the evaluation result of the relevance of product information. Based on the first score and the second score, the evaluation result of the screen quality of the explanation segment is obtained. The screen quality of the explanation segment is quantified by the integral method, which is more efficient and more objective than traditional manual evaluation.
[0197] Figure 7 This is a flowchart illustrating the process of obtaining a target keyframe from a received explanatory slice, provided in an embodiment of this application. It may include steps 701 to 703, which are described in detail below.
[0198] 701. Using the preset extraction tool, extract the first keyframe image after a preset time from the start of the explanation slice;
[0199] Preset extraction tools can be used to extract keyframe images from the explanatory slices.
[0200] Since audio and video development tools generally have the function of converting digital audio and video and converting them into streaming data, the preset extraction tool can be an audio and video development tool, such as FF (Fast Forward) or MPEG (Moving Pictures Experts Group).
[0201] In the process of extracting keyframe images from the narration slice, the preset extraction tool can obtain keyframe images while maintaining the original quality parameters of the narration slice, without re-encoding the narration slice, so as to maintain the image quality of the narration slice as much as possible.
[0202] Furthermore, since each frame in the narration slice undergoes lossy compression, the keyframe images are saved at the highest quality after extraction to ensure that the image quality in the extracted images is consistent with that in the narration slice.
[0203] Since the appearance of keyframes is determined by the streaming method during live broadcasting, the streaming processing function of the FFMPEG tool can be used to extract keyframes. This can be achieved by adding video filter parameters to the FFMPEG command, filtering out keyframes, and using the first filtered keyframe as the first keyframe image.
[0204] Since the narration clip is extracted from the live stream video, it's possible that the product isn't fully displayed at the beginning of the clip, or that the live streamer isn't near the product, or that the product isn't even in the live stream frame. As the narration progresses, the product will be placed in the narration frame, and the live streamer will be near the product while narrating. Therefore, keyframe images can be extracted after a preset time period at the beginning of the narration clip.
[0205] In one possible implementation, a preset time can be set to trigger a preset extraction tool to extract keyframe images after a preset time has elapsed since the start of the narration slice.
[0206] The keyframe image extracted for this target is the image of the first keyframe after a preset time.
[0207] As an example, the preset time can be a relatively long time, such as 30 seconds. Of course, the value of the preset time is not limited to this and can be set according to the actual situation. This application does not impose any restrictions.
[0208] 702. Based on the extracted first keyframe image, use the first keyframe image as the target keyframe image;
[0209] Since the explanation slice focuses on a single product and needs to get straight to the point, the first keyframe in the explanation slice should contain all the product information. Therefore, only one keyframe image can be used as a basis to evaluate the image quality of the explanation slice.
[0210] In one possible implementation, if the first keyframe image is extracted, it is used as the target keyframe image for subsequent processing, and there is no need to perform subsequent step 703.
[0211] In one possible implementation, determining whether the first keyframe image has been extracted includes:
[0212] Based on the preset extraction tool, key frame images are extracted from the explanatory slices within a preset time period, and it is determined that the first key frame has been extracted.
[0213] If the preset extraction tool fails to extract keyframe images from the explanatory slices within a preset time period, it is determined that the first keyframe has not been extracted.
[0214] Live streaming may be interrupted due to network fluctuations, server issues, or the stream ending. Therefore, the explanatory segments extracted from the live video may also be affected. Setting a timeout prevents the program from waiting indefinitely. This application uses a preset time period to represent the timeout duration. If a keyframe image is extracted within the preset time period, that keyframe image is used as the target keyframe image, indicating that the extraction did not time out; otherwise, the extraction timed out.
[0215] In one possible implementation, the process of extracting keyframe images for the preset extraction tool can be set as a task, and an extraction timeout can be set to avoid situations where abnormalities cause the preset extraction tool to block the execution of the task.
[0216] If the preset extraction tool extracts the first keyframe image from the explanatory slice within the timeout period, the extracted first keyframe image can be used as the target keyframe image. If the first keyframe image is not extracted from the explanatory slice within the timeout period, it can be determined that the extraction has failed, and the keyframe image is extracted from the subsequent image frames in the explanatory slice, and step 703 is executed.
[0217] 703. Based on the fact that the first keyframe image was not extracted, extract the second keyframe image and use the second keyframe image as the target keyframe image. The second keyframe image is the one in the slice whose time is later than the first keyframe image. The second keyframe is a keyframe that is directly adjacent to or indirectly adjacent to the first keyframe.
[0218] If the first keyframe image is not extracted, the second keyframe image is extracted from the subsequent image frames in this explanatory slice. The extraction method is the same as that for the first keyframe image, and will not be repeated here.
[0219] If a second keyframe image is extracted, it is used as the target keyframe image. The second keyframe image can be the first keyframe image after the first keyframe image, or it can be the second or third keyframe image after the first keyframe image.
[0220] In one possible implementation, the number of extraction attempts is counted. If a keyframe image is successfully extracted before the preset number of attempts is reached, the extracted keyframe image is used as the target keyframe image. If the keyframe image is not successfully extracted after the preset number of attempts is reached, the image quality of the narration segment can be judged to be poor, and a low score, such as 0 points, can be assigned to the narration segment.
[0221] After extracting the target keyframe image, it can be uploaded to the media library for storage. Subsequently, the image links and prompts in the media library can be used to generate a request, which is then sent to the large model. The large model uses the image links to obtain the target keyframe image from the media library and uses the prompts to process the target keyframe image to obtain the corresponding screen content evaluation results and product quality evaluation results.
[0222] In this embodiment, a preset extraction tool is used to extract a first keyframe image after a preset time from the start of the explanatory slice. Based on the extraction of the first keyframe image, it is used as the target keyframe image. If the first keyframe image is not extracted, a second keyframe image is extracted from subsequent image frames in the explanatory slice, and this second keyframe image is used as the target keyframe image. Using a preset extraction tool to extract keyframe images from the explanatory slice improves extraction efficiency.
[0223] Figure 8 This is a flowchart illustrating an application scenario of a method for evaluating the quality of a live-streamed narration segment, as provided in this application embodiment. In this scenario, a separate program for this evaluation method is created. This program accepts a link to a live-streamed narration segment to be evaluated. The program also determines the length of the narration segment; if the length is less than a set duration threshold (e.g., 30 seconds), the program terminates and does not trigger subsequent tasks. The program consists of a screen extraction program, a screen recognition program, and a scoring program. The screen extraction program uses FFMPEG to extract keyframes from the narration segment. The screen recognition program includes a basic evaluation task and a correlation evaluation task. The basic evaluation task triggers a large model to evaluate the human figures in the keyframe images, and the correlation evaluation task triggers a large model to evaluate the products in the keyframe images. The scoring program records the scores based on the evaluation results.
[0224] In this application scenario, after determining that the explanatory segments are suitable for evaluation, the following evaluation process is included:
[0225] 801. Capture the frame;
[0226] FFMPEG is used to capture the target keyframe from the narration slice. In this application scenario, the capture is only performed once.
[0227] 802. Determine if it was successful;
[0228] If the capture is successful, proceed to step 803; otherwise, end.
[0229] 803. Upload to media library;
[0230] The captured keyframes of the target image are uploaded to the media library for storage.
[0231] The target keyframes from the media library are used to perform basic evaluation and correlation evaluation tasks, respectively.
[0232] The basic assessment task execution process includes steps 804-808, and the correlation assessment task execution process includes steps 809-815.
[0233] 804. Create basic assessment tasks;
[0234] During the creation of this task, a prompt word can be determined, which represents the analysis rules for analyzing the human image in the target keyframe image. A request is then generated to link the prompt word with the target keyframe image in the media library.
[0235] 805. Large-scale model reasoning begins;
[0236] The above request is sent to the large model, which then begins inference based on the request.
[0237] 806. Determine if standard one has been passed;
[0238] The first standard is used to determine whether a human figure appears in the image. If a human figure appears in the image, the first standard passes and the subsequent step 807 is executed. Otherwise, the first standard fails, the reasoning process ends, and the step 816 is executed.
[0239] 807. Determine whether Standard 2 has been passed;
[0240] Standard 2 is used to determine whether the portrait is located in the preset core area of the image. If the portrait is located in the core area, Standard 2 passes and the subsequent step 808 is executed; otherwise, Standard 2 fails, the reasoning process ends, and step 816 is executed.
[0241] 808. Determine if the standard three-pass test is passed;
[0242] Standard 3 is used to determine whether the face of the portrait meets the preset clarity condition. If the portrait is clear, Standard 3 passes, the reasoning process ends, and step 816 is executed. Otherwise, Standard 3 fails, the reasoning process ends, and step 816 is executed.
[0243] 809. Create a correlation assessment task;
[0244] During the creation of this task, prompt words can be determined, which may include consistency judgment prompt words, etc.
[0245] During the creation of this task, prompt words can be determined, which may include product attribute information prompt words, category prompt words, and consistency judgment prompt words, etc., and a request is generated to link the prompt words with the target keyframe images in the media library.
[0246] 810. Obtain product information;
[0247] You can query the corresponding product ID from the slice list in the database according to the ID of the slice, and then retrieve the product information corresponding to that product ID from the database.
[0248] 811. Determine if the acquisition was successful;
[0249] To determine if the acquisition was successful, you can check if you can obtain the product information of the product in the explanation slice from the database. If you can obtain it, proceed to the next step 812; if you cannot obtain it, end the process.
[0250] After obtaining product information, prompt words can also be determined. These prompt words can include product attribute information prompt words and category prompt words.
[0251] Requests can be generated using product attribute information prompts, category prompts, consistency judgment prompts, and links to target keyframe images in the media library.
[0252] 812. Large-scale model reasoning begins;
[0253] The above request is sent to the large model, which then begins to evaluate the relevance of the product to the target keyframe image based on the request and the product information.
[0254] 813. Extract product information;
[0255] The large model extracts product information from the target keyframe images.
[0256] 814. Product information classification;
[0257] The large model uses category clue words to classify the extracted product information, resulting in product information in multiple categories.
[0258] 815. Consistency comparison;
[0259] The large model uses consistency discrimination prompts and product attribute information prompts to compare the product information extracted from the target keyframe with the attribute information obtained from the database according to the category, and determines whether they are consistent, thus obtaining the evaluation result of the correlation between the target keyframe and the attribute information of the product in the database.
[0260] 816. The large-scale model reasoning has ended;
[0261] The evaluation process based on the aforementioned two dimensions has concluded, yielding evaluation results for both dimensions. This concludes the image quality evaluation process for the explanatory slice.
[0262] 817. Record the results.
[0263] The evaluation results of the two dimensions are quantified by integral, and the evaluation results and corresponding integrals of the two dimensions are recorded. The integrals of the two dimensions are added together to obtain the integral of the explanation slice. The evaluation process for this explanation slice is now complete.
[0264] The above describes a method for evaluating the quality of a narration slice image provided by an embodiment of this application. The following will describe an electronic device that performs the above-described method for evaluating the quality of a narration slice image.
[0265] Please see Figure 9 , Figure 9 This is a schematic diagram of an electronic device illustrating a method for evaluating the quality of sliced images, as provided in an embodiment of this application. For example... Figure 9 As shown, the electronic device 900 includes:
[0266] Memory 901 is used to store large models;
[0267] The storage device can be a data storage structure in an electronic device, such as a hard drive or memory, but it is not limited to this. The storage location of large models can be set according to the actual situation.
[0268] Processor 901 is used to receive the narration slice to be evaluated, which is a video clip explaining the target product; extract the target keyframe image from the narration slice; and perform product information relevance evaluation and image content evaluation on the target keyframe image using a large model to obtain the product information relevance evaluation result and the image content evaluation result. Product information relevance represents the correlation between product information in the target keyframe image and the attribute information of the target product, while image content is the relevant information of the human image in the target keyframe image; based on the image content evaluation result and the product information relevance evaluation result, the image quality evaluation result of the narration slice is obtained.
[0269] In one possible implementation, the specific process by which the processor performs its functions and related explanations can be found in the explanations in the foregoing method embodiments, and will not be repeated here.
[0270] In this embodiment, the electronic device includes: a memory for storing a large model; a processor for receiving a narration slice to be evaluated, which is a video clip explaining a target product; extracting target keyframe images from the narration slice; using the large model to perform product information correlation evaluation and image content evaluation on the target keyframe images, obtaining product information correlation evaluation results and image content evaluation results. Product information correlation characterizes the correlation between product information in the target keyframe image and the attribute information of the target product, and image content is the relevant information of the human image in the target keyframe image; based on the image content evaluation results and product information correlation evaluation results, the image quality evaluation result of the narration slice is obtained. Using the large model to evaluate the target keyframe images in the narration slice from two dimensions—the human image dimension and the product information correlation dimension—results in high accuracy. Moreover, because the processing of the large model is stable, its evaluation of the input target keyframe images has a unified evaluation standard, is not affected by personal bias, and can obtain standardized evaluation results.
[0271] This application also provides an electronic device in its embodiments. (See reference...) Figure 10 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 10 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0272] like Figure 10 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. When the electronic device is powered on, the RAM 1003 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0273] Typically, the following devices can be connected to the I / O interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1008 including, for example, memory card, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to exchange data via wireless or wired communication with other devices. Although Figure 10 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0274] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the methods for evaluating the quality of explanatory slice images provided in this application.
[0275] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the methods for evaluating the quality of explanatory slice images provided in this application.
[0276] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0277] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0278] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0279] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0280] In the above embodiments, the implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product.
[0281] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A method of evaluating the quality of a slice picture, characterized by, The method comprises the following steps: receiving a to-be-evaluated explanation clip, the explanation clip being a video segment for explaining a target commodity; extracting a target key frame image from the explanation clip; performing commodity information correlation evaluation and picture content evaluation on the target key frame image by a large model to obtain commodity information correlation evaluation results and picture content evaluation results, the commodity information correlation representing the correlation between commodity information in the target key frame image and attribute information of the target commodity, and the picture content being related information of a person image in the target key frame image; obtaining an evaluation result of the picture quality of the explanation clip according to the picture content evaluation result and the commodity information correlation evaluation result.
2. The method of claim 1, wherein, performing picture content evaluation on the target key frame image by a large model to obtain a picture content evaluation result, comprising: obtaining a preset prompt word, the preset prompt word being used to indicate an analysis rule of the large model for analyzing the picture quality of the person image; sending the preset prompt word and the target key frame image to the large model to enable the large model to determine a corresponding analysis rule according to the preset prompt word and analyze the target key frame image by using the analysis rule to obtain a picture content evaluation result of the target key frame image, the picture content evaluation result representing whether the person image in the target key frame image passes the evaluation or not.
3. The method of claim 2, wherein, The analysis rule comprises: the picture appearing a person image, the person image being located in a preset core area of the picture, and a face of the person image satisfying a preset clearness condition.
4. The method of claim 1, wherein, Performing commodity information correlation evaluation on the target key frame image by a large model to obtain commodity information correlation evaluation results, comprising: obtaining attribute information of a target commodity from a database according to a clip identifier of the explanation clip; performing commodity information correlation evaluation on the target key frame image by a large model according to the attribute information to obtain commodity information correlation evaluation results.
5. The method of claim 4, wherein, The performing commodity information correlation evaluation on the target key frame image by a large model according to the attribute information to obtain commodity information correlation evaluation results comprises: extracting commodity information of the target key frame image by a large model to obtain commodity information existing in the picture of the target key frame image; performing semantic comparison between the attribute information and the commodity information by a large model to obtain matching information of the attribute information and the commodity information; obtaining the commodity information correlation evaluation result according to the matching information.
6. The method of claim 5, wherein, The performing semantic comparison between the attribute information and the commodity information by a large model to obtain matching information of the attribute information and the commodity information comprises: dividing the commodity information into at least one category according to at least one category of the attribute information by a large model; analyzing the matching information of the commodity information and the attribute information by a large model to obtain matching information of each category according to the category as a reference; determining the matching information of the attribute information and the commodity information according to the matching information of each category.
7. The method of claim 1, wherein, The obtaining an evaluation result of the picture quality of the explanation clip according to the picture content evaluation result and the commodity information correlation evaluation result comprises: based on the picture content evaluation result, a first score is obtained, the picture content evaluation result represents passing the evaluation, the first score is a first score value, the picture content evaluation result represents failing the evaluation, the first score is a second score value, the first score value is higher than the second score value; based on the commodity information relevance evaluation result, a second score is obtained; according to the first score and the second score, an evaluation result of the quality of the explanation slice picture is obtained.
8. The method of claim 1, wherein, the target key frame is obtained from the received explanation slice, including: extracting a first key frame image after a preset time from the beginning of the explanation slice through a preset extraction tool; based on the first key frame image being extracted, the first key frame image is taken as a target key frame image; based on the first key frame image not being extracted, a second key frame image is extracted, and the second key frame image is taken as a target key frame image, the second key frame image is later in time than the first key frame image in the explanation slice, and the second key frame is a key frame directly or indirectly adjacent to the first key frame.
9. The method of claim 8, wherein, determining whether the first key frame image is extracted, including: based on the key frame image being extracted from the explanation slice within a preset time period by the preset extraction tool, it is determined that the first key frame is extracted; based on the key frame image not being extracted from the explanation slice within a preset time period by the preset extraction tool, it is determined that the first key frame is not extracted.
10. An electronic device, comprising: a memory for storing a large model; a processor for receiving an explanation slice to be evaluated, the explanation slice being a video segment for explaining a target commodity; extracting a target key frame image from the explanation slice; through a large model, commodity information relevance evaluation and picture content evaluation are performed on the target key frame image respectively, and commodity information relevance evaluation result and picture content evaluation result are obtained, the commodity information relevance represents the relevance between the commodity information in the target key frame image and the attribute information of the target commodity, and the picture content is related information of a portrait in the target key frame image; according to the picture content evaluation result and the commodity information relevance evaluation result, an evaluation result of the quality of the explanation slice picture is obtained.