Method, device and equipment for quality evaluation of cataloging information and medium
By using a large language model to perform multi-dimensional semantic evaluation of the cataloging information of video content platforms, the problem of insufficient accuracy of cataloging information in existing technologies is solved, and efficient and accurate quality evaluation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING QIYI CENTURY SCI & TECH CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies are insufficient to effectively assess the quality of cataloging information on video content platforms, resulting in inaccurate cataloging information.
By acquiring the cataloging information of multimedia content, determining the cataloging content to be evaluated and the comparison cataloging content, filling them into a preset prompt word template to generate evaluation prompt words, and inputting them into a large language model for processing, multi-dimensional evaluation is carried out based on semantic understanding to generate quality evaluation results.
It enables multi-dimensional quality assessment of cataloging information, improves assessment accuracy and processing efficiency, and ensures the accuracy of cataloging information.
Smart Images

Figure CN122113936A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and medium for quality assessment of cataloging information. Background Technology
[0002] In video content platforms, each video needs to be cataloged. Cataloging information includes, for example, the content name, a brief description, tags, and actor information. This cataloging information directly affects user experience, the accuracy of content recommendation algorithms, and the platform's compliance.
[0003] Currently, quality checks on cataloging information of video content platforms employ rule-based methods, classifier-based methods, and regression model-based methods to check for format and keyword errors. However, with the increasing complexity and diversity of cataloging information, these methods are insufficient to meet the quality requirements of business scenarios, and the accuracy of cataloging information quality assessment needs to be improved. Summary of the Invention
[0004] To address the aforementioned technical problems, this disclosure provides a method, apparatus, equipment, and medium for quality assessment of cataloging information.
[0005] In a first aspect, embodiments of this disclosure provide a method for quality assessment of cataloging information, including: Obtain cataloging information of multimedia content, and determine from the cataloging information the cataloging content to be evaluated and the corresponding reference cataloging content; the reference cataloging content includes other cataloging content in the cataloging information used to evaluate the cataloging content to be evaluated; The cataloging content to be evaluated and the comparison cataloging content are filled into a preset prompt word template to generate evaluation prompt words for the cataloging information; wherein, the evaluation prompt words include the comparison cataloging content, the cataloging content to be evaluated, and the first prompt information; The evaluation prompt words are input into a large language model for processing to generate a quality evaluation result for the cataloging content to be evaluated; wherein, the first prompt information is used to enable the large language model to evaluate the potential risks of the cataloging content to be evaluated based on semantic understanding to determine the quality evaluation result, and the evaluation of the potential risks includes a combined evaluation based on several dimensions of data and / or the correlation between data in the cataloging information.
[0006] Secondly, embodiments of this disclosure provide a quality assessment device for cataloging information, comprising: The acquisition module is used to acquire cataloging information of multimedia content, and determine the cataloging content to be evaluated and the corresponding reference cataloging content from the cataloging information; the reference cataloging content includes other cataloging content in the cataloging information used to evaluate the cataloging content to be evaluated; The generation module is used to fill the cataloging content to be evaluated and the comparison cataloging content into a preset prompt word template to generate evaluation prompt words for the cataloging information; wherein, the evaluation prompt words include the comparison cataloging content, the cataloging content to be evaluated, and the first prompt information; The processing module is used to input the evaluation prompt words into a large language model for processing, and generate a quality evaluation result of the cataloging content to be evaluated; wherein, the first prompt information is used to enable the large language model to evaluate the potential risks of the cataloging content to be evaluated based on semantic understanding to determine the quality evaluation result, and the evaluation of the potential risks includes a combined evaluation based on several dimensions of data and / or the correlation between data in the cataloging information.
[0007] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the cataloging information quality assessment method described in the first aspect above.
[0008] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the cataloging information quality assessment method described in the first aspect.
[0009] Compared with the prior art, the technical solution provided in this disclosure has the following advantages: By determining the cataloging content to be evaluated and the corresponding reference cataloging content from the cataloging information of multimedia content, the cataloging content to be evaluated and the reference cataloging content are filled into a preset prompt word template to generate evaluation prompt words for the cataloging information. Then, the evaluation prompt words are input into a large language model for processing. The first prompt information in the evaluation prompt words enables the large language model to evaluate the potential risks of the cataloging content to be evaluated based on semantic understanding, so as to generate the quality evaluation result of the cataloging content to be evaluated. Thus, the potential risks of the cataloging information can be analyzed by the large language model based on semantic understanding, realizing multi-dimensional quality evaluation of the cataloging information. A single call can cover multiple evaluation dimensions, improving processing efficiency, and the evaluation accuracy is higher than that of rule-based methods, ensuring the accuracy of the cataloging information. Attached Figure Description
[0010] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0011] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart illustrating a method for quality assessment of cataloging information provided in an embodiment of this disclosure; Figure 2 A flowchart illustrating another method for quality assessment of cataloging information provided in this embodiment of the present disclosure; Figure 3 A flowchart illustrating another method for quality assessment of cataloging information provided in this embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of a cataloging information quality assessment device provided in an embodiment of this disclosure. Detailed Implementation
[0013] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0014] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0015] Figure 1 This is a flowchart illustrating a method for quality assessment of cataloging information provided in an embodiment of this disclosure. The method provided in this embodiment can be executed by a device for quality assessment of cataloging information. This device can be implemented using software and / or hardware and can be integrated into any electronic device with computing capabilities.
[0016] like Figure 1 As shown, the quality assessment method for cataloging information provided in this disclosure embodiment may include: Step 101: Obtain the cataloging information of the multimedia content, and determine the cataloging content to be evaluated and the corresponding reference cataloging content from the cataloging information.
[0017] The method disclosed in this embodiment can be applied to quality assessment of cataloging information of multimedia content on a network platform. The network platform includes various video content platforms, content management system providers, e-commerce platforms, news media platforms, knowledge management platforms, enterprise content management systems, etc. The multimedia content includes, but is not limited to, video content, user-generated content, product content, news content, etc. The cataloging information is descriptive information of the multimedia content. Based on the cataloging information of the multimedia content, further content recommendation, search algorithm optimization, etc., can be performed.
[0018] In this embodiment, taking video content as an example, the video content includes, for example, movies, TV series, variety shows, etc. The cataloging information includes one or more of the following: content description information, tag information, actor information, category information, popularity data, and content rating information. Among them, the content description information includes, but is not limited to, content name, brief description, channel category, online time, release year, etc. The tag information includes, for example, explicit tags, main tags, track tags, V2 tags, etc. The actor information includes, for example, the names of the lead actors, supporting actors, etc. Taking the ranking information as an example, the category information is the ranking category, the popularity data is the ranking popularity, the ranking category indicates the name of the ranking list in which the content is located, and the ranking category includes, for example, the ancient costume ranking list, the romance ranking list, the mythology ranking list, etc. The ranking popularity indicates the ranking position of the content in the ranking list, and the content rating information includes the level label, and the level label includes, for example, "mainstream", "positive energy", "ordinary", etc.
[0019] In this embodiment, the cataloging content to be evaluated is determined from the cataloging information according to the quality assessment requirements, and the corresponding reference cataloging content is determined from the cataloging information. The reference cataloging content includes other cataloging content in the cataloging information used to evaluate the cataloging content to be evaluated. Different cataloging contents to be evaluated may correspond to the same reference cataloging content or different reference cataloging content. A piece of cataloging information may be the cataloging content to be evaluated or the reference cataloging content of another cataloging content to be evaluated. For example, the cataloging content to be evaluated includes tag information and ranking information, and the reference cataloging content includes content description information. Another example is that the cataloging content to be evaluated includes content rating information, and the reference cataloging content includes tag information.
[0020] The steps for obtaining cataloging information are explained below.
[0021] In one embodiment of this disclosure, cataloging information can be obtained from a content management system. Obtaining cataloging information includes: receiving content update notifications via a message queue, or retrieving a content list via a scheduled task; and then, for multimedia content whose cataloging information is to be evaluated, calling the application programming interface (API) of the content management system to obtain complete cataloging information.
[0022] Optionally, after obtaining the cataloging information, the cataloging information is preprocessed. Preprocessing includes data cleaning and format conversion, and evaluation scope filtering. The data cleaning and format conversion steps include: cleaning the cataloging information to remove invalid characters and unify the format; then, converting data of different formats into a unified data structure; and further extracting key fields to construct the input data required for quality assessment. Evaluation scope filtering is used to filter out content types that are not evaluated based on business rules. Taking video content as an example, evaluation scope filtering includes excluding user-generated content, excluding content from specific channels (such as unknown channels), excluding content with empty names or descriptions, and excluding content with empty tags.
[0023] Step 102: Fill the cataloging content to be evaluated and the comparison cataloging content into the preset prompt word template to generate evaluation prompt words for the cataloging information.
[0024] In this embodiment, a Large Language Model (LLM) is used to perform a multi-dimensional structured evaluation of cataloging information. The LLM possesses natural language understanding and generation capabilities, enabling it to understand the semantics, contextual relationships, and logical connections of text. This solution leverages the characteristics of the LLM to construct structured evaluation prompts, performing a semantically understanding-based structured evaluation of the cataloging information from multiple dimensions. Optionally, corresponding prompt templates can be set for different business rules and content types. After obtaining the cataloging information of multimedia content, the cataloging information is filled into the prompt template to generate evaluation prompts.
[0025] The assessment prompts include the reference catalog content, the catalog content to be assessed, and the first prompt information. The first prompt information guides the large language model to understand the reference catalog content, the catalog content to be assessed, and the assessment rules and requirements for quality assessment.
[0026] As an example, taking the evaluation prompts for video content as an example, the cataloged content includes content description information and tag information, while the cataloged content to be evaluated includes tag information, ranking information, and content classification information. The first prompt information includes evaluation rules and evaluation requirements. The evaluation rules include content classification risk assessment rules, tag classification rules, etc. The evaluation requirements, for example, require the large language model to give a score of 0 to 10 for each cataloged content to be evaluated and an evaluation reason of no more than 25 words. In this example, the cataloged content to be evaluated in the evaluation prompts can be displayed in the form of an evaluation list, which may include, for example, an explicit tag list, a ranking list, a main tag list, a track tag list, etc.
[0027] Step 103: Input the evaluation prompts into the large language model for processing to generate the quality evaluation results of the cataloged content to be evaluated.
[0028] In this embodiment, the first prompt information is used to enable the large language model to assess the potential risks of the cataloged content to be evaluated based on semantic understanding to determine the quality assessment result. The assessment of potential risks includes a combined assessment based on several dimensions of data and / or the correlations between data in the cataloging information. Categories of potential risks include, for example, semantic correlation, logical consistency, and compliance with business rules. The adopted dimensional data includes, for example, tag information, category information, and popularity data.
[0029] The following example illustrates the assessment of potential risks.
[0030] As an example, the large language model uses semantic understanding to determine the semantics of the cataloging content to be evaluated and the reference cataloging content, thereby determining the semantic relevance between them. Based on this semantic relevance, it judges whether the cataloging content to be evaluated has potential risks. For example, if the semantic relevance is lower than a preset threshold, the cataloging content to be evaluated has potential risks. In this example, taking a scoring format for quality assessment results, the scoring rules can be as follows: establish a mapping relationship between semantic relevance and scores, with different semantic relevance levels corresponding to different scores; or compare the semantic relevance with a preset threshold. If the semantic relevance is lower than the preset threshold, the first score is used; otherwise, the second score is used, where the second score is greater than the first score. It should be noted that the above scoring rules are an example and can be adjusted according to the actual application scenario; no restrictions are imposed here.
[0031] As another example, the large language model uses semantic understanding to determine the semantics of the cataloging content to be evaluated and the comparison cataloging content separately. Based on semantics, it judges whether the cataloging content to be evaluated and the comparison cataloging content involve content of a preset dimension. The preset dimension can be set according to the needs of actual application scenarios, such as including popularity ranking. Then, when the preset dimension is involved, it judges the degree of consistency of the meaning of the cataloging content to be evaluated and the comparison cataloging content on that preset dimension based on semantics, to determine whether the cataloging content to be evaluated has potential risks. For example, if the degree of consistency of meaning on that preset dimension is lower than a preset threshold, then the cataloging content to be evaluated has potential risks. In this example, the scoring rules for the quality assessment results can be as follows: establish a mapping relationship between the degree of consistency and the score, with different degrees of consistency corresponding to different scores; or compare the degree of consistency with a preset threshold. If the degree of consistency is lower than the preset threshold, the third score is used; otherwise, the fourth score is used, where the fourth score is greater than the third score.
[0032] As another example, the large language model uses semantic understanding to determine the semantics of the cataloging content to be evaluated and the reference cataloging content separately. Based on the category of the cataloging content to be evaluated, it determines whether business risk assessment is involved. This can be achieved by pre-setting the categories of cataloging content involved in business risk assessment and the corresponding business risk assessment rules. Business risk assessment includes, but is not limited to, sensitive content assessment. Furthermore, when business risk assessment is involved, the number of business risky contents present in the reference cataloging content is determined based on semantics. The number of business risky contents is then used to determine whether the cataloging content to be evaluated has potential risks. The business risky contents can be determined through semantics and business risk assessment rules. In this example, the scoring rules for the quality assessment results can be as follows: establish a mapping relationship between the number of business risky contents and the score. Different numbers of business risky contents can correspond to different scores. For example, the more business risky contents, the lower the score; or, if the number of business risky contents is zero, the sixth score is used; if the number of business risky contents is greater than zero, the fifth score is used. The sixth score is greater than the fifth score.
[0033] The above illustrates an example of potential risk assessment. For any cataloging content in the cataloging information, potential risk assessment can be conducted using one or more of the above methods in combination, without any restrictions.
[0034] Optionally, the application programming interface (API) of the large language model is invoked, with the constructed evaluation prompts as input. The large language model returns a structured JSON-formatted quality evaluation result through the API, which may include, for example, a score and the reason for the score.
[0035] In one embodiment of this disclosure, the cataloging content to be evaluated includes tag information, and the reference cataloging content includes content description information. Evaluation prompts are input into a large language model for processing to generate a quality evaluation result for the cataloging content to be evaluated. This includes: evaluating a first matching degree between each tag and the content description information using the large language model to determine a tag quality score for each tag. The first matching degree is positively correlated with the tag quality score; for example, the higher the semantic relevance, logical consistency, and business rule compliance between the tag and the content description information, the higher the tag quality score, and the higher the accuracy and reasonableness of the tag.
[0036] As an example, if the tag information includes the tag "historical drama" and the content description information includes the content name "historical legend," the first match degree of the tag is evaluated using a large language model, and the tag quality score is determined to be 'a'. In another case, if the tag information includes the tag "historical drama" and the content description information includes the brief description "...a story from a certain period...", the first match degree of the tag is evaluated using a large language model, and the tag quality score is determined to be 'b'. In this example, the semantic relevance between the tag "historical drama" and the content name "historical legend" is greater than a preset threshold, and the tag quality score 'a' is greater than the preset score. Conversely, the semantic relevance between the tag "historical drama" and the brief description "...a story from a certain period..." is less than the preset threshold, and the tag quality score 'b' is less than the preset score. For example, the evaluation requires the large language model to give a score from 0 to 10, with a preset score of 5. For each tag quality score, the large language model can also output the reason corresponding to the tag quality score.
[0037] Optionally, a weighted sum can be performed based on the tag quality score of each tag to obtain the overall tag quality score and the reasons corresponding to the overall tag quality score. For example, an overall tag quality score and processing suggestions can be generated for all explicit tags.
[0038] Optionally, if the tag information includes actor tags, and the catalog content also includes actor information, a large language model is used to evaluate the matching degree between each tag information and the actor information to determine the tag quality score for each tag information. This matching degree is positively correlated with the tag quality score.
[0039] In one embodiment of this disclosure, the cataloging content to be evaluated includes category information, and the comparison cataloging content includes content description information. Evaluation prompts are input into a large language model for processing to generate a quality evaluation result for the cataloging content to be evaluated. This includes: obtaining the category information from the cataloging content to be evaluated; evaluating the second matching degree between the category information and the content description information using the large language model to determine a category matching degree score. The second matching degree is positively correlated with the category matching degree score; for example, the higher the semantic relevance, logical consistency, and business rule compliance between the category information and the content description information, the higher the category matching degree score.
[0040] As an example, let's take a ranking list as an example. The ranking list information is "5th on the list of most popular historical dramas." By identifying the ranking list information, we obtain the ranking category "historical drama." We then use a large language model to evaluate the second matching degree of this ranking category. If the content description information includes the content name "historical legend," the ranking category matching degree score is determined to be c. In another case, the content name is "modern...story," and the ranking category matching degree score is d. In this example, the semantic relevance between the ranking category "historical drama" and the content name "historical legend" is greater than a preset threshold, and the ranking category matching degree score c is greater than a preset score. Conversely, the semantic relevance between the ranking category "historical drama" and the content name "modern...story" is less than a preset threshold, and the ranking category matching degree score d is less than a preset score. For example, the evaluation requires the large language model to give a score from 0 to 10, and the preset score is, for example, 5.
[0041] In one embodiment of this disclosure, reference cataloging content for comparison can also be obtained. The reference cataloging content can be cataloging content of the same type as the cataloging content to be evaluated. The first prompt information is also used to prompt the large language model to evaluate the degree of matching between the cataloging content to be evaluated and the reference cataloging content to determine the quality evaluation result. The reference cataloging content includes reference popularity data, reference introduction descriptions, reference actor information, etc.
[0042] In this embodiment, taking popularity data as an example, the evaluation prompts are input into a large language model for processing to generate a quality evaluation result for the cataloged content to be evaluated. This includes: obtaining target popularity data from the cataloged content to be evaluated; evaluating the third degree of matching between the target popularity data and the reference popularity data of the multimedia content using the large language model to determine the popularity matching score. The third degree of matching is positively correlated with the popularity matching score; for example, the higher the semantic relevance, logical consistency, and business rule compliance between the target popularity data and the reference popularity data, the higher the popularity matching score.
[0043] As an example, the list information is "5th place on the list of most popular historical dramas". By identifying the list information, the popularity data of the list is "5th place". The third matching degree of the popularity data of the list is evaluated by the large language model. If the list is not on the list in the reference popularity data, the consistency between the list popularity data and the reference popularity data is lower than the preset threshold. The matching degree of the list popularity is determined to be e. e is less than the preset score. For example, the evaluation requires the large language model to give a score of 0 to 10. The preset score is, for example, 5 points.
[0044] Optionally, a weighted sum can be calculated based on the ranking category matching score and the ranking popularity matching score to obtain the ranking information matching score and the reason corresponding to the ranking matching score.
[0045] In one embodiment of this disclosure, the cataloging content to be evaluated includes content classification information, and the corresponding cataloging content includes tag information. Evaluation prompts are input into a large language model for processing to generate a quality assessment result for the cataloging content to be evaluated. This includes: evaluating the fourth matching degree between the content classification information and the tag information using the large language model to determine the risk score of the content classification information. The fourth matching degree is negatively correlated with the risk score; for example, the lower the semantic relevance, logical consistency, and business rule compliance between the content classification information and the tag information, the higher the risk score, indicating a greater risk.
[0046] As an example, the content classification information includes the classification label "mainstream". The fourth matching degree of the content classification information is evaluated by the big language model. If the label information contains negative labels, that is, the number of business risk content in the label information is greater than zero, the risk score is determined to be f. If f is greater than the preset score, it indicates that there is public opinion risk. The big language model can also output the reasons corresponding to the risk score.
[0047] In this embodiment, a multi-level, multi-dimensional multimedia content cataloging information quality inspection system is implemented through a large language model. After determining the quality assessment results of the cataloging information, the system can perform multimedia content cataloging quality monitoring and correction, content recommendation and search algorithm optimization, compliance risk prevention and control, and operation strategy formulation based on the quality assessment results of the cataloging information.
[0048] According to the technical solution of this disclosure, by determining the cataloging content to be evaluated and the corresponding reference cataloging content from the cataloging information of multimedia content, the cataloging content to be evaluated and the reference cataloging content are filled into a preset prompt word template to generate evaluation prompt words for the cataloging information. Then, the evaluation prompt words are input into a large language model for processing. Based on semantic understanding, the first prompt information in the evaluation prompt words generates a quality evaluation result for the cataloging content to be evaluated. Thus, the semantic relationship, logical consistency, and business rule compliance between cataloging information can be analyzed by the large language model based on semantic understanding. It can be applied to various types of information such as tags and lists in the cataloging information to achieve multi-dimensional quality evaluation of the cataloging information. A single call can cover multiple evaluation dimensions, improving processing efficiency and ensuring higher evaluation accuracy than rule-based methods, thus guaranteeing the accuracy of the cataloging information.
[0049] Based on the above embodiments, Figure 2 This is a flowchart illustrating another method for assessing the quality of cataloging information provided in an embodiment of this disclosure, as shown below. Figure 2 As shown, after generating the quality assessment results for the cataloging content to be evaluated, the method further includes: Step 201: If the quality assessment result of the cataloging content to be evaluated does not meet the preset accuracy conditions, generate review prompt words based on the cataloging content to be evaluated.
[0050] In this embodiment, the quality assessment result includes a score and a reason. Whether the quality assessment result meets the preset accuracy condition includes comparing the score of the quality assessment result with a preset score to determine whether the preset accuracy condition is met. For example, if the label quality score is less than the preset score, the preset accuracy condition is not met.
[0051] As an example, taking tag information as an example, the multimedia content cataloging information quality inspection system traverses the tag quality scores of each tag and identifies the first tag whose quality score is lower than a preset score. Then, for each first tag, it checks whether a review task already exists. If a review task already exists and the review status is "successful," the current review is not executed, and the existing quality review result is used. If no review task exists or the review status is "failed," a review task for the first tag is created. In this example, the created review task includes the following: multimedia content identifier, content name, tag name, tag type, task status (initialized to "pending"), creation time, update time, etc. Optionally, the system uses an asynchronous task processing mechanism to submit the review task to a thread pool for asynchronous execution, uses atomic operations to ensure that the same review task is not processed repeatedly by multiple threads, and updates the task status to "processing" to increase the number of attempts.
[0052] In this embodiment, an artificial intelligence agent is invoked for review, constructing review prompts. These prompts include the cataloged content to be evaluated, a preset data source, and a second prompt message. For example, a review prompt might be: "Please search [preset data source] and determine, based on the following: Is it reasonable to tag '[tag name]' in the video album '[content name]' on [channel]? Brief description: [first 50 words of the brief description]. If reasonable, the score should be higher than 5 points; otherwise, lower than 5 points. The last line outputs only an integer score between 0 and 10."
[0053] Step 202: Input the review prompts into the AI system for processing and generate the quality review results of the cataloging content to be evaluated.
[0054] The second prompt information is used to enable the AI agent to query the review cataloging content of multimedia content from a preset data source, and to evaluate the consistency of the cataloging content to be evaluated based on the review cataloging content to determine the quality review result.
[0055] In this embodiment, an artificial intelligence agent is invoked. Based on the review prompts, the AI agent automatically accesses external data sources and searches for multimedia content information. It extracts the cataloging content of the multimedia content from the search results as the review cataloging content, which includes, for example, tags, types, and brief descriptions. Furthermore, the AI agent evaluates the consistency of the cataloging content to be evaluated based on the review cataloging content to determine the quality review result. The quality review result includes, for example, a review score and the corresponding analytical basis.
[0056] As an example, the multimedia content cataloging information quality inspection system parses the quality review results returned by the AI agent, extracts the review score from the results, and updates the review status of the task to "successful" if the parsing is successful and the score is valid, recording the score, review time, and analysis basis. If the parsing fails or the score is invalid, the review status is updated to "failed" and the error message is recorded. Then, the review results are written back to the multimedia content data record and stored. Optionally, the review score can overwrite the original quality assessment score; for example, the review score of a tag can overwrite the original tag quality score, and the field can be labeled "AI Review X Score".
[0057] In one embodiment of this disclosure, the evaluation prompts further include feedback information, which may be input by a user. The first prompt is also used to enable the large language model to refer to the feedback information when determining the quality evaluation result.
[0058] After determining the quality assessment and / or quality review results of the cataloging information, these results can be displayed on the front-end interface for experts to view. Experts can then input feedback in the feedback input area provided on the front-end interface. For example, in cases of false alarms or omissions, the feedback information can explain the correct assessment result, the reason why the current assessment result is incorrect, and the rules or standards that should be followed in the assessment. Furthermore, the system stores the feedback information in the data record of the corresponding multimedia content and associates the feedback information with the multimedia content identifier for persistent storage.
[0059] As an example, when evaluating the cataloging information of this multimedia content again, the system automatically reads the saved feedback information and adds it to the evaluation prompts. For instance, the evaluation prompts might include "Historical evaluation feedback: The tag 'historical drama' for this content is reasonable because the plot does indeed take place in ancient times. Please re-evaluate." The large language model refers to the feedback information during the evaluation and corrects the quality evaluation results. In this example, the effectiveness of the feedback can be verified. For example, the system records the quality evaluation results after applying the feedback information and compares the results before and after the application to verify whether the feedback is effective. If the feedback is effective, the prompt template is further optimized, and the rules in the feedback information are abstracted into general rules.
[0060] In this embodiment, for quality assessment results that do not meet the accuracy criteria, external data sources from the artificial intelligence entity are used for comparison and verification. This confirms the accuracy of the LLM assessment results, reduces false alarms, and thus triggers alarms more accurately. The system supports inputting feedback information, and through a feedback learning mechanism, it continuously optimizes and improves the accuracy and reliability of the assessment, reducing the recurrence of similar false alarms. This allows the system to adapt to different business scenarios and content types, improving its adaptability and flexibility.
[0061] Based on the above embodiments, Figure 3 This is a flowchart illustrating another method for assessing the quality of cataloging information provided in an embodiment of this disclosure, as shown below. Figure 3 As shown, the method also includes: Step 301: If the cataloging content to be evaluated does not meet the preset review conditions, generate an alarm message for the cataloging content to be evaluated.
[0062] In this embodiment, the review conditions are set according to the alarm requirements of the cataloging information quality assessment scenario. Optionally, the cataloging content to be evaluated not meeting the preset review conditions includes: the quality review result of the cataloging content to be evaluated does not meet the preset accuracy conditions.
[0063] The quality review results include a review score and the corresponding analytical basis. Whether the quality review results meet the preset accuracy conditions involves comparing the review score with a preset score to determine if the preset accuracy conditions are met. For example, if the label's review score is less than the preset score, the preset accuracy conditions are not met. The specific determination of these accuracy conditions can be the same as or different from the accuracy conditions of the quality assessment results; no specific restrictions are imposed here.
[0064] In one embodiment of this disclosure, generating an alarm message for the cataloging content to be evaluated includes: generating a target evaluation result for the cataloging content to be evaluated based on the quality evaluation result and quality review result; querying the historical evaluation results of the cataloging content to be evaluated; and determining the range of change between the target evaluation result and the historical evaluation results. Furthermore, if the range of change is greater than or equal to a preset threshold, an alarm message for the cataloging content to be evaluated is generated; if the range of change is less than the preset threshold, the alarm flag information for the cataloging content to be evaluated is queried. If the alarm flag information is a no-alarm flag, an alarm message for the cataloging content to be evaluated is generated; if the alarm flag information is an alarm flag, the current alarm is not executed to avoid duplicate alarms.
[0065] The system integrates the quality assessment results and quality review results of the cataloging content to be evaluated into a target assessment result. This target assessment result includes a list of tag scores for each dimension, a list of ranking matching scores, a risk score list, and an overall tag quality score. The system calls the alarm optimization module to perform alarm judgment for each cataloging content to be evaluated, determining whether to trigger an alarm based on historical assessment results in historical data. Historical assessment results include historical scores, which are records of the cataloging information after quality assessment and review.
[0066] Step 302: For each multimedia content, generate an aggregated alarm message for the multimedia content based on the alarm message of the cataloging content to be evaluated corresponding to the multimedia content.
[0067] In this embodiment, if a multimedia content has a cataloging item that needs to be evaluated and requires an alarm, an alarm message for that cataloging item is sent. If there are multiple cataloging items that need to be evaluated and require an alarm, the multiple alarm messages are aggregated into one message to reduce the number of alarms and improve alarm readability.
[0068] As an example, the system collects all cataloged content requiring alerts and forms a list of issue tags. If the same multimedia content has multiple issues, such as multiple tags failing to meet quality standards or failing to match ranking criteria, the alert messages for these multiple issues are aggregated into a single aggregated alert message. This aggregated alert message includes the multimedia content's name, multimedia content identifier, number of issues, and a description of each issue. The issue description includes the tag name, rating, reason, and alert type. Alert types include tag quality alerts, ranking matching alerts, and content classification risk alerts. In this example, the system sends alert messages through the alert service. These alert messages or aggregated alert messages are sent to designated groups with suggested actions, such as "Please check and fix the tags."
[0069] In one embodiment of this disclosure, after determining the quality assessment results and / or quality review results of the cataloging information, the quality assessment results and / or quality review results can be saved as historical assessment results of the multimedia content.
[0070] As an example, after generating an alarm message for the cataloging content to be evaluated, the historical evaluation results of the cataloging content to be evaluated are updated based on the quality evaluation results and quality review results. Furthermore, the historical evaluation results of the cataloging content to be evaluated can be added to the evaluation prompts in the cataloging information. When the cataloging information of this multimedia content is evaluated next time, the first prompt will guide the large language model to compare the quality evaluation result with the historical evaluation results.
[0071] In this embodiment, by comparing the current evaluation results with historical evaluation results, it is determined whether an alarm should be triggered. This avoids repeated alarms for issues that have already been alarmed but not improved, reducing the number of alarms and improving their effectiveness. By aggregating alarms to integrate multiple issues with the same content, alarm messages become clearer and easier to read, allowing operations personnel to quickly understand the full picture of the problem, reducing invalid alarms, and saving processing resources and time.
[0072] The following examples illustrate this with specific business scenarios.
[0073] Evaluation Phase: Cataloging information is obtained, including the content title "Ancient Costume Legend," a brief description "Telling a story of ancient court intrigue," tags "Ancient Costume Drama," "Court," and "Power Struggle," ranking information "Ranked 5th on the Popular Ancient Costume Drama Chart," and content rating "Ordinary." Evaluation prompts are constructed and input into a large language model for processing. The quality evaluation results output by the large language model include: a score of 8 for the tag "Ancient Costume Drama" (reasonable), a score of 7 for the tag "Court" (reasonable), a score of 6 for the tag "Power Struggle" (reasonable), a ranking matching score of 7 (reasonable), and an overall tag quality score of 7. Verification Phase: All scores are higher than the preset scores, requiring no verification by the AI. Alarm Judgment Phase: Comparing with historical data, this multimedia content is being evaluated for the first time, and all scores meet the accuracy criteria; no alarm is triggered. Feedback optimization phase: Experts rated the tag "political intrigue" at 6 points, which is too low. They can enter feedback information such as "This drama does contain a lot of political intrigue plots, and the tag 'political intrigue' should be rated higher." In the next evaluation of this multimedia content, the feedback information will be added to the evaluation prompts so that the large language model can refer to the feedback information for evaluation.
[0074] The Multimedia Content Cataloging Information Quality Inspection System includes: a data source module, a data acquisition module, an LLM evaluation module, a review module, an alarm optimization module, a feedback management module, a business knowledge base, and a data storage module. The data source module includes a content management system and a message queue, used to provide content update notifications and content information. The data acquisition module includes a message listening service, a data transformation service, and a filtering service. The message listening service listens for content update notifications in the message queue; the data transformation service converts data from external systems to the system's internal format; and the filtering service filters cataloging content that does not require evaluation. The LLM evaluation module includes a prompt word builder, an application programming interface (API), a result parser, and a score updater. The prompt word builder constructs evaluation prompt words based on business rules, feedback information, and cataloging information. The API calls a large language model for structured evaluation. The result parser parses the quality evaluation results returned by the LLM, and the score updater updates the quality evaluation results to the database. The review module includes a review task manager, an AI agent client, a review result parser, and a result write-back unit. The review task manager manages the creation, execution, and status updates of review tasks. The AI agent client calls the AI agent to perform information capture and review. The review result parser parses the quality review results returned by the AI agent, and the result write-back unit writes the quality review results back to the database. The alarm optimization module includes a historical data comparer, an alarm judge, an alarm aggregator, an alarm service, and an application robot service. The historical data comparer compares the current evaluation results with historical evaluation data. The alarm judge intelligently determines whether an alarm is needed based on historical data. The alarm aggregator aggregates multiple issues related to the same multimedia content into a single alarm message. The alarm service sends alarm messages, and the application robot service receives alarm messages and sends them to designated groups. The feedback management module includes a feedback storage and a prompt word optimizer. The feedback storage stores expert feedback information, and the prompt word optimizer integrates feedback information into evaluation prompt words. The business knowledge base includes rule templates and rule configurations. Rule templates store configurable prompt word templates, while rule configurations support dynamic configuration of business rules. The data storage module uses a database to store data, including content information, evaluation results, review tasks, alarm records, etc. Configurable prompt word templates allow for dynamic adjustments to adapt to changing business needs. Prompt word templates can be directly modified, resulting in low rule maintenance costs.
[0075] Figure 4 This is a schematic diagram of the structure of a quality assessment device for cataloging information provided in an embodiment of this disclosure, as shown below. Figure 4 As shown, the quality assessment device for cataloging information includes: an acquisition module 41, a generation module 42, and a processing module 43.
[0076] The acquisition module 41 is used to acquire cataloging information of multimedia content, and to determine the cataloging content to be evaluated and the corresponding reference cataloging content from the cataloging information; the reference cataloging content includes other cataloging content in the cataloging information used to evaluate the cataloging content to be evaluated. The generation module 42 is used to fill the cataloging content to be evaluated and the comparison cataloging content into a preset prompt word template to generate evaluation prompt words for the cataloging information; wherein, the evaluation prompt words include the comparison cataloging content, the cataloging content to be evaluated, and the first prompt information; The processing module 43 is used to input the evaluation prompt words into the large language model for processing and generate the quality evaluation result of the cataloging content to be evaluated; wherein, the first prompt information is used to enable the large language model to evaluate the potential risks of the cataloging content to be evaluated based on semantic understanding to determine the quality evaluation result, and the evaluation of potential risks includes a combined evaluation based on several dimensions of data in the cataloging information and / or the correlation between data.
[0077] In one embodiment of this disclosure, the cataloging content to be evaluated includes tag information, the cataloging content to be compared includes content description information, and the processing module 43 is specifically used for: The first matching degree between each tag information and the content description information is evaluated using a large language model to determine the tag quality score of each tag information; where the first matching degree is positively correlated with the tag quality score. And / or, Obtain category information from the cataloging content to be evaluated; The second matching degree between category information and content description information is evaluated using a large language model to determine the category matching degree score of the cataloging content to be evaluated; the second matching degree is positively correlated with the category matching degree score.
[0078] In one embodiment of this disclosure, the processing module 43 is specifically used for: Obtain target popularity data from the catalog content to be evaluated; The third degree of matching between the target popularity data and the reference popularity data of multimedia content is evaluated using a large language model to determine the popularity matching score of the cataloging content to be evaluated; among which, the third degree of matching is positively correlated with the popularity matching score.
[0079] In one embodiment of this disclosure, the cataloging content to be evaluated includes content classification information, and the cataloging content to be compared includes tag information. The processing module 43 is specifically used for: The fourth matching degree between content classification information and tag information is evaluated using a large language model to determine the risk score of content classification information; the fourth matching degree is negatively correlated with the risk score.
[0080] In one embodiment of this disclosure, the device further includes: The review module is used to generate review prompts based on the cataloging content to be evaluated when the quality assessment results do not meet the preset accuracy conditions. The review prompts include the cataloging content to be evaluated, the preset data source, and the second prompt information. The review prompts are input into the AI agent for processing, generating a quality review result for the cataloged content to be evaluated. The second prompt is used to enable the AI agent to query the review cataloged content of the multimedia content from a preset data source, and to evaluate the consistency of the cataloged content to be evaluated based on the review cataloged content to determine the quality review result.
[0081] In one embodiment of this disclosure, the device further includes: The alarm module is used to generate alarm messages for the cataloging content to be evaluated if the content does not meet the preset review conditions. For each multimedia content, an aggregated alarm message for the multimedia content is generated based on the alarm messages of the cataloging content to be evaluated corresponding to the multimedia content. The aggregated alarm message includes the content name, content identifier, number of issues, and description information of each issue.
[0082] In one embodiment of this disclosure, the alarm module is specifically used for: Based on the quality assessment results and quality review results of the cataloging content to be evaluated, generate the target assessment results of the cataloging content to be evaluated; Query the historical evaluation results of the cataloging content to be evaluated to determine the range of change between the target evaluation results and the historical evaluation results; If the change is greater than or equal to a preset threshold, an alarm message is generated for the cataloging content to be evaluated; If the change is less than a preset threshold, query the alarm flag information of the cataloging content to be evaluated; If the alarm flag is "no alarm", then an alarm message for the cataloged content to be evaluated is generated.
[0083] In one embodiment of this disclosure, the device further includes: The history module is used to update the historical evaluation results of the cataloging content to be evaluated based on the quality evaluation results and quality review results. Add the historical evaluation results of the cataloging content to be evaluated to the evaluation prompts in the cataloging information; The first prompt information is also used to enable the large language model to compare the quality assessment results with historical assessment results when determining the quality assessment results.
[0084] In one embodiment of this disclosure, the evaluation prompt words further include feedback information, and the first prompt information is also used to enable the large language model to refer to the feedback information when determining the quality evaluation result.
[0085] The cataloging information quality assessment device provided in this disclosure can execute any cataloging information quality assessment method provided in this disclosure, and has the corresponding functional modules and beneficial effects for executing the method. Content not described in detail in the device embodiments of this disclosure can be referred to the description in any method embodiment of this disclosure.
[0086] This disclosure also provides an electronic device including one or more processors and a memory. The processor may be a central processing unit (CPU) or other processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute the program instructions to implement the methods of the embodiments of this disclosure above and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.
[0087] In one example, the electronic device may also include input and output devices, which are interconnected via a bus system and / or other forms of connection. Furthermore, the input device may include, for example, a keyboard, a mouse, etc. The output device can output various information to the outside, including determined distance information, direction information, etc. The output device may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc. In addition, depending on the specific application, the electronic device may include any other suitable components such as a bus, input / output interfaces, etc.
[0088] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform any of the methods provided in the embodiments of this disclosure.
[0089] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0090] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform any of the methods provided in the embodiments of this disclosure.
[0091] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0092] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0093] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for quality assessment of cataloging information, characterized in that, The method includes: Obtain cataloging information of multimedia content, and determine from the cataloging information the cataloging content to be evaluated and the corresponding reference cataloging content; the reference cataloging content includes other cataloging content in the cataloging information used to evaluate the cataloging content to be evaluated; The cataloging content to be evaluated and the comparison cataloging content are filled into a preset prompt word template to generate evaluation prompt words for the cataloging information; wherein, the evaluation prompt words include the comparison cataloging content, the cataloging content to be evaluated, and the first prompt information; The evaluation prompts are input into a large language model for processing to generate a quality evaluation result for the cataloging content to be evaluated. The first prompt is used to enable the large language model to assess the potential risks of the cataloging content to be evaluated based on semantic understanding to determine the quality evaluation result. The assessment of the potential risks includes a combined assessment based on several dimensions of data and / or the correlation between data in the cataloging information.
2. The method as described in claim 1, characterized in that, The cataloging content to be evaluated includes tag information, the comparison cataloging content includes content description information, and the step of inputting the evaluation prompts into a large language model for processing to generate a quality evaluation result for the cataloging content to be evaluated includes: The first matching degree between each tag information and the content description information is evaluated using the large language model to determine the tag quality score for each tag information; wherein, the first matching degree is positively correlated with the tag quality score; And / or, Obtain the category information from the cataloging content to be evaluated; The second matching degree between the category information and the content description information is evaluated using the large language model to determine the category matching degree score of the cataloging content to be evaluated; wherein the second matching degree is positively correlated with the category matching degree score.
3. The method as described in claim 2, characterized in that, The process of inputting the evaluation prompts into a large language model for processing to generate quality evaluation results for the cataloging content to be evaluated includes: Obtain the target popularity data of the cataloging content to be evaluated; The third matching degree between the target popularity data and the reference popularity data of the multimedia content is evaluated using the large language model to determine the popularity matching degree score of the cataloging content to be evaluated; wherein the third matching degree is positively correlated with the popularity matching degree score.
4. The method as described in claim 1, characterized in that, The cataloging content to be evaluated includes content classification information, the comparison cataloging content includes tag information, and the step of inputting the evaluation prompts into a large language model for processing to generate a quality evaluation result for the cataloging content to be evaluated includes: The fourth matching degree between the content classification information and the tag information is evaluated using the large language model to determine the risk score of the content classification information; wherein the fourth matching degree is negatively correlated with the risk score.
5. The method as described in claim 1, characterized in that, After generating the quality assessment results for the cataloging content to be evaluated, the method further includes: If the quality assessment result of the cataloging content to be evaluated does not meet the preset accuracy conditions, a review prompt word is generated based on the cataloging content to be evaluated; wherein, the review prompt word includes the cataloging content to be evaluated, the preset data source, and the second prompt information; The verification prompt is input into the artificial intelligence agent for processing, generating a quality verification result for the cataloged content to be evaluated; wherein, the second prompt information is used to enable the artificial intelligence agent to query the verification cataloged content of the multimedia content from the preset data source, and to evaluate the consistency of the cataloged content to be evaluated based on the verification cataloged content to determine the quality verification result.
6. The method as described in claim 1, characterized in that, The method further includes: If the cataloging content to be evaluated does not meet the preset review conditions, an alarm message is generated for the cataloging content to be evaluated. For each multimedia content, an aggregated alarm message for the multimedia content is generated based on the alarm message of the cataloging content to be evaluated corresponding to the multimedia content; wherein, the aggregated alarm message includes the content name, content identifier, number of issues, and description information of each issue of the multimedia content.
7. The method as described in claim 6, characterized in that, The alarm message generated for the cataloging content to be evaluated includes: Based on the quality assessment results and quality review results of the cataloging content to be evaluated, the target assessment results of the cataloging content to be evaluated are generated; Query the historical evaluation results of the cataloging content to be evaluated, and determine the range of change between the target evaluation result and the historical evaluation results; If the change is greater than or equal to a preset threshold, an alarm message is generated for the cataloging content to be evaluated. If the change is less than the preset threshold, query the alarm marker information of the cataloging content to be evaluated; If the alarm flag information is a no-alarm flag, then an alarm message for the catalog content to be evaluated is generated.
8. The method as described in claim 6, characterized in that, After generating the alarm message for the cataloging content to be evaluated, the method further includes: Update the historical evaluation results of the cataloging content to be evaluated based on the quality evaluation results and quality review results of the cataloging content to be evaluated; Add the historical evaluation results of the cataloging content to be evaluated to the evaluation prompts in the cataloging information; The first prompt information is also used to enable the large language model to compare the quality assessment result with the historical assessment result when determining the quality assessment result.
9. The method as described in claim 1, characterized in that, The evaluation prompts also include feedback information, and the first prompt information is further used to enable the large language model to refer to the feedback information when determining the quality evaluation result.
10. A device for quality assessment of cataloging information, characterized in that, include: The acquisition module is used to acquire cataloging information of multimedia content, and determine the cataloging content to be evaluated and the corresponding reference cataloging content from the cataloging information; The comparison cataloging content includes other cataloging content in the cataloging information used to evaluate the cataloging content to be evaluated; The generation module is used to fill the cataloging content to be evaluated and the comparison cataloging content into a preset prompt word template to generate evaluation prompt words for the cataloging information; wherein, the evaluation prompt words include the comparison cataloging content, the cataloging content to be evaluated, and the first prompt information; The processing module is used to input the evaluation prompt words into a large language model for processing, and generate a quality evaluation result of the cataloging content to be evaluated; wherein, the first prompt information is used to enable the large language model to evaluate the potential risks of the cataloging content to be evaluated based on semantic understanding to determine the quality evaluation result, and the evaluation of the potential risks includes a combined evaluation based on several dimensions of data and / or the correlation between data in the cataloging information.
11. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the quality assessment method for cataloging information as described in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the quality assessment method for cataloging information as described in any one of claims 1-9.