Resource processing method and device based on large model, intelligent agent, equipment, medium and product
By using a resource processing method based on a large model, multi-dimensional prompts are generated and integrated for evaluation, solving the problem of lagging resource quality assessment in the fields of e-commerce live streaming and short videos, and realizing comprehensive quantification and efficient distribution of resource quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies in the e-commerce live streaming and short video fields lack real-time and refined evaluation of the intrinsic quality of resources, resulting in low-quality content wasting traffic and damaging user experience, insufficient identification and incentives for high-quality content, and a lack of systematic technical standards.
By using a resource processing method based on a large model, prompts are generated for the resource to be processed in multiple dimensions. The large model is used to obtain the dimensional information, and the information is integrated into integrated information through integration instructions to execute the matching processing operations.
It enables comprehensive, objective, quantitative, and standardized judgment of resource quality, improves the targeting and distribution efficiency of resource processing, enhances the identification and incentive of high-quality content, and improves user experience.
Smart Images

Figure CN121743589A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to fields such as smart e-commerce, digital humans, large models, and intelligent agents. More specifically, it relates to a resource processing method, device, intelligent agent, equipment, medium, and product based on a large model. Background Technology
[0002] In digital content platforms, especially in the fields of e-commerce live streaming and short videos, the quality of resources is related to user experience, platform reputation, and long-term commercial value. Summary of the Invention
[0003] This disclosure provides a resource processing method, apparatus, intelligent agent, device, medium, and product based on a large model.
[0004] According to one aspect of this disclosure, a resource processing method based on a large model is provided, comprising: in response to receiving a resource to be processed, generating prompt information based on multiple dimensions of a subject associated with the resource to be processed; inputting the prompt information and the resource to be processed into a large model to obtain dimension information of each of the multiple dimensions; integrating the dimension information of the multiple dimensions according to an integration instruction to obtain integrated information of the resource to be processed, wherein the integration instruction indicates the integration method for the dimension information; and performing processing operations on the resource to be processed according to a processing instruction matching the integrated information.
[0005] According to another aspect of this disclosure, a resource processing apparatus based on a large model is provided, comprising: a generation module, configured to generate prompt information based on multiple dimensions of a subject associated with the resource to be processed in response to receiving a resource to be processed; an acquisition module, configured to input the prompt information and the resource to be processed into a large model to obtain dimensional information of each of the multiple dimensions; an integration module, configured to integrate the dimensional information of the multiple dimensions according to an integration instruction to obtain integrated information of the resource to be processed, wherein the integration instruction indicates the integration method for the dimensional information; and a processing module, configured to perform processing operations on the resource to be processed according to a processing instruction matching the integrated information.
[0006] According to another aspect of this disclosure, a large-model-based intelligent agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a target large model based on the target task, and executing a resource processing method based on the large model by calling the target large model to obtain output information; and an output module for outputting the output information obtained by the processing module.
[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0008] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program or instructions thereon, which, when executed by a processor, implement the steps of the above-described method.
[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0012] Figure 1 The illustration schematically shows a system architecture to which a large-model-based resource processing method can be applied according to embodiments of the present disclosure;
[0013] Figure 2 A flowchart illustrating a resource processing method based on a large model according to an embodiment of the present disclosure is shown schematically.
[0014] Figure 3A The illustration shows an example diagram of a process based on multiple dimensions of a subject associated with a resource to be processed, according to an embodiment of the present disclosure, wherein each of the multiple dimensions has its own dimension information.
[0015] Figure 3B The illustration shows an example diagram of a process based on multiple dimensions of a subject associated with a resource to be processed, according to an embodiment of the present disclosure, wherein each of the multiple dimensions has its own dimension information.
[0016] Figure 4 This illustration schematically shows an example diagram of a dimension determination process according to an embodiment of the present disclosure;
[0017] Figure 5 The illustration shows an example diagram of a resource processing procedure based on a large model according to an embodiment of the present disclosure;
[0018] Figure 6A block diagram of a large-model-based resource processing apparatus according to an embodiment of the present disclosure is shown schematically.
[0019] Figure 7 A schematic diagram illustrating the structure of a large-model-based intelligent agent according to embodiments of the present disclosure; and
[0020] Figure 8 A block diagram of an electronic device suitable for implementing a large-model-based resource processing method according to embodiments of the present disclosure is shown schematically. Detailed Implementation
[0021] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0024] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0025] In one example, resources can be sorted and traffic allocated based on post-interaction data such as user click-through rate and conversion rate.
[0026] However, this method is lagging and inaccurate in its characterization of the long-term value of resources. By the time the platform identifies that the content has low long-term value, a lot of traffic has already been wasted and the user experience has been damaged.
[0027] Furthermore, due to the lack of an effective mechanism for real-time, refined, and pre-emptive evaluation of the intrinsic quality of resources, especially for emerging content formats such as digital humans and virtual live streaming, there is a lack of targeted technical quality standards.
[0028] Furthermore, low-quality content is often dealt with by simple demotion or traffic restriction, but no specific improvement guidance is provided; while the identification and incentives for high-quality content are not systematic enough, resulting in insufficient motivation for merchants to produce high-quality content.
[0029] To address this, embodiments of this disclosure propose a resource processing scheme based on a large model. For example, in response to receiving a resource to be processed, prompt information is generated based on multiple dimensions of the subject associated with the resource; the prompt information and the resource to be processed are input into the large model to obtain the dimension information of each of the multiple dimensions; according to an integration instruction, the dimension information of the multiple dimensions is integrated to obtain integrated information of the resource to be processed, wherein the integration instruction indicates the integration method for the dimension information; and processing operations on the resource to be processed are performed according to processing instructions matching the integrated information.
[0030] According to embodiments of this disclosure, by generating prompt information based on multiple dimensions of the subject associated with the resource to be processed, and using this prompt information to guide a large model in processing the resource to obtain dimensional information for each of the multiple dimensions, the intrinsic content quality of the resource can be quantified more comprehensively and objectively. By employing integration instructions to integrate the dimensional information of each dimension into unified integrated information, standardized judgment of the overall quality of the resource is achieved. Based on this, by executing matching processing instructions based on this integrated information, processing operations on the resource to be processed are performed, thereby making resource processing more targeted and resource distribution more efficient.
[0031] The collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, involved in the technical solution of this invention comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0032] In the technical solution of the present invention, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.
[0033] Figure 1 The illustration schematically depicts a system architecture to which a large-model-based resource processing method can be applied according to embodiments of the present disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0034] like Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0035] Users can interact with server 105 via network 104 using at least one of the first terminal device 101, second terminal device 102, and third terminal device 103 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, second terminal device 102, and third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0036] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0037] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0038] It should be noted that the resource processing method based on a large model provided in this disclosure can generally be executed by server 105. Correspondingly, the resource processing device based on a large model provided in this disclosure can generally be located in server 105. The resource processing method based on a large model provided in this disclosure can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the resource processing device based on a large model provided in this disclosure can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0039] Alternatively, the resource processing method based on a large model provided in this disclosure can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the resource processing apparatus based on a large model provided in this disclosure can also be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0040] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0041] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.
[0042] Figure 2 A flowchart illustrating a resource processing method based on a large model according to an embodiment of the present disclosure is shown schematically.
[0043] like Figure 2 As shown, the resource processing method 200 based on the large model includes operations S210~S240.
[0044] In operation S210, in response to receiving a resource to be processed, a prompt message is generated based on multiple dimensions of the subject in the resource to be processed.
[0045] In operation S220, the prompt information and resources to be processed are input into the large model to obtain the dimensional information of each of the multiple dimensions.
[0046] In operation S230, according to the integration instruction, the dimensional information of multiple dimensions is integrated to obtain the integrated information of the resource to be processed. The integration instruction defines the integration method for the dimensional information.
[0047] In operation S240, processing operations on the resources to be processed are executed according to the processing instructions that match the integrated information.
[0048] Resources awaiting processing refer to digital content resources that require quality assessment to determine their distribution strategy. For example, a resource awaiting processing could be a live video stream featuring a digital human, a short video promoting a product, or a live-streaming session. The subject refers to the object within the resource awaiting processing that needs to be evaluated. For example, in a digital human or live-streaming scenario, the subject could refer to a real person, a digital human, or an object.
[0049] A dimension refers to a specific aspect used to measure the quality of a subject. Dimensional information refers to information generated after a quantitative or qualitative assessment of a single dimension. For example, dimensional information can be numerical; for instance, the dimensional information for the clarity dimension could be a score of 1, representing slight blur. Alternatively, dimensional information can also be a probability value; for instance, the dimensional information for the clarity dimension could be 95%, representing a 95% confidence level in clarity.
[0050] The method for obtaining dimensional information can be configured according to actual business needs and is not limited here. For example, a multi-task neural network can be trained, taking the resource to be processed as input and directly outputting the dimensional information of each dimension. Alternatively, a general feature extraction network can be used to extract high-dimensional features of the resource to be processed, and then a set of threshold judgment rules based on feature values can be designed for each dimension to generate dimensional information.
[0051] Integration directives are rules that instruct how to combine multiple independent dimensions of information into a single comprehensive conclusion. Integration directives can be configured according to actual business needs and are not limited here. For example, an integration directive could specify a weighted summation method, where each dimension is assigned a weight and a weighted total score is calculated. Alternatively, an integration directive could specify a logical judgment method, where if any dimension fails to meet the standard, the entire result is judged as low quality. Another alternative is a decision tree method, where the scores from multiple dimensions are mapped to a predefined quality level.
[0052] Integrated information refers to comprehensive information representing the overall quality level of the resources to be processed, obtained after integration. The specific form of integrated information can be configured according to actual business needs and is not limited here. For example, integrated information can be a score, such as 85 points; alternatively, integrated information can be a grade label, such as excellent; alternatively, integrated information can be a binary judgment, such as pass or fail.
[0053] Processing instructions refer to specific operational instructions for the resources to be processed, which are associated with the integrated information. In other words, they can trigger corresponding traffic or distribution control actions based on the final quality label. Different processing instruction classification methods can be used depending on the form of the integrated information.
[0054] For example, when the integrated information is a score, integrated information of 0-60 points can be associated with processing instruction 1, integrated information of 60-80 points with processing instruction 2, and integrated information of 80-100 points with processing instruction 3. Alternatively, when the integrated information is a level label, low-quality integrated information can be associated with processing instruction 1, medium-quality integrated information with processing instruction 2, and high-quality integrated information with processing instruction 3; this is not limited here.
[0055] In the embodiments of this disclosure, by generating prompt information based on multiple dimensions of the subject associated with the resource to be processed, and using this prompt information to guide the large model in processing the resource to obtain dimensional information for each of the multiple dimensions, the intrinsic content quality of the resource can be quantified more comprehensively and objectively. By employing integration instructions to integrate the dimensional information of each dimension into unified integrated information, a standardized judgment of the overall quality of the resource is achieved. Based on this, by executing matching processing instructions based on this integrated information, the processing operation of the resource to be processed is performed, thereby making the processing of the resource more targeted and the distribution of the resource more efficient.
[0056] Figure 3A The illustration shows an example diagram of a process for obtaining dimensional information of multiple dimensions based on multiple dimensions of a subject associated with a resource to be processed, according to an embodiment of the present disclosure.
[0057] like Figure 3A As shown, in embodiment 300A for obtaining dimension information, in response to receiving a resource to be processed 301, dimension information 306 for each of the multiple dimensions 301 can be obtained based on multiple dimensions 303 of the subject 302 associated with the resource to be processed 301.
[0058] In one embodiment, the target metric for dimension 303 can be added to the prompt template to obtain prompt information 305. The prompt template includes instruction information for indicating that the resource 301 to be processed should be processed based on the target metric. The prompt information 305 and the resource 301 to be processed are input into the large model M1 to obtain the dimension information 306 of dimension 303.
[0059] A target indicator refers to a specific, actionable criterion or question used to measure a particular dimension 303. It can be tailored to the specific dimension 303. For example, target indicators tailored to dimension 303 can include target indicator 1, target indicator 2, ..., target indicator M, where M is a positive integer.
[0060] The prompt template may include, but is not limited to, instructions for processing resource 301 based on target metrics. In one embodiment, processing of the resource may refer to quality assessment of the resource. The prompt template may also include identity information for instructing the execution of the task. For example, the prompt template may include: identity information "You are an expert in resource quality assessment," and instructions "View resource content, focusing primarily on subject 302 in the resource, determine if the following problems exist, and please return dimension information 306 in a preset format," where the preset format may include dimension values and assessment descriptions.
[0061] In addition to the above, the prompt template 304 may also include instructions guiding the model's reasoning steps. For example, "Please think step by step: 1. Identify the main subject 302 and foreground items in the image; 2. Determine whether there are any actual goods in the foreground; 3. Answer the following questions based on your judgment: {indicator}". Furthermore, the prompt template 304 can also be a script containing conditional judgments and loop logic, that is, it can dynamically generate prompt information 305 containing multiple sub-questions and logical jumps, etc., depending on the complexity of dimension 303, without limitation here.
[0062] Prompt Message 305 is a complete and executable text instruction formed by filling specific target metrics into placeholders in the prompt template. By generating prompt Message 305 by combining target metrics with prompt template 304, the operation method and expected results of the task performed by the large model M1 can be clearly indicated, thereby improving the accuracy and effectiveness of dimensional information 306.
[0063] After obtaining the prompt information 305, the prompt information 305 and the resource to be processed 301 can be input into the large model M1 to obtain the dimensional information 306 of dimension 303. The large model M1 can be a general large model M1, or a large model M1 finely tuned on a large amount of resource quality assessment paired data, etc., without limitation here. By using the large model M1 to process the resource to be processed 301 based on the target index, its powerful processing capabilities can be fully utilized to improve the efficiency of quality assessment.
[0064] In the embodiments of this disclosure, by dynamically filling target metrics into preset prompt templates to generate prompt information, the same large model can flexibly adapt to various dimensions that can be added or removed at any time by changing text instructions, without needing to retrain or deploy dedicated models for each dimension, thereby improving the efficiency of the evaluation process in adapting to new metrics. Furthermore, by inputting prompt information containing clear evaluation instructions along with the resources to be processed into the large model, and leveraging the powerful multimodal understanding and reasoning capabilities of the large model, structured dimensional information is directly output, improving the comprehensiveness and efficiency of processing the resources to be processed, and contributing to more efficient subsequent resource distribution decisions.
[0065] According to embodiments of this disclosure, the above method may further include: generating a prompt template based on instruction information and condition information; matching the condition information with dimension 303, wherein the condition information is used to indicate the limitation of the target indicator on obtaining dimension information 306.
[0066] Conditional information is a set of rules used to limit or constrain how to obtain dimensional information 306 based on the target metric. Conditional information can define how to do it and what to output, and may include evaluation criteria, output format, and logical judgments. For example, the evaluation criteria could be "if occlusion exists, score=0; if not, score=2; only two values, 0 and 2," the output format could be "please return in the following JSON (JavaScript Object Notation) format: {"score": ,"reason":}," and the logical judgments could be "first check A; if A is true, then output result X directly without checking B; otherwise, check B and C..."
[0067] In the embodiments of this disclosure, by decoupling instruction information from condition information, constraints such as evaluation logic and output format can be managed and flexibly adjusted independently without redesigning the entire task instruction, thereby improving the efficiency and flexibility of evaluation rule updates, and thus improving the automation level and efficiency of resource processing, as well as the accuracy of the obtained dimensional information.
[0068] Figure 3B The illustration shows an example diagram of a process for obtaining dimensional information of multiple dimensions based on multiple dimensions of a subject associated with a resource to be processed, according to an embodiment of the present disclosure.
[0069] like Figure 3B As shown, in embodiment 300B for obtaining dimensional information, for each frame in the resource to be processed 307, the image of the subject 308 region is extracted from the frame based on the position 309 of the subject 308 obtained by object detection of the frame; the image of the subject 308 region is processed based on the dimension 311 of the subject 308 to obtain the frame evaluation information 312 of each frame; and the dimensional information 313 is determined based on the frame evaluation information 312 of each frame.
[0070] A frame refers to a single image frame in the resource 307 to be processed. For example, if the resource 307 to be processed is a live video, then a frame can be a still image extracted from the live video at a rate of 1 frame per second. Object detection refers to the process of identifying the position 309 of a specific subject 308 in the frame using a computer vision model. For example, the position 309 of the subject 308 in the frame can be represented using a bounding box.
[0071] After obtaining the position 309 of the subject 308 in the image, an image region containing only the subject 308 can be cropped from the image based on the coordinates of the position 309 obtained by object detection, to obtain the image of the subject 308 region. For example, if the position 309 of the half-body image of the subject 308 in the image is obtained through object detection, the rectangular area corresponding to that position 309 can be extracted from the image to obtain a half-body image without the background.
[0072] In addition to the methods mentioned above, a semantic segmentation model can be used to obtain a pixel-level precise mask of the subject 308, and then the image of the subject 308 region can be extracted based on the mask, thereby better handling irregular shapes and avoiding the inclusion of background edges. Furthermore, if there are multiple subjects 308 in the image, the subject 308 with the largest area and the center position 309 can be selected as the subject 308 to be cropped, without any limitation here.
[0073] After obtaining the image of the subject region 308, each dimension 311 of the image of the subject region 308 can be processed separately to obtain sub-evaluation information of each dimension 311, and the sub-evaluation information of each dimension 311 can be integrated into the image evaluation information 312 of the image corresponding to the image of the subject region 308.
[0074] After obtaining the individual image evaluation information 312 for each image in the resource 307 to be processed, the image evaluation results of all images on the timeline for a certain dimension 311 can be summarized to form the dimensional information 313 for that dimension 311. This process can adopt an aggregation method based on time series analysis, that is, treating the image evaluation information 312 as a time series and using sequence models such as LSTM to learn the dynamic changes of the image evaluation information 312, thereby outputting comprehensive dimensional information 313 that considers time consistency. Alternatively, this process can also adopt a rule-based weighted fusion method, that is, by defining aggregation rules, such as determining the extreme values in each image evaluation information 312 as dimensional information 313, etc., which is not limited here.
[0075] In the embodiments of this disclosure, by performing target detection on each frame and cropping the subject area image, the interference of complex backgrounds on the evaluation is eliminated, allowing for a more focused analysis of the subject's detailed quality, thereby improving the accuracy of identifying the subject's own quality defects. Furthermore, by independently evaluating the subject area image of each frame based on dimensions, and by comprehensively considering the evaluation information from each frame, the final dimensional information is determined. This reflects the overall quality level and stability of the resource to be processed throughout its entire lifecycle, improving the accuracy of the dimensional information and enabling more efficient subsequent resource distribution.
[0076] According to embodiments of this disclosure, determining dimension information 313 based on the image evaluation information 312 of each image may include: determining a fusion instruction for dimension 311, wherein the fusion instruction indicates a fusion method for the image evaluation information 312; and fusing the image evaluation information 312 of each image according to the fusion instruction to obtain dimension information 313.
[0077] Fusion instructions define how to aggregate the evaluation results of multiple screens under the same dimension 311 into dimension information 313. The method for determining fusion instructions can be configured according to actual business needs and is not limited here. For example, fusion instructions can be determined based on the characteristics of dimension 311, business objectives, or resource types.
[0078] For example, a dimensional attribute library can be pre-established, in which each dimension 311 is labeled with a corresponding fusion method, such as fusion method 1 for dimension 311, fusion method 2 for dimension 311, etc., so that the corresponding fusion instruction can be matched from the dimensional attribute library based on the current dimension 311. Alternatively, the description and contextual features used for the current dimension 311 can be input into the deep learning model, so that the model outputs the recommended fusion instruction for that dimension 311. The contextual features may include resource type, resource duration, etc.
[0079] After receiving the fusion instruction, the image evaluation information 312 of all images in this dimension 311 can be processed according to the fusion instruction to obtain dimension information 313. For example, if the fusion instruction indicates the use of a fusion method based on a time series model, the image evaluation information 312 can be regarded as a time series, and its dynamic changes can be modeled using a time series analysis model. The dimension information 313 that integrates time dependencies can be directly output. For example, for the sharpness dimension, the dimension information 313 can be "the opening blur is more severe than the mid-process blur", etc.
[0080] Alternatively, if the fusion instruction specifies a comprehensive fusion method based on bucketing statistics and rules, then instead of directly calculating a single value, the image evaluation information 312 can be statistically bucketed (e.g., the proportion of images with scores of 0, 1, and 2 is calculated), and then these proportions are converted into dimensional information 313 according to preset rules. For example, the preset rule could be: "If the proportion of images with scores of 0 is greater than 10%, then the overall quality is judged as poor; otherwise, if the proportion of images with scores of 2 is greater than 70%, then the quality is judged as good; the rest are considered average."
[0081] In the embodiments of this disclosure, by determining corresponding fusion instructions based on the characteristics of different dimensions, the derivation process from instantaneous image evaluation to overall quality is made more consistent with the inherent logic of each dimension, thereby improving the rationality and accuracy of subsequent dimensional information. Furthermore, by integrating the dispersed image evaluation information using a clearly defined fusion method, the random errors in single-frame evaluation can be effectively smoothed, resulting in a more stable and representative overall quality characterization, improving the accuracy of dimensional information, and thus contributing to more efficient subsequent resource distribution decisions.
[0082] Figure 4 The illustration shows an example schematic diagram of a dimension determination process according to an embodiment of the present disclosure.
[0083] like Figure 4 As shown, in embodiment 400 of determining dimensions, the dimensions may be determined in the following manner: by performing content understanding on the resource to be processed 401 to classify the subject and obtain subject category 402; based on subject category 402, multiple dimensions are determined from multiple candidate dimensions.
[0084] Category 402 refers to a category label that classifies entities according to their nature, form, or function. For example, Category 402 can include live streamers, digital human streamers, food products, electronic products, and clothing products. The specific method of classifying entities can be configured according to actual business needs and is not limited here.
[0085] For example, a hierarchical classification model can be used. The first-level network distinguishes between "people" and "objects." The first branch of the second-level network distinguishes between real people and digital people under the "people" branch. The second branch of the second-level network further subdivides the "objects" branch into product leaf categories. Alternatively, the visual attributes of the subject can be identified first (such as the presence of human figures, skin texture, the presence of virtual background edges, and the presence of product labels), and then the subject category 402 can be inferred according to preset rules. For example, the preset rule can be to classify subjects with both human features and obvious CG (Computer Graphics) rendering texture as digital people.
[0086] When the subject is a digital human, the dimension can be any one or more of the candidate dimensions for the digital human; when the subject is an item, the dimension can be any one or more of the candidate dimensions for the item. Since different categories of subjects focus on different information, the dimension can be determined based on subject category 402.
[0087] For example, if the subject is an item, and the item is clothing, then the focus is more on the visual effect of the clothing when worn, so the dimension can include the completeness of the item; if the item is food, then the focus is more on the ingredients, shelf life, etc., so the dimension can include the completeness of the item information.
[0088] The method for determining dimensions can be configured according to actual business needs and is not limited here. For example, a lookup table (LUT) or configuration file can be pre-configured to specify which dimensions correspond to each subject category 402. After determining the subject category 402, the dimensions can be determined directly by looking up the lookup table. Alternatively, a lightweight knowledge graph containing subject category 402, dimensions, and context can be pre-built, allowing the dimensions to be derived through graph relationships based on the identified subject category 402.
[0089] For example, taking multiple candidate dimensions including candidate dimension 1, candidate dimension 2, ..., candidate dimension P, where P is a positive integer, we can determine dimension 1, dimension 2, ..., dimension Q from candidate dimension 1, candidate dimension 2, ..., candidate dimension P based on subject category 402, where Q is a positive integer and less than or equal to P.
[0090] In the embodiments of this disclosure, by accurately classifying subjects based on subject region images, the essence of the subjects can be accurately identified, laying the foundation for subsequent differentiated evaluation. Furthermore, by dynamically selecting dimensions from candidate dimensions according to the determined subject category, unnecessary computational overhead is reduced, thereby improving the processing efficiency of the evaluation process. Simultaneously, it ensures that the evaluation closely focuses on the core quality concerns of the current subject category, thus improving the relevance and accuracy of the final integrated information.
[0091] According to embodiments of this disclosure, the subject may include a digital human, and the candidate dimensions for the digital human may include at least one of the following: the completeness of the digital human, the sharpness of the digital human, the edge roughness of the digital human, the color anomaly of the digital human, and the proportion of the digital human in the image.
[0092] The completeness of a digital human can be used to assess whether the digital human image is presented completely in the image, and whether key body parts (such as the head and arms) are missing due to reasons such as image cropping or obstruction by foreground objects. For example, if the digital human's head is partially obscured by the floating icon of the live broadcast room, or if the arms are not fully displayed due to image cropping, the completeness is insufficient.
[0093] For example, target metrics for the completeness of a digital human may include at least one of the following: "whether the face is obscured by a text box, foreground image, or object, and the full facial features are not visible", "whether the body is heavily obscured by a foreground image or object", and "whether the arms and hands of the human body are not fully displayed and are obviously truncated".
[0094] The sharpness of a digital human can be used to assess the image sharpness and detail discernibility of the digital human figure itself, primarily focusing on whether there is overall or localized blurring. For example, insufficient model rendering resolution or transmission compression can cause the facial features of a digital human to be blurry.
[0095] For example, the target indicators for the sharpness of a digital human may include at least one of the following: "whether the overall image is obviously blurry and the facial features are not clear", "whether key facial features (such as eyes, nose, and mouth) are partially blurry and details are difficult to identify", and "whether there is motion blur or focus failure that causes the edges of the image to be blurred".
[0096] Edge roughness of digital humans can be used to evaluate whether there are technical defects in the contour edges of digital humans synthesized using green screen keying technology. For example, the edges may have obvious jagged lines, irregular black or white borders, etc.
[0097] For example, the target indicators for edge roughness of digital humans may include at least one of the following: "whether there are obvious lines", "whether there are black edges", "whether there are white edges", "whether there are jagged lines", "whether there are blurred edges and lines that are not visible".
[0098] Color anomaly in digital humans can be used to assess the naturalness of skin tone and overall color reproduction, and to determine whether there are any illogical color casts or brightness anomalies. For example, a digital human's skin tone may appear severely grayish, too dark overall (e.g., blackish), or exhibit unnatural overexposure (e.g., too white), yellowish, or reddish.
[0099] For example, the target indicators for color anomalies of digital humans may include at least one of the following: "whether the portrait is grayish", "whether the portrait is dull", "whether the portrait is too white", "whether the portrait is black", "whether the portrait is yellowish", "whether the portrait is reddish", and "whether the portrait is overexposed".
[0100] The proportion of a digital human in the frame can be used to assess whether the size of the digital human subject within the overall image meets preset visual comfort or content display standards. For example, a digital human that is too large will crowd out the product display space (e.g., >80%), or a digital human that is too small will make it difficult for viewers to discern details (e.g., <30%).
[0101] For example, the target indicators for the proportion of digital human in the image may include at least one of the following: "Is the human image too large?", "Is the human image too small?", "Is the proportion of the human image greater than or equal to 80% of the image?", "Is the proportion of the human image less than or equal to 30% of the image?".
[0102] In the embodiments of this disclosure, by explicitly defining a dedicated dimension for digital humans, the one-sidedness of a single dimension is overcome, and the targeting, effectiveness, and operability of processing digital human resources are enhanced, thereby improving the accuracy of processing digital humans.
[0103] According to embodiments of this disclosure, the subject may include an item, and the candidate dimensions for the item may include at least one of the following: the completeness of the item, the completeness of the item information.
[0104] The integrity of an item can be used to assess whether the item itself is visually complete and undamaged in the image. For example, severely damaged product packaging, half of the product being obscured, or shoes being displayed with only the upper showing and the sole or other key parts not visible.
[0105] For example, the target indicators for the completeness of an item may include at least one of the following: "whether the foreground of the image lacks actual product display", "whether the foreground of the image lacks stacked product display", "whether the host is holding a product".
[0106] The completeness of item information can be used to assess whether descriptive and identifying information related to the item is clearly and completely presented in the image. For example, product labels may be blurry and illegible, price tags may be obscured, brand logos may be incomplete, or necessary product certification marks may be missing.
[0107] For example, the target indicators for the completeness of item information may include at least one of the following: "Is the decoration rudimentary?", "Are there any missing stickers?", "Are there any missing merchandise cards?".
[0108] It should be noted that the completeness of an item focuses on the completeness of its physical form, while the completeness of the item's information focuses on the completeness of its informational aspects.
[0109] In the embodiments of this disclosure, by clearly defining the specific dimensions for items, the one-sidedness of a single dimension is overcome, and the targeting, effectiveness and operability of the processing of item resources are enhanced, thereby improving the accuracy of the processing of items.
[0110] According to embodiments of this disclosure, the dimension information includes dimension values; operation S220 may include: determining the evaluation level of a dimension based on the dimension values, with a first evaluation threshold for the dimension as a reference; and determining integrated information based on the evaluation levels of each dimension.
[0111] After obtaining the dimensional information, the dimensional values can be compared with a preset first evaluation threshold to classify them into the corresponding level range. The first evaluation threshold is a critical value used to map dimensional values to evaluation levels. The evaluation level is a qualitative level assigned to each dimension based on the comparison result between the dimensional value and the first evaluation threshold.
[0112] For example, for the "edge roughness" dimension, the relationship between the first evaluation threshold and the evaluation level can be set as follows: 0 points corresponds to poor, 1 point corresponds to medium, and 2 points corresponds to excellent. If the dimension value is 1 point, the evaluation level of the edge roughness dimension can be determined to be medium.
[0113] Alternatively, the first evaluation threshold can not be a fixed value, but can be dynamically adjusted based on the resource category, resource context, or historical data distribution. For example, the "excellent" threshold for "color naturalness" can be different for "digital humans" and "real people." Alternatively, the first evaluation threshold can also be automatically adjusted periodically based on improvements in the overall quality level of the platform.
[0114] After obtaining the evaluation levels for each dimension, integrated information can be determined based on these levels. For example, the evaluation levels for each dimension can be synthesized using a yes / no logic approach; that is, dimensions with poor evaluation levels do not belong to high-quality resources. Alternatively, the evaluation level can be determined by the proportion of excellent evaluation levels in each dimension out of the total number of evaluation levels. For instance, if this proportion is greater than 50%, then it can be determined that the resource belongs to high-quality resources.
[0115] In the embodiments of this disclosure, the evaluation level is determined by comparing the dimensional values of each dimension with a first evaluation threshold, thus giving the multi-dimensional evaluation results a unified and intuitive semantic scale, thereby improving the understandability and comparability of the evaluation results. Based on this, by determining integrated information according to the evaluation levels of each dimension, the logic of comprehensive judgment is simplified, and the flexibility and maintainability of determining integrated information are improved, making subsequent resource distribution decisions more efficient.
[0116] According to an embodiment of this disclosure, the dimension information includes dimension values; operation S220 may include: integrating the dimension values of multiple dimensions according to the weights used for each dimension to obtain an intermediate evaluation value, wherein the weights characterize the degree of influence of the dimension on the quality of the resource to be processed; and determining integrated information based on the intermediate evaluation value with a second evaluation threshold as a reference.
[0117] Weights are coefficients used to characterize the degree to which different dimensions affect the overall quality of resources. For example, a higher weight indicates that the dimensional value carries greater weight in the final decision. For instance, in the case of a digital human, edge roughness might be given a higher weight because severe image defects significantly impact visual appeal.
[0118] The intermediate evaluation value is a numerical value obtained by weighted summation based on the dimensional values and their corresponding weights of each dimension. For example, taking a digital human as the subject and dimensions including the digital human's completeness, clarity, and color anomaly, if the dimension value of clarity is 0.8 with a weight of 0.3, the dimension value of color anomaly is 0.9 with a weight of 0.3, and the dimension value of completeness is 0.6 with a weight of 0.4, then the integrated intermediate evaluation value can be 0.8*0.3+0.9*0.3+0.6*0.4=0.75.
[0119] It should be noted that the weights used for each dimension do not have to be statically configured, but are learned by a machine learning model based on historical data. The goal of training this model can be to maximize long-term user retention rate, repurchase rate, etc., thereby enabling dynamic allocation of weights based on business objectives.
[0120] Alternatively, the integration method is not limited to linear weighted summation. For example, geometric mean or a penalty term can be used. Geometric mean can place greater emphasis on the balance of each dimension, while penalty term can avoid significantly lowering the intermediate evaluation value for dimensions with serious problems, even if the weight is not high.
[0121] After obtaining the intermediate evaluation value, the integrated information can be determined based on the intermediate evaluation value, using the second evaluation threshold as a reference. The second evaluation threshold is a critical value used to determine the quality level of the intermediate evaluation value. This second evaluation threshold is not necessarily a fixed value, but can be dynamically adjusted according to the type of resource, the scenario of the resource, or the distribution of historical data, etc., and is not limited here.
[0122] In the embodiments of this disclosure, by assigning corresponding weights to different dimensions based on their varying degrees of influence on resource quality and then weighted and integrating them, the dimensions with a greater impact on quality occupy a more important position in the final evaluation. This results in intermediate evaluation values that more reasonably reflect the overall quality level of resources and improve the business relevance of the intermediate evaluation values. Furthermore, by comparing the integrated intermediate evaluation values with a clearly defined second evaluation threshold, the accuracy and efficiency of the judgment process are improved, thereby contributing to more efficient subsequent resource allocation decisions.
[0123] According to an embodiment of this disclosure, operation S230 may include: positively adjusting the recommendation weight for the resource to be processed when the integrated information indicates that the resource to be processed belongs to the first level; and performing at least one of the following operations on the resource to be processed when the integrated information indicates that the resource to be processed belongs to the second level: negatively adjusting the recommendation weight for the resource to be processed and limiting the recommendation range of the resource to be processed, wherein the first level is higher than the second level.
[0124] The first and second levels are tiers of resource quality based on integrated information, with the first level being of higher quality than the second. For example, the first level is high-quality and the second level is low-quality.
[0125] Recommendation weight refers to an adjustable parameter factor in a recommendation system's ranking model that influences the final ranking position of a resource. For example, positive adjustment means increasing the recommendation weight, which can increase the exposure probability, such as increasing the recommendation weight from 1.0 to 1.2; negative adjustment means decreasing the recommendation weight, which can decrease the exposure probability, such as decreasing the recommendation weight from 1.0 to 0.6.
[0126] The recommendation scope refers to the set of scenarios, channels, or groups within the platform where the resources to be processed may be distributed to users. For example, the recommendation scope may include the homepage recommendation stream, related recommendation lists, specific vertical channels, and content visible to all users or only to followers.
[0127] In one embodiment, when a resource is determined to belong to the first tier (e.g., high quality), its recommendation weight in the recommendation algorithm can be increased to give it an advantage in ranking competition with other resources, thereby gaining more exposure. It should be noted that the magnitude of the positive adjustment to the recommendation weight is not fixed, but dynamically calculated based on the specific score of the integrated information. For example, if the quality score increases from 80 to 100, the recommendation weight can increase from 1.1 to 1.5.
[0128] In addition to the methods mentioned above, high-quality resources awaiting processing can be injected into specific high-value traffic pools or recall channels. For example, a separate high-quality content queue can be created for high-quality resources awaiting processing, and the recommendation system can prioritize selecting and distributing content from this queue. Furthermore, other potential high-quality resources that are similar to the current high-quality resources in terms of content or subject matter can be identified, and the recommendation weight of these resources can be adjusted to a certain extent.
[0129] In another embodiment, when a resource is determined to be in the second level (e.g., low quality), its negative impact on users and the platform can be reduced by decreasing its recommendation weight in the recommendation algorithm or restricting the scenarios in which it can be distributed. It should be noted that the magnitude of the negative adjustment to the recommendation weight is not fixed, but rather dynamically calculated based on the severity of the low quality; this is not limited here.
[0130] In the embodiments of this disclosure, by positively adjusting the recommendation weight when a resource is identified as high quality, and negatively adjusting the recommendation weight or limiting its recommendation range when a resource is identified as low quality, the accuracy of resource distribution can be improved to be more efficient, thereby improving the overall user experience.
[0131] According to embodiments of this disclosure, the dimensional information includes an evaluation description that characterizes the reason why the resource to be processed has a dimensional value under a dimensionality.
[0132] An evaluation description refers to the textual explanation provided alongside a quantitative score when evaluating a specific dimension, explaining the reasons for that score. For example, for the clarity dimension of a digital human, the evaluation description when the dimension value is 0 might be "the facial features are severely blurred"; for the edge roughness dimension of a digital human, the evaluation description when the dimension value is 0 might be "the human image has obvious jagged white edges".
[0133] If the integrated information characterizes the resource to be processed as belonging to the second level, an adjustment instruction can be sent to the producer corresponding to the resource to be processed, so that the producer can adjust the resource to be processed and return the adjusted resource. The adjustment instruction is determined based on the evaluation description of each dimension. In response to the integrated information characterizing the adjusted resource as belonging to the first level, the recommended weight for the resource to be processed is positively adjusted.
[0134] When the integrated information characterizes the resource to be processed as belonging to the second level, in addition to imposing penalties, the specific source can be analyzed, and adjustment instructions can be generated and sent to the producer to guide them in making technical or content-level improvements. The producer refers to the creator or provider of the resource to be processed. For example, the producer could be a merchant conducting a live stream, a content creator publishing a video, or their operations team.
[0135] Adjustment instructions are generated based on the assessment description and are designed to guide producers in targeted optimization of resources. For example, for the assessment description "obvious jagged white edges exist," the adjustment instruction might be "Please optimize the green screen keying algorithm parameters to eliminate jagged edges and white borders on portraits." The generation method of adjustment instructions can be configured according to actual business needs and is not limited here.
[0136] For example, a knowledge base can be pre-built to map common evaluation description keywords (such as blurry, black border, occlusion, etc.) to standardized adjustment instruction templates. When generating adjustment instructions, keywords are extracted from the evaluation description, and the corresponding adjustment instruction templates are queried and combined to obtain the adjustment instructions.
[0137] Alternatively, the assessment descriptions for each dimension can be input into a text generation model, which can be trained to transform the assessment descriptions into guiding text and obtain adjustment instructions.
[0138] Alternatively, adjustment instructions can include not only text, but also processed screenshots that highlight the problem areas (such as circling blurry areas), or screenshots of high-quality resources in the same category as reference examples to make the guidance more intuitive.
[0139] Adjusted resources refer to resource versions resubmitted by the producer after modifying and optimizing the original resources according to adjustment instructions. For example, a live stream where the streamer adjusts lighting and green screen settings before restarting, or a video re-uploaded after replacing the source material with clearer footage. Once the producer has rectified and resubmitted the adjusted resources according to the instructions, the resources can be re-evaluated. If the adjusted resources meet the high-quality standard, they will receive a positive traffic reward.
[0140] It should be noted that a separate incentive weight pool can be established for resources that successfully optimize from level two to level one. This weighting coefficient can be higher than that of ordinary high-quality content to significantly reward improvement. Alternatively, different levels of incentives can be given based on the degree of optimization. For example, optimizing from 1 point to 2 points would result in a small positive adjustment, while optimizing from 1 point to 3 points would result in a larger positive adjustment.
[0141] In the embodiments of this disclosure, by sending adjustment instructions based on evaluation descriptions to producers of low-quality resources, producers can quickly understand the root cause of the problem and make effective rectifications. This reduces the trial-and-error costs and time costs for producers optimizing content, making the optimization process more targeted and improving the efficiency of resource optimization. Furthermore, by giving positive weight incentives to resources successfully optimized to high quality, subsequent resource distribution decisions become more efficient.
[0142] Figure 5 The illustration shows an example schematic diagram of a resource processing procedure based on a large model according to an embodiment of the present disclosure.
[0143] like Figure 5 As shown, in embodiment 500 of the resource processing procedure based on a large model, in response to receiving the resource to be processed 501, the resource processing method based on the large model provided in this disclosure can be executed to obtain integration information 502. After obtaining the integration information 502, operation S510 can be executed.
[0144] In operation S510, does the integrated information 502 indicate that the resource to be processed 501 belongs to the first level? If so, the recommended weight 503 for the resource to be processed 501 can be positively adjusted. If not, operation S520 can be executed.
[0145] In operation S520, does the integrated information 502 indicate that the resource to be processed 501 belongs to the second level? If not, the recommendation weight 503 can remain unchanged. If yes, at least one of the following operations can be performed on the resource to be processed 501: negatively adjust the recommendation weight 503 for the resource to be processed 501, or limit the recommendation range of the resource to be processed 501.
[0146] Furthermore, an adjustment instruction can be sent to the producer corresponding to the resource to be processed, so that the producer adjusts the resource to be processed and returns the adjusted resource 505. After receiving the adjusted resource 505, the resource processing method based on the large model provided in this disclosure can be executed again to obtain the integrated information 506 of the adjusted resource 505. Based on this, operation S510 can be executed again.
[0147] In operation S510, the integrated information 506 of the adjusted resource 505 indicates whether the adjusted resource 505 belongs to the first level? If so, the recommended weight 503 used for the resource to be processed can be positively adjusted.
[0148] Based on the above-described resource processing method based on a large model, this invention also provides a resource processing device based on a large model. The following will be combined with... Figure 6 The device is described in detail.
[0149] Figure 6 A block diagram of a large-model-based resource processing apparatus according to an embodiment of the present disclosure is shown schematically.
[0150] like Figure 6 As shown, the resource processing device 600 based on a large model may include a generation module 610, an acquisition module 620, an integration module 630, and a processing module 640.
[0151] The generation module 610 is used to generate prompt information based on multiple dimensions of the subject associated with the resource to be processed in response to receiving the resource to be processed.
[0152] The module 620 is used to input the prompt information and the resources to be processed into the large model to obtain the dimensional information of each of the multiple dimensions.
[0153] The integration module 630 is used to integrate dimensional information from multiple dimensions according to the integration instructions to obtain the integrated information of the resource to be processed. The integration instructions specify the integration method for the dimensional information.
[0154] The processing module 640 is used to perform processing operations on the resources to be processed according to processing instructions that match the integrated information.
[0155] According to embodiments of this disclosure, the processing module 640 may include a first processing unit and a second processing unit.
[0156] The first processing unit is used to positively adjust the recommendation weights for the resource to be processed when the integrated information indicates that the resource belongs to the first level.
[0157] The second processing unit is used to perform at least one of the following operations on the resource to be processed when the integrated information characterizes that the resource to be processed belongs to the second level: negatively adjusting the recommendation weight of the resource to be processed and limiting the recommendation range of the resource to be processed, wherein the first level is higher than the second level.
[0158] According to embodiments of this disclosure, dimensional information includes dimensional values and an evaluation description, the evaluation description characterizing the reason why the resource to be processed has dimensional values under a given dimension.
[0159] When the integrated information representation of the resource to be processed belongs to the second level, the processing module 640 may also include a sending unit and a third processing unit.
[0160] The sending unit is used to send adjustment instructions to the producer corresponding to the resource to be processed, so that the producer can adjust the resource to be processed and return the adjusted resource. The adjustment instructions are determined based on the evaluation description of each dimension.
[0161] The third processing unit is used to respond to the received integrated information of the adjusted resources, which indicates that the adjusted resources belong to the first level, and to positively adjust the recommendation weights used for the resources to be processed.
[0162] According to embodiments of this disclosure, the generation module 610 may include a first obtaining unit.
[0163] The first obtaining unit is used to add the target metrics for the dimension to the prompt template to obtain prompt information. The prompt template includes instruction information for indicating how to process the resource to be processed based on the target metrics.
[0164] According to embodiments of this disclosure, the generation module 610 may further include a second obtaining unit.
[0165] The second acquisition unit is used to generate a prompt template based on instruction information and condition information; the condition information is matched with the dimension, and the condition information is used to indicate the restrictions of the target indicator on the acquisition of dimension information.
[0166] According to embodiments of this disclosure, the generation module 610 may further include an interception unit, an evaluation unit, and a first determination unit.
[0167] The cropping unit is used to crop the subject area image from each frame in the resource to be processed, based on the position of the subject in the frame obtained by object detection.
[0168] The evaluation unit is used to process the image of the subject area based on the dimension of the subject to obtain the image evaluation information of each frame.
[0169] The first determining unit is used to determine the dimensional information based on the image evaluation information of each image.
[0170] According to embodiments of this disclosure, the first determining unit may include a determining subunit and a fusion subunit.
[0171] A sub-unit is defined to determine the fusion instructions for the dimensions, wherein the fusion instructions indicate the fusion method for the image evaluation information.
[0172] The fusion subunit is used to fuse the image evaluation information of each frame according to the fusion command to obtain dimensional information.
[0173] According to embodiments of this disclosure, dimensions are determined by: performing content understanding on the resource to be processed to classify the subject and obtain subject categories; and determining multiple dimensions from multiple candidate dimensions based on the subject categories.
[0174] According to embodiments of this disclosure, the subject includes a digital human, and the candidate dimensions for the digital human include at least one of the following: the completeness of the digital human, the sharpness of the digital human, the edge roughness of the digital human, the color anomaly of the digital human, and the proportion of the digital human in the image.
[0175] According to embodiments of this disclosure, the subject includes an article, and the candidate dimensions for the article include at least one of the following: the completeness of the article, the completeness of the article information.
[0176] According to embodiments of this disclosure, the dimension information includes dimension values; the integration module 630 may include a second determining unit and a third determining unit.
[0177] The second determining unit is used to determine the evaluation level of a dimension based on the dimension value, with reference to the first evaluation threshold for the dimension.
[0178] The third determining unit is used to determine the integrated information based on the evaluation level of each dimension.
[0179] According to embodiments of this disclosure, the dimension information includes dimension values; the integration module 630 may further include an integration unit and a fourth determination unit.
[0180] The integration unit is used to integrate the dimensional values of multiple dimensions according to the weights assigned to each dimension to obtain an intermediate evaluation value. The weights represent the degree of influence of the dimension on the quality of the resource to be processed.
[0181] The fourth determining unit is used to determine integrated information based on intermediate evaluation values, with the second evaluation threshold as a reference.
[0182] Figure 7 A schematic diagram illustrating the structure of a large-model-based intelligent agent according to an embodiment of the present disclosure is shown.
[0183] In embodiments of this disclosure, the von Neumann architecture in modern computer theory is inspired, such as... Figure 7 As shown, the AI agent 700 may include five core modules: an input module 710, a processing module 720, and an output module 730.
[0184] The input module 710 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the AI agent 700 can understand and process. The input module 710 is the primary link for the AI agent 700 to interact with the outside world, enabling it to efficiently and accurately acquire necessary "sensory" information and respond to it. In the example, the input module 710 can input the resource to be processed described above.
[0185] In embodiments of this disclosure, the processing module 720 may include a control module 721, a storage module 722, and a computation module 723. The processing module 720 is configured to determine a target task based on the input information received by the input module 710, determine a target large model based on the target task, and obtain output information by calling the corresponding method executed by the target large model.
[0186] The control module 721 is the core support for the AI agent 700's ability to handle complex tasks. The control module 721 can execute the resource processing methods based on large models described above.
[0187] In the example, the control module 721 will continuously interact with the storage module 722, the arithmetic module 723, and / or the output module 730 during operation. However, it should be noted that in the embodiments of this disclosure, the control module 721 initiates communication with the storage module 722, the arithmetic module 723, and / or the output module 730 as a single initiator, and there is no communication coupling between the storage module 722, the arithmetic module 723, and the output module 730.
[0188] In the example, the performance of the control module 721 is closely related to the large model on which the AI agent 700 is based. To fully leverage the capabilities of the large language model, the internal structure of the control module 721 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.
[0189] The storage module 722 is responsible for remembering the generated target sample set. Candidate dimensions, as mentioned earlier, can be included in the storage module 722.
[0190] In the example, after receiving the resource to be processed, the AI agent 700 can invoke candidate dimensions and the resource processing procedure based on the large model to obtain integrated information, and then feed it back to the control module 721. The control module 721 can then pass the integrated information back to the output module 730.
[0191] The arithmetic module 723 can be viewed as a predefined tool library. Tools used for integration, as described above, can be included in the arithmetic module 723.
[0192] In the example, when the AI agent 700 needs to perform a linear transformation on features, it can invoke relevant tools from the computation module 723 and feed them back to the control module 721. The control module 721 can then use the fed-back tools to process the relevant information. It's understandable that while large language models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks without any tools is limited. When the AI agent 700 is given the ability to invoke tools, it can perform tasks such as integration using tools designed for integration.
[0193] The output module 730 can output the integrated information described above.
[0194] The AI agent 700 according to embodiments of this disclosure can simply and effectively improve the level of intelligence, as well as enhance flexibility and versatility.
[0195] Figure 8 A block diagram schematically illustrates an electronic device suitable for implementing a large-model-based resource processing method according to embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0196] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0197] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0198] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the large model-based resource processing method. For example, in some embodiments, the large model-based resource processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the large model-based resource processing method described above may be performed. Alternatively, in other embodiments, computing unit 801 may be configured to perform resource processing methods based on large models by any other suitable means (e.g., by means of firmware).
[0199] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0200] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0201] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0202] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0203] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0204] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0205] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0206] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A large model-based resource processing method, comprising: in response to receiving a to-be-processed resource, generating prompt information based on a plurality of dimensions for a subject associated with the to-be-processed resource; inputting the prompt information and the to-be-processed resource into a large model to obtain dimension information of each of the plurality of dimensions; integrating the dimension information of the plurality of dimensions according to integration instructions to obtain integrated information of the to-be-processed resource, wherein the integration instructions indicate an integration manner for the dimension information; and performing a processing operation on the to-be-processed resource according to a processing instruction matched with the integrated information.
2. The method of claim 1, wherein, The performing of the processing operation on the to-be-processed resource according to the processing instruction matched with the integrated information comprises: in a case where the integrated information indicates that the to-be-processed resource belongs to a first level, positively adjusting a recommendation weight for the to-be-processed resource; and in a case where the integrated information indicates that the to-be-processed resource belongs to a second level, performing at least one of the following operations on the to-be-processed resource: negatively adjusting the recommendation weight for the to-be-processed resource, limiting a recommendation range of the to-be-processed resource, wherein the first level is higher than the second level.
3. The method of claim 2, wherein, The dimension information comprises a dimension value and an evaluation description, and the evaluation description indicates a reason why the to-be-processed resource has the dimension value under the dimension; in a case where the integrated information indicates that the to-be-processed resource belongs to a second level, the method further comprises: sending an adjustment instruction to a production party corresponding to the to-be-processed resource, so that the production party adjusts the to-be-processed resource and returns an adjusted resource, wherein the adjustment instruction is determined according to the evaluation description of each dimension; and in response to integrated information of the adjusted resource indicating that the adjusted resource belongs to a first level, positively adjusting a recommendation weight for the to-be-processed resource.
4. The method of any one of claims 1 to 3, wherein, The generating of the prompt information based on the plurality of dimensions for the subject associated with the to-be-processed resource comprises: adding a target indicator for the dimension into a prompt template to obtain the prompt information, wherein the prompt template comprises instruction information indicating that the to-be-processed resource is processed based on the target indicator.
5. The method of claim 4, further comprising: generating the prompt template based on the instruction information and condition information; wherein the condition information matches the dimension, and the condition information is used to indicate a limitation of the target indicator on obtaining the dimension information.
6. The method of any one of claims 1 to 5, further comprising: for each picture in the to-be-processed resource, cutting a subject region image from the picture according to a position of the subject in the picture obtained by performing target detection on the picture; processing the subject region image based on a dimension for the subject to obtain picture evaluation information of each picture; and determining the dimension information according to the picture evaluation information of each picture.
7. The method of claim 6, wherein, The determining of the dimension information according to the picture evaluation information of each picture comprises: determine a fusion instruction for the dimension, wherein the fusion instruction indicates a fusion manner for the picture evaluation information; and fuse the picture evaluation information of each of the pictures according to the fusion instruction to obtain the dimension information.
8. The method of any one of claims 1 to 7, wherein, The dimension is determined in the following manner: classify the subject to obtain a subject category by performing content understanding on the to-be-processed resource; and determine the plurality of dimensions from a plurality of candidate dimensions according to the subject category.
9. The method of claim 8, wherein, The subject includes a digital person, and the candidate dimensions for the digital person include at least one of the following: completeness of the digital person, clarity of the digital person, edge roughness of the digital person, color abnormality of the digital person, and proportion of the digital person in the picture.
10. The method of claim 8, wherein, The subject includes an article, and the candidate dimensions for the article include at least one of the following: completeness of the article, and completeness of article information of the article.
11. The method of any one of claims 1 to 10, wherein, The dimension information includes a dimension value. The integrating the dimension information of the plurality of dimensions according to the integration instruction to obtain the integration information of the to-be-processed resource includes: determining an evaluation level of the dimension based on the dimension value with a first evaluation threshold for the dimension as a reference; and determining the integration information according to the evaluation level of each of the dimensions.
12. The method of claim 11, wherein, The integrating the dimension information of the plurality of dimensions according to the integration instruction to obtain the integration information of the to-be-processed resource includes: integrating the dimension values of the plurality of dimensions to obtain an intermediate evaluation value according to a weight for each of the dimensions, wherein the weight represents an influence degree of the dimension on the quality of the to-be-processed resource; and determining the integration information based on the intermediate evaluation value with a second evaluation threshold as a reference.
13. A resource processing apparatus based on a large model, comprising: a generation module configured to generate prompt information based on a plurality of dimensions for a subject associated with a to-be-processed resource in response to receiving the to-be-processed resource; an obtaining module configured to input the prompt information and the to-be-processed resource into a large model to obtain dimension information of each of the plurality of dimensions; an integrating module configured to integrate the dimension information of the plurality of dimensions according to an integration instruction to obtain integration information of the to-be-processed resource, wherein the integration instruction indicates an integration manner for the dimension information; and a processing module configured to perform a processing operation on the to-be-processed resource according to a processing instruction matched with the integration information.
14. An intelligent agent based on a large model, comprising: an input module configured to receive input information; a processing module configured to determine a target task based on the input information received by the input module, determine a target large model based on the target task, execute the method of claims 1-12 by calling the target large model, and obtain output information; an output module configured to output the output information obtained by the processing module.
15. An electronic device, comprising: one or more processors; a memory configured to store one or more computer programs, characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 12.
16. A computer readable storage medium having stored thereon a computer program or instructions, characterized in that, The computer program or instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 12.
17. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 12.