Resource processing method and device, electronic equipment and storage medium
Through domain perception retrieval and dynamic analysis of multimodal large language model, the target cover image of multimedia resources is determined, which solves the problem of retraining the model in the existing technology, and realizes efficient and high-quality cover image determination.
Patent Information
- Application Number
- CN202510150042.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art requires retraining the classification model when processing multimedia resources in different fields, resulting in low efficiency in determining the cover image.
By obtaining multimedia resources, performing field classification, retrieving candidate cover images that meet the resource field, conducting business feature detection, building prompt instructions, input candidate cover images and business feature information into the multimodal large language model for content generation, and obtaining prediction scores to determine the target cover image.
While ensuring the high quality of the target cover image of the multimedia resource, it effectively reduces the workload, improves the efficiency of the cover image determination, and has strong migration. It can handle multimedia resources in different fields without retraining the model.
Smart Images

Figure CN120067349A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and particularly to a resource processing method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of network technology, multimedia resources such as videos, picture sets, and live broadcasts recommended based on the network have the advantages of large coverage and strong timeliness, and are widely used. When recommending resources, it is mostly necessary to display the cover image of the multimedia resource, and the cover image is crucial for the recommendation effect of the multimedia resource.
[0003] In the related art, for multimedia resources in a specific field, a classification model trained for that field is usually used to determine the cover image. However, when processing multimedia resources in other fields, it is necessary to retrain the classification model, and the workload of retraining the model is relatively large, resulting in low efficiency in determining the cover image. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in the present disclosure. This overview is not intended to limit the scope of protection of the claims.
[0005] Embodiments of the present disclosure provide a resource processing method, apparatus, electronic device, and storage medium, which can effectively reduce the workload while ensuring that the target cover image of the multimedia resource has high quality, thereby improving the efficiency of determining the target cover image.
[0006] On the one hand, an embodiment of the present disclosure provides a resource processing method, including:
[0007] Obtain a multimedia resource, perform domain classification on the multimedia resource to obtain the resource domain of the multimedia resource, and retrieve multiple candidate cover images that conform to the resource domain from the multimedia resource;
[0008] Perform business feature detection on each of the candidate cover images to obtain the business feature information corresponding to each of the candidate cover images;
[0009] Construct a prompt instruction for prompting to score the candidate cover images, input the prompt instruction, multiple candidate cover images, and the corresponding business feature information into a multi-modal large language model for content generation, and obtain the predicted scores of the multiple candidate cover images;
[0010] Based on the predicted scores, determine the target cover image corresponding to the multimedia resource from the multiple candidate cover images.
[0011] On the other hand, an embodiment of the present disclosure further provides a resource processing apparatus, including:
[0012] An acquisition module, configured to acquire a multimedia resource, classify the multimedia resource by field to obtain the resource field of the multimedia resource, and retrieve a plurality of candidate cover images that conform to the resource field from the multimedia resource;
[0013] A detection module, configured to respectively perform business feature detection on each of the candidate cover images to obtain business feature information corresponding to each of the candidate cover images;
[0014] A scoring module, configured to construct a prompt instruction for prompting to score the candidate cover images, input the prompt instruction, a plurality of the candidate cover images, and the corresponding business feature information into a multi-modal large language model for content generation, and obtain predicted scores of the plurality of candidate cover images;
[0015] A cover determination module, configured to determine a target cover image corresponding to the multimedia resource from the plurality of candidate cover images based on the predicted scores.
[0016] Further, the multimedia resource includes a plurality of original images, and the acquisition module is specifically configured to:
[0017] Construct a query text that conforms to the resource field;
[0018] Respectively determine a first similarity between each of the original images and the query text;
[0019] For each of the original images, when the first similarity is greater than or equal to a preset first threshold, determine the original image as a candidate cover image that conforms to the resource field.
[0020] Further, the multimedia resource includes meta information text, and the acquisition module is specifically configured to:
[0021] Respectively perform target detection on each of the original images to obtain key regions corresponding to each of the original images;
[0022] Extract keywords from the meta information text to obtain first keywords;
[0023] Respectively determine a second similarity between each of the key regions and the first keywords;
[0024] For each of the original images, when the first similarity is greater than or equal to a preset first threshold and the second similarity is greater than or equal to a preset second threshold, determine the original image as a candidate cover image that conforms to the resource field.
[0025] Further, the acquisition module is specifically configured to:
[0026] Extract keywords from the meta-information text to obtain multiple second keywords;
[0027] Retrieve the first keyword that matches the resource field from among the multiple second keywords.
[0028] Furthermore, the multimedia resource includes meta-information text, and the above resource processing device further includes a filtering module. The filtering module is specifically configured to:
[0029] Determine the third similarity between each candidate cover image and the meta-information text respectively;
[0030] Filter the multiple candidate cover images based on the third similarity.
[0031] Furthermore, the above filtering module is specifically configured to:
[0032] Sort the candidate cover images in descending order based on the third similarity;
[0033] For the first n - 1 candidate cover images after sorting, determine the adjacent ratio corresponding to the i-th candidate cover image according to the ratio between the third similarity corresponding to the i-th candidate cover image and the third similarity corresponding to the (i + 1)-th candidate cover image, where n is the number of candidate cover images, i ≤ n, and i is a positive integer;
[0034] When the adjacent ratio corresponding to the k-th candidate cover image is greater than a preset third threshold, retain the first k candidate cover images after sorting and eliminate the remaining candidate cover images, where k ≤ n and k is a positive integer.
[0035] Furthermore, the multimedia resource includes meta-information text and multiple original images, and one of the original images is configured as the initial cover image. The above acquisition module is specifically configured to:
[0036] Encode the initial cover image based on a multimodal encoding model to obtain a first image feature;
[0037] Encode the meta-information text based on the multimodal encoding model to obtain a first text feature;
[0038] Concatenate the first image feature and the first text feature and input them into a vertical domain classification model for classification to obtain the resource field of the multimedia resource.
[0039] Furthermore, the multimedia resource includes original audio, and the above acquisition module is specifically configured to:
[0040] After splicing the first image feature and the first text feature, input them into the vertical domain classification model for classification;
[0041] When the confidence of the classification result is less than or equal to a preset fourth threshold, perform speech recognition on the original audio to obtain a recognized text;
[0042] Extract keywords from the recognized text to obtain a third keyword;
[0043] Input the third keyword into the first large language model for domain prediction to determine the resource domain of the multimedia resource.
[0044] Further, the above detection module is specifically used for:
[0045] Input each of the candidate cover images into the attractiveness prediction model for prediction to obtain the attractiveness score corresponding to each candidate cover image;
[0046] Input each of the candidate cover images into the clarity prediction model for prediction to obtain the clarity score corresponding to each candidate cover image;
[0047] Based on the attractiveness score and the clarity score, determine the service feature information corresponding to each candidate cover image.
[0048] Further, the above detection module is specifically used for:
[0049] Input each of the candidate cover images into the emotion prediction model for prediction to obtain the predicted emotion corresponding to each candidate cover image;
[0050] Determine the expected emotion of the resource domain, and respectively determine the fourth similarity between each predicted emotion and the expected emotion;
[0051] Based on the attractiveness score, the clarity score, and the fourth similarity, determine the service feature information corresponding to each candidate cover image.
[0052] Further, both the attractiveness prediction model and the clarity prediction model match the audience object of the multimedia resource. The above resource processing device further includes an adjustment module, and the adjustment module is specifically used for:
[0053] Construct an object description text for describing the audience object;
[0054] Add the object description text to the prompt instruction.
[0055] On the other hand, an embodiment of the present disclosure also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the above-mentioned resource processing method is implemented.
[0056] On the other hand, an embodiment of the present disclosure also provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned resource processing method is implemented.
[0057] On the other hand, an embodiment of the present disclosure also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned resource processing method.
[0058] The embodiments of the present disclosure at least include the following beneficial effects: By obtaining multimedia resources, then performing domain classification on the multimedia resources to obtain the resource domain of the multimedia resources, then retrieving multiple candidate cover images that meet the resource domain from the multimedia resources, and then respectively performing business feature detection on each candidate cover image to obtain the business feature information corresponding to each candidate cover image. The business feature information provides the context closely related to the business requirements. Then, a prompt instruction is constructed, and the prompt instruction, multiple candidate cover images, and the corresponding business feature information are input into a multi-modal large language model for content generation to obtain the prediction scores of the multiple candidate cover images. Then, based on the prediction scores, high-quality target cover images are selected as the cover images of the multimedia resources. It realizes first extracting highly relevant candidate cover images from the multimedia resources through domain-aware retrieval, and then using the multi-modal large language model to dynamically analyze the retrieved candidate cover images. This not only narrows the analysis scope of the multi-modal large language model to improve the reasoning efficiency, but also avoids the interference of irrelevant information or low-quality information to improve the accuracy of the prediction scores. Moreover, under the guidance of the prompt instruction, the multi-modal large language model can utilize rich knowledge reserves and the external knowledge contained in the business feature information to infer accurate prediction scores for the candidate cover images and determine the target cover images, ensuring that the target cover images have high visual and semantic quality and meet the business requirements. On this basis, when processing multimedia resources in different fields, it can be directly applied without retraining, that is, directly using domain-aware retrieval and dynamically analyzing using the multi-modal large language model, making this method have strong transferability, capable of effectively reducing the workload while ensuring that the target cover images of the multimedia resources have high quality, thereby improving the determination efficiency of the cover images.
[0059] Other features and advantages of the present disclosure will be set forth in the following description, and in part will be obvious from the description, or can be learned by practicing the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0061] Figure 1 Schematic diagram of an optional implementation environment provided for an embodiment of the present disclosure;
[0062] Figure 2 Schematic diagram of an optional process flow of a resource processing method provided for an embodiment of the present disclosure;
[0063] Figure 3 Schematic diagram of an optional architecture of a multimodal large language model provided for an embodiment of the present disclosure;
[0064] Figure 4 Schematic diagram of an optional process flow for determining a first similarity provided for an embodiment of the present disclosure;
[0065] Figure 5 Schematic diagram of an optional process flow for determining a target cover image provided for an embodiment of the present disclosure;
[0066] Figure 6 Schematic diagram of an optional process flow for determining a resource field provided for an embodiment of the present disclosure;
[0067] Figure 7 Schematic diagram of an optional process flow for determining business feature information provided for an embodiment of the present disclosure;
[0068] Figure 8 Schematic diagram of an optional process flow for determining a prediction result provided for an embodiment of the present disclosure;
[0069] Figure 9 Schematic diagram of another optional process flow for determining a prediction result provided for an embodiment of the present disclosure;
[0070] Figure 10 Schematic diagram of another optional process flow for determining a prediction result provided for an embodiment of the present disclosure;
[0071] Figure 11 Schematic diagram of an optional process flow for determining a fourth similarity provided for an embodiment of the present disclosure;
[0072] Figure 12 Schematic diagram of an optional structural diagram of a resource processing device provided for an embodiment of the present disclosure;
[0073] Figure 13 Partial structural block diagram of the terminal provided by the embodiments of the present disclosure;
[0074] Figure 14 Partial structural block diagram of the server provided by the embodiments of the present disclosure. Detailed implementation manners
[0075] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure, and are not used to limit the present disclosure.
[0076] It should be noted that in each specific implementation manner of the present disclosure, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or attribute information set, etc., the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. Among them, the target object can be a user. In addition, when the embodiments of the present disclosure need to obtain the target object attribute information, the separate permission or separate consent of the target object will be obtained by means of a pop-up window or jumping to a confirmation page, etc. After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for the embodiments of the present disclosure to operate normally will be obtained.
[0077] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0078] In the related art, for multimedia resources in a specific field, a classification model trained for that field is usually used to determine the cover image. However, when processing multimedia resources in other fields, the classification model needs to be retrained, and the workload of retraining the model is relatively large, resulting in a low efficiency of determining the cover image.
[0079] Based on this, the embodiments of the present disclosure provide a resource processing method, device, electronic device, and storage medium, which can effectively reduce the workload while ensuring that the target cover image of the multimedia resource has high quality, thereby improving the efficiency of determining the target cover image.
[0080] Refer to Figure 1 , Figure 1Schematic diagram of an optional implementation environment provided by an embodiment of the present disclosure. The implementation environment includes a terminal 101 and a server 102. Among them, the terminal 101 and the server 102 are connected through a communication network.
[0081] Exemplarily, the server 102 can obtain multimedia resources, classify the multimedia resources by domain to obtain the resource domain of the multimedia resources, retrieve multiple candidate cover images that conform to the resource domain from the multimedia resources; respectively perform business feature detection on each candidate cover image to obtain the business feature information corresponding to each candidate cover image; construct a prompt instruction for prompting to score the candidate cover images, input the prompt instruction, multiple candidate cover images, and the corresponding business feature information into a multi-modal large language model for content generation to obtain the predicted scores of the multiple candidate cover images; based on the predicted scores, determine the target cover image corresponding to the multimedia resources from the multiple candidate cover images; the server 102 sends the target cover image for recommending the multimedia resources to the terminal 101.
[0082] The server 102 obtains multimedia resources, then classifies the multimedia resources by domain to obtain the resource domain of the multimedia resources, then retrieves multiple candidate cover images that conform to the resource domain from the multimedia resources, then respectively performs business feature detection on each candidate cover image to obtain the business feature information corresponding to each candidate cover image. The business feature information provides context closely related to business requirements. Then, a prompt instruction is constructed, and the prompt instruction, multiple candidate cover images, and the corresponding business feature information are input into a multi-modal large language model for content generation to obtain the predicted scores of the multiple candidate cover images. Then, based on the predicted scores, high-quality target cover images are selected as the cover images of the multimedia resources, realizing first extracting highly relevant candidate cover images from the multimedia resources through domain-aware retrieval, and then using the multi-modal large language model to dynamically analyze the retrieved candidate cover images. This not only narrows the analysis scope of the multi-modal large language model to improve the inference efficiency, but also avoids the interference of irrelevant information or low-quality information to improve the accuracy of the predicted scores. Moreover, under the guidance of the prompt instruction, the multi-modal large language model can utilize rich knowledge reserves and external knowledge contained in the business feature information to infer accurate predicted scores for the candidate cover images and determine the target cover image, ensuring that the target cover image has high visual and semantic quality and meets business requirements. On this basis, when processing multimedia resources in different domains, it can be directly applied without retraining, that is, directly using domain-aware retrieval and using the multi-modal large language model for dynamic analysis, making this method have strong transferability, capable of ensuring high quality of the target cover image of the multimedia resources while effectively reducing the workload, thereby improving the determination efficiency of the cover image.
[0083] The server 102 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Additionally, the server 102 can also be a node server in a blockchain network.
[0084] The terminal 101 can be a mobile phone, a computer, an intelligent voice interaction device, an intelligent home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication means, and the embodiments of the present disclosure do not limit this here.
[0085] Referring to Figure 2 , Figure 2 FIG. is an optional flowchart of a resource processing method provided by an embodiment of the present disclosure. The resource processing method can be executed by the server, or can also be executed by the terminal, or can also be executed by the server in cooperation with the terminal. The resource processing method includes, but is not limited to, the following steps 201 to step 204.
[0086] Step 201: Obtain a multimedia resource, perform domain classification on the multimedia resource to obtain the resource domain of the multimedia resource, and retrieve multiple candidate cover images that conform to the resource domain from the multimedia resource.
[0087] Among them, the multimedia resource is a resource including multiple images. For example, the multimedia resource can be a video, a picture set, an article, etc. Taking the multimedia resource as a video as an example, the images included in the multimedia resource can be each video frame of the video. Taking the multimedia resource as a picture set as an example, the images included in the multimedia resource can be the images in the picture set. Taking the multimedia resource as an article as an example, the images included in the multimedia resource can be the images carried by the article.
[0088] It should be noted that the multimedia resource can be divided into two categories: non-real-time and real-time. The non-real-time multimedia resource refers to a resource that is prepared, processed, and transmitted in advance. For example, a pre-recorded video, and such a multimedia resource can be accessed at any time after the transmission is completed. The real-time multimedia resource refers to a resource that is captured and transmitted simultaneously when an event occurs. For example, a live video, and such a multimedia resource can be accessed in real time.
[0089] Specifically, for the domain classification of multimedia resources, it can be based on the content, usage, and characteristics of the multimedia resources to classify the multimedia resources into one or more specific domains. The classified resource domains are used to represent the knowledge domain or application domain to which the multimedia resources belong, etc. For example, the resource domains include the fashion domain, the education domain, the technology domain, the entertainment domain, the food domain, etc. Taking a video for programming learning as an example of multimedia resources, the resource domain of this multimedia resource can be the education domain. Taking a movie clip as an example of multimedia resources, the resource domain of this multimedia resource can be the entertainment domain.
[0090] It can be understood that by classifying the multimedia resources by domain, the accurate resource domain of the multimedia resources can be obtained, and then through domain-aware retrieval, multiple candidate cover images that meet the resource domain can be retrieved from the multimedia resources, which is equivalent to extracting candidate cover images highly relevant to the resource domain from the images included in the multimedia resources.
[0091] For example, taking the resource domain as the food domain, domain-aware retrieval can use the images related to food in the multimedia resources as candidate cover images, rather than using the images not related to food in the multimedia resources as candidate cover images, which can avoid the interference of irrelevant information or low-quality information.
[0092] Step 202: Perform business feature detection on each candidate cover image respectively to obtain the business feature information corresponding to each candidate cover image.
[0093] It can be understood that business feature detection is used to determine the information associated with business requirements. Business feature information is the information associated with business requirements. Business requirements refer to the requirements for the cover image to meet specific business goals. For example, business requirements can include requirements such as high attractiveness of the cover image, high clarity of the cover image, or the cover image having a title, etc. Business feature information can be used as the context for evaluating candidate cover images. By using business feature information as the external knowledge of the multimodal large language model, the higher the prediction score of the candidate cover image that better meets the business requirements, and the lower the prediction score of the candidate cover image that less meets the business requirements, which can improve the evaluation accuracy of candidate cover images.
[0094] In addition, since different business requirements usually associate with different business feature information, and there are a wide variety of business requirement types, business feature information has the characteristic of diversification. For example, business feature information can include information in dimensions such as physical, information, quality, perception, etc. In actual application, all business feature information can be detected simultaneously, or only the business feature information associated with the current business requirements can be detected. The embodiments of the present disclosure do not limit this here.
[0095] Specifically, the business feature information in the physical dimension may include texture, edge, brightness, contrast, etc. The business feature information in the information dimension may include subject information, person information, emotion information, information entropy, etc. The business feature information in the quality dimension may include information for characterizing whether an image is a spliced image or a text image, and may also include information on whether it contains mosaics, QR codes, sensitive information, as well as information on whether it has been compressed or stretched. The business feature information in the perception dimension may include attractiveness, clarity, etc.
[0096] It should be noted that different business requirements usually need to analyze the business feature information in different dimensions. For example, if the business requirement is an image with high attractiveness and high clarity, then it is necessary to detect the attractiveness and clarity of the candidate cover image through business feature detection, and then use the attractiveness and clarity as external knowledge for the multimodal large language model, so that the multimodal large language model can generate accurate prediction scores based on the attractiveness and clarity, thereby determining the target cover image that meets the business requirements.
[0097] Step 203: Construct a prompt instruction for prompting the scoring of candidate cover images, and input the prompt instruction, multiple candidate cover images, and the corresponding business feature information into the multimodal large language model for content generation to obtain the prediction scores of the multiple candidate cover images.
[0098] Among them, multimodal refers to information from different senses or sources, such as vision, hearing, touch, etc. The multimodal large language model (Multimodal Large Language Models, MLLMs) is an extended form of the large language model. Compared with the large language model that can only process input data corresponding to the text modality, the multimodal large language model can not only process input data corresponding to the text modality, but also process input data corresponding to other modalities except the text modality. For example, other modalities include the visual modality, the audio modality, and so on. Therefore, the multimodal large language model can process the prompt instruction and business feature information in the text modality, and can also process the candidate cover images in the visual modality.
[0099] Among them, the prompt can be understood as a way to start a multi-modal large language model. Since the prompt can prompt the multi-modal large language model to score candidate cover images. For example, the prompt can be "Please generate the predicted score of the following cover image", and the user can also give a more specific prompt according to specific business requirements. For example, the prompt can also be "Require the cover image to be as clear as possible, please generate the predicted score of the following cover image", or "Require the cover image to include a title, please generate the predicted score of the following cover image". At this time, the prompt is used to prompt to score the candidate cover image with reference to the corresponding business feature information. Therefore, under the guidance of the prompt text, the multi-modal large language model can more deeply understand the information related to scoring in the candidate cover image by combining the business feature information, so as to generate an accurate predicted score.
[0100] In a possible implementation, referring to Figure 3 , Figure 3 is an optional architecture diagram of the multi-modal large language model provided by the embodiments of the present disclosure. The prompt, multiple candidate cover images, and corresponding business feature information are input into the multi-modal large language model, specifically referring to inputting the prompt, multiple candidate cover images, and the business feature information corresponding to each candidate cover image into the multi-modal large language model simultaneously or sequentially. The multi-modal large language model can determine the embeddings of the corresponding modal data through multiple embedding modules respectively, that is, the multi-modal large language model provides a general multi-modal embedding space and can effectively combine multi-modal information.
[0101] For example, the multi-modal large language model performs embedding processing on the prompt through the word embedding module to obtain the embedding of the prompt, performs embedding processing on the candidate cover image through the graph embedding module to obtain the embedding of the candidate cover image, and performs embedding processing on the business feature information through the context embedding module to obtain the embedding of the business feature information. The context embedding module can specifically be a word embedding module.
[0102] Among them, the graph embedding module can include a visual encoder and an adapter. By encoding the candidate cover image through the visual encoder, visual features in the visual feature space can be extracted. By mapping the visual features through the adapter, the embedding of the candidate cover image in the text embedding space can be obtained. The embedding of the candidate cover image has language information, enabling the multi-modal large language model to effectively process the input data.
[0103] Then, the multi-modal large language model can determine the input sequence obtained by splicing the embeddings of the prompt instruction, the embeddings of multiple candidate cover images, and the embeddings of business feature information. Then, the multi-modal large language model generates prediction scores for multiple candidate cover images based on the input sequence. The internal architecture of the multi-modal large language model can be in various forms.
[0104] For example, a transformer decoding layer and a regression layer can be set inside the multi-modal large language model. After inputting the input sequence into the transformer decoding layer, the transformer decoding layer can perform self-attention processing on the input sequence, which is equivalent to weighted fusion of multi-modal data, aligning each modality, generating the first prediction features corresponding to multiple candidate cover images, and then inputting the first prediction features into the regression layer for regression to obtain the prediction scores corresponding to multiple candidate cover images.
[0105] Specifically, the fusion formula for multi-modal data is as follows:
[0106] E fused = α·E t (x t ) + β·E r (X′ r ) + γ·E c (X c )
[0107] Among them, E fused is the first prediction feature, x t is the prompt instruction, X′ r is the image set of multiple candidate cover images, X c is the information set of business feature information corresponding to each candidate cover image, E t is the word embedding module, E r is the graph embedding module, E c is the context embedding module. α, β, and γ are all modality weighting coefficients determined by the multi-modal large language model after training. The modality weighting coefficients can be represented as learnable attention matrices in the attention mechanism.
[0108] Then, the formula for determining the prediction score is as follows:
[0109] score i = f score (E fused , x ri )
[0110] Among them, score i is the prediction score corresponding to the i-th candidate cover image in the image set, f score is the regression layer, E fused is the first prediction feature, xri is the i-th candidate cover image in the image set, where i ≤ k and k is the number of candidate cover images.
[0111] For another example, in another possible implementation, a multi-modal large language model may be internally provided with a transformer decoding layer and a classification layer. After inputting the input sequence into the transformer decoding layer, the transformer decoding layer can perform self-attention processing on the input sequence to sequentially generate multiple second prediction features. Inputting the second prediction features into the classification layer for classification can obtain the probability distribution of each candidate token in the vocabulary, and then sampling the vocabulary through the probability distribution to obtain the target token. The prediction scores corresponding to multiple candidate cover images are determined by the target tokens obtained through sequential sampling. Assuming that adjacent two prediction scores are separated by [], if the target tokens obtained through sequential sampling are 9, 8, [], 9, 0, [], then the first prediction score can be 98 and the second prediction score can be 90.
[0112] Exemplarily, assuming that a prompt instruction, 5 candidate cover images, and corresponding service feature information are input into the multi-modal large language model, the multi-modal large language model can output 5 prediction scores. Assuming the output is "98 90 86 70 65", then it can be determined that the prediction score of the first candidate cover image is 98, the prediction score of the second candidate cover image is 90, the prediction score of the third candidate cover image is 86, the prediction score of the fourth candidate cover image is 70, and the prediction score of the fifth candidate cover image is 65.
[0113] In a possible implementation, before generating the prediction scores of multiple candidate cover images based on the input sequence, the resource processing method further includes: in the input sequence, adding an image marker to the head of the embedding of each candidate cover image and adding a reference marker to the head of the embedding of the service feature information; based on this, by adding the image marker and the reference marker, the multi-modal large language model can understand and focus on each candidate cover image and the service feature information, thereby improving the accuracy of the prediction scores.
[0114] Step 204: Based on the prediction scores, determine the target cover image corresponding to the multimedia resource from multiple candidate cover images.
[0115] Among them, the predicted score is used to characterize the quality level of the candidate cover image. The higher the predicted score, the higher the quality of the candidate cover image, and the lower the predicted score, the lower the quality of the candidate cover image. Therefore, among multiple candidate cover images, a suitable target cover image can be selected through the predicted score. For example, when the number of candidate cover images with the highest predicted score is one, the candidate cover image with the highest predicted score is used as the target cover image. When the number of candidate cover images with the highest predicted score is multiple, one can be randomly selected from the multiple candidate cover images with the highest predicted score as the target cover image.
[0116] Based on this, by obtaining the multimedia resource, then classifying the multimedia resource by domain to obtain the resource domain of the multimedia resource, then retrieving multiple candidate cover images that conform to the resource domain from the multimedia resource, then respectively detecting the business characteristics of each candidate cover image to obtain the business characteristic information corresponding to each candidate cover image, the business characteristic information provides the context closely related to the business requirements, then constructing a prompt instruction, and inputting the prompt instruction, multiple candidate cover images, and the corresponding business characteristic information into the multi-modal large language model for content generation to obtain the predicted scores of the multiple candidate cover images, and then screening out high-quality target cover images based on the predicted scores as the cover image of the multimedia resource, it realizes first extracting highly relevant candidate cover images from the multimedia resource through domain-aware retrieval, and then using the multi-modal large language model to dynamically analyze the retrieved candidate cover images, which not only narrows the analysis scope of the multi-modal large language model to improve the reasoning efficiency, but also avoids the interference of irrelevant information or low-quality information to improve the accuracy of the predicted score. Moreover, under the guidance of the prompt instruction, the multi-modal large language model can utilize the rich knowledge reserve and the external knowledge contained in the business characteristic information to infer an accurate predicted score for the candidate cover image and determine the target cover image, ensuring that the target cover image has high visual and semantic quality and meets the business requirements. On this basis, when processing multimedia resources in different domains, it can be directly applied without retraining, that is, directly using domain-aware retrieval and dynamically analyzing using the multi-modal large language model, making this method have strong transferability, capable of ensuring the high quality of the target cover image of the multimedia resource while effectively reducing the workload, thereby improving the determination efficiency of the cover image.
[0117] In a possible implementation manner, the multimedia resource includes multiple original images. Retrieving multiple candidate cover images that conform to the resource domain from the multimedia resource can specifically be constructing a query text that conforms to the resource domain; respectively determining the first similarity between each original image and the query text; for each original image, when the first similarity is greater than or equal to a preset first threshold, determining the original image as a candidate cover image that conforms to the resource domain.
[0118] Among them, the original image is the image in the original state in the multimedia resource. Taking the multimedia resource as a video as an example, the original image can be the video frame in the original state in the video. Taking the multimedia resource as an atlas as an example, the original image can be the image in the original state in the atlas; the query text is a natural language description used to represent the resource domain. Through the query text, the features, attributes, content, etc. of the resource domain can be captured. Assuming the resource domain is the food domain, the query text can be a short phrase. For example, the query text is "food", and the query text can be a long sentence. For example, the query text is "This is a picture showing food".
[0119] Among them, the first similarity is used to characterize the semantic consistency between the original image and the query text. When the first similarity is higher, the semantic consistency between the original image and the query text is higher, indicating that the original image is more likely to be an image highly relevant to the resource domain. When the first similarity is lower, the semantic consistency between the original image and the query text is lower, indicating that the original image is less likely to be an image highly relevant to the resource domain.
[0120] It can be understood that in domain-aware retrieval, by constructing the query text of the resource domain, and then respectively determining the first similarity between each original image and the query text. For each original image, when the first similarity is greater than or equal to the first threshold, it can be considered that the original image is an image highly relevant to the resource domain, and the highly relevant original images are determined as candidate cover images that meet the resource domain, realizing the accurate retrieval of highly relevant candidate cover images from the multimedia resource, which not only narrows the analysis scope of the multimodal large language model to improve the inference efficiency, but also avoids the interference of irrelevant information or low-quality information to improve the accuracy of the prediction score.
[0121] It should be noted that the number of query texts can be one or more. When the number of query texts is multiple, for each original image, it is necessary to respectively determine the reference similarity between the original image and each query text, and then take the maximum reference similarity as the first similarity of the original image, so that when the semantic consistency between the original image and any one of the query texts in the resource domain is relatively high, it is considered that the original image is an image highly relevant to the resource domain, and the appropriate candidate cover images can be quickly and accurately screened out.
[0122] It should be noted that the value range of the first similarity can be from 0 to 1, and the first threshold is a real number between 0 and 1. For example, the first threshold is 0.8. The appropriate first threshold can be obtained through existing experience from multiple experiments or selected based on other strategies. The embodiments of the present disclosure do not limit this here.
[0123] Specifically, the original image is data in the image modality, the query text is data in the text modality, and the first similarity is the similarity between the image modality data and the text modality data. Specifically, the first similarity between the original image and the query text can be determined by a multimodal retriever, that is, retrieval-augmented generation (RAG) is performed through the multimodal retriever to achieve domain-aware retrieval of each original image in the multimedia resource, and the multimodal retriever can output candidate cover images. In the multimodal retriever, the first similarity can be determined in various ways, and the following describes in detail various ways to determine the first similarity.
[0124] In the first way to determine the first similarity, refer to Figure 4 , Figure 4 which is an optional flowchart for determining the first similarity provided by the embodiments of the present disclosure. Each original image is encoded by the image encoder in the Contrastive Language–Image Pre-training (CLIP) to obtain the second image feature corresponding to each original image, and the query text is encoded by the text encoder in the CLIP to obtain the second text feature, so that the second image feature and the second text feature are in the same feature space. Then, the similarity between each second image feature and the second text feature is determined respectively to obtain the first similarity between each second image feature and the query text. The first similarity can be a cosine similarity or other similarities.
[0125] In the second way to determine the first similarity, content analysis can be performed on each original image to determine the image description text of each original image, and then the first similarity between each image description text and the query text is determined respectively to obtain the first similarity between each second image feature and the query text.
[0126] In a possible implementation, the multimedia resource includes meta-information text. For each original image, when the first similarity is greater than or equal to a preset first threshold, the original image is determined as a candidate cover image that conforms to the resource domain. Specifically, object detection can be performed on each original image to obtain the key region corresponding to each original image; keyword extraction is performed on the meta-information text to obtain the first keyword; the second similarity between each key region and the first keyword is determined respectively; for each original image, when the first similarity is greater than or equal to the preset first threshold and the second similarity is greater than or equal to the preset second threshold, the original image is determined as a candidate cover image that conforms to the resource domain.
[0127] Among them, the meta-information text is the text containing the basic information of the multimedia resource. For example, the meta-information text may include the title, abstract, or tags of the multimedia resource. The meta-information text can be the text input by the user when uploading the multimedia resource, or it can also be the content analysis result of the multimedia resource. For example, when the user does not input the meta-information text, the multimedia resource can be automatically analyzed for content to obtain the meta-information text of the multimedia resource. The first keyword refers to the word that can summarize the core content, features, or uses of the meta-information text, etc. For example, if the meta-information text is "10 simple and delicious home-cooked dishes recipes", then the extracted first keywords may include: home-cooked dishes, delicious, recipes, simple.
[0128] Among them, the second similarity is used to characterize the semantic consistency between the key region and the first keyword. When the second similarity is higher, the semantic consistency between the key region and the first keyword is stronger, and when the second similarity is lower, the semantic consistency between the key region and the first keyword is weaker.
[0129] It should be noted that through object detection, the target objects in the original image can be detected, and the positions of the target objects are indicated by bounding boxes. Since the key region can refer to the region where the target object is located in the original image, the region within the bounding box determined by object detection can be used as the key region. One or more bounding boxes can be determined through object detection, indicating that the original image can include one or more target objects. At this time, one or more key regions can be obtained. In addition, through object detection, it may also be impossible to determine any bounding box, indicating that the original image may not include target objects, or the target objects cannot be detected by this detection method. At this time, the key region is empty. When the key region is empty, the second similarity between the key region and any first keyword can be set to 0.
[0130] It can be understood that during domain-aware retrieval, first, the first similarity between each original image and the query text is determined separately. Then, object detection is performed on each original image to obtain the key regions of each original image, and keyword extraction is performed on the meta-information text to obtain the first keywords. Next, the second similarity between each key region and the first keyword is determined separately. Since the range of the key region is small and the range of the first keyword is small, the determination efficiency of the second similarity is faster. Also, since the key region is within the original image and the first keyword is within the meta-information text, the second similarity can represent the consistency between the original image and the meta-information text from another dimension. When the second similarity is greater than or equal to the second threshold, it can be considered that the consistency between the original image and the meta-information text is high. Based on this, the second similarity is determined quickly. Then, on the premise that the first similarity is greater than or equal to the first threshold, it is also necessary to satisfy the condition that the second similarity is greater than or equal to the second threshold to determine that the original image is an image highly relevant to the resource domain, which can further improve the retrieval accuracy of candidate cover images.
[0131] It should be noted that the value range of the second similarity can be from 0 to 1, and the second threshold is a real number between 0 and 1. For example, the second threshold is 0.8. The appropriate second threshold can be obtained through existing experience from multiple experiments or selected based on other strategies. The embodiments of the present disclosure do not limit this here.
[0132] Specifically, the key region is data in the image modality, and the first keyword is data in the text modality. Therefore, the second similarity determined by the key region and the first keyword is the similarity between two modalities of data. Referring to the aforementioned method for determining the first similarity, the second similarity can also be determined in a similar way.
[0133] In a possible implementation, keyword extraction is performed on the meta-information text to obtain the first keyword. Specifically, keyword extraction can be performed on the meta-information text to obtain multiple second keywords, and the first keyword that conforms to the resource domain is retrieved from the multiple second keywords.
[0134] Based on this, keyword extraction is performed on the meta-information text to obtain multiple second keywords, and then the first keyword that conforms to the resource domain is retrieved from the multiple second keywords, ensuring that the first keyword can indicate relevant objects in the resource domain, which can improve the reliability of the second similarity and thus further improve the retrieval accuracy of candidate cover images.
[0135] Exemplarily, assume that the resource domain is the food domain, and the meta-information text is "How to promote local special cuisine through short videos to let more people know the taste of their hometown". Then, the second keywords extracted can include: short videos, promotion, local special cuisine, taste of hometown, etc. Further, the first keywords that match the resource domain retrieved can include: local special cuisine and taste of hometown.
[0136] In a possible implementation, the multimedia resource includes the meta-information text. After retrieving multiple candidate cover images that match the resource domain from the multimedia resource, the resource processing method further includes: respectively determining the third similarity between each candidate cover image and the meta-information text; filtering the multiple candidate cover images based on the third similarity.
[0137] Based on this, by determining the third similarity between the candidate cover image and the meta-information text, and then filtering the multiple candidate cover images based on the third similarity, inappropriate candidate cover images can be filtered out, thereby further narrowing the analysis scope of the multimodal large language model and avoiding the interference of irrelevant information or low-quality information.
[0138] Specifically, the candidate cover image is data in the image modality, and the meta-information text is data in the text modality. Therefore, the third similarity determined by the candidate cover image and the meta-information text is the similarity between two modality data. Referring to the aforementioned method for determining the first similarity, the third similarity can also be determined in a similar way.
[0139] It should be noted that based on the third similarity, the multiple candidate cover images can be filtered in various ways. The following describes in detail various ways of filtering candidate cover images.
[0140] In the first way of filtering candidate cover images, filtering the multiple candidate cover images based on the third similarity can specifically be sorting each candidate cover image in descending order based on the third similarity; for the first n - 1 candidate cover images after sorting, determining the adjacent ratio corresponding to the i-th candidate cover image according to the ratio between the third similarity corresponding to the i-th candidate cover image and the third similarity corresponding to the (i + 1)-th candidate cover image, where n is the number of candidate cover images, i ≤ n, and i is a positive integer; when the adjacent ratio corresponding to the k-th candidate cover image is greater than a preset third threshold, retaining the first k candidate cover images after sorting and excluding the remaining candidate cover images, where k ≤ n and k is a positive integer.
[0141] Among them, the adjacent ratio is used to indicate the deviation degree between the third similarities corresponding to two candidate cover images before and after sorting. When the adjacent ratio is larger, it means that the deviation degree between the third similarities corresponding to the two candidate cover images before and after sorting is larger, that is, the qualities of the two candidate cover images before and after sorting are less similar. Assuming that the candidate cover image ranked first belongs to an image with better quality, then the candidate cover image ranked later is more likely to belong to an image with worse quality. When the adjacent ratio is smaller, it means that the deviation degree between the third similarities corresponding to the two candidate cover images before and after sorting is smaller, that is, the qualities of the two candidate cover images before and after sorting are more similar.
[0142] Based on this, sort the candidate cover images obtained by domain-aware retrieval in descending order of the third similarity. Then, sort each candidate cover image in descending order of the third similarity, and then determine the adjacent ratio corresponding to each candidate cover image. Then, find the k-th candidate cover image whose adjacent ratio is greater than the third threshold. The third similarities of the first k candidate cover images after sorting are within a relatively large range. It can be considered that the first k candidate cover images after sorting are images with a relatively high alignment degree with the meta-information text, that is, they belong to images with better quality. By retaining the first k candidate cover images after sorting and excluding the remaining candidate cover images arranged after the k-th candidate cover image, the number of high-quality candidate cover images is dynamically determined by analyzing the distribution of the third similarity, and then the n candidate cover images are truncated. It can automatically filter out low-quality candidate cover images, thereby further narrowing the analysis range of the multimodal large language model and avoiding the interference of irrelevant information or low-quality information, further improving the accuracy of the prediction score, and then quickly and effectively determining the target cover image.
[0143] Specifically, referring to Figure 5 , Figure 5 is an optional flowchart for determining the target cover image provided by the embodiments of the present disclosure. Assume that the retrieval result of the multimedia resource is n candidate cover images. First, determine the third similarity corresponding to each candidate cover image, and then sort each candidate cover image in descending order of the third similarity. The first image set composed of the n candidate cover images after sorting is:
[0144] X r ={x r1 ,x r2 ,…,x rn}
[0145] Among them, X r is the first image set, x r1 is the first candidate cover image after sorting, x r2 is the second candidate cover image after sorting, x rnis the nth candidate cover image after sorting;
[0146] Then, the similarity set composed of the third similarities corresponding to the n sorted candidate cover images is:
[0147] S = {s 1 , s 2 , …, s n}
[0148] where S is the similarity set, s 1 is the third similarity corresponding to the first candidate cover image after sorting, s 2 is the third similarity corresponding to the second candidate cover image after sorting, s n is the third similarity corresponding to the nth candidate cover image after sorting;
[0149] Then, the calculation formula for the adjacent ratio is as follows:
[0150]
[0151] where u i is the adjacent ratio corresponding to the i-th candidate cover image, s i is the third similarity corresponding to the i-th candidate cover image after sorting, s i+1 is the third similarity corresponding to the (i + 1)-th candidate cover image after sorting, is used to take the logarithm of the ratio between s i and s i+1 , and i ∈ [1, n - 1] means that the i-th candidate cover image is one of the first n - 1 candidate cover images;
[0152] Then, based on the adjacent ratio, the first k candidate cover images are retained among the n candidate cover images and the remaining candidate cover images are removed to obtain the sorted filtering result, and then the target cover image is determined among the k candidate cover images.
[0153] In a possible implementation, the k-th candidate cover image is the image that is ranked the most forward among all candidate cover images with adjacent ratios greater than the third threshold, which can ensure that the retained candidate cover images are all of high quality and similar quality. At this time, the second image set composed of the retained candidate cover images is:
[0154] X′ r = {x r1 , x r2 , …, x rk}, k = min{i | u i > γ}
[0155] where X′ ris the second image set, x rk is the k-th candidate cover image after sorting, γ is the third threshold, u i is the adjacent ratio corresponding to the i-th candidate cover image. For each adjacent ratio, when u i > γ, the arrangement order i of the candidate cover image is divided into the same order set, and the order set {i|u i > γ} is obtained. min{i|u i > γ} is used to take the minimum value in the order set, so that x rk is the candidate cover image with the largest adjacent ratio greater than the third threshold and the earliest arrangement.
[0156] Exemplarily, assume that the third similarity is 7, and they are 0.9, 0.89, 0.88, 0.55, 0.5, 0.2, 0.19 from large to small in turn. Then the respective adjacent ratios can be determined as follows: 0.0049, 0.0049, 0.204, 0.0413, 0.397, 0.022. Assume γ = 0.1, u 3 and u 5 are both greater than γ. Since 3 is less than 5, then k = 3 can be determined.
[0157] In the second method of filtering candidate cover images, based on the third similarity, multiple candidate cover images are filtered. Specifically, it can be based on the order from large to small of the third similarity to sort each candidate cover image; based on the number of candidate cover images and the preset filtering ratio, the truncation number is determined; the first truncation number of candidate cover images after sorting is retained and the remaining candidate cover images are excluded. Based on this, first sort each candidate cover image based on the order from large to small of the third similarity, and then determine the truncation number through the number of candidate cover images and the preset filtering ratio. For example, multiply the number of candidate cover images by the filtering ratio to obtain the truncation number, and then quickly filter the candidate cover images based on the truncation number, realizing the efficient filtering of candidate cover images. Since the quality of the candidate cover images with a more forward sorting is usually higher, by retaining the first truncation number of candidate cover images after sorting and excluding the remaining candidate cover images, candidate cover images with higher quality can be retained, thereby further narrowing the analysis scope of the multi-modal large language model and avoiding the interference of irrelevant information or low-quality information, further improving the accuracy of the prediction score, and then quickly and effectively determining the target cover image.
[0158] In a possible implementation, refer to Figure 6 , Figure 6FIG. 0 is an alternative process diagram for determining the resource domain provided by an embodiment of the present disclosure. The multimedia resource includes a meta-information text and multiple original images, and one of the original images is configured as the initial cover image. The resource domain of the multimedia resource is obtained by performing domain classification on the multimedia resource. Specifically, the initial cover image is encoded based on a multimodal encoding model to obtain a first image feature; the meta-information text is encoded based on the multimodal encoding model to obtain a first text feature; the first image feature and the first text feature are concatenated and then input into a vertical domain classification model for classification to obtain the resource domain of the multimedia resource.
[0159] Among them, the initial cover image can be the cover image specified by the user when uploading the multimedia resource. When the user does not specify a specific cover image, one of the original images can be automatically selected as the initial cover image. For example, the first original image is used as the initial cover image.
[0160] Among them, the multimodal encoding model can map inputs of multiple modalities to the same feature space. Therefore, the multimodal encoding model can map the initial cover image and the meta-information text to the same feature space. For example, the multimodal encoding model can be a contrastive language-image pre-training model, a LanguageBind model, etc. Exemplarily, the image encoder in CLIP is used to encode the initial cover image in the image modality to obtain a first image feature, and the text encoder in CLIP is used to encode the meta-information text in the text modality to obtain a first text feature.
[0161] Based on this, through encoding by the multimodal encoding model, the first image feature encoded from the initial cover image and the first text feature encoded from the meta-information text are in the same feature space, enabling better alignment of the features between the initial cover image and the meta-information text. The first image feature and the first text feature are concatenated and then input into the vertical domain classification model, enabling the vertical domain classification model to make full use of and combine the information of both modalities for classification prediction, thereby predicting the accurate resource domain.
[0162] In a possible implementation, the multimedia resource includes original audio. The first image feature and the first text feature are concatenated and then input into the vertical domain classification model for classification to obtain the resource domain of the multimedia resource. Specifically, the first image feature and the first text feature are concatenated and then input into the vertical domain classification model for classification; when the confidence level of the classification result is less than or equal to a preset fourth threshold, speech recognition is performed on the original audio to obtain a recognized text; keywords are extracted from the recognized text to obtain a third keyword; the third keyword is input into a first large language model for domain prediction to determine the resource domain of the multimedia resource.
[0163] Among them, the vertical domain classification model is a classification model used to determine vertical domain categories. For example, vertical domain categories include fields such as fashion, education, technology, entertainment, and food. After splicing the first image feature and the first text feature, they are input into the vertical domain classification model, which can predict the scores of each vertical domain category, and then determine the resource domain of the multimedia resource as the vertical domain category with the highest score.
[0164] Among them, the first large language model belongs to the large language model (LLM). The large language model is a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. Generally, the large language model adopts a recurrent neural network (RNN) or variants such as long short-term memory network (LSTM) and gated recurrent unit (GRU) to capture the context information in the text sequence, so as to realize tasks such as natural language text generation, language model evaluation, text classification, and sentiment analysis. In the field of natural language processing, large language models have been widely used, such as speech recognition, machine translation, automatic summarization, dialogue systems, intelligent question answering, etc.
[0165] Based on this, after splicing the first image feature and the first text feature, they are input into the vertical domain classification model for classification. When the confidence level of the classification result is less than or equal to the fourth threshold, it can be considered that the prediction uncertainty of the vertical domain classification model for the current input is relatively high, indicating that the classification result may not be the actual resource domain of the multimedia resource. At this time, speech recognition is performed on the original audio of the multimedia resource to obtain the recognized text, and then keyword extraction is performed on the recognized text to obtain the third keyword, which can extract the key part of the recognized text. Then, the third keyword is input into the first large language model for domain prediction. The first large language model can use its rich knowledge reserve to quickly and accurately determine the resource domain of the multimedia resource, realizing the determination of the resource domain through audio modality information when the resource domain cannot be reliably determined through the information of image modality and text modality, ensuring the reliability of the resource domain.
[0166] In a possible implementation, refer to Figure 7 , Figure 7An optional process schematic diagram for determining service feature information provided by an embodiment of the present disclosure, which respectively performs service feature detection on each candidate cover image to obtain the service feature information corresponding to each candidate cover image. Specifically, each candidate cover image can be input into an attraction prediction model for prediction to obtain the attraction score corresponding to each candidate cover image; each candidate cover image can be input into a clarity prediction model for prediction to obtain the clarity score corresponding to each candidate cover image; based on the attraction score and the clarity score, the service feature information corresponding to each candidate cover image is determined.
[0167] Among them, the attraction prediction model can be a regression model pre-trained, and through the attraction prediction model, the attraction score can be accurately predicted. The clarity prediction model can also be a regression model pre-trained, and through the clarity prediction model, the clarity score can be accurately predicted.
[0168] Based on this, during service feature detection, the candidate cover image is input into the attraction prediction model, enabling the attraction prediction model to make a prediction based on the visual information related to attraction in the candidate cover image, thereby predicting an accurate attraction score. At the same time, the candidate cover image is input into the clarity prediction model, enabling the clarity prediction model to make a prediction based on the visual information related to clarity in the candidate cover image, thereby predicting an accurate clarity score. Both the attraction score and the clarity score can be used as service metrics for service requirements. Determining the service feature information based on the attraction score and the clarity score is equivalent to constructing comprehensive service feature information through service metrics in multiple dimensions, which can improve the comprehensiveness, reliability, and accuracy of the service feature information, thereby providing more comprehensive and effective external knowledge for the multi-modal large language model, improving the accuracy of the prediction score, and being able to determine the target cover image whose attraction score and clarity score meet the service requirements, that is, screening out a better target cover image.
[0169] It should be noted that when the service feature information is determined by the attraction score and the clarity score, the prompt instruction can be "It is required that the attraction and clarity of the cover image be as high as possible. Please generate the prediction scores for the following cover images."
[0170] It should be noted that the attraction prediction model or the clarity prediction model can be a prediction model for image modal data. For example, a convolutional neural network or a vision Transformer. The attraction prediction model or the clarity prediction model can also be a prediction model for multi-modal data. The prediction model for multi-modal data will be described in detail below.
[0171] In the prediction model of multi-modal data with the first architecture, refer to Figure 8 , Figure 8FIG. 0 is an alternative process flow diagram for determining a prediction result provided by an embodiment of the present disclosure. First, a plurality of modality data are obtained, that is, modality data 1, modality data 2, …, modality data M are obtained. For example, image modality data, text modality data, or video modality data are obtained. Among them, the image modality data may be a candidate cover image, the text modality data may be meta-information text, and the video modality data may be a multimedia resource. Then, the plurality of modality data are fused and then input into a prediction model for prediction to obtain a prediction result. For example, the prediction result may be an attractiveness score or a clarity score.
[0172] In the prediction model of multi-modal data with the second architecture, refer to Figure 9 , Figure 9 FIG. 7 is another alternative process flow diagram for determining a prediction result provided by an embodiment of the present disclosure. First, a plurality of modality data are obtained, that is, modality data 1, modality data 2, …, modality data M are obtained. For example, image modality data, text modality data, or video modality data are obtained. Among them, the image modality data may be a candidate cover image, the text modality data may be meta-information text, and the video modality data may be a multimedia resource. Then, feature extraction is performed on each modality data respectively to obtain respective corresponding data features, that is, data feature 1, data feature 2, …, data feature M are obtained. Then, each data feature is input into a corresponding sub-model for prediction, that is, data feature 1 is input into sub-model 1, data feature 2 is input into sub-model 2, and data feature 3 is input into sub-model M to obtain respective corresponding prediction scores. Then, decision fusion is performed on each prediction score to obtain a prediction result. For example, each prediction score is weighted and added to obtain an attractiveness score or a clarity score.
[0173] In the prediction model of multi-modal data with the third architecture, refer to Figure 10 , Figure 10 FIG. 14 is another alternative process flow diagram for determining a prediction result provided by an embodiment of the present disclosure. First, a plurality of modality data are obtained, that is, modality data 1, modality data 2, and modality data 3 are obtained. For example, image modality data, text modality data, or video modality data are obtained. Among them, the image modality data may be a candidate cover image, the text modality data may be meta-information text, and the video modality data may be a multimedia resource. Then, feature extraction is performed on each modality data respectively to obtain respective corresponding data features. Then, in layer 1 of the network architecture, two of the data features are fused to obtain a first fused feature, and the remaining one data feature is mapped to obtain a first mapped feature. Then, in layer 2 of the network architecture, the first fused feature and the first mapped feature are fused to obtain a second fused feature. Then, in layer L of the network architecture, the second fused feature is mapped to obtain a prediction result.
[0174] In a possible implementation, refer to Figure 11 , Figure 11 is an alternative flowchart for determining the fourth similarity provided by the embodiments of the present disclosure. Based on the attractiveness score and the clarity score, the business feature information corresponding to each candidate cover image is determined. Specifically, each candidate cover image can be input into an emotion prediction model for prediction to obtain the predicted emotion corresponding to each candidate cover image; the expected emotion in the resource field is determined, and the fourth similarity between each predicted emotion and the expected emotion is determined respectively; based on the attractiveness score, the clarity score, and the fourth similarity, the business feature information corresponding to each candidate cover image is determined.
[0175] Among them, the expected emotion can indicate the emotion demand in the resource field, and the fourth similarity is used to characterize the semantic consistency between the predicted emotion and the expected emotion. When the fourth similarity is higher, the semantic consistency between the predicted emotion and the expected emotion is higher, indicating that the emotion of the candidate cover image more conforms to the emotion demand in the resource field. When the fourth similarity is lower, the semantic consistency between the predicted emotion and the expected emotion is lower, indicating that the emotion of the candidate cover image less conforms to the emotion demand in the resource field.
[0176] Based on this, during business feature detection, in addition to predicting the attractiveness score and the clarity score corresponding to the candidate cover image, it is also necessary to input the candidate cover image into the emotion prediction model, so that the emotion prediction model can make a prediction based on the emotion-related content information in the candidate cover image, thereby predicting an accurate predicted emotion. Then, the fourth similarity between the predicted emotion and the expected emotion in the resource field is determined, and the business feature information is determined based on the attractiveness score, the clarity score, and the fourth similarity, further improving the comprehensiveness, reliability, and accuracy of the business feature information, and being able to determine the target cover image whose attractiveness score, clarity score, and fourth similarity meet the business requirements, that is, screening out a better target cover image.
[0177] It should be noted that when the business feature information is determined by the attractiveness score, the clarity score, and the fourth similarity, the prompt instruction can be "It is required that the attractiveness and clarity of the cover image be as high as possible, and the emotion of the cover image be as consistent as possible with the emotion in the resource field. Please generate the predicted scores for the following cover images."
[0178] In a possible implementation, both the attractiveness prediction model and the clarity prediction model match the audience object of the multimedia resource. After constructing a prompt instruction for prompting the scoring of the candidate cover image, the resource processing method further includes: constructing an object description text for describing the audience object; adding the object description text to the prompt instruction.
[0179] Among them, the audience refers to the object that receives content push. For example, the audience can be a user. The audience can usually be divided according to preferences. By constructing corresponding object description texts for different types of audiences, the object description texts are used to represent the preferences of the objects.
[0180] Exemplarily, assume that the audience likes food. Then the object description text can be: The audience of the cover image likes food. Add the object description text to the prompt instruction. The prompt instruction can be: The audience of the cover image likes food. Please generate the predicted score of the following cover image.
[0181] Based on this, by constructing object description texts for describing the audience, and then adding the object description texts to the prompt instruction, the multi-modal large language model can capture relevant information of the audience, guide the multi-modal large language model to make predictions for the audience, further improve the accuracy of the predicted score, make the predicted score more suitable for the audience, so as to screen out the target cover images that the audience is interested in. When pushing the target cover images of the multimedia resources to the audience subsequently, it can effectively improve the degree of interest of the audience, thereby increasing the click-through rate of the audience for the multimedia resources and improving the delivery effect of the multimedia resources.
[0182] It can be seen that the resource processing method provided by the embodiments of the present disclosure can be applied to a variety of scenarios.
[0183] For example, in the scenario of cover selection, through the resource processing method provided by the embodiments of the present disclosure, the target cover image of the multimedia resource can be quickly determined, the efficient selection of the cover image can be realized, and it can also ensure that the target cover image of the multimedia resource has high quality, which can effectively improve the degree of interest of the audience, thereby increasing the click-through rate of the audience for the multimedia resources and improving the delivery effect of the multimedia resources.
[0184] It should be noted that the resource processing method provided by the embodiments of the present disclosure can be deployed to the business pipeline and encapsulated into a reusable plug-in for wide use by multiple downstream business departments. Through efficient interface design, the plug-in can be quickly integrated into different business scenarios to meet various high-quality cover mining requirements.
[0185] Again, for example, in the scenario of cover review, the multi-modal large language model used in the resource processing method provided by the embodiments of the present disclosure is equivalent to a cover screening model. The precision rate and recall rate of this cover screening model are relatively high, and it can screen out high-quality target cover images and display the target cover images as the bottom result of the review, which can effectively improve the efficiency of the reviewer and thus significantly reduce the review cost.
[0186] The complete process of the resource processing method will be described in detail below.
[0187] First, multimedia resources are obtained, and the initial cover image is encoded based on the multimodal coding model to obtain a first image feature; the meta-information text is encoded based on the multimodal coding model to obtain a first text feature; the first image feature and the first text feature are spliced and input into the vertical domain classification model for classification; when the confidence of the classification result is less than or equal to the preset fourth threshold, speech recognition is performed on the original audio to obtain a recognized text; keyword extraction is performed on the recognized text to obtain a third keyword; the third keyword is input into the first language model for domain prediction to determine the resource domain of the multimedia resource;
[0188] Then, construct a query text that conforms to the resource field; respectively determine the first similarity between each original image and the query text; respectively perform target detection on each original image to obtain the key area corresponding to each original image; perform keyword extraction on the meta information text to obtain multiple second keywords; retrieve the first keyword that conforms to the resource field from the multiple second keywords; respectively determine the second similarity between each key area and the first keyword; for each original image, when the first similarity is greater than or equal to a preset first threshold value, and the second similarity is greater than or equal to a preset second threshold value, determine the original image as a candidate cover image that conforms to the resource field;
[0189] Then, the third similarity between each candidate cover image and the meta-information text is determined respectively; each candidate cover image is sorted in descending order based on the third similarity; for the first n-1 candidate cover images after sorting, the adjacent ratio corresponding to the i-th candidate cover image is determined according to the ratio between the third similarity corresponding to the i-th candidate cover image and the third similarity corresponding to the i+1-th candidate cover image, where n is the number of candidate cover images, i≤n, and i is a positive integer; when the adjacent ratio corresponding to the k-th candidate cover image is greater than a preset third threshold, the first k candidate cover images after sorting are retained and the remaining candidate cover images are eliminated, where k≤n, and k is a positive integer.
[0190] Then, each candidate cover image is input into the attractiveness prediction model for prediction, and the attractiveness score corresponding to each candidate cover image is obtained; each candidate cover image is input into the clarity prediction model for prediction, and the clarity score corresponding to each candidate cover image is obtained; each candidate cover image is input into the emotion prediction model for prediction, and the predicted emotion corresponding to each candidate cover image is obtained; the expected emotion of the resource field is determined, and the fourth similarity between each predicted emotion and the expected emotion is determined; based on the attractiveness score, the clarity score and the fourth similarity, the business feature information corresponding to each candidate cover image is determined;
[0191] Then, construct a prompting instruction for prompting the scoring of candidate cover images, and input the prompting instruction, multiple candidate cover images, and the corresponding business feature information into a multi-modal large language model for content generation to obtain the predicted scores of the multiple candidate cover images;
[0192] Then, based on the predicted scores, determine the target cover image corresponding to the multimedia resource from the multiple candidate cover images.
[0193] Based on this, by obtaining the multimedia resource, then classifying the multimedia resource by domain to obtain the resource domain of the multimedia resource, then retrieving multiple candidate cover images that meet the resource domain from the multimedia resource, and then respectively performing business feature detection on each candidate cover image to obtain the business feature information corresponding to each candidate cover image. The business feature information provides the context closely related to the business requirements. Then, construct a prompting instruction, and input the prompting instruction, multiple candidate cover images, and the corresponding business feature information into a multi-modal large language model for content generation to obtain the predicted scores of the multiple candidate cover images. Then, based on the predicted scores, screen out high-quality target cover images as the cover image of the multimedia resource, which realizes first extracting highly relevant candidate cover images from the multimedia resource through domain-aware retrieval, and then using the multi-modal large language model to dynamically analyze the retrieved candidate cover images. This not only narrows the analysis scope of the multi-modal large language model to improve the reasoning efficiency, but also avoids the interference of irrelevant or low-quality information to improve the accuracy of the predicted scores. Moreover, under the guidance of the prompting instruction, the multi-modal large language model can utilize the rich knowledge reserve and the external knowledge contained in the business feature information to infer accurate predicted scores for the candidate cover images and determine the target cover image, ensuring that the target cover image has high visual and semantic quality and meets the business requirements. On this basis, when processing multimedia resources in different domains, it can be directly applied without retraining, that is, directly use domain-aware retrieval and dynamic analysis using the multi-modal large language model, making this method have strong transferability, capable of effectively reducing the workload while ensuring that the target cover image of the multimedia resource has high quality, thereby improving the determination efficiency of the cover image.
[0194] It can be understood that although the steps in each of the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this embodiment, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0195] Referring to Figure 12 , Figure 12 FIG. is an optional structural schematic diagram of a resource processing device provided by an embodiment of the present disclosure. The resource processing device 1200 includes:
[0196] An acquisition module 1201, configured to acquire a multimedia resource, perform domain classification on the multimedia resource to obtain the resource domain of the multimedia resource, and retrieve multiple candidate cover images that conform to the resource domain from the multimedia resource;
[0197] A detection module 1202, configured to respectively perform service feature detection on each candidate cover image to obtain service feature information corresponding to each candidate cover image;
[0198] A scoring module 1203, configured to construct a prompt instruction for prompting to score the candidate cover images, input the prompt instruction, multiple candidate cover images, and corresponding service feature information into a multi-modal large language model for content generation, and obtain predicted scores of the multiple candidate cover images;
[0199] A cover determination module 1204, configured to determine a target cover image corresponding to the multimedia resource from multiple candidate cover images based on the predicted scores.
[0200] Further, the multimedia resource includes multiple original images. The acquisition module 1201 is specifically configured to:
[0201] Construct a query text that conforms to the resource domain;
[0202] Respectively determine the first similarity between each original image and the query text;
[0203] For each original image, when the first similarity is greater than or equal to a preset first threshold, determine the original image as a candidate cover image that conforms to the resource domain.
[0204] Further, the multimedia resource includes meta-information text. The acquisition module 1201 is specifically configured to:
[0205] Perform object detection on each original image to obtain the key regions corresponding to each original image;
[0206] Extract keywords from the meta-information text to obtain the first keywords;
[0207] Determine the second similarity between each key region and the first keywords respectively;
[0208] For each original image, when the first similarity is greater than or equal to a preset first threshold and the second similarity is greater than or equal to a preset second threshold, determine the original image as a candidate cover image that conforms to the resource field.
[0209] Furthermore, the obtaining module 1201 is specifically configured to:
[0210] Extract multiple second keywords from the meta-information text;
[0211] Retrieve the first keywords that conform to the resource field from the multiple second keywords.
[0212] Furthermore, the multimedia resource includes meta-information text, and the resource processing device further includes a filtering module (not shown in the figure), and the filtering module is specifically configured to:
[0213] Determine the third similarity between each candidate cover image and the meta-information text respectively;
[0214] Filter the multiple candidate cover images based on the third similarity.
[0215] Furthermore, the filtering module is specifically configured to:
[0216] Sort each candidate cover image based on the order from large to small of the third similarity;
[0217] For the first n - 1 candidate cover images after sorting, determine the adjacent ratio corresponding to the i-th candidate cover image according to the ratio between the third similarity corresponding to the i-th candidate cover image and the third similarity corresponding to the i + 1-th candidate cover image, where n is the number of candidate cover images, i ≤ n, and i is a positive integer;
[0218] When the adjacent ratio corresponding to the k-th candidate cover image is greater than a preset third threshold, retain the first k candidate cover images after sorting and eliminate the remaining candidate cover images, where k ≤ n and k is a positive integer.
[0219] Furthermore, the multimedia resource includes meta-information text and multiple original images, and one of the original images is configured as the initial cover image, and the obtaining module 1201 is specifically configured to:
[0220] Encode the initial cover image based on the multimodal encoding model to obtain the first image feature;
[0221] Encode the meta-information text based on the multimodal encoding model to obtain the first text feature;
[0222] Concatenate the first image feature and the first text feature and input them into the vertical domain classification model for classification to obtain the resource domain of the multimedia resource.
[0223] Furthermore, the multimedia resource includes the original audio, and the above-mentioned acquisition module 1201 is specifically used for:
[0224] Concatenate the first image feature and the first text feature and input them into the vertical domain classification model for classification;
[0225] When the confidence level of the classification result is less than or equal to the preset fourth threshold, perform speech recognition on the original audio to obtain the recognized text;
[0226] Extract keywords from the recognized text to obtain the third keyword;
[0227] Input the third keyword into the first large language model for domain prediction to determine the resource domain of the multimedia resource.
[0228] Furthermore, the above-mentioned detection module 1202 is specifically used for:
[0229] Input each candidate cover image into the attractiveness prediction model for prediction to obtain the attractiveness score corresponding to each candidate cover image;
[0230] Input each candidate cover image into the clarity prediction model for prediction to obtain the clarity score corresponding to each candidate cover image;
[0231] Based on the attractiveness score and the clarity score, determine the business feature information corresponding to each candidate cover image.
[0232] Furthermore, the above-mentioned detection module 1202 is specifically used for:
[0233] Input each candidate cover image into the emotion prediction model for prediction to obtain the predicted emotion corresponding to each candidate cover image;
[0234] Determine the expected emotion of the resource domain, and respectively determine the fourth similarity between each predicted emotion and the expected emotion;
[0235] Based on the attractiveness score, the clarity score, and the fourth similarity, determine the business feature information corresponding to each candidate cover image.
[0236] Further, both the attraction prediction model and the clarity prediction model match the audience of the multimedia resource. The above resource processing device further includes an adjustment module (not shown in the figure), and the adjustment module is specifically configured to:
[0237] Construct an object description text for describing the audience;
[0238] Add the object description text to the prompt instruction.
[0239] The above resource processing device 1200 and the resource processing method are based on the same inventive concept. By obtaining a multimedia resource, then performing domain classification on the multimedia resource to obtain the resource domain of the multimedia resource, then retrieving multiple candidate cover images that meet the resource domain from the multimedia resource, and then respectively performing business feature detection on each candidate cover image to obtain the business feature information corresponding to each candidate cover image. The business feature information provides context closely related to the business requirements. Then, a prompt instruction is constructed, and the prompt instruction, multiple candidate cover images, and the corresponding business feature information are input into a multi-modal large language model for content generation to obtain the predicted scores of the multiple candidate cover images. Then, based on the predicted scores, high-quality target cover images are selected as the cover images of the multimedia resource, realizing first extracting highly relevant candidate cover images from the multimedia resource through domain-aware retrieval, and then using the multi-modal large language model to dynamically analyze the retrieved candidate cover images, which not only reduces the analysis scope of the multi-modal large language model to improve the reasoning efficiency, but also avoids the interference of irrelevant information or low-quality information to improve the accuracy of the predicted scores. Moreover, under the guidance of the prompt instruction, the multi-modal large language model can utilize rich knowledge reserves and external knowledge contained in the business feature information to infer accurate predicted scores for the candidate cover images and determine the target cover images, ensuring that the target cover images have high visual and semantic quality and meet the business requirements. On this basis, when processing multimedia resources in different fields, it can be directly applied without retraining, that is, directly using domain-aware retrieval and dynamically analyzing using the multi-modal large language model, making this method have strong transferability, capable of ensuring high quality of the target cover images of multimedia resources while effectively reducing the workload, thereby improving the determination efficiency of the cover images.
[0240] The electronic device provided in the embodiments of the present disclosure for executing the above resource processing method may be a terminal. Refer to Figure 13 , Figure 13It is a partial structural block diagram of the terminal provided by the embodiments of the present disclosure. The terminal includes components such as a camera assembly 1310, a first memory 1320, an input unit 1330, a display unit 1340, a sensor 1350, an audio circuit 1360, a wireless fidelity (WiFi) module 1370, a first processor 1380, and a first power supply 1390. Those skilled in the art can understand that Figure 13 the terminal structure shown in
[0241] does not limit the terminal, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. The camera assembly 1310 can be used to collect images or videos. Optionally, the camera assembly 1310 includes a front camera and a rear camera. Generally, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to realize functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions.
[0242] The first memory 1320 can be used to store software programs and modules. The first processor 1380 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the first memory 1320.
[0243] The input unit 1330 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 1330 may include a touch panel 1331 and other input devices 1332.
[0244] The display unit 1340 can be used to display the input information or provided information and various menus of the terminal. The display unit 1340 may include a display panel 1341.
[0245] The audio circuit 1360, the speaker 1361, and the microphone 1362 can provide an audio interface.
[0246] The first power supply 1390 can be alternating current, direct current, a primary battery, or a rechargeable battery.
[0247] The number of sensors 1350 can be one or more. The one or more sensors 1350 include, but are not limited to: an acceleration sensor, a gyroscope sensor, a pressure sensor, an optical sensor, etc. Among them:
[0248] The acceleration sensor can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor can be used to detect the components of the gravitational acceleration on the three coordinate axes. The first processor 1380 can control the display unit 1340 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor. The acceleration sensor can also be used for games or the collection of the user's motion data.
[0249] The gyroscope sensor can detect the body direction and rotation angle of the terminal. The gyroscope sensor can cooperate with the acceleration sensor to collect the user's 3D actions on the terminal. Based on the data collected by the gyroscope sensor, the first processor 1380 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0250] The pressure sensor can be disposed on the side frame of the terminal and / or the lower layer of the display unit 1340. When the pressure sensor is disposed on the side frame of the terminal, it can detect the user's holding signal, and the first processor 1380 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor. When the pressure sensor is disposed on the lower layer of the display unit 1340, the first processor 1380 can control the operable controls on the UI interface according to the user's pressure operation on the display unit 1340. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0251] The optical sensor is used to collect the ambient light intensity. In one embodiment, the first processor 1380 can control the display brightness of the display unit 1340 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 1340 is increased; when the ambient light intensity is low, the display brightness of the display unit 1340 is decreased. In another embodiment, the first processor 1380 can also dynamically adjust the shooting parameters of the camera assembly 1310 according to the ambient light intensity collected by the optical sensor.
[0252] In this embodiment, the first processor 1380 included in the terminal can execute the resource processing method of the previous embodiment.
[0253] The electronic device provided by the embodiments of the present disclosure for executing the above resource processing method can also be a server. Refer to Figure 14 , Figure 14This is a partial structural block diagram of the server provided by the embodiments of the present disclosure. The server may vary greatly due to configuration or performance differences, and may include one or more second processors 1410 and a second memory 1430, and one or more storage media 1440 (such as one or more mass storage devices) for storing application programs 1443 or data 1442. Among them, the second memory 1430 and the storage media 1440 may be transient storage or persistent storage. The program stored in the storage media 1440 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the second processor 1410 may be configured to communicate with the storage media 1440 and execute a series of instruction operations in the storage media 1440 on the server.
[0254] The server may further include one or more second power supplies 1420, one or more wired or wireless network interfaces 1450, one or more input / output interfaces 1460, and / or one or more operating systems 1441, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0255] The second processor 1410 in the server may be used to execute the resource processing method.
[0256] The embodiments of the present disclosure further provide a computer-readable storage medium for storing a computer program, and the computer program is used to execute the resource processing methods of the foregoing various embodiments.
[0257] The embodiments of the present disclosure further provide a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the resource processing method described above.
[0258] In the description of the present disclosure and the above-mentioned accompanying drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0259] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0260] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", and "exceeding" do not include the present number, and understandings such as "above", "below", and "within" include the present number.
[0261] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0262] The unit described as a separation component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0263] In addition, each functional unit in various embodiments of the present disclosure may be integrated in a processing unit, may exist physically alone for each unit, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0264] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0265] It should also be understood that the various embodiments provided in the present disclosure can be combined arbitrarily to achieve different technical effects.
[0266] The above is a specific description of the preferred embodiments of the present disclosure, but the present disclosure is not limited to the above-mentioned embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure. These equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.
Claims
1. A resource processing method, characterized in that: include: Acquire multimedia resources, classify the multimedia resources by fields, obtain resource fields of the multimedia resources, and retrieve multiple candidate cover images that meet the resource fields from the multimedia resources; Performing business feature detection on each of the candidate cover images respectively to obtain business feature information corresponding to each of the candidate cover images; Constructing a prompt instruction for prompting to score the candidate cover images, inputting the prompt instruction, the plurality of candidate cover images, and the corresponding business feature information into a multimodal large language model for content generation, and obtaining predicted scores of the plurality of candidate cover images; Based on the prediction score, a target cover image corresponding to the multimedia resource is determined from the plurality of candidate cover images.
2. The resource processing method according to claim 1, characterized in that: The multimedia resource includes a plurality of original images, and the step of retrieving a plurality of candidate cover images that meet the resource domain from the multimedia resource includes: Constructing query text that matches the domain of the resource; respectively determining a first similarity between each of the original images and the query text; For each of the original images, when the first similarity is greater than or equal to a preset first threshold, the original image is determined as a candidate cover image that conforms to the resource field.
3. The resource processing method according to claim 2, characterized in that: The multimedia resource includes meta information text, and for each of the original images, when the first similarity is greater than or equal to a preset first threshold, determining the original image as a candidate cover image that meets the resource domain includes: Performing target detection on each of the original images respectively to obtain a key area corresponding to each of the original images; Extracting keywords from the meta information text to obtain a first keyword; respectively determining a second similarity between each of the key areas and the first keyword; For each of the original images, when the first similarity is greater than or equal to a preset first threshold and the second similarity is greater than or equal to a preset second threshold, the original image is determined as a candidate cover image that meets the resource field.
4. The resource processing method according to claim 3, characterized in that: The step of extracting keywords from the meta information text to obtain a first keyword includes: Extracting keywords from the meta information text to obtain a plurality of second keywords; A first keyword matching the resource field is retrieved from the plurality of second keywords.
5. The resource processing method according to claim 1, characterized in that: The multimedia resource includes meta information text. After retrieving a plurality of candidate cover images that meet the resource domain from the multimedia resource, the resource processing method further includes: respectively determining a third similarity between each of the candidate cover images and the meta information text; Based on the third similarity, the plurality of candidate cover images are filtered.
6. The resource processing method according to claim 5, characterized in that: The filtering of the plurality of candidate cover images based on the third similarity comprises: Sorting the candidate cover images based on the third similarity from large to small; For the first n-1 candidate cover images after sorting, determine the adjacent ratio corresponding to the i-th candidate cover image according to the ratio between the third similarity corresponding to the i-th candidate cover image and the third similarity corresponding to the i+1-th candidate cover image, where n is the number of the candidate cover images, i≤n, and i is a positive integer; When the adjacent ratio corresponding to the kth candidate cover image is greater than a preset third threshold, the first k sorted candidate cover images are retained and the remaining candidate cover images are eliminated, where k≤n and k is a positive integer.
7. The resource processing method according to claim 1, characterized in that: The multimedia resource includes meta information text and a plurality of original images, one of the original images being configured as an initial cover image, and the field classification of the multimedia resource to obtain the resource field of the multimedia resource includes: Encoding the initial cover image based on a multimodal coding model to obtain a first image feature; Encoding the meta information text based on the multimodal encoding model to obtain a first text feature; The first image feature and the first text feature are spliced and input into a vertical domain classification model for classification to obtain the resource domain of the multimedia resource.
8. The resource processing method according to claim 7, characterized in that: The multimedia resource includes original audio, and the first image feature and the first text feature are spliced and input into the vertical domain classification model for classification to obtain the resource domain of the multimedia resource, including: splicing the first image feature and the first text feature and inputting them into a vertical domain classification model for classification; When the confidence level of the classification result is less than or equal to a preset fourth threshold, performing speech recognition on the original audio to obtain a recognized text; Extracting keywords from the recognized text to obtain a third keyword; The third keyword is input into the first language model for domain prediction to determine the resource domain of the multimedia resource.
9. The resource processing method according to claim 1, characterized in that: The performing business feature detection on each of the candidate cover images to obtain business feature information corresponding to each of the candidate cover images includes: Inputting each of the candidate cover images into an attractiveness prediction model for prediction, and obtaining an attractiveness score corresponding to each of the candidate cover images; Inputting each of the candidate cover images into a clarity prediction model for prediction, and obtaining a clarity score corresponding to each of the candidate cover images; Based on the attractiveness score and the clarity score, business feature information corresponding to each of the candidate cover images is determined.
10. The resource processing method according to claim 9, characterized in that: The determining, based on the attractiveness score and the clarity score, the business feature information corresponding to each of the candidate cover images includes: Inputting each of the candidate cover images into a sentiment prediction model for prediction, respectively, to obtain a predicted sentiment corresponding to each of the candidate cover images; Determining the expected emotion of the resource domain, and respectively determining a fourth similarity between each of the predicted emotions and the expected emotion; Based on the attractiveness score, the clarity score and the fourth similarity, the business feature information corresponding to each of the candidate cover images is determined.
11. The resource processing method according to claim 9, characterized in that: The attractiveness prediction model and the clarity prediction model are both matched with the audience object of the multimedia resource. After the prompt instruction for prompting to score the candidate cover image is constructed, the resource processing method further includes: Constructing an object description text for describing the audience object; The object description text is added to the prompt instruction.
12. A resource processing device, characterized in that: include: An acquisition module is used to acquire multimedia resources, classify the multimedia resources by fields, obtain resource fields of the multimedia resources, and retrieve multiple candidate cover images that meet the resource fields from the multimedia resources; A detection module, used to perform business feature detection on each of the candidate cover images respectively to obtain business feature information corresponding to each of the candidate cover images; A scoring module is used to construct a prompt instruction for prompting to score the candidate cover images, input the prompt instruction, the plurality of candidate cover images and the corresponding business feature information into a multimodal large language model for content generation, and obtain predicted scores of the plurality of candidate cover images; A cover determination module is used to determine a target cover image corresponding to the multimedia resource from a plurality of candidate cover images based on the predicted score.
13. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the resource processing method described in any one of claims 1 to 11 is implemented.
14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the resource processing method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the resource processing method according to any one of claims 1 to 11 is implemented.