Material salvage method based on modal fusion matching

Through the method based on modal fusion matching, historical materials and popular materials are vectorized and similarity calculations are solved, and the problems of low efficiency and poor accuracy of material retrieval in the existing technology are achieved, and efficient and accurate material retrieval effect is achieved.

CN120047186APending Publication Date: 2025-05-27GUANGZHOU TAIDONG TECH CO LTD

Patent Information

Application Number
CN202510112516.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing material recovery method has the problems of high manual labeling cost, low efficiency and strong subjectivity of the results. The direct matching method is difficult to deal with fuzzy or noisy materials, and the calculation cost is high.

Method used

The material recovery method based on modal fusion matching is adopted. By obtaining historical material and popular material, it performs vector processing, constructs a material vector library, and calculates the similarity through traversal matching, and filters back recovery materials.

Benefits of technology

It improves the efficiency and accuracy of material recovery, reduces labor costs and calculation costs, and ensures consistency and high quality of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047186A_ABST
    Figure CN120047186A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a material salvage method based on modal fusion matching. The method comprises the following steps: acquiring a plurality of historical delivery materials to construct a historical delivery material library; vectorizing each historical delivery material in the historical delivery material library to obtain a corresponding first material vector so as to construct a first material vector library; obtaining a plurality of hot delivery materials to construct a hot delivery material library; vectorizing each hot delivery material in the hot delivery material library to obtain a corresponding second material vector so as to construct a second material vector library; performing traversal matching on the first material vector in the first material vector library and the second material vector in the second material vector library to calculate the similarity between the corresponding historical put materials and the hot put materials; and based on the similarity, screening historical putting materials from the historical putting material library to serve as salvaged putting materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of data processing, and in particular, to a method for retrieving materials based on modal fusion matching. Background Art

[0002] In the current digital marketing and content creation fields, the technology of retrieving materials is crucial for improving work efficiency and resource utilization. The effective management and accurate retrieval of placed materials can help enterprises quickly locate and use materials that meet their needs, reduce costs, and improve the advertising placement effect and content creation quality.

[0003] Currently, the methods for retrieving materials are mainly divided into two types: manual annotation and direct matching.

[0004] The manual annotation method is to manually mark and classify materials for subsequent retrieval. However, this method has many drawbacks. On the one hand, manual annotation requires a large amount of labor costs. Enterprises need to hire professional personnel to spend a lot of time annotating materials one by one, and the efficiency is extremely low. On the other hand, due to the dependence on personal subjective judgment in the annotation process, different annotators may have differences in the understanding and annotation of the same material, resulting in strong subjectivity of the annotation results and difficulty in ensuring consistency. This makes it impossible to accurately and reliably obtain the required materials during material retrieval, seriously affecting the utilization efficiency and quality of materials.

[0005] The direct matching method attempts to achieve material retrieval by directly matching materials. However, this method also has a series of problems. First, when the materials are unclear or noisy, it is difficult for this method to accurately identify them, resulting in the inability to effectively extract the features of the effective visual elements therein, thus affecting the accuracy of matching. Second, when processing materials, it cannot accurately identify the key elements in them, and a large number of redundant features will be mixed in the matching calculation process, not only resulting in low matching efficiency, but also difficult to guarantee the matching quality. Moreover, the direct matching method has limited semantic understanding ability for materials and cannot accurately identify semantically similar materials, making it difficult to effectively retrieve some materials that are semantically similar but slightly different in visual performance. In addition, direct matching requires complex calculations on a large amount of data, with a huge amount of calculations and high costs. Summary of the Invention

[0006] In view of this, embodiments of the present invention provide a method for retrieving materials based on modal fusion matching to at least partially solve the above problems.

[0007] According to the first aspect of the embodiments of the present invention, a method for retrieving materials based on modal fusion matching is provided, which includes:

[0008] Obtaining a plurality of historical placed materials to construct a historical placed material library;

[0009] Vectorize each historical placement material in the historical placement material library to obtain a corresponding first material vector to construct a first material vector library;

[0010] Obtain multiple popular placement materials to construct a popular placement material library;

[0011] Vectorize each popular placement material in the popular placement material library to obtain a corresponding second material vector to construct a second material vector library;

[0012] Traverse and match the first material vectors in the first material vector library and the second material vectors in the second material vector library to calculate the similarity between the corresponding historical placement materials and popular placement materials;

[0013] Based on the similarity, screen historical placement materials from the historical placement material library as the placement materials to be retrieved.

[0014] Optionally, the vectorizing each historical placement material in the historical placement material library to obtain a corresponding first material vector to construct a first material vector library includes:

[0015] Determine the evaluation labels of each historical placement material in the historical placement material library;

[0016] Vectorize the evaluation labels of each historical placement material to obtain a corresponding first material vector to construct a first material vector library.

[0017] Optionally, before determining the evaluation labels of each historical placement material in the historical placement material library, it includes:

[0018] Separate the content of each historical placement material in the historical placement material library to obtain the corresponding main-modal original content and sub-modal original content;

[0019] Separate the features of the main-modal original content corresponding to each historical placement material to obtain the key elements of the main-modal content;

[0020] Determine the evaluation labels of the key elements of the main-modal content and use them as the evaluation labels of the corresponding historical placement materials.

[0021] Optionally, the vectorizing each historical placement material in the historical placement material library to obtain a corresponding first material vector to construct a first material vector library includes:

[0022] Vectorize the evaluation labels of the key elements of the main-modal content corresponding to each historical placement material in the historical placement material library to generate the main-modal content evaluation vector of the corresponding historical placement material;

[0023] Vectorize the sub-modal original content corresponding to each historical placement material in the historical placement material library to generate a sub-modal content evaluation vector for the corresponding historical placement material;

[0024] Fuse the main-modal content evaluation vector and the sub-modal content evaluation vector of each historical placement material to generate the first material vector to construct the first material vector library.

[0025] Optionally, the vectorizing each popular placement material in the popular placement material library to obtain a corresponding second material vector to construct the second material vector library includes:

[0026] Determine the evaluation tags of each popular placement material in the popular placement material library;

[0027] Vectorize the evaluation tags of each popular placement material to generate a corresponding second material vector to construct the second material vector library.

[0028] Optionally, before determining the evaluation tags of each popular placement material in the popular placement material library, it includes:

[0029] Separate the content of each popular placement material in the popular placement material library to obtain the corresponding main-modal original content and sub-modal original content;

[0030] Separate the features of the main-modal original content corresponding to each popular placement material to obtain the corresponding key elements of the main-modal content;

[0031] Determine the evaluation tags of the key elements of the main-modal content corresponding to each popular placement material and use them as the evaluation tags of the corresponding popular placement material.

[0032] Optionally, the vectorizing each popular placement material in the popular placement material library to obtain a corresponding second material vector to construct the second material vector library includes:

[0033] Vectorize the evaluation tags of the key elements of the main-modal content corresponding to each popular placement material in the popular placement material library to generate a main-modal content evaluation vector for the corresponding popular placement material;

[0034] Vectorize the sub-modal original content corresponding to each popular placement material in the popular placement material library to generate a sub-modal content evaluation vector for the corresponding popular placement material;

[0035] Fuse the main-modal content evaluation vector and the sub-modal content evaluation vector of each popular placement material to generate the second material vector to construct the second material vector library.

[0036] Optionally, the evaluation tags of the historical placement materials or the evaluation tags of the popular placement materials include at least one of the following: evaluation index tags, content semantic tags.

[0037] Optionally, the evaluation index tags include at least one of a rhythm tag, a product brand tag introduced in a video advertisement, and a product brand type tag, and the content semantic tags include an advertisement category tag, and the advertisement category tag includes at least one of a text advertisement tag, an animation advertisement tag, and a real machine advertisement tag.

[0038] Optionally, the historical placement materials and the popular placement materials are visual and audio materials.

[0039] Optionally, traversing and matching the first material vectors in the first material vector library and the second material vectors in the second material vector library to calculate the similarity between the corresponding historical placement materials and the popular placement materials includes:

[0040] Invoking a matching model trained to completion based on set reinforcement learning constraints, traversing the first material vector library and the second material vector library to calculate the similarity between each of the first material vectors and each of the second material vectors, and using this as the similarity between the corresponding historical placement materials and the popular placement materials.

[0041] Optionally, the method further includes:

[0042] Determining placement material samples and matching result tags between the samples;

[0043] Separating the content of each of the placement material samples to obtain the corresponding primary modality sample original content and secondary modality sample original content;

[0044] Training the matching model to be trained for several rounds based on the following steps until a trained matching model is obtained:

[0045] Vectorizing the secondary modality sample original content participating in the current round of training to obtain secondary modality sample vectors;

[0046] The matching model to be trained performs forward propagation on the primary modality sample original content to generate a predicted matching result between the placement material samples;

[0047] Invoking the objective function that performs reinforcement learning constraints on the predicted matching result based on the secondary modality sample vectors, and calculating the vector similarity between the secondary modality sample vectors in the current round of training and the secondary modality sample vectors in the previous round of training;

[0048] Adjust the network parameters of the to-be-trained matching model in the gradient direction that increases the similarity until a predetermined condition for completing model training is reached. The predetermined condition is at least that the similarity reaches a set similarity threshold and the loss of the predicted matching result relative to the label of the sample matching result is less than a set loss value threshold.

[0049] In the solution of the embodiment of the present invention, the material retrieval method based on modal fusion matching has the following technical advantages:

[0050] 1. Compared with the manual annotation method, which requires a large amount of human input and time cost and has low efficiency. The present invention uses an automated vectorization and matching process, eliminating the need for a large amount of manual participation in the one-by-one labeling and classification of materials. The solution of this application can quickly vectorize multiple historical delivery materials and popular delivery materials, construct a vector library, and calculate the similarity through traversal matching, and then screen out the retrieved materials, greatly saving labor costs and significantly improving the efficiency of material retrieval.

[0051] 2. Compared with manual annotation, due to the influence of subjective judgment, the annotation results are difficult to ensure consistency, affecting the accuracy and reliability of material retrieval. The present invention is based on objective vectorization and similarity calculation methods, and processes materials according to a unified mathematical model and algorithm. Whether it is historical delivery materials or popular delivery materials, they are vectorized and matched according to the same rules, avoiding the interference of human factors, making the material retrieval results more accurate and reliable, and effectively ensuring the consistency of the results.

[0052] 3. Compared with the direct matching method, when facing unclear or noisy materials, it is difficult to accurately identify and extract effective visual element features. The present invention uses a vectorization processing method, which can perform more in-depth feature extraction and analysis of materials. Through the modal fusion technology, different modal information (such as vision, audio, etc.) of the materials is integrated. Even if the materials are fuzzy or noisy, key features of the materials can be obtained from multiple dimensions, thereby improving the recognition ability of such complex materials and enhancing the accuracy of matching.

[0053] 4. Compared with the direct matching method, it is impossible to accurately identify the key elements in the materials, resulting in a large amount of redundant features in the matching calculation, affecting efficiency and quality. In the vectorization process of the present invention, attention is paid to the feature extraction of the key elements of the materials. Through detailed feature analysis of historical delivery materials and popular delivery materials, the core content of the materials can be accurately grasped. In the matching process, calculations are based on these accurately extracted key features, avoiding the interference of redundant features, thereby improving the matching efficiency and quality.

[0054] 5. Compared with the direct matching method, the ability to understand the semantics of materials is limited, and it is difficult to identify materials with similar semantics. Through the vectorization process of materials, the present invention can convert the semantic information of materials into vector form, and measure the similarity degree of materials at the semantic level by calculating the similarity between vectors. This method can capture the semantic associations between materials more accurately, effectively identify materials with similar semantics but slightly different visual performances, broaden the scope of material retrieval, and improve the utilization rate of materials.

[0055] 6. Compared with direct matching, which requires complex calculations on a large amount of data and has high costs, the present invention is based on the method of vectorization and modality fusion, and effectively extracts features and simplifies the representation of materials during the data processing. By constructing a vector library and performing vector matching, the material retrieval operation can be completed with relatively low computational complexity, greatly reducing the computational cost, and improving the operating efficiency and economy of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0057] Figure 1 It is a schematic flowchart of a method for retrieving materials based on modality fusion matching provided by an embodiment of the present invention.

[0058] Figure 2 It is a schematic diagram of a device for retrieving materials based on modality fusion matching provided by an embodiment of the present invention.

[0059] Figure 3 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following will clearly and detailedly describe the technical solutions in the embodiments of the present invention in combination with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present invention.

[0061] It should be understood that the terms "first", "second", "third", etc. in the claims, specification and drawings of this disclosure are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" used in the specification and claims of this disclosure indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0062] It should also be understood that the terms used in the specification of this disclosure herein are only for the purpose of describing specific embodiments and are not intended to limit this disclosure. As used in the specification and claims of this disclosure, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in the specification and claims of this disclosure refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0063] Figure 1 Schematic flow diagram of a material retrieval method based on modal fusion matching provided for an embodiment of the present invention. As Figure 1 shown, it includes:

[0064] Obtain a plurality of historical delivered materials to construct a historical delivered material library;

[0065] Vectorize each historical delivered material in the historical delivered material library to obtain a corresponding first material vector to construct a first material vector library;

[0066] Obtain a plurality of popular delivered materials to construct a popular delivered material library;

[0067] Vectorize each popular delivered material in the popular delivered material library to obtain a corresponding second material vector to construct a second material vector library;

[0068] Traverse and match the first material vectors in the first material vector library and the second material vectors in the second material vector library to calculate the similarity between the corresponding historical delivered materials and popular delivered materials;

[0069] Based on the similarity, screen historical delivered materials from the historical delivered material library as the retrieved delivered materials.

[0070] Optionally, the vectorizing each historical delivered material in the historical delivered material library to obtain a corresponding first material vector to construct a first material vector library includes:

[0071] Determine the evaluation labels of each historical delivered material in the historical delivered material library;

[0072] Vectorize the evaluation tags of each of the historical delivery materials to obtain corresponding first material vectors, thereby constructing a first material vector library.

[0073] Optionally, before determining the evaluation tags of each historical delivery material in the historical delivery material library, it includes:

[0074] Perform content separation on each historical delivery material in the historical delivery material library to obtain corresponding primary modality original content and secondary modality original content;

[0075] Perform feature separation on the primary modality original content corresponding to each historical delivery material to obtain key elements of the primary modality content;

[0076] Determine the evaluation tags of the key elements of the primary modality content and use them as the evaluation tags of the corresponding historical delivery materials.

[0077] Specifically, the technical implementation details of each step are elaborated in detail below.

[0078] 1. Perform content separation on each historical delivery material in the historical delivery material library

[0079] Let the historical delivery material library be where S i represents the i-th historical delivery material, and n is the total number of materials in the material library.

[0080] In an actual application scenario, such as advertising delivery, these materials may contain multiple modality information, such as vision (image, video), audition (audio), etc. To perform content separation, a multi-modal decomposition network based on deep learning is adopted

[0081] For each material S i , process it through the network to obtain the primary modality original content and the secondary modality original content which can be expressed as:

[0082] where θ sep is the parameter set of the network and these parameters are optimized through training on a large-scale multi-modal material data set. In the advertising delivery scenario, the primary modality original content may be information such as the core product display in the advertisement, the key advertising slogan, etc.; the secondary modality original content may be secondary information such as background music, auxiliary picture elements, product introduction music, etc.

[0083] 2. Feature separation is performed on the original content of the main modality corresponding to each historical delivery material

[0084] For the original content of the main modality Use a deep convolutional attention network (DCAN) to perform feature separation to extract the features of the key elements of the main modality content.

[0085] DCAN combines the powerful feature extraction ability of the convolutional neural network (CNN) and the attention mechanism, and can focus more on the key features. The network input is The output is the feature vector F i .

[0086]

[0087] where θ feat is the parameter set of the network The convolutional layer inside the network performs a convolutional operation on the input through the convolutional kernel K j (j = 1, 2, …, J, J is the number of convolutional layers) to extract features at different levels: X j = σ(K j *X j-1 + b j )

[0088] where X j represents the output feature map of the j-th layer, σ is the activation function (such as ReLU), b j is the bias vector of the j-th layer, and * represents the convolutional operation.

[0089] The attention mechanism calculates the attention weight A i for the feature map to highlight the key features:

[0090] A i = Softmax(W a X L + b a )

[0091]

[0092] where W a and b a are the parameters of the attention mechanism, X L is the output of the last convolutional layer, and are the attention weight and the value of the feature map at the k-th position respectively.

[0093] In the scenario of advertising materials, these feature vectors F iIt can accurately represent the key elements in the main modal content, such as the unique appearance and functional features of the product, etc.

[0094] 3. Determine the evaluation labels for the key elements of the main modal content

[0095] To determine the evaluation labels, a multi-classification model based on Support Vector Machine (SVM) is adopted

[0096] Given the feature vector F i , the evaluation label y i is calculated by the following formula:

[0097]

[0098] where w k and b k are the weight vector and bias of the k-th class respectively, k = 1, 2, …, K, and K is the number of categories of the evaluation labels.

[0099] In the training stage, the parameters of the SVM model are optimized by maximizing the classification margin, and the objective function is:

[0100]

[0101]

[0102] where C is the penalty parameter, used to balance the classification margin and the penalty for misclassified samples, ξ ik is the slack variable, y ik is the true label of sample i belonging to the k-th class (1 if it belongs, otherwise -1). In the determination of the evaluation labels for advertising materials, these labels can represent information such as the target audience, product type, and advertising style of the materials.

[0103] 4. Vectorize the evaluation labels for each historical delivery material

[0104] Adopt a pre-trained language model based on Transformer to vectorize the evaluation labels. Let the evaluation label y i be a text sequence, which is input into the Transformer model after word segmentation and word embedding.

[0105]

[0106] where θ vec is the set of parameters of the model . The self-attention mechanism in the Transformer model encodes the input sequence by calculating the attention weights between different positions:

[0107]

[0108] where Q, R, and K are the query, key, and value matrices respectively, and d k is the dimension of the key matrix.

[0109] Through the processing of multi-layer self-attention and feed-forward neural networks, the vector representation V of the evaluation label is finally obtained i , and these vectors constitute the first material vector library.

[0110] 5. Construction of the first material vector library

[0111] The vectors V 1 , V 2 , …, V n obtained by vectorizing the evaluation labels of all historical delivery materials are collected to form the first material vector library

[0112] Therefore, the above technical process for constructing the first material vector library has the following technical advantages:

[0113] Regarding the content separation formula in 1

[0114] This formula represents the use of a deep learning-based multi-modal decomposition network to decompose the historical delivery material S i into the main-modal original content and the secondary-modal original content The network parameter θ sep is obtained by training on a large-scale multi-modal material dataset to achieve effective separation of different modal contents.

[0115] Therefore, in the advertising delivery scenario, materials often contain various types of information. Through this separation method, the core information (main-modal original content) and auxiliary information (secondary-modal original content) can be accurately distinguished. This helps to focus on the core information for subsequent processing, improve the accuracy and efficiency of feature extraction, and avoid interference from auxiliary information in key feature extraction. For example, in a car advertising material, the main-modal contents such as the display screen of the car and the key advertising slogans can be separated from the secondary-modal contents such as the background music and background decorations, providing cleaner data for subsequent analysis.

[0116] 2. Feature separation formula

[0117] Convolution operation: X j = σ(K j * X j-1 + b j )

[0118] Attention mechanism: A i= Softmax(W a X L + b a )

[0119]

[0120] The convolution operation is performed on the input feature map X through the convolution kernel K j to extract features at different levels. The activation function σ increases the non-linear expression ability of the model. The attention mechanism calculates the attention weight A j-1 to weight the feature map X i and highlight key features, enabling the model to focus on important information. Finally, the feature vector F L of the key elements of the main modal content is obtained. i .

[0121] Therefore, in the processing of advertising materials, the convolution operation can automatically learn local features in advertising images or videos, such as the appearance details and logos of products. The attention mechanism further improves the sensitivity of the model to key features and can accurately capture information related to the core selling points of products in complex advertising content. For example, in a cosmetics advertisement, it can highlight key features such as the product packaging and usage effects while ignoring some irrelevant background elements, providing a more representative feature vector for subsequent determination of evaluation labels.

[0122] 3. Determine the evaluation label formula

[0123]

[0124] Objective function:

[0125]

[0126] The first formula maps the feature vector F i to different evaluation label categories y i through a multi-classification model of support vector machine (SVM). The objective function is to optimize the parameters w k and b k of the SVM model by minimizing the objective function during the training phase, where is used to control the model complexity, and is used to penalize misclassified samples to maximize the classification margin.

[0127] Therefore, in the advertising placement scenario, the model can accurately assign evaluation tags to advertising materials according to the extracted feature vectors, such as target audience, product type, etc. By optimizing the classification margin, the model has strong generalization ability and can stably classify on different advertising material datasets, improving the accuracy and reliability of determining evaluation tags. This helps to perform more precise management and matching of materials subsequently. For example, appropriate advertising materials can be screened out for placement according to different target audience tags.

[0128] 4. Evaluation Tag Vectorization Formula

[0129]

[0130] Self-attention Mechanism:

[0131] Utilize the pre-trained language model based on Transformer Convert the evaluation tag text sequence y i into the vector representation V i . The self-attention mechanism encodes the input sequence by calculating the relationships between the query, key, and value matrices, capturing the semantic information in the text.

[0132] Therefore, in the processing of advertising materials, after vectorizing the evaluation tags, the text information can be converted into a vector form that is easy for computers to process, facilitating subsequent similarity calculation and matching operations. The self-attention mechanism can effectively capture the semantic relationships in the text, making the vector representation more accurately reflect the meaning of the evaluation tags. For example, for the two evaluation tags "Fashion electronic product advertisement suitable for young people" and "Trendy digital product promotion targeting young groups", the semantic similarity between them can be accurately captured through vectorization, providing a more precise basis for material matching and retrieval.

[0133] 5. Construct the First Material Vector Library

[0134] Collect all the V i to form the first material vector library

[0135] Integrate the vector representations of the evaluation tags of each historical placement material together to form a vector library that can be queried and matched.

[0136] Therefore, in advertising placement, this vector library provides basic data for material retrieval and matching. By calculating the similarity of the vectors in the vector library, historical placement materials similar to the current requirements can be quickly found, improving the reuse rate of materials and the placement effect. For example, when making a new electronic product advertisement, advertising materials of previous similar products can be quickly found from the vector library, learning from their successful experiences and saving production costs and time.

[0137] Optionally, vectorize each historical delivery material in the historical delivery material library to obtain a corresponding first material vector to construct a first material vector library, including:

[0138] Vectorize the evaluation tags of the key elements of the main modal content corresponding to each historical delivery material in the historical delivery material library to generate a main modal content evaluation vector for the corresponding historical delivery material;

[0139] Vectorize the secondary modal original content corresponding to each historical delivery material in the historical delivery material library to generate a secondary modal content evaluation vector for the corresponding historical delivery material;

[0140] Fuse the main modal content evaluation vector and the secondary modal content evaluation vector of each historical delivery material to generate the first material vector to construct a first material vector library.

[0141] Specifically, the following details the technical implementation details of each step of constructing the first material vector library from the historical delivery material library.

[0142] 1. Vectorize the evaluation tags of the key elements of the main modal content

[0143] Let the historical delivery material library be For each material S i , the evaluation tag of the key element of its main modal content is denoted as t i .

[0144] To vectorize the evaluation tags, a method based on a variational autoencoder (VAE) is adopted, combined with an improved attention mechanism to capture the semantic information in the tags.

[0145] The variational autoencoder consists of an encoder and a decoder. The encoder E maps the input evaluation tag t i to a latent space z i , and here a multi-layer perceptron (MLP) is used to implement the encoder.

[0146] z i = E(t i ; θ E ) = MLP(t i ; θ E )

[0147] where θ E is the parameter set of the encoder, including the weights and biases of the multi-layer perceptron. In the advertising delivery scenario, the evaluation tag t iIt may contain information such as product type, target audience, advertising style, etc. By learning the characteristics of this information, the encoder transforms it into a vector representation in the latent space.

[0148] To better capture the important information in the evaluation tags, an improved attention mechanism is introduced. Calculate the attention weight α i :

[0149] α i = Softmax(W a z i + b a )

[0150] where W a and b a are the parameters of the attention mechanism. The attention weight α i represents the degree of importance of different parts in the tag.

[0151] Then, the main modal content evaluation vector is obtained through attention weighting

[0152]

[0153] where is the j-th element of z i of.

[0154] 2. Vectorize the secondary modal original content

[0155] For the secondary modal original content of each historical placement material S i adopt a hybrid model based on convolutional neural network (CNN) and recurrent neural network (RNN) for vectorization.

[0156]

[0157]

[0158]

[0159] cnn is the parameter set of the convolutional neural network, including the weights and biases of the convolutional kernels. In advertising placement, the secondary modal original content may contain information such as background music and background images, and the convolutional neural network can extract the local features of this information.

[0159] Then, the feature map is flattened and input into the recurrent neural network to obtain the secondary modal content evaluation vector

[0160]

[0161] where θ rnn is the parameter set of the recurrent neural network. The recurrent neural network can process sequence information and capture the temporal or spatial order features in the raw content of the sub-modalities.

[0162] 3. Vector Fusion

[0163] Fuse the main-modal content evaluation vector and the sub-modal content evaluation vector to generate the first material vector v i . Adopt a fusion method based on a gating mechanism, which can adaptively control the fusion ratio of the main-modal and sub-modal information.

[0164] Define the gating vector g i :

[0165]

[0166] where W g and b g are the parameters of the gating mechanism, denotes concatenating the main-modal and sub-modal content evaluation vectors, and σ is the Sigmoid activation function.

[0167] The first material vector v i is:

[0168]

[0169] where ⊙ represents element-wise multiplication. Through the gating mechanism, the fusion ratio is dynamically adjusted according to the importance of the main-modal and sub-modal information, so that the generated first material vector can better reflect the overall characteristics of the material.

[0170] 4. Construct the First Material Vector Library

[0171] Collect all the first material vectors v i obtained from the historical delivered materials S i through the above steps to form the first material vector library

[0172] Therefore, the technical benefits of the specific algorithm in the above process of constructing the first material vector library are as follows:

[0173] 1. Vectorization of the evaluation labels of the key elements in the main-modal content

[0174] Use the variational autoencoder (VAE) to vectorize the evaluation label t iMapped to the latent space z i , the formula z i = E(t i ; θ E ) = MLP(t i ; θ E ), where the multi - layer perceptron (MLP) serves as the encoder E, learning the feature representation of the evaluation label through the parameter θ E , converting the high - dimensional label data into a low - dimensional latent vector for subsequent processing.

[0175] An improved attention mechanism is introduced. The attention weight α i = Softmax(W a z i + b a ) is calculated, which reflects the importance degree of each part in the latent vector z i . Then, through i the main - modality content evaluation vector is obtained such that the vector can highlight the key information in the evaluation label. For this reason, in the advertising placement scenario, the evaluation label contains rich semantic information, such as product type, target audience, etc. This vectorization method can convert this semantic information into a vector form that can be processed by a computer, and at the same time, the attention mechanism can capture the importance of different parts in the label, making the generated main - modality content evaluation vector more representative. For example, for an evaluation label of "a fashion beauty product advertisement targeting young women", this method can accurately incorporate key information related to "young women" and "fashion beauty" into the vector, providing a precise feature representation for subsequent material matching and analysis.

[0176]

[0177] 2. Vectorization of the secondary - modality original content

[0178] First, a convolutional neural network (CNN) is used to extract features from the secondary - modality original content . In the formula , θ cnn is the parameter of the CNN. Through convolutional operations, local features in the secondary - modality original content are extracted to obtain the feature map

[0179] Then, after flattening the feature map, it is input into a recurrent neural network (RNN). Using the secondary - modality content evaluation vector is obtained The RNN can process sequential information and capture temporal or spatial order features in the secondary - modality original content, thus representing the secondary - modality information more comprehensively.

[0180] ​Therefore, in advertising placement, although secondary modal original content such as background music and background images is relatively less important, it still contains information that affects the advertising effect. CNN can extract local features of this information, such as the rhythm features of background music, and the color and texture features of background images. RNN can further process these features, capture their sequential relationships in time or space, so that the generated evaluation vector of secondary modal content can fully reflect the characteristics of the secondary modal original content. This helps to fully consider the impact of secondary modal information on the overall features of the material in subsequent vector fusion, and improve the comprehensiveness of material representation.

[0181] 3. Vector Fusion

[0182] Generate the gating vector g through the gating mechanism i , in the formula , after concatenating the evaluation vectors of the primary modal and secondary modal content, calculate the gating vector through the parameters W g and b g and the Sigmoid activation function σ. Its value ranges from 0 to 1 and is used to control the fusion ratio of the primary modal and secondary modal information.

[0183] Finally, through generate the first material vector v i , adaptively adjust the fusion degree of the primary and secondary modal information according to the gating vector, so that the generated vector can better reflect the overall characteristics of the material.

[0184] Therefore, in the advertising placement scenario, the importance of the primary modal and secondary modal information of different advertising materials may vary. Through the vector fusion method of the gating mechanism, the proportion of the primary and secondary modal information in the final vector can be dynamically adjusted according to the characteristics of the specific material. For example, for some advertisements that emphasize the product itself, the primary modal information may be more important, and the gating vector will make the evaluation vector of the primary modal content account for a larger proportion in the final fusion vector; while for some advertisements that focus on creating an atmosphere, the secondary modal information may be more critical, and the gating vector will adjust the fusion ratio accordingly. The first material vector generated in this way can more accurately reflect the core characteristics of the material, improve the accuracy and efficiency of material matching, and provide more powerful support for the formulation of advertising placement strategies.

[0185] 4. Construct the First Material Vector Library

[0186] Collect all the first material vectors v i obtained from the above steps for all historical placement materials to form the first material vector library to provide a data basis for subsequent material matching and analysis.

[0187] Therefore, in advertising placement, it is crucial to have a comprehensive and accurate material vector library. The vectors in this vector library can accurately represent the characteristics of each historical placement material. By calculating the similarity of the vectors in the vector library, historical materials similar to the current requirements can be quickly found. For example, when planning a new advertising placement campaign, historical materials similar in terms of target audience, product type, advertising style, etc. can be retrieved from the vector library, and their successful experiences can be referred to optimize the production and placement strategies of the new advertisement, thereby saving time and costs and improving the effectiveness and return on investment of advertising placement.

[0188] Optionally, the process of vectorizing each popular placement material in the popular placement material library to obtain corresponding second material vectors for constructing a second material vector library includes:

[0189] Determine the evaluation tags for each popular placement material in the popular placement material library;

[0190] Vectorize the evaluation tags for each popular placement material to generate corresponding second material vectors for constructing a second material vector library.

[0191] Optionally, before determining the evaluation tags for each popular placement material in the popular placement material library, it includes:

[0192] Separate the content of each popular placement material in the popular placement material library to obtain corresponding primary modal original content and secondary modal original content;

[0193] Separate the features of the primary modal original content corresponding to each popular placement material to obtain corresponding key elements of the primary modal content;

[0194] Determine the evaluation tags for the key elements of the primary modal content corresponding to each popular placement material and use them as the evaluation tags for the corresponding popular placement materials.

[0195] Optionally, the process of vectorizing each popular placement material in the popular placement material library to obtain corresponding second material vectors for constructing a second material vector library includes:

[0196] Vectorize the evaluation tags for the key elements of the primary modal content corresponding to each popular placement material in the popular placement material library to generate the primary modal content evaluation vectors for the corresponding popular placement materials;

[0197] Vectorize the secondary modal original content corresponding to each popular placement material in the popular placement material library to generate the secondary modal content evaluation vectors for the corresponding popular placement materials;

[0198] Fuse the main modal content evaluation vector and the secondary modal content evaluation vector of each of the said popular advertising materials to generate the second material vector to construct a second material vector library.

[0199] Specifically, the following elaborates in detail the technical implementation details for constructing a second material vector library for popular advertising materials.

[0200] 1. Separate the content of popular advertising materials

[0201] Let the popular advertising material library be where H j (j = 1, 2, …, m, m being the total number of popular advertising materials) represents the j-th popular advertising material.

[0202] Adopt a multi-modal decomposition network based on deep learning to separate the content of each popular advertising material H j from it.

[0203]

[0204] Among them, represents the main modal original content of the j-th popular advertising material, represents the secondary modal original content. θ sep-hot is the parameter set of the multi-modal decomposition network , and these parameters are optimized through training on a large number of popular advertising material data sets. In the advertising placement scenario, the main modal original content may be information such as the prominent product features and core selling points in popular advertisements, and the secondary modal original content may be auxiliary elements in the advertisement, such as background pictures, sound effects, etc.

[0205] 2. Separate the features of the main modal original content

[0206] For the main modal original content of each popular advertising material use a deep network that combines a convolutional neural network (CNN) and a self-attention mechanism (Self-Attention) to separate the features and obtain the features of the key elements of the main modal content.

[0207] First, perform feature extraction through the CNN layer:

[0208] Let X l-1 be the output of the (l - 1)-th layer (initially ), and after the l-th layer of convolutional operation, it is obtained:

[0209] X l = σ(K l * Xl-1 +b l )

[0210] Among them, K l is the convolutional kernel of the l-th layer, and its dimension determines the receptive field size of the convolutional operation. In the feature extraction of advertising images or videos, convolutional kernels of different sizes can capture features at different scales. For example, small convolutional kernels capture detailed features, and large convolutional kernels capture overall structural features; b l is the bias vector of the l-th layer; σ is the activation function, such as the ReLU (Rectified Linear Unit) function, σ(x) = max(0, x), which introduces non-linearity into the network and enables the network to learn more complex feature relationships.

[0211] After L layers of convolution, the feature map X L is obtained. Next, the self-attention mechanism is used to calculate the attention weights:

[0212]

[0213] Among them, Q j , K j are the query matrix and key matrix obtained by linearly transforming X L , and d k is the dimension of the key matrix. The attention weight A j represents the degree of association between different positions in the feature map, and in this way, the key elements in the main modality content can be highlighted.

[0214] Finally, the feature vector F j of the key elements of the main modality content is obtained through attention weighting:

[0215] F j ×A j V j

[0216] Among them, V j is the value matrix obtained by performing another linear transformation on X L .

[0217] 3. Determine the evaluation labels

[0218] A multi-classifier based on the support vector machine (SVM) is used to determine the evaluation labels of each popular advertising material.

[0219] Given the feature vector F j of the key elements of the main modality content, the evaluation label y j is calculated through the following formula:

[0220]

[0221] where w k and b k are the weight vector and bias of the k-th class respectively, where k = 1, 2, …, K, and K is the number of categories of evaluation labels. In the advertising placement scenario, K can represent different classifications such as different types of advertisements, target audience groups, product categories, etc.

[0222] In the training stage, the parameters of the SVM model are optimized by minimizing the following objective function:

[0223]

[0224]

[0225] where C is the penalty parameter, used to balance the classification margin and the penalty degree of misclassified samples; ξ jk is the slack variable, used to allow a certain degree of misclassification; y jk is the true label of sample j belonging to the k-th class (1 if it belongs, otherwise -1).

[0226] 4. Vectorize the evaluation labels

[0227] For the evaluation label y j of each popular placement material, a pre-trained language model based on Transformer is used for vectorization.

[0228]

[0229] where θ vec-hot is the parameter set of the model . Inside the Transformer model, the attention weights at different positions are calculated through the self-attention mechanism to capture the semantic information in the evaluation labels.

[0230] 5. Vectorize the sub-modal original content

[0231] For the sub-modal original content a hybrid model based on the variational autoencoder (VAE) and the long short-term memory network (LSTM) is used for vectorization.

[0232] First, the sub-modal original content is input into the encoder E vae of the VAE:

[0233]

[0234] where z j is the encoded latent vector, and θ e-vaeIt is a set of parameters of the encoder.

[0235] Then, input the latent vector z j into the LSTM network as follows:

[0236]

[0237] where θ lstm is a set of parameters of the LSTM network. The LSTM network can process sequence information and effectively capture the time series features or structural information in the original content of the sub-modal through memory units and gating mechanisms.

[0238] 6. Vector Fusion

[0239] Fuse the main-modal content evaluation vector and the sub-modal content evaluation vector to generate the second material vector V j . A fusion method based on weighted summation is adopted, and the weights are learned through a small multi-layer perceptron (MLP).

[0240] First, calculate the weights and

[0241]

[0242] where θ mlp is a set of parameters of the MLP, represents concatenating the main-modal and sub-modal content evaluation vectors.

[0243] Then, obtain the second material vector V j through weighted summation:

[0244]

[0245] 7. Construct the Second Material Vector Library

[0246] Collect all the popular placement materials H j and the second material vector V j obtained through the above steps to form the second material vector library

[0247] Therefore, the technical benefits of the above process of constructing the second material vector library are as follows:

[0248] 1. Content Separation

[0249]

[0250] This formula utilizes a deep learning-based multi-modal decomposition network The popular delivery material H j Separate the original content from the main modal and the submodal original content Network parameters θ sep-hot The method is determined by training on a large number of popular delivery material data sets, aiming to achieve effective distinction and extraction of different modal information.

[0251] For this reason, in the advertising delivery scenario, popular delivery materials often contain a variety of information. This separation method can clearly divide the core information (primary modal original content) and auxiliary information (secondary modal original content). This helps to carry out targeted processing of different modal contents separately in the future, avoid information confusion, and improve the accuracy and efficiency of feature extraction. For example, in a popular electronic product advertisement, the primary modal original content may be key information such as the product's function demonstration and core selling point introduction, while the secondary modal original content may be auxiliary elements such as the background environment and background music in the advertisement. Accurately separating these contents lays the foundation for the subsequent in-depth analysis of material features.

[0252] 2. Separation of main modal original content features

[0253] Convolution operation: X l =σ(K l *X l-1 +b l )

[0254] Self-Attention Mechanism:

[0255] F j ×A j V j

[0256] The convolution operation is performed through the convolution kernel K l For the input feature map X l-1 Convolution operations are performed to extract local features at different levels, and the activation function σ increases the nonlinear expression ability of the network, enabling the network to learn more complex patterns.

[0257] The self-attention mechanism calculates the query matrix Q j , key matrix K j Sum value matrix V j The relationship between the two is used to obtain the attention weight A j , thereby highlighting the key elements in the main modal content, allowing the model to focus on important information and generate a more representative feature vector F of the key elements of the main modal content j .

[0258] Therefore, in the processing of advertising materials, convolutional operations can automatically extract various local features in advertising images or videos, such as the appearance details of products, unique identifiers, etc. The self-attention mechanism further enhances the model's attention to key features and can accurately capture information closely related to the core selling points of products in complex advertising content. For example, in a popular advertisement for a smartphone, this method can highlight key features such as the high-definition screen display effect and fast charging interface of the phone, while ignoring some irrelevant background decorations, providing an accurate feature basis for subsequent determination of evaluation tags and helping to more accurately describe and classify advertising materials.

[0259] 3. Determine evaluation tags

[0260]

[0261] Objective function:

[0262]

[0263] The first formula is based on the multi-classification principle of Support Vector Machine (SVM), which maps the feature vector F of the key elements of the main modal content j to the corresponding evaluation tag category y j . The objective function is to optimize the parameters w k and b k of the SVM model by minimizing this objective function during the training phase, where is used to control the model complexity and prevent overfitting, is used to penalize misclassified samples to maximize the classification margin and improve the classification accuracy.

[0264] Therefore, in the advertising placement scenario, the model can accurately assign evaluation tags to popular placement materials based on the extracted feature vectors. These tags can cover important information such as the type of advertisement, target audience, product features, etc. By optimizing the classification margin, the model has strong generalization ability and can stably classify on different popular advertising material datasets, improving the accuracy and reliability of the determination of evaluation tags. This provides a clear classification basis for subsequent material management, screening, and matching. For example, according to different target audience tags, popular advertising materials suitable for specific audience groups can be quickly screened for placement, improving the targeting and effectiveness of advertising placement.

[0265] 4. Evaluation tag vectorization formula

[0266]

[0267] Using a pre-trained language model based on Transformer to vectorize the evaluation tag yj Converted to vector representation The Transformer model captures the semantic information in the evaluation labels through the self-attention mechanism, converts the labels in text form into a vector form that is easy for computers to process, and facilitates subsequent operations such as similarity calculation and fusion.

[0268] Therefore, in the field of advertising placement, after vectorizing the evaluation labels, the text information can be converted into numerical vectors, facilitating various mathematical operations and analyses. Through the self-attention mechanism of the Transformer model, the vector representation can accurately reflect the semantic connotations of the evaluation labels, such as the semantic similarity or difference between different labels. This helps to more precisely find popular materials that are semantically similar to the current requirements during the material matching process, improving the accuracy and efficiency of material screening. For example, when looking for popular materials similar to "electronic product advertisements suitable for high-end business people", the vector representation based on the vectorization of evaluation labels can quickly calculate the similarity between other materials and this label, thereby screening out the most suitable materials.

[0269] 5. Vectorization of sub-modal original content

[0270] Encoder:

[0271] LSTM network:

[0272] First, through the encoder E of the variational autoencoder (VAE) vae Encode the sub-modal original content Into the latent vector z j , The VAE can learn the latent distribution of the data, map the high-dimensional sub-modal original content to a low-dimensional latent space, and facilitate subsequent processing.

[0273] Then input the latent vector z j Into the long short-term memory network (LSTM) The LSTM network can process sequence information. Through its unique gating mechanism, it can effectively capture the time series features or structural information in the sub-modal original content and generate the sub-modal content evaluation vector

[0274] Therefore, in advertising placement, although secondary modal original content such as background music and dynamic background images is not the core information, it has an important impact on the overall effect of the advertisement. The VAE encoder can compress and extract features from these complex secondary modal information to obtain representative latent vectors. The LSTM network further processes these latent vectors to capture the time series features therein, such as the rhythm change of background music and the transition effect of dynamic background images. The secondary modal content evaluation vector generated in this way can comprehensively reflect the characteristics of the secondary modal original content, providing rich information for subsequent vector fusion, enabling the finally generated material vector to more completely reflect the overall characteristics of the advertising material, and improving the accuracy and comprehensiveness of material matching.

[0275] 6. Vector Fusion

[0276]

[0277]

[0278] Calculate the main modal content evaluation vector through a small multi-layer perceptron (MLP) and the secondary modal content evaluation vector fusion weights and The MLP takes the concatenated vector as input and determines the appropriate weights by learning the parameter θ mlp Finally, according to the calculated weights, generate the second material vector V j through weighted summation to achieve the fusion of main and secondary modal information.

[0279] Therefore, in the advertising placement scenario, the importance of the main and secondary modal information of different popular advertising materials varies. This vector fusion method based on MLP learning weights can adaptively adjust the fusion ratio of the main and secondary modal information in the final vector according to the specific characteristics of each material. For example, for some advertisements emphasizing product functions, the main modal information may be more important, and the fusion weights will make the main modal content evaluation vector account for a larger proportion in the final second material vector; while for some advertisements focusing on creating an atmosphere, the weight of the secondary modal information will increase accordingly. The second material vector generated in this way can more accurately reflect the unique characteristics of each popular advertising material, provide more powerful support for the accurate matching and retrieval of materials, and further improve the effect and resource utilization rate of advertising placement.

[0280] 7. Construct the Second Material Vector Library

[0281] All popular placement materials H j The second material vector V obtained through the above steps jCollect them to form a second material vector library

[0282] Integrate the second material vectors obtained after a series of processes on each popular advertising material together to form a vector library that can be queried and compared, providing a data basis for subsequent material matching and analysis.

[0283] Therefore, in the advertising business, having a high-quality second material vector library is of great significance. The vectors in this vector library accurately represent the comprehensive characteristics of each popular advertising material. By calculating the similarity of the vectors in the vector library, popular materials similar to the current requirements can be quickly found. For example, when planning a new advertising campaign, popular materials similar to the target audience, product type, advertising style, etc. can be retrieved from the vector library, and their successful experiences can be referred to optimize the production and placement strategies of the new advertisement. This can not only save a large amount of time and resources, but also improve the accuracy and effectiveness of advertising placement, and enhance the market competitiveness of the advertisement.

[0284] Optionally, the evaluation labels of the historical advertising materials or the evaluation labels of the popular advertising materials include at least one of the following: evaluation index labels, content semantic labels.

[0285] Optionally, the evaluation index labels include at least one of the rhythm label, the product brand label introduced in the video advertisement, and the product brand type label, and the content semantic label includes an advertisement category label, and the advertisement category label includes at least one of a text advertisement label, an animated advertisement label, and a real machine advertisement label.

[0286] The rhythm label is mainly used to describe the characteristics of the advertising material in terms of rhythm. In a video advertisement, the rhythm covers elements such as the speed of picture switching, the beat of the music, and the narrative rhythm of the entire advertisement. For example, an advertisement with a strong rhythm may have fast picture switching and dynamic music, while an advertisement with a slow rhythm may have slower picture switching and a more stable music rhythm.

[0287] Therefore, by labeling the advertising material with a rhythm label, technicians can screen out advertising materials with specific rhythm characteristics according to different needs during the material management and matching process. For example, for advertising placement targeting young and energetic audiences, materials with a strong rhythm may be more preferred; while for advertisements of some high-end and steady products, materials with a slow rhythm may be more suitable. This helps to improve the matching degree between the advertisement and the target audience, and enhance the attractiveness and communication effect of the advertisement.

[0288] The product brand label introduced in the video advertisement clearly indicates the product brand advertised in the advertisement. It is a direct identification of the core promotional object of the advertisement, such as brand names like "Apple", "Huawei", "Coca-Cola", etc.

[0289] Therefore, in the process of material management and delivery, product brand tags are an important basis for classification and retrieval. The advertising delivery platform can quickly screen out historical or popular delivery materials related to a specific brand according to brand requirements. This is very helpful for brand owners to carry out brand promotion and market analysis. They can understand the advertising materials of their own brand in different periods, and it is also convenient to distinguish and compare with the materials of other brands, so as to formulate more targeted advertising strategies.

[0290] Product brand type tags are used to classify the brand types to which products belong, such as electronic product brands, food and beverage brands, fashion brands, automotive brands, etc. It describes the attributes of the brand from a more macroscopic level.

[0291] Therefore, such tags help to conduct a more extensive classification and management of advertising materials. When conducting advertising delivery, select appropriate advertising materials of brand types according to different market research and target audience positioning. For example, for advertising delivery targeting technology enthusiasts, materials of electronic product brand types can be preferred; while for consumers who pay attention to quality of life, advertising materials of fashion brands or high-end food and beverage brands may be more attractive. This makes the advertising delivery more accurate and improves the utilization efficiency of advertising resources.

[0292] Text advertising tags are used to identify advertisements with text as the main communication element. Such advertisements may mainly convey product information, brand concepts, promotional activities, etc. through text descriptions, such as newspaper advertisements, text promotions on posters, and pure text advertisements on web pages.

[0293] Therefore, in the material management system, by marking text advertising tags, it is convenient to separately manage and retrieve such text-based advertising materials. For some advertising delivery requirements that need to accurately convey information through text, appropriate materials can be quickly found. At the same time, for analyzing the effects and audience feedback of different advertising forms, text advertising tags also provide an important basis for classification, which helps to optimize advertising strategies.

[0294] Animation advertising tags are used to mark advertisements presented in the form of animation. Animation advertisements can vividly display product features, brand images, or storylines through various animation styles (such as 2D animation, 3D animation) and creative techniques, attracting the attention of the audience.

[0295] Therefore, labeling animated advertising tags enables quick screening of animated advertising materials in the material library. For products or brands that are suitable for promotion in the form of animation, such as children's products, creative brands, etc., relevant animated advertising materials can be conveniently found for reference or direct placement. In addition, through the analysis of animated advertising materials, the market response of different animation styles and creativity can be understood, providing reference for the subsequent production of more attractive animated advertisements.

[0296] The real machine advertising tag is mainly used to identify advertisements with real product objects as the main display. Such advertisements usually let consumers more intuitively understand the features and advantages of the product by showing the actual appearance of the product, function demonstration, etc. For example, the real machine operation demonstration advertisement of a mobile phone, the real car display advertisement of a car, etc.

[0297] Therefore, in the process of advertising placement, the real machine advertising tag helps to quickly locate and screen out advertising materials that can directly display the real product objects. For some categories where consumers are more concerned about the actual appearance and functions of the product, such as electronic products, household products, etc., real machine advertising materials can provide more intuitive information and enhance consumers' willingness to purchase. At the same time, through the management and analysis of real machine advertising materials, the effects of different products in real machine display can be evaluated, and the product display strategy can be optimized.

[0298] Optionally, the historical placement materials and the popular placement materials are visual and audio materials.

[0299] Optionally, the traversing and matching of the first material vectors in the first material vector library and the second material vectors in the second material vector library to calculate the similarity between the corresponding historical placement materials and the popular placement materials includes:

[0300] Invoking a matching model trained based on the set reinforcement learning constraints, traversing the first material vector library and the second material vector library to calculate the similarity between each first material vector and each second material vector, and using this as the similarity between the corresponding historical placement materials and the popular placement materials.

[0301] Optionally, the method further includes:

[0302] Determining placement material samples and matching result labels between samples;

[0303] Separating the content of each placement material sample to obtain the corresponding main modal sample original content and secondary modal sample original content;

[0304] Training the matching model to be trained for several rounds based on the following steps until a trained matching model is obtained:

[0305] Vectorize the original content of the sub-modal samples participating in the current round of training to obtain sub-modal sample vectors;

[0306] The to-be-trained matching model performs forward propagation on the original content of the main-modal samples to generate predicted matching results between the placement material samples;

[0307] Call the objective function that performs reinforcement learning constraints on the predicted matching results based on the sub-modal sample vectors, and calculate the vector similarity between the sub-modal sample vectors in the current round of training and the sub-modal sample vectors in the previous round of training;

[0308] Adjust the network parameters of the to-be-trained matching model in the gradient direction that increases the similarity until a predetermined condition for completing model training is reached. The predetermined condition is at least: the similarity reaches a set similarity threshold and the loss of the predicted matching result relative to the sample-inter matching result label is less than a set loss value threshold.

[0309] Specifically, the technical details of the above similarity calculation and model training are as follows:

[0310] 1. Calculate the vector similarity between the first material vector library and the second material vector library

[0311] The matching model used is based on a deep neural network architecture, and is illustrated by taking a multi-layer perceptron as an example. Let the matching model be M(·; θ), where θ is the set of model parameters, including the weight matrix W and bias vector b of each layer.

[0312] Vector similarity calculation

[0313] For the first material vector library and the second material vector library Traverse and calculate the similarity between each pair of vectors. Here, cosine similarity is used to measure the similarity between vectors. The cosine similarity formula is:

[0314]

[0315] where is the dot product of vectors and and are the L2 norms of vectors and respectively.

[0316] In the advertising placement scenario, the first material vector represents the feature vector of historical placement materials, and the second material vector ​The feature vectors representing popular ad materials. By calculating the cosine similarity between them, the similarity degree of historical ad materials and popular ad materials at the feature level can be measured, providing a quantitative basis for retrieving materials.

[0317] 2. Train the matching model

[0318] Determine the ad material samples and labels

[0319] Let the set of ad material samples be Each sample S k has a corresponding matching result label L k , L k represents the actual matching situation between sample S k and other samples, such as "match" or "no match", which can be encoded and represented by 0 and 1.

[0320] For each ad material sample S k , perform content separation to obtain the original content of the main modality sample and the original content of the secondary modality sample which can be expressed as:

[0321]

[0322] where is the content separation network, and θ sep is its parameter set. In the ad placement scenario, the original content of the main modality sample may be the product information, core selling points, etc. prominently displayed in the ad, and the original content of the secondary modality sample may be the background music, auxiliary pictures, etc. of the ad.

[0323] Generate the secondary modality sample vector

[0324] Vectorize the original content of the secondary modality sample participating in the current round of training and use the method based on variational autoencoder (VAE) to generate the secondary modality sample vector z k .

[0325] First, the encoder E maps to the latent space:

[0326]

[0327] where θ E is the parameter set of the encoder. In the ad scenario, converting the original content of the secondary modality into a low-dimensional latent vector through the VAE encoder can extract the key features therein for subsequent processing.

[0328] Predict the matching result

[0329] The original content of the main modal sample is forward propagated by the matching model \(M(\cdot;\theta)\) to be trained, generating a predicted matching result between the placement material samples.

[0330]

[0331] Define an objective function \(J(\theta)\) based on the sub-modal sample vectors for reinforcement learning constraints on the predicted matching results. The objective function considers the sub-modal sample vector \(z\) in the current round of training k and the sub-modal sample vector \(z\) k-1 in the previous round of training.

[0332] Suppose the vector similarity is calculated using cosine similarity, that is:

[0333]

[0334] The objective function \(J(\theta)\) can be expressed as:

[0335]

[0336] where \(\lambda\) is a balance coefficient used to balance the vector similarity and the loss between the predicted matching result and the true label. is the predicted matching result and the loss function between the matching result label \(L\) k of the samples, such as the cross-entropy loss function:

[0337]

[0338] Adjust the network parameters \(\theta\) of the matching model \(M(\cdot;\theta)\) to be trained in the gradient direction that increases the vector similarity. Using the gradient descent algorithm, the parameter update formula is:

[0339]

[0340] where \(\alpha\) is the learning rate that controls the step size of each parameter update.

[0341] The training process continues until the predetermined conditions for completing the model training are reached. The predetermined conditions are at least:

[0342] 1. The vector similarity reaches the set similarity threshold \(\tau\), s i.e., Similarity\((z\) k , \(z\) k-1 ) \(\geq \tau\) s .

[0343] 2. The loss of the predicted matching result relative to the matching result label between samples is less than the set loss value threshold τ l , that is

[0344] Therefore, each specific processing step involved in the above calculation of similarity and model training has the following technical advantages:

[0345] 1. Calculate vector similarity

[0346]

[0347] This formula is the cosine similarity calculation formula, which is used to measure the similarity between two vectors and . The numerator is the dot product of the two vectors, reflecting their consistency in direction; the denominator is the product of the L2 norms of the two vectors, which is used to normalize the result so that the similarity value ranges between [-1, 1]. The closer the value is to 1, the more similar the two vectors are.

[0348] Therefore, in the advertising placement scenario, the historical placement material vector and the popular placement material vector can be quantified for their similarity through cosine similarity calculation. This helps to quickly screen out historical materials similar to the characteristics of popular placement materials, providing reference for advertising planning and placement. For example, advertising planners can find past successful similar advertising materials based on the similarity, draw on their creativity and strategies, and improve the production quality and placement effect of new advertisements. At the same time, this vector-based similarity calculation method can comprehensively consider the characteristics of materials from multiple dimensions, avoiding the limitations of single-dimensional evaluation and more comprehensively reflecting the similarity relationship between materials.

[0349] 2. Content separation

[0350]

[0351] This formula uses the content separation network to separate the placement material sample S k into the original content of the main modality sample and the original content of the secondary modality sample The network parameters θ sep are obtained through training and learning, enabling the network to accurately distinguish different modality information in the material.

[0352] Therefore, in advertising, a single piece of material often contains multiple types of information. By separating the content, the core information (main modality) and auxiliary information (secondary modality) can be processed separately. This helps to perform targeted feature extraction and analysis on different modality content in subsequent steps, improving processing efficiency and accuracy. For example, the original content of the main modality sample may contain key information such as the core selling points of a product and brand logos, while the original content of the secondary modality sample may contain auxiliary elements such as background music and background images. After separating them, we can more focusedly analyze the impact of the main modality content on the advertising effect, and at the same time consider how the secondary modality content enhances the overall attractiveness of the advertisement, providing a more detailed basis for advertising optimization.

[0353] 3. Generation of Secondary Modality Sample Vectors

[0354]

[0355] This formula uses the encoder E of the Variational Autoencoder (VAE) to map the original content of the secondary modality sample to the latent space, obtaining the secondary modality sample vector z k . The parameters θ E of the encoder E are learned through training, such that the generated vector can effectively represent the characteristics of the original secondary modality content.

[0356] Therefore, in an advertising scenario, secondary modality information such as background music and dynamic background images usually has high dimensionality and complexity. By converting it into a low-dimensional latent vector through the VAE encoder, on the one hand, this can compress this information, reducing the amount of data for subsequent processing; on the other hand, the latent vector can extract the key features in the original secondary modality content and capture the implicit information therein. For example, through feature extraction of the background music, key information such as its rhythm and melody can be incorporated into the vector, providing more valuable information for subsequent model training and similarity calculation, and helping to more accurately evaluate the overall characteristics and effects of the advertising material.

[0357] 4. Prediction of Matching Results

[0358]

[0359] The to-be-trained matching model M(·; θ) takes the original content of the main modality sample as input, performs forward propagation calculation through the parameters θ of the model, and outputs the predicted matching result The model M(·; θ) takes the original content of the main modality sample as input, performs forward propagation calculation through the parameters θ of the model, and outputs the predicted matching result What the model M learns is the mapping relationship between the main modality content and the matching result.

[0360] Therefore, in advertising placement, by predicting the matching results through a model, the matching situation between different materials can be estimated. This helps to quickly screen out possible matching material combinations from a large number of materials, improving the efficiency of material matching. For example, when selecting materials for a new advertising campaign, the model can predict which historical materials may match the current requirements based on the main modal content, reducing the workload of manual screening while improving the accuracy and rationality of matching.

[0361] 5. Reinforcement Learning Constraints and Objective Function

[0362]

[0363]

[0364]

[0365] The first formula calculates the cosine similarity between the sub-modal sample vectors of the current round and the previous round, which is used to measure the stability of the model's representation of sub-modal information in different rounds.

[0366] The second formula is the objective function, which combines the vector similarity and the loss between the predicted matching result and the true label. The negative vector similarity term encourages the model to maintain the consistency of the sub-modal information representation during training, that is, the similarity of the sub-modal sample vectors in adjacent rounds should be as high as possible; λ is the balance coefficient, which is used to adjust the relative importance of the two terms; is the cross-entropy loss function, which is used to measure the predicted matching result and the true label L k to measure the difference between them, prompting the model to improve the accuracy of prediction.

[0367] Therefore, in the advertising placement scenario, through this objective function with reinforcement learning constraints, the model not only has to accurately predict the matching results between materials, but also has to ensure the stability of the representation of sub-modal information. This helps to improve the generalization ability and robustness of the model, enabling it to better adapt to different advertising materials and scenarios. For example, when faced with advertising materials of different styles and contents, the model can, while accurately predicting the matching results, maintain a stable understanding and representation of sub-modal information (such as background music style, picture color tone, etc.), so as to more comprehensively and accurately evaluate the similarity between materials and improve the quality and effect of advertising material matching.

[0368] 6. Model Parameter Adjustment

[0369]

[0370] This formula uses the gradient descent algorithm to adjust the model parameters θ. The gradient It represents the rate of change of the objective function J(θ) with respect to the parameter θ, and the learning rate α controls the step size of each parameter update. By updating the parameters in the opposite direction of the gradient, the value of the objective function gradually decreases, that is, the model is gradually adjusted towards a better direction.

[0371] Therefore, in the model training of advertising placement, the gradient descent algorithm can effectively optimize the model parameters, enabling the model to continuously learn and adapt to the characteristics of advertising material data. By adjusting the parameters, the model can better capture the complex relationship between the main modality and sub-modality content and the matching result, improving the accuracy and stability of prediction. For example, during the training process, the model can automatically adjust the parameters according to different advertising material data to adapt to different advertising styles, product types and other factors, thereby improving the application effect of the model in the actual advertising placement scenario.

[0372] 7. Training termination conditions

[0373] Similarity(z k ,z k-1 )≥τ s

[0374]

[0375] The first condition requires that the similarity between the sub-modality sample vectors in adjacent rounds reaches the set threshold τ s , which means that the model's representation of the sub-modality information has tended to be stable and there are no longer large fluctuations.

[0376] The second condition requires that the loss between the predicted matching result and the true label is less than the set threshold τ l , indicating that the prediction accuracy of the model has reached an acceptable level.

[0377] Therefore, in the advertising placement scenario, these two training termination conditions ensure the training effect of the model. When the model meets these two conditions, it means that it can already stably and accurately predict the matching results between materials, and the representation of the sub-modality information is also consistent. The model trained in this way can reliably calculate the similarity between historical placement materials and popular placement materials in actual applications, providing accurate decision-making basis for advertising placement, improving the efficiency and effect of advertising placement, and avoiding the problems of over-training or under-training of the model.

[0378] Optionally, when screening historical placement materials from the historical placement material library as the placement materials for redelivery based on the similarity, multiple or one historical placement material with a similarity greater than the set similarity threshold is selected as the placement material for redelivery.

[0379] Based on the above embodiments, an embodiment of the present application further provides a material recommendation method, which includes:

[0380] Based on the above material retrieval method, the retrieved placement materials are obtained;

[0381] Based on the set recommendation conditions, the recommended materials are screened out from the retrieved placement materials.

[0382] Specifically, the process of screening out the recommended materials from the retrieved placement materials based on the set recommendation conditions is described in detail:

[0383] The recommendation conditions can be set according to specific business requirements and goals. For example, the recommendation conditions may include the theme relevance of the materials, the matching degree with the target audience, the adaptability to the placement channels, the timeliness of the materials, etc. These conditions can be represented by some quantitative indicators. For example:

[0384] Theme relevance: It can be measured by calculating the text similarity of the text description of the materials (such as the content semantic tags in the evaluation labels). For example, use the cosine similarity to calculate the similarity between the material text and the target theme text, and set a threshold. Only the materials with a similarity higher than this threshold meet the requirements.

[0385] Target audience matching degree: Match according to the evaluation index labels of the materials (such as product brand type labels, target audience labels, etc.) and the characteristics of the target audience. For example, if the target audience is young men, then screen out the materials with product brand type labels related to young men and relevant pointers in the evaluation labels.

[0386] Placement channel adaptability: Consider the requirements of different placement channels for the materials. For example, video platforms may have requirements for video duration and resolution; social media platforms may have requirements for the format and interactivity of the materials. Screen out the materials that meet the corresponding placement channels according to these requirements.

[0387] Timeliness of the materials: If promoting seasonal products or time-limited activities, it is necessary to screen out the materials that match the current time and the activity period.

[0388] When screening and recommending materials:

[0389] For each retrieved placement material, evaluate it one by one according to the above recommendation conditions. Different weights can be set for each recommendation condition, and then the scores of each material are calculated comprehensively. For example, the weight of theme relevance is w 1 , the weight of target audience matching degree is w 2 , the weight of placement channel adaptability is w 3 , the weight of material timeliness is w 4 , and w 1 + w 2 + w3 +w 4 = 1.

[0390] For each piece of material, calculate the score S according to its performance under various conditions: S = w 1 × theme_score + w 2 × audience_score + w 3 + channel_scor + w 4 × timeliness_score.

[0391] For example, in a specific scenario:

[0392] 1. Calculate the theme relevance score:

[0393] Let the material vector be The target theme vector be The formula for calculating the theme relevance score theme_score is:

[0394]

[0395] 2. Calculate the target audience matching score:

[0396] The material vector The 0th dimension represents the target audience characteristics. The formula for calculating the target audience matching score audience_score is:

[0397]

[0398] 3. Calculate the adaptability score for the placement channel:

[0399] The material vector The 1st dimension represents whether it is suitable for the placement channel (1 means suitable). The formula for calculating the adaptability score for the placement channel channel_score is:

[0400]

[0401] 4. Calculate the timeliness score for the material:

[0402] The material vector The 2nd dimension represents the timeliness. The larger the value, the newer. The formula for calculating the timeliness score for the material timeliness_score is:

[0403]

[0404] 5. Calculate the comprehensive score S

[0405] Let the theme relevance weight be w 1= 0.3, the weight of the target audience matching degree is w 2 = 0.3, the weight of the adaptability of the delivery channel is w 3 = 0.2, the weight of the timeliness of the material is w 4 = 0.2. The specific formula for calculating the comprehensive score S is as follows:

[0406]

[0407] On the above basis, according to the calculated scores, all the retrieved materials are sorted, and a certain number of materials with higher scores are selected as the recommended materials.

[0408] Figure 2 The figure is a schematic diagram of a material retrieval device based on modal fusion matching provided by an embodiment of the present invention. As Figure 2 shown, it includes:

[0409] A first program unit, configured to obtain a plurality of historical delivery materials to construct a historical delivery material library;

[0410] A second program unit, configured to vectorize each historical delivery material in the historical delivery material library to obtain a corresponding first material vector to construct a first material vector library;

[0411] A third program unit, configured to obtain a plurality of popular delivery materials to construct a popular delivery material library;

[0412] A fourth program unit, configured to vectorize each popular delivery material in the popular delivery material library to obtain a corresponding second material vector to construct a second material vector library;

[0413] A fifth program unit, configured to traverse and match the first material vectors in the first material vector library and the second material vectors in the second material vector library to calculate the similarity between the corresponding historical delivery materials and popular delivery materials;

[0414] A sixth program unit, configured to screen historical delivery materials from the historical delivery material library as the retrieved delivery materials based on the similarity.

[0415] The above Figure 2 In the embodiment, an exemplary explanation of the technical processing process executed by each program unit can be found in the above Figure 1 record.

[0416] Figure 3 The figure is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. As Figure 3 shown, the electronic device includes a memory and a processor. A computer-executable program is stored on the memory, and the processor runs the computer-executable program to perform the following steps:

[0417] Obtain multiple historical delivery materials to construct a historical delivery material library;

[0418] Vectorize each historical delivery material in the historical delivery material library to obtain a corresponding first material vector to construct a first material vector library;

[0419] Obtain multiple popular delivery materials to construct a popular delivery material library;

[0420] Vectorize each popular delivery material in the popular delivery material library to obtain a corresponding second material vector to construct a second material vector library;

[0421] Traverse and match the first material vectors in the first material vector library and the second material vectors in the second material vector library to calculate the similarity between the corresponding historical delivery materials and popular delivery materials;

[0422] Based on the similarity, screen historical delivery materials from the historical delivery material library as the delivery materials to be retrieved.

[0423] The above Figure 3 In the embodiments, the exemplary explanations of the technical processing procedures for each step can be referred to the above Figure 1 records.

[0424] This application also provides a method for retrieving materials, which includes:

[0425] Obtain multiple first delivery materials to construct a first delivery material library;

[0426] Vectorize each first delivery material in the first delivery material library to obtain a corresponding first material vector to construct a first material vector library;

[0427] Obtain to construct a second delivery material library;

[0428] Vectorize each second delivery material in the second delivery material library to obtain a corresponding second material vector to construct a second material vector library;

[0429] Traverse and match the first material vectors in the first material vector library and the second material vectors in the second material vector library to calculate the similarity between the corresponding first delivery materials and second delivery materials;

[0430] Based on the similarity, screen first delivery materials from the first delivery material library as the delivery materials to be retrieved.

[0431] The first delivery materials are, for example, the above historical delivery materials, and the second delivery materials are, for example, the above popular delivery materials.

[0432] In this embodiment, for the exemplary explanations of the technical processing procedures for each step, reference may be made to the above Figure 1 description.

[0433] This application also provides a method for retrieving materials, which includes:

[0434] Traverse and match the first material vector and the second material vector to calculate the similarity between the corresponding first placed material and the second placed material;

[0435] Based on the similarity, screen the first placed materials from the first placed material library to serve as the retrieved placed materials.

[0436] The first placed material is, for example, the above-mentioned historical placed material, and the second placed material is, for example, the above-mentioned popular placed material.

[0437] In this embodiment, for the exemplary explanations of the technical processing procedures for each step, reference may be made to the above Figure 1 description.

[0438] Optionally, in an embodiment, considering that the first material vector and the second material vector may not be generated, therefore, before traversing and matching the first material vector and the second material vector, it further includes:

[0439] Obtain a plurality of first placed materials to construct a first placed material library;

[0440] Vectorize each first placed material in the first placed material library to obtain the corresponding first material vector to construct a first material vector library;

[0441] Obtain to construct a second placed material library;

[0442] Vectorize each second placed material in the second placed material library to obtain the corresponding second material vector to construct a second material vector library;

[0443] The above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention. The patent protection scope of the embodiments of the present invention shall be defined by the claims. The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions.

[0444] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

Claims

1. A material recovery method based on modal fusion matching, characterized in that: include: Obtain multiple historical delivery materials to build a historical delivery material library; Vectorizing each historical delivery material in the historical delivery material library to obtain a corresponding first material vector to construct a first material vector library; Obtain multiple popular delivery materials to build a popular delivery material library; Vectorizing each hot delivery material in the hot delivery material library to obtain a corresponding second material vector to construct a second material vector library; Performing traversal matching on the first material vector in the first material vector library and the second material vector in the second material vector library to calculate the similarity between the corresponding historical delivery material and the popular delivery material; Based on the similarity, historical delivery materials are screened from the historical delivery material library to serve as the recovered delivery materials.

2. A material recovery method based on modal fusion matching according to claim 1, characterized in that: The vectorizing each historical delivery material in the historical delivery material library to obtain a corresponding first material vector to construct a first material vector library comprises: Determining an evaluation label of each historical delivery material in the historical delivery material library; The evaluation label of each of the historical delivery materials is vectorized to obtain a corresponding first material vector to construct a first material vector library.

3. The material recovery method based on modal fusion matching according to claim 2 is characterized in that: The step of determining the evaluation label of each historical delivery material in the historical delivery material library comprises: Separating the content of each historical delivery material in the historical delivery material library to obtain the corresponding main modal original content and sub-modal original content; Performing feature separation on the main modal original content corresponding to each of the historical delivery materials to obtain key elements of the main modal content; The evaluation labels of the key elements of the main modal content are determined and used as the evaluation labels of the corresponding historical delivery materials.

4. The material recovery method based on modal fusion matching according to claim 3 is characterized in that: The vectorizing each historical delivery material in the historical delivery material library to obtain a corresponding first material vector to construct a first material vector library includes: Vectorizing the evaluation labels of the key elements of the main modal content corresponding to each historical delivery material in the historical delivery material library to generate a main modal content evaluation vector corresponding to the historical delivery material; Vectorizing the submodal original content corresponding to each historical delivery material in the historical delivery material library to generate a submodal content evaluation vector corresponding to the historical delivery material; The primary modal content evaluation vector and the secondary modal content evaluation vector of each of the historical delivery materials are fused to generate the first material vector to construct a first material vector library.

5. The material recovery method based on modal fusion matching according to claim 1 is characterized in that: The vectorizing each hot delivery material in the hot delivery material library to obtain a corresponding second material vector to construct a second material vector library comprises: Determine an evaluation label of each hot delivery material in the hot delivery material library; The evaluation label of each of the popular delivery materials is vectorized to generate a corresponding second material vector to construct a second material vector library.

6. The material recovery method based on modality fusion matching according to claim 5 is characterized in that: The step of determining the evaluation label of each hot delivery material in the hot delivery material library comprises: Separating the content of each hot delivery material in the hot delivery material library to obtain the corresponding main modal original content and sub-modal original content; Performing feature separation on the main modal original content corresponding to each of the popular delivery materials to obtain the corresponding main modal content key elements; An evaluation label of a key element of the main modal content corresponding to each of the hot delivery materials is determined and used as the evaluation label of the corresponding hot delivery material.

7. The material recovery method based on modality fusion matching according to claim 6 is characterized in that: The step of vectorizing each hot delivery material in the hot delivery material library to obtain a corresponding second material vector to construct a second material vector library includes: Vectorizing the evaluation labels of the key elements of the main modal content corresponding to each hot delivery material in the hot delivery material library to generate a main modal content evaluation vector corresponding to the hot delivery material; Vectorizing the submodal original content corresponding to each hot delivery material in the hot delivery material library to generate a submodal content evaluation vector corresponding to the hot delivery material; The primary modal content evaluation vector and the secondary modal content evaluation vector of each of the hot delivery materials are fused to generate the second material vector to construct a second material vector library.

8. A material recovery method based on modal fusion matching according to any one of claims 1 to 7, characterized in that: The traversing and matching the first material vector in the first material vector library and the second material vector in the second material vector library to calculate the similarity between the corresponding historical delivery material and the popular delivery material includes: Call the matching model trained based on the set reinforcement learning constraints, traverse the first material vector library and the second material vector library to calculate the similarity between each of the first material vectors and each of the second material vectors, and use this as the similarity between the corresponding historical delivery materials and the popular delivery materials.

9. The material recovery method based on modality fusion matching according to claim 8 is characterized in that: The method further comprises: Determine the delivery material samples and the labels of the matching results between the samples; Separating the content of each of the delivery material samples to obtain the corresponding original content of the main modality sample and the original content of the sub-modality sample; The training matching model is trained for several rounds based on the following steps until a training matching model is obtained: Vectorize the original content of the submodal samples participating in the current round of training to obtain the submodal sample vector; The matching model to be trained performs forward propagation on the original content of the main modality sample to generate a predicted matching result between the delivery material samples; Calling the objective function for performing reinforcement learning constraints on the predicted matching results based on the submodal sample vectors, and calculating the vector similarity between the submodal sample vectors in the current round of training and the submodal sample vectors in the previous round of training; According to the gradient direction that increases the similarity, the network parameters of the matching model to be trained are adjusted until a predetermined condition for completing the model training is met, and the predetermined condition is at least: the similarity reaches a set similarity threshold and the loss of the predicted matching result relative to the matching result label between the samples is less than a set loss value threshold.

Citation Information

Patent Citations

  • Personalized recommendation method and system based on historical materials

    CN110020200A

  • Bid text generation method and system based on large language model and storage medium

    CN118982007A

Cited By

  • Mask screening-based material salvage system

    CN120929621A