Product label extraction method based on Internet big data and AI big language model

By combining Internet big data and AI big language model, the problem of insufficient performance of traditional product label extraction methods when facing language diversity and complex semantics is solved, and more accurate vocabulary recognition and deep semantic capture are achieved, and the generated product labels are more in line with actual needs.

CN120012774AActive Publication Date: 2025-05-16BEIJING TAOMI TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510083215.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Traditional product label extraction methods are insufficient in the face of language diversity and complex semantics, difficult to capture the deep semantic relationships between contexts, and are inefficient and inadaptive.

Method used

The product label extraction method based on Internet big data and AI large language models is adopted, and Internet text data is captured through crawling technology, combined with TF-IDF algorithm and Skip-Gram model to capture semantic associations between vocabulary, and large-scale pre-trained language models are used to generate product labels, and finally position and classify them through BERT and CRF models.

Benefits of technology

It realizes more accurately identifying and extracting important vocabulary, capturing deep semantic relationships, and the generated product labels are more in line with actual needs, improving the accuracy and semantic richness of labels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012774A_ABST
    Figure CN120012774A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of product label extraction, in particular to a product label extraction method based on Internet big data and an AI big language model. The method comprises the following steps: S1, capturing text data of a product on the Internet by using a crawler technology; s2, determining important vocabularies in the text data by adopting a TF-IDF algorithm, capturing semantic association among the vocabularies in combination with a Skip-Gram model, and introducing a weight reflecting user browsing frequency and a behavior feature vector of a user in a process of capturing the semantic association among the vocabularies to optimize a capturing process; s3, based on the extracted important vocabularies and semantic association information between the vocabularies, generating product labels by using a large-scale pre-trained language model; and S4, positioning and classifying product labels by combining a sequence labeling model BERT and a conditional random field CRF, and outputting the finally extracted product labels. According to the technology, the BERT model and the conditional random field (CRF) layer are combined, so that product labels can be effectively positioned and classified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of product label extraction, and in particular to a product label extraction method based on Internet big data and an AI big language model. Background Art

[0002] The keyword matching and rule-based algorithms of traditional label extraction methods rely on predefined lexicons and language rules. Although simple and direct, they often perform poorly when faced with language diversity and complex semantics. Secondly, traditional machine learning models have limited ability to express features and are difficult to capture deep semantic relationships between contexts, resulting in extraction effects that are difficult to meet actual needs. In addition, Internet data is large in scale, diverse in form (such as structured data and unstructured text), and has significant differences in semantic expression. At the same time, hot spots and user needs change rapidly, and traditional methods appear to be inefficient and lack adaptability. Therefore, a product label extraction method based on Internet big data and AI big language model is provided. Summary of the invention

[0003] The purpose of the present invention is to provide a product label extraction method based on Internet big data and AI big language model to solve the limitations of the traditional product label extraction process mentioned in the above background technology.

[0004] To achieve the above object, the present invention aims to provide a product label extraction method based on Internet big data and AI big language model, comprising the following steps:

[0005] S1. Use crawler technology to capture text data of products on the Internet;

[0006] S2. Use the TF-IDF algorithm to determine the important words in the text data, and combine it with the Skip-Gram model to capture the semantic associations between words. In the process of capturing the semantic associations between words, the weights reflecting the user's browsing frequency and the user's behavioral feature vector are introduced to optimize the capture process.

[0007] S3, based on the extracted important words and the semantic association information between words, generate product labels using a large-scale pre-trained language model;

[0008] S4. Combine the sequence annotation model BERT and conditional random field CRF to locate and classify product labels, and output the final extracted product labels.

[0009] As a further improvement of the technical solution, in S2, the TF-IDF algorithm is used to determine the important words in the text information, including the following steps:

[0010] S2.1. Preprocess the text data of the product and organize the captured text data into a corpus D, where D contains N documents d, D = {d1, d2, d3, ..., d N};

[0011] S2.2. Calculate the term frequency TF(t, d) of each word t in document d:

[0012]

[0013] Where cou(t, d) represents the number of occurrences of word t in document d, ∑ w∈d (w, d) represents the total number of occurrences of all words in document d; w represents each word in document d;

[0014] S2.3. Calculate the inverse document frequency IDF(t, D) of word t in the entire corpus:

[0015]

[0016] Where N is the total number of documents in the corpus, |{d∈D:t∈d}| is the number of documents containing word t;

[0017] S2.4. Combine the term frequency and the inverse document frequency to get the importance weight of term t in document d: TF-IDF(t, d, D):

[0018] TF-IDF(t,d,D)=TF(t,d)·IDF(t,D).

[0019] As a further improvement of the technical solution, in S2, the Skip-Gram model is combined to capture the semantic association between words, including the following steps:

[0020] S2.5. Build a vocabulary based on the preprocessed text data to record all unique words and their frequencies that appear in the text;

[0021] S2.6, set a fixed window size c, for each central word, take c words on the left and right of the central word as context words to form a training pair;

[0022] S2.7. Randomly initialize a low-dimensional vector for each word in the vocabulary and set hyperparameters.

[0023] S2.8. Set the objective function to train the Skip-Gram model, assign weights to product word vectors that reflect browsing frequency, and transform the user behavior feature vector u i and u o Introduce it into the objective function for optimization, and introduce the related products as additional positive examples into the objective function for further optimization;

[0024] S2.9. The trained Skip-Gram model maps each word into a high-dimensional vector space to obtain the word vector v for each word. t ;

[0025] S2.10. Combining word frequency, inverse document frequency and word vector v t , forming a new product word vector v d :

[0026]

[0027] S2.11. Calculate the cosine similarity of word vectors to measure the semantic similarity between words. For words with similar contexts, introduce a clustering method in the calculation process of the cosine similarity of word vectors. Through clustering, semantically similar words are classified into the same category.

[0028] As a further improvement of the technical solution, in S2.8, the objective function is:

[0029]

[0030] Among them, L neg represents the optimized objective function; σ(*) represents the Sigmoid function; v o The vector representation of the context word; v i represents the vector representation of the central word; k represents the number of negative samples; v j Represents the vector representation of negative sample words; w j Represents negative sample words; represents the expected value of a random draw from a vocabulary according to a certain vocabulary distribution P(w); P(w) represents the vocabulary distribution; T represents the transpose of the vector; i represents the index of the center word; o represents the index of the context word; j represents the index of the negative sample word;

[0031] When a product frequently appears in the user's browsing path, the product word vector is given a weight that reflects the browsing frequency, and a user behavior matrix is ​​established. Two user behavior feature vectors u are extracted from the user behavior matrix. i and u o , and the user behavior feature vector u i and u o Introduce into the objective function for optimization:

[0032]

[0033]

[0034] Among them, L neg,uesr represents the objective function after optimizing the user's browsing path; u iRepresents the user behavior feature vector corresponding to the target word; u o Represents the user behavior feature vector corresponding to the context word; bp u,r represents the browsing frequency of user u on the central product r; bp u,r1 represents the browsing frequency of user u on negative sample product r1;

[0035] If multiple products often appear in the browsing path of the same user, then the products are considered to be related. For each center word v i , find products that are often viewed by the same user group based on the user's browsing history, and introduce the associated products as additional positive examples into the objective function for further optimization:

[0036]

[0037] Among them, L neg,be represents the objective function after further optimization; B represents a set of additional positive examples selected based on the user's browsing path; v b Represents the word vector corresponding to the additional positive example; u b represents the user behavior feature vector corresponding to the additional positive example; α represents the hyperparameter that controls the contribution of the user behavior positive example; b represents the index of the additional positive example.

[0038] As a further improvement of the technical solution, in S2.11, the cosine similarity of the word vectors is calculated to measure the semantic similarity between words:

[0039]

[0040] Among them, Sim(v c , v e ) represents the cosine similarity of word vectors; v c The word vector representation of word c; v e The word vector representation of word e;

[0041] For words with similar contexts, a clustering method is introduced in the calculation process of the cosine similarity of word vectors. Through clustering, semantically similar words are classified into the same category:

[0042]

[0043] Among them, Sim1(v c , v e ) represents the cosine similarity of the optimized word vector; CL(v c , v e ) represents the similarity calculated based on vocabulary clustering information; γ represents the weighting coefficient of clustering information.

[0044] As a further improvement of the technical solution, in S3, based on the extracted important words and the semantic association information between the words, a large-scale pre-trained language model is used to generate product labels, including the following steps:

[0045] S3.1. Construct an input sequence X=(x1, x2, ..., x3) based on the extracted important words and the semantic association information between words. n ), and input the input sequence into the pre-trained language model;

[0046] S3.2. The pre-trained language model predicts the next word at each position through autoregression to generate a complete text:

[0047]

[0048] in, Indicates the generated product label; y n+1 represents the next word generated; P(y n+1 |X) indicates that the model calculates the next word y by autoregression n+1 The conditional probability distribution of ;

[0049] S3.3, generating a probability distribution of product labels based on the input sequence and the trained model parameters θ;

[0050] S3.4. Use temperature sampling strategy to extract label candidates from conditional probability distribution.

[0051] As a further improvement of the technical solution, in S3.3, the probability distribution of the product label is:

[0052]

[0053] in, Represents the generated product label The conditional probability distribution of ; n represents the length of the input sequence; t represents the position index in the process of product label generation.

[0054] As a further improvement of the technical solution, in S4, the sequence annotation model BERT and the conditional random field CRF are combined to locate and classify product labels, including the following steps:

[0055] S4.1. Arrange the product label candidates generated by the pre-trained language model;

[0056] S4.2. Create a corresponding feature vector T' for each product label candidate h ={t' h,1 , t' h,2 ,…,t' h,m}, where T' hrepresents the product label set after preliminary screening, t' h,m represents the mth product label candidate;

[0057] S4.3. Product label candidate sequence T' h Input into the BERT model, the BERT model encodes each product label candidate according to the context and captures the semantic features of the product label candidate:

[0058]

[0059] Among them, R m Represents product label candidate t' h,m BERT encoding result, Represents the parameters of the BERT model;

[0060] S4.4. Define product label space L = {l1, l2, ..., l k}, where k represents the index of the product label category;

[0061] S4.5. Apply the conditional random field layer to calculate the transition probability matrix A between labels. The transition probability matrix represents the probability of transferring from one product label to another product label:

[0062] A=[a hm ];

[0063]

[0064] Among them, [a hm ] indicates product label space l g Transfer to product label space k The transition probability, represents the parameters of the conditional random field layer; g represents the index of the product label category;

[0065] S4.6. BERT encoding result R m Combined with the transition probability of the conditional random field layer, the optimal label sequence is determined:

[0066]

[0067] Among them, T h Represents the finalized product label set; t" h,m represents the mth final product label; m1 represents the total length of the product label sequence;

[0068] S4.7. Use the Viterbi algorithm to find the product label sequence with the highest probability.

[0069] As a further improvement of the technical solution, in S4.7, the Viterbi algorithm is used to find the product label sequence with the highest probability, including the following steps:

[0070] S4.71. For the first product label candidate, calculate the initial probability of each possible product label:

[0071]

[0072] Where l represents one of the product label categories; δ1(l) represents the first product label candidate t” h,1 The initial probability of product label l under the condition of ;

[0073] S4.72. Initialize the backtracking pointer for each product tag

[0074]

[0075] S4.73. For the e1th product label candidate in the sequence, calculate the maximum cumulative probability of each possible product label and update the backtracking pointer;

[0076] S4.74. At the end of the sequence, select the product label with the highest cumulative probability as the last product label z m :

[0077] z m = argmax l δ m (l);

[0078] S4.75, starting from the last product tag, according to the backtracking pointer Go back step by step and reconstruct the entire optimal product label sequence:

[0079] S4.76. After the reconstruction is completed, the optimal product label sequence T is obtained. h , which is the final extracted product label:

[0080] T” h =(z1, z2, ..., z m ).

[0081] As a further improvement of the technical solution, in S4.73, the maximum cumulative probability of each possible product label is calculated as:

[0082]

[0083] Among them, δ e1 (l) represents the probability of the best path to the e1th position with product label l given the observation condition of the e1th position; δe1-1 (l') represents the probability of the best path to the e1-1th position with product label l' given the observation of the e1-1th position; Indicates that given the e1th product label candidate t" h,e1 The maximum cumulative probability of product label l under the condition of; Represents a pointer to the previous best product label; a l',l represents the transition probability from product label l' to product label l; t' h,e1 Indicates the product label candidate at the e1th position; t" h,e1 Indicates the final product label for the e1th position.

[0084] Compared with the prior art, the present invention has the following beneficial effects:

[0085] 1. In this product label extraction method based on Internet big data and AI big language model, by combining the TF-IDF algorithm, Skip-Gram model and large-scale pre-trained language models (such as the GPT series), this method can identify and extract important words from a large amount of text data, and capture the semantic associations between these words. This method not only considers the frequency of vocabulary, but also pays attention to the distribution of vocabulary in different documents and the context, so as to more accurately reflect the characteristics of the product and the user's concerns. In addition, the introduction of user behavior data as additional positive examples to optimize the objective function further enhances the model's understanding of the potential relationship between products, making the generated product labels more in line with actual needs and improving the accuracy and semantic richness of the labels.

[0086] 2. In the product label extraction method based on Internet big data and AI big language model, the BERT model and the conditional random field (CRF) layer are combined to effectively locate and classify product labels. The BERT model can provide a context-sensitive representation for each label candidate based on the context, while the CRF layer can calculate the transition probability between labels, ensuring that the final output label sequence not only conforms to the definition of a single label, but also has logical consistency and coherence in the entire sequence. The application of the Viterbi algorithm ensures that the label sequence with the highest probability is selected, which helps to improve the intelligence level of label classification, while ensuring the consistency and professionalism of the label system, facilitating subsequent data analysis and business decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 The figure is a flow chart of the overall method of the present invention. DETAILED DESCRIPTION

[0088] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0089] Please refer to Figure 1 As shown, this embodiment provides a method for extracting product tags based on Internet big data and AI large language models, including the following steps:

[0090] S1. Use web crawler technology to capture the text data of products on the Internet. The text data includes e-commerce platforms, product reviews, user feedback, product descriptions, etc.;

[0091] S2. Adopt the TF-IDF algorithm to determine the important words in the text data, and combine the Skip-Gram model to capture the semantic associations between words. During the process of capturing the semantic associations between words, introduce the weight reflecting the user browsing frequency and the user behavior feature vector to optimize the capture process;

[0092] In this embodiment, TF-IDF is a statistical method used to evaluate the importance of a word in a document or a set of documents. It calculates by combining the term frequency (the frequency of a word appearing in a document) and the inverse document frequency (the rarity of a word appearing in the entire document collection); the main purpose of using the TF-IDF algorithm to determine the important words in the text information is to identify those words that appear frequently in a specific document but are relatively uncommon in the entire corpus. These words often have a high discrimination degree and can well represent the theme or content of the document; by calculating the term frequency (TF) and inverse document frequency (IDF) of each word, a weight value can be assigned to each word. This weight reflects the importance of the word for a certain document, that is, its significance in the document; some very common words (such as stop words like "of", "is", "in", etc.) may appear frequently in most documents, but they contribute little to understanding the document theme. Through the IDF part, TF-IDF effectively reduces the weights of these general words, thereby reducing their impact on the analysis results;

[0093] Adopting the TF-IDF algorithm to determine the important words in the text information includes the following steps:

[0094] S2.1. Preprocess the text data of the product, and organize the captured text data into a corpus D, where D contains N documents d, D = {d1, d2, d3, …, d N};

[0095] S2.2. Calculate the term frequency TF(t, d) of each word t in document d:

[0096]

[0097] Among them, cou(t, d) represents the number of occurrences of word t in document d, Σ w∈d (w, d) represents the total number of occurrences of all words in document d; w represents each word in document d;

[0098] S2.3. Calculate the inverse document frequency IDF(t, D) of word t in the entire corpus:

[0099]

[0100] Where N is the total number of documents in the corpus, |{d∈D:t∈d}| is the number of documents containing word t;

[0101] S2.4. Combine the term frequency and the inverse document frequency to get the importance weight of term t in document d: TF-IDF(t, d, D):

[0102] TF-IDF(t,d,D)=TF(t,d)·IDF(t,D).

[0103] Among them, the Skip-Gram model is a neural network model for training word vectors. It learns the distributed representation of words by predicting the context words around a given word. By building and training a word vector model, Skip-Gram can map semantically similar or related words to similar positions in a high-dimensional space, even if the words are literally completely different. This helps to identify synonyms, hyponyms, and words with similar contexts. Compared with methods that only use word frequency statistics, word vectors provide richer feature representations that include contextual information of words. This is very important for tasks such as product label extraction because it allows the system to infer the meaning of words based on the patterns in which they are used in different contexts. For some uncommon but potentially very important words, traditional statistical methods may underestimate their importance. The Skip-Gram model is able to use surrounding context information to enhance understanding of these words, even if they do not appear many times.

[0104] Combining the Skip-Gram model to capture the semantic associations between words includes the following steps:

[0105] S2.5. Build a vocabulary based on the preprocessed text data to record all unique words and their frequencies that appear in the text;

[0106] S2.6, set a fixed window size c to determine the context range around the current word. For each central word, take c words on the left and right of the central word as context words to form a training pair;

[0107] S2.7. Randomly initialize a low-dimensional vector for each word in the vocabulary, usually with a dimension between 50 and 300, and set hyperparameters, including window size, vector dimension, minimum word frequency, learning rate, etc.

[0108] S2.8. Set the objective function to train the Skip-Gram model. The objective function is used to maximize the probability of predicting the correct context word under the given central word condition, give the product word vector a weight that reflects the browsing frequency, and convert the user behavior feature vector u i and u o Introduce it into the objective function for optimization, and introduce the related products as additional positive examples into the objective function for further optimization;

[0109] Furthermore, the objective function is:

[0110]

[0111] Among them, L neg represents the objective (loss) function of optimization, and the specific output is a scalar value, which represents the degree of error between the model's prediction result on a given training sample and the actual label; σ(*) represents the Sigmoid function, which is used to map real values ​​to between (0,1); v o The vector representation of the context word; v i represents the vector representation of the central word; k represents the number of negative samples; v j Represents the vector representation of negative sample words; w j Represents a negative sample word, which is extracted from the vocabulary distribution P(w) and is not the central word v i The real context word of It represents the expected value of a random sample from a vocabulary according to a certain vocabulary distribution P(w), which is actually approximated by random sampling during the training process; P(w) represents the vocabulary distribution; T represents the transpose of the vector; i represents the index of the center word; o represents the index of the context word; j represents the index of the negative sample word;

[0112] When users frequently visit certain products in their browsing paths, this indicates that these products may have some kind of intrinsic connection, for example, they may be complements, substitutes, or belong to the same category. This co-occurrence pattern can help us discover semantic associations that are not explicitly stated in the text data but actually exist; through browsing paths, we can identify which products usually appear in the same scenario (such as a shopping session), thereby providing more contextual information for the labels of these products, which helps to more accurately capture their semantic features; users' browsing behavior often reveals their current needs or interests. Even if this information is not directly expressed in the comments or descriptions, it can serve as a supplementary clue to enrich the semantic representation of related vocabulary. For example, if many users are looking at a brand's new mobile phone, then words such as "new model", "mobile phone", and the brand's name may gain additional importance weights;

[0113] When a product frequently appears in the user's browsing path, the product word vector is given a weight that reflects the browsing frequency, and a user behavior matrix is ​​established (each row of the user behavior matrix represents a user's behavior pattern, and each column corresponds to a product). Two user behavior feature vectors u are extracted from the user behavior matrix. i and u o , and the user behavior feature vector u i and u o Introduce into the objective function for optimization:

[0114]

[0115] Among them, L neg,uesr represents the objective (loss) function after optimizing the user's browsing path; u i Represents the user behavior feature vector corresponding to the target word; u o Represents the user behavior feature vector corresponding to the context word; bp u,r represents the browsing frequency of user u on the central product r; bp u,r1 represents the browsing frequency of user u on negative sample product r1;

[0116] Additional positive examples refer to products (or commodities) that are viewed, clicked, or browsed by users within a period of time, and are considered as "additional positive examples" related to the current product (or commodity). Additional positive examples generated in combination with user browsing paths refer to using user browsing behavior data (such as clickstreams, browsing history) to identify products or pages that are often viewed by the same user or similar user groups, and consider these products or pages as positive examples of each other. In other words, if two or more products often appear in the browsing path of the same user, they can be considered to be associated or have similar attributes, even if this association is not explicitly stated in the text description.

[0117] If multiple products often appear in the browsing path of the same user, then the products are considered to be related or have similar attributes. For each central word v i , find products that are often viewed by the same user group based on the user's browsing history, and introduce the associated products as additional positive examples into the objective function for further optimization:

[0118]

[0119] Among them, L neg,be represents the objective (loss) function after further optimization; B represents a set of additional positive examples selected based on the user's browsing path; v b Represents the word vector corresponding to the additional positive example; u b represents the user behavior feature vector corresponding to the additional positive example; α represents the hyperparameter that controls the contribution of the user behavior positive example, which is used to balance the weight between the traditional positive example and the user behavior positive example; b represents the index of the additional positive example.

[0120] S2.9. The trained Skip-Gram model maps each word into a high-dimensional vector space to obtain the word vector v for each word. t , each word vector contains the potential semantic features of the word in the corpus. Through these word vectors, the semantic relationship between words can be captured, such as synonyms, hyponyms, and so on;

[0121] S2.10. Combining word frequency, inverse document frequency and word vector v t , forming a new product word vector v d :

[0122]

[0123] S2.11. Calculate the cosine similarity of word vectors to measure the semantic similarity between words. For words with similar contexts, introduce a clustering method in the calculation process of the cosine similarity of word vectors. Through clustering, semantically similar words are grouped into the same category.

[0124] Furthermore, the main purpose of calculating the cosine similarity of word vectors to measure the semantic similarity between words is to quantify and compare the semantic relationship between different words, so as to more accurately understand and utilize the meaning of words in various natural language processing tasks; by calculating the cosine value of the angle between two word vectors, their relative position in the high-dimensional space can be obtained, which reflects the degree of semantic similarity between the two words. The higher the similarity score, the more likely the two words appear in similar contexts in the corpus, and therefore may have similar or related meanings; in tasks such as product label extraction, using cosine similarity can help identify those words that are most closely related to the target concept, ensuring that the generated labels are both accurate and representative, while avoiding the limitations of relying solely on frequency statistics;

[0125] The cosine similarity of word vectors is calculated to measure the semantic similarity between words:

[0126]

[0127] Among them, Sim(v c , v e ) represents the cosine similarity of word vectors; v c The word vector representation of word c; v e The word vector representation of word e;

[0128] In natural language processing, traditional cosine similarity calculations judge the similarity of words based on the angular difference between word vectors. However, although many words are semantically similar, their contexts may be very similar, resulting in the inability of cosine similarity calculations to accurately distinguish the subtle differences between them. In particular, the vectors of synonyms, words with similar hyponyms or contexts may be very close, resulting in the inability of traditional cosine similarity to effectively identify them. For example, "smartphone" and "mobile phone" are very close in semantics, but may have different applications or importance in different contexts. Traditional cosine similarity calculations will consider them to be extremely similar, while ignoring the differences in their usage scenarios. Clustering methods are introduced to group semantically similar words into the same category through clustering, so that clustering information between words is considered when calculating similarity. In this way, clustering results can be used to enhance the ability to distinguish the semantics of words, and more accurately reflect the relationship between words when calculating similarity.

[0129] For words with similar contexts, a clustering method is introduced in the calculation process of the cosine similarity of word vectors. Through clustering, semantically similar words are classified into the same category:

[0130]

[0131] Among them, Sim1(v c , v e ) represents the cosine similarity of the optimized word vector; CL(v c , v e ) represents the similarity calculated based on vocabulary clustering information; γ represents the weighting coefficient of clustering information; G represents the set of all clusters, and each g is a cluster; represents the indicator function, indicating whether word c and word e belong to the same cluster g; S g (c, e) represents the similarity measure between word c and word e in cluster.

[0132] S3, based on the extracted important words and the semantic association information between words, generate product labels using a large-scale pre-trained language model;

[0133] In this embodiment, the main purpose of using a large-scale pre-trained language model (GPT series) to generate product labels based on the extracted important words and the semantic association information between words is to automatically create high-quality, representative labels that can accurately reflect the characteristics of the product and match the needs of users and market trends; by combining important words and the semantic associations between them, the pre-trained language model can generate labels that are more in line with the actual content of product descriptions and user reviews. This helps to ensure that the labels not only cover the key topics of the document, but also express deeper product characteristics or user preferences; the pre-trained language model has been trained with a large amount of text data and has rich language rules, contextual relationships, and world knowledge. Therefore, it can generate diverse and informative labels, rather than just simple keyword extraction, thereby providing a more detailed and detailed description for each product;

[0134] Based on the extracted important words and the semantic association information between words, a large-scale pre-trained language model (GPT series) is used to generate product labels, including the following steps:

[0135] S3.1. Construct an input sequence X=(x1, x2, ..., x3) based on the extracted important words and the semantic association information between words. n ), and input the input sequence into the pre-trained language model so that the model can generate appropriate product labels. The large-scale pre-trained language model has learned rich language rules, contextual relationships and world knowledge through a large corpus, and can generate labels that are highly relevant to the input text;

[0136] S3.2. The pre-trained language model predicts the next word at each position through autoregression to generate a complete text:

[0137]

[0138] in, Indicates the generated product label; y n+1 represents the next word generated; P(y n+1 |X) indicates that the model calculates the next word y by autoregression n+1 The conditional probability distribution of ;

[0139] S3.3, generating a probability distribution of product labels based on the input sequence and the trained model parameters θ;

[0140] Furthermore, the probability distribution of product labels is:

[0141]

[0142] in, Represents the generated product label The conditional probability distribution of ; n represents the length of the input sequence; t represents the position index in the product label generation process, which is used to represent each step of the model generating labels;

[0143] S3.4. A temperature sampling strategy is used to extract label candidates from the conditional probability distribution. This is a strategy for controlling the smoothness of vocabulary distribution when generating text. Temperature sampling controls randomness by adjusting the "temperature" parameter of the generated probability distribution.

[0144] S4, combining the sequence annotation model BERT and the conditional random field CRF to locate and classify product labels, and output the final extracted product labels;

[0145] In this embodiment, the main purpose of combining the sequence annotation model BERT and the conditional random field (CRF) to locate and classify product labels is to achieve high-precision label recognition and classification, ensuring that the generated labels not only accurately reflect the characteristics of the product, but also maintain the consistency and logic of the label sequence; the BERT model can provide a context-sensitive representation for each label candidate by capturing context information, which helps to more accurately understand the meaning of vocabulary in a specific context, thereby improving the accuracy of label classification; the CRF layer uses the transition probability matrix, and CRF can consider the dependencies between labels to ensure the overall coherence and rationality of the label sequence, avoid processing each label in isolation, and thus improve the quality of the classification results; the CRF layer calculates the transition probability between labels, ensuring that the label sequence not only conforms to the definition of a single label, but also has logical consistency and coherence in the entire sequence. This is very important for building a structured label system, especially when multiple labels jointly describe a complex product or concept; by combining the deep learning ability of BERT and the sequence modeling ability of CRF, the long dependency problem in label generation can be effectively dealt with, that is, the choice of label depends not only on the current word, but also on the context before and after. This method improves the overall efficiency and effectiveness of the label generation process;

[0146] Combining the sequence tagging model BERT and the conditional random field CRF to locate and classify product labels includes the following steps:

[0147] S4.1. Arrange the product label candidates generated by the pre-trained language model to form a serialized input format;

[0148] S4.2. Create a corresponding feature vector T' for each product label candidate h ={t' h,1 , t' h,2 ,…,t' h,m}, including word vectors, position information, etc., where T' h represents the product label set after preliminary screening, t' h,m represents the mth product label candidate;

[0149] S4.3. Product label candidate sequence T' h After being input into the BERT model, the BERT model encodes each product label candidate according to the context, capturing the semantic features of the product label candidate. The BERT model outputs a context-sensitive representation of each label candidate, which contains the complex relationship between the label candidate and its surrounding environment:

[0150]

[0151] Among them, R mRepresents product label candidate t' h,m BERT encoding result, Represents the parameters of the BERT model;

[0152] S4.4. Define product label space L = {l1, l2, ..., l k}, where k represents the index of the product label category (such as "brand", "color", "size", etc.);

[0153] S4.5. Apply the conditional random field layer to calculate the transition probability matrix A between labels. The transition probability matrix represents the probability of transferring from one product label to another product label:

[0154] A=[a hm ];

[0155]

[0156] Among them, [a hm ] indicates product label space l g Transfer to product label space k The transition probability, represents the parameters of the conditional random field layer; g represents the index of the product label category;

[0157] S4.6. BERT encoding result R m Combined with the transition probability of the conditional random field layer, the optimal label sequence is determined:

[0158]

[0159] Among them, T h Represents the finalized product label set; t" h,m represents the mth final product label; m1 represents the total length of the product label sequence;

[0160] S4.7. Use the Viterbi algorithm to find the product label sequence with the highest probability, that is, to maximize the probability of the label sequence.

[0161] Furthermore, the main purpose of using the Viterbi algorithm to find the product label sequence with the highest probability is to ensure that the final generated label sequence is not only the optimal choice locally (each word or phrase), but also has the highest overall coherence and logical consistency in the entire sequence; the Viterbi algorithm searches for a path that maximizes the probability of the entire sequence among all possible label sequences through dynamic programming. This ensures that the selected label sequence is not simply based on the best label at each position, but takes into account the optimality of the entire sequence; because the Viterbi algorithm takes into account the transition probability between labels (calculated by the CRF layer), it can ensure that there is a reasonable conversion relationship between adjacent labels in the label sequence, avoiding the inconsistency problem that may be caused by selecting each label in isolation;

[0162] Using the Viterbi algorithm to find the product label sequence with the highest probability includes the following steps:

[0163] S4.71. For the first product label candidate, calculate the initial probability of each possible product label:

[0164]

[0165] Where l represents one of the product label categories; δ1(l) represents the first product label candidate t” h,1 The initial probability of product label l under the condition of ;

[0166] S4.72. Initialize the backtracking pointer for each product tag For recording the best path:

[0167]

[0168] S4.73. For the e1th product label candidate in the sequence (2≤e1≤m), calculate the maximum cumulative probability of each possible product label and update the backtracking pointer;

[0169] Furthermore, by calculating the maximum cumulative probability, the global optimization of the entire label sequence can be achieved, rather than just the local optimum. This means that the label sequence finally selected is the best solution obtained after considering all possible paths, ensuring the overall quality and accuracy of the label sequence;

[0170] Calculate the maximum cumulative probability for each possible product label as:

[0171]

[0172] Among them, δ e1 (l) represents the observation at the given e1th position (i.e., product label candidate t” h,e1) condition, the probability of the best path to the e1th position with product label l; δ e1-1 (l') represents the probability of the best path to the e1-1th position with product label l' given the observation of the e1-1th position; Indicates that given the e1th product label candidate t" h,e1 The maximum cumulative probability of product label l under the condition of; Represents a pointer to the previous best product label; a l',l represents the transition probability from product label l' to product label l, calculated by the CRF layer; t' h,e1 Indicates the product label candidate at the e1th position; t" h,e1 Indicates the final product label of the e1th position;

[0173] S4.74. At the end of the sequence, select the product label with the highest cumulative probability as the last product label z m :

[0174] z m = argmax l δ m (l);

[0175] S4.75, starting from the last product tag, according to the backtracking pointer Go back step by step and reconstruct the entire optimal product label sequence:

[0176]

[0177] S4.76. After the reconstruction is completed, the optimal product label sequence T is obtained. h , which is the final extracted product label:

[0178] T” h =(z1, z2, ..., z m ).

[0179] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and descriptions are only preferred examples of the present invention, and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements all fall within the scope of the present invention to be protected.

Claims

1. A product label extraction method based on Internet big data and AI big language model, characterized in that: The following steps are involved: S1. Use crawler technology to capture text data of products on the Internet; S2. Use the TF-IDF algorithm to determine the important words in the text data, and combine it with the Skip-Gram model to capture the semantic associations between words. In the process of capturing the semantic associations between words, the weights reflecting the user's browsing frequency and the user's behavioral feature vector are introduced to optimize the capture process. S3, based on the extracted important words and the semantic association information between words, generate product labels using a large-scale pre-trained language model; S4. Combine the sequence annotation model BERT and conditional random field CRF to locate and classify product labels, and output the final extracted product labels.

2. The product label extraction method based on Internet big data and AI large language model according to claim 1 is characterized by: In S2, the TF-IDF algorithm is used to determine the important words in the text information, including the following steps: S2.

1. Preprocess the text data of the product and organize the captured text data into a corpus D, where D contains N documents d, D = {d1, d2, d3, ..., d N }; S2.

2. Calculate the term frequency TF(t, d) of each word t in document d: Where cou(t, d) represents the number of occurrences of word t in document d, ∑ w∈d (w, d) represents the total number of occurrences of all words in document d; w represents each word in document d; S2.

3. Calculate the inverse document frequency IDF(t, D) of word t in the entire corpus: Where N is the total number of documents in the corpus, |{d∈D:t∈d}| is the number of documents containing word t; S2.

4. Combine the term frequency and the inverse document frequency to get the importance weight of term t in document d: TF-IDF(t, d, D): TF-IDF(t,d,D)=TF(t,d)·IDF(t,D).

3. The product label extraction method based on Internet big data and AI big language model according to claim 2 is characterized by: In S2, the Skip-Gram model is combined to capture the semantic association between words, including the following steps: S2.

5. Build a vocabulary based on the preprocessed text data to record all unique words and their frequencies that appear in the text; S2.6, set a fixed window size c, for each central word, take c words on the left and right of the central word as context words to form a training pair; S2.

7. Randomly initialize a low-dimensional vector for each word in the vocabulary and set hyperparameters. S2.

8. Set the objective function to train the Skip-Gram model, assign weights to product word vectors that reflect browsing frequency, and transform the user behavior feature vector u i and u o Introduce it into the objective function for optimization, and introduce the related products as additional positive examples into the objective function for further optimization; S2.

9. The trained Skip-Gram model maps each word into a high-dimensional vector space to obtain the word vector v for each word. t ; S2.

10. Combining word frequency, inverse document frequency and word vector v t , forming a new product word vector v d : S2.

11. Calculate the cosine similarity of word vectors to measure the semantic similarity between words. For words with similar contexts, introduce a clustering method in the calculation process of the cosine similarity of word vectors. Through clustering, semantically similar words are classified into the same category.

4. The product label extraction method based on Internet big data and AI big language model according to claim 3 is characterized by: In S2.8, the objective function is: Among them, L neg represents the optimized objective function; σ(*) represents the Sigmoid function; v o The vector representation of the context word; v i represents the vector representation of the central word; k represents the number of negative samples; v j Represents the vector representation of negative sample words; w j Represents negative sample words; represents the expected value of a random draw from a vocabulary according to a certain vocabulary distribution P(w); P(w) represents the vocabulary distribution; T represents the transpose of the vector; i represents the index of the center word; o represents the index of the context word; j represents the index of the negative sample word; When a product frequently appears in the user's browsing path, a weight reflecting the browsing frequency is given to the product word vector, a user behavior matrix is ​​established, and two user behavior feature vectors u are extracted from the user behavior matrix. i and u o , and the user behavior feature vector u i and u o Introduce into the objective function for optimization: Among them, L neg,uesr represents the objective function after optimizing the user's browsing path; u i Represents the user behavior feature vector corresponding to the target word; u o Represents the user behavior feature vector corresponding to the context word; bp u,r represents the browsing frequency of user u on the central product r; bp u,r1 represents the browsing frequency of user u on negative sample product r1; If multiple products often appear in the browsing path of the same user, then the products are considered to be related. For each center word v i , find products that are often viewed by the same user group based on the user's browsing history, and introduce the associated products as additional positive examples into the objective function for further optimization: Among them, L neg,be represents the objective function after further optimization; B represents a set of additional positive examples selected based on the user's browsing path; v b Represents the word vector corresponding to the additional positive example; u b represents the user behavior feature vector corresponding to the additional positive example; α represents the hyperparameter that controls the contribution of the user behavior positive example; b represents the index of the additional positive example.

5. The product label extraction method based on Internet big data and AI big language model according to claim 4 is characterized by: In S2.11, the cosine similarity of word vectors is calculated to measure the semantic similarity between words: Among them, Sim(v c , v e ) represents the cosine similarity of word vectors; v c The word vector representation of word c; v e The word vector representation of word e; For words with similar contexts, a clustering method is introduced in the calculation process of the cosine similarity of word vectors. Through clustering, semantically similar words are classified into the same category: Among them, Sim1(v c , v e ) represents the cosine similarity of the optimized word vector; CL(v c , v e ) represents the similarity calculated based on vocabulary clustering information; γ represents the weighting coefficient of clustering information.

6. The product label extraction method based on Internet big data and AI big language model according to claim 5 is characterized by: In S3, based on the extracted important words and the semantic association information between the words, a large-scale pre-trained language model is used to generate product labels, including the following steps: S3.

1. Construct an input sequence X=(x1, x2, ..., x3) based on the extracted important words and the semantic association information between words. n ), and input the input sequence into the pre-trained language model; S3.

2. The pre-trained language model predicts the next word at each position through autoregression to generate a complete text: in, Indicates the generated product label; y n+1 represents the next word generated; P(y n+1 |X) indicates that the model calculates the next word y by autoregression n+1 The conditional probability distribution of ; S3.3, generating a probability distribution of product labels based on the input sequence and the trained model parameters θ; S3.

4. Use temperature sampling strategy to extract label candidates from conditional probability distribution.

7. The product label extraction method based on Internet big data and AI big language model according to claim 6 is characterized by: In S3.3, the probability distribution of product labels is: in, Represents the generated product label The conditional probability distribution of ; n represents the length of the input sequence; t represents the position index in the product label generation process.

8. The product label extraction method based on Internet big data and AI big language model according to claim 7 is characterized by: In S4, the sequence tagging model BERT and the conditional random field CRF are combined to locate and classify product labels, including the following steps: S4.

1. Arrange the product label candidates generated by the pre-trained language model; S4.

2. Create a corresponding feature vector T′ for each product label candidate h ={t″ h,1 , t′ h,2 ,…,t′ h,m }, where T′ h represents the product label set after preliminary screening, t′ h,m represents the mth product label candidate; S4.

3. Product label candidate sequence T′ h Input into the BERT model, the BERT model encodes each product label candidate according to the context and captures the semantic features of the product label candidate: Among them, R m represents the product label candidate t′ h,m BERT encoding result, Represents the parameters of the BERT model; S4.

4. Define product label space L = {l1, l2, ..., l k }, where k represents the index of the product label category; S4.

5. Apply the conditional random field layer to calculate the transition probability matrix A between labels. The transition probability matrix represents the probability of transferring from one product label to another product label: A=[a hm ]; Among them, [a hm ] indicates product label space l g Transfer to product label space k The transition probability, represents the parameters of the conditional random field layer; g represents the index of the product label category; S4.

6. BERT encoding result R m Combined with the transition probability of the conditional random field layer, the optimal label sequence is determined: Among them, T″ h Represents the finalized product label set; t″ h,m represents the mth final product label; m1 represents the total length of the product label sequence; S4.

7. Use the Viterbi algorithm to find the product label sequence with the highest probability.

9. The product label extraction method based on Internet big data and AI big language model according to claim 8 is characterized by: In S4.7, the Viterbi algorithm is used to find the product label sequence with the highest probability, including the following steps: S4.

71. For the first product label candidate, calculate the initial probability of each possible product label: Where l represents one of the product label categories; δ1(l) represents the first product label candidate t″ h,1 The initial probability of product label l under the condition of ; S4.

72. Initialize the backtracking pointer for each product tag S4.

73. For the e1th product label candidate in the sequence, calculate the maximum cumulative probability of each possible product label and update the backtracking pointer; S4.

74. At the end of the sequence, select the product label with the highest cumulative probability as the last product label z m : With m =argmax l δ m (l); S4.75, starting from the last product tag, according to the backtracking pointer Go back step by step and reconstruct the entire optimal product label sequence: S4.

76. After the reconstruction is completed, the optimal product label sequence T″ is obtained h , which is the final extracted product label: T″ h =(z1,z2,…,z m )。 10. The product label extraction method based on Internet big data and AI big language model according to claim 9 is characterized by: In S4.73, the maximum cumulative probability of each possible product label is calculated as: Among them, δ e1 (l) represents the probability of the best path to the e1th position with product label l given the observation condition of the e1th position; δ e1-1 (l′) represents the probability of the best path to the e1-1th position with product label l′ given the observation of the e1-1th position; Indicates that given the e1th product label candidate t″ h,e1 The maximum cumulative probability of product label l under the condition of; Represents a pointer to the previous best product label; a l′,l represents the transition probability from product label l′ to product label l; t′ h,e1 Indicates the product label candidate at the e1th position; t″ h,e1 Indicates the final product label for the e1th position.

Citation Information

Patent Citations

  • Intelligent customer service system text matching method based on Word2Vector and TF-IDF

    CN115168559A

  • Intelligent traffic text analysis method based on natural language processing

    CN115934936A

  • BERT-BiLSTM-CRF-based ship named entity identification method

    CN117744658A

  • Text classification method based on particle swarm optimization and CNN (Convolutional Neural Network)

    CN117891939A

  • Untoward drug reaction monitoring and early warning method

    CN118280603A

Cited By

  • Label generation method, device, equipment and product based on large model and configuration

    CN120508659A