Content processing method and device, readable medium, electronic equipment and program product

Through the machine learning model combining category information and multiple rounds of clustering algorithms, the problem of inaccurate classification and summary of attribute words is solved, more efficient and accurate object analysis is achieved, and the content promotion effect of the brand or industry is improved.

CN120353935APending Publication Date: 2025-07-22BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510489137.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When extracting attribute words from object-related content in the prior art, words with high semantic similarity cannot be accurately classified and summarized, resulting in low classification precision, affecting the efficiency and accuracy of subsequent object analysis.

Method used

Through the machine learning model, multiple attribute words are classified and summarized through the category information of the category to which the target object belongs. High-dimensional coding model and multi-round clustering algorithm are adopted, including LP unsupervised clustering, China Unicom component unsupervised clustering and large language model to conduct semantic understanding and similarity verification, and optimize the classification and induction process of attribute words.

Benefits of technology

It improves the classification precision and accuracy of attribute words, enhances the efficiency and accuracy of object analysis, helps brands or industries understand consumer cognition, and optimizes content promotion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353935A_ABST
    Figure CN120353935A_ABST
Patent Text Reader

Abstract

The invention discloses a content processing method and device, a readable medium, electronic equipment and a program product. The content processing method comprises the steps of obtaining a plurality of first attribute words used for describing a target object; classifying and summarizing the plurality of first attribute words at least based on category information of a category to which the target object belongs through a machine learning model to obtain a plurality of second attribute words used for describing the target object, determining a target attribute word group based on the plurality of second attribute words, and processing at least based on semantics of each attribute word in the target attribute word group to obtain a target attribute word group; and obtaining a target attribute word for describing the target object. The attribute words are subjected to finer-dimension semantic understanding in combination with the categories to which the objects belong, so that complex semantics of the attribute words are obtained, accurate classification and semantic induction of the attribute words are facilitated, and the attribute words in the classified and induced attribute word groups are subjected to semantic processing to obtain target attribute words. Complex semantics in the attribute words can be further captured, so that more accurate and more representative attribute words are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of large language models, machine learning models, and computer technologies. Specifically, it relates to a content processing method, apparatus, readable medium, electronic device, and program product. Background Art

[0002] In the related art, it is possible to deeply analyze an object itself based on the vocabulary extracted from the object-associated content. However, among the vast number of extracted vocabulary, there may be a large number of semantically similar words, which need to be classified and summarized. For example, when extracting the attribute words of the content-associated object from the object-associated content, usually a single clustering model is used to classify and summarize the attribute words, which may not be able to capture the complex semantics of the attribute words, resulting in low classification fineness, and easily leading to low accuracy of the summarized attribute words, thereby affecting the efficiency and accuracy of subsequent object analysis. Summary of the Invention

[0003] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the subsequent Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.

[0004] In a first aspect, the present disclosure provides a content processing method, the content processing method comprising: obtaining a plurality of first attribute words for describing a target object, the plurality of first attribute words being obtained based on target content associated with the target object; classifying and summarizing the plurality of first attribute words by a machine learning model at least based on category information of the category to which the target object belongs, obtaining a plurality of second attribute words for describing the target object, determining a target attribute word group based on the plurality of second attribute words, and processing at least based on the semantics of each attribute word in the target attribute word group to obtain a target attribute word for describing the target object.

[0005] In a second aspect, the present disclosure provides a content processing apparatus, the content processing apparatus comprising: an obtaining module, configured to obtain a plurality of first attribute words for describing a target object, the plurality of first attribute words being obtained based on target content associated with the target object; a processing module, configured to classify and summarize the plurality of first attribute words by a machine learning model at least based on category information of the category to which the target object belongs, obtain a plurality of second attribute words for describing the target object, determine a target attribute word group based on the plurality of second attribute words, and process at least based on the semantics of each attribute word in the target attribute word group to obtain a target attribute word for describing the target object.

[0006] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, and when the program is executed by a processing device, the steps of the method described in the first aspect are implemented.

[0007] In a fourth aspect, the present disclosure provides an electronic device, including: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the method described in the first aspect.

[0008] In a fifth aspect, the present disclosure provides a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0009] Through the above technical solutions, multiple first attribute words can be classified and summarized at least based on the category information of the target object's category by a machine learning model, and multiple second attribute words for describing the target object are obtained. In this way, the semantic understanding of the attribute words can be carried out in a finer dimension in combination with the object's category to obtain the complex semantics of the attribute words, which is convenient for accurate classification and semantic induction of the attribute words. Furthermore, a target attribute word group is determined based on the multiple second attribute words, and processing is performed at least based on the semantics of each attribute word in the target attribute word group to obtain a target attribute word for describing the target object. By performing semantic processing on the attribute words in the classified and summarized attribute word group to obtain the target attribute word, the complex semantics in the attribute words can be further captured, and then more accurate and representative attribute words can be obtained. In an object analysis scenario, based on the induced target attribute words, the efficiency and accuracy of object analysis can be improved.

[0010] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Combined with the drawings and referring to the following specific implementation manners, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale. In the drawings: Figure 1 is a flowchart of a content processing method shown according to an exemplary embodiment of the present disclosure; Figure 2 is a schematic diagram of the process of a secondary clustering method shown according to an exemplary embodiment of the present disclosure; Figure 3 is a schematic diagram of the process of a split-cluster clustering method shown according to an exemplary embodiment of the present disclosure; Figure 4 is a schematic diagram of a multi-round clustering method shown according to an exemplary embodiment of the present disclosure; Figure 5 is a schematic diagram of a preset attribute definition shown according to an exemplary embodiment of the present disclosure; Figure 6 is a schematic diagram of a similarity verification process shown according to an exemplary embodiment of the present disclosure; Figure 7 is a schematic diagram of an architecture for classification and induction shown according to an exemplary embodiment of the present disclosure; Figure 8 is a block diagram of a content processing device shown according to an exemplary embodiment of the present disclosure; Figure 9 is a schematic diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0012] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0013] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0014] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0015] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0016] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0017] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of these messages or information.

[0018] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0019] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operations of the technical solutions of the present disclosure according to the prompt message.

[0020] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of the present disclosure, and other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present disclosure.

[0022] At the same time, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0023] In the related art, with the user's authorization, it is possible to deeply analyze the object itself based on the vocabulary extracted from the object-associated content. However, among the massive extracted vocabulary, there may be a large number of words with similar semantics, and it is necessary to classify and summarize them. However, since different words may have different meanings in different fields, in the face of complex and diverse semantics, a single clustering model is difficult to capture the complex meanings of words such as polysemous words and synonyms in different contexts, and thus it is impossible to accurately classify and summarize them.

[0024] For example, when extracting the attribute words of the content-associated object from the object-associated content, a large number of attribute words with high semantic similarity appear. The clustering model usually classifies and summarizes the attribute words based on the semantic similarity between the attribute words. For example, "good complexion" and "nice complexion" have a relatively large semantic distance determined based on word vectors, but in the actual business scenario, they describe the same attribute. In addition, the clustering model lacks the ability of context awareness and cannot capture the specific meanings of polysemous words and synonyms. Therefore, it may cause similar words not to be clustered together, forming miscellaneous clustering clusters with low classification fineness, which easily leads to low accuracy of the induced attribute words, thereby affecting the efficiency and accuracy of subsequent object analysis.

[0025] In view of this, the present disclosure provides a content processing method, apparatus, readable medium, electronic device, and program product to solve the above technical problems.

[0026] The following further explains the embodiments of the present disclosure with reference to the accompanying drawings.

[0027] Figure 1 is a flowchart of a content processing method shown according to an exemplary embodiment of the present disclosure. Referring to Figure 1 The method may include the following steps: S101: Obtain a plurality of first attribute words for describing a target object, where the plurality of first attribute words are obtained based on target content associated with the target object.

[0028] In this embodiment, the target object may be a brand or a category to which the brand belongs, or it may be a commodity or a live broadcast room. Of course, it may also be other objects, and the embodiments of the present disclosure do not impose any restrictions on this. The target content may be content in the form of short videos, long videos, graphics, texts, audios, etc. obtained under the authorization of the user. Taking the content associated with the brand as an example, it may include promotional videos, promotional articles, and other promotional content published by the brand side or creators such as anchors and bloggers who cooperate with the brand side on social media platforms, as well as comment content of users on the promotional content or evaluation content of users after purchasing the brand's products and search content of users on social media platforms. The embodiments of the present disclosure do not impose any restrictions on this.

[0029] S102: Classify and summarize the plurality of first attribute words through a machine learning model based at least on the category information of the category to which the target object belongs to obtain a plurality of second attribute words for describing the target object, determine a target attribute word group based on the plurality of second attribute words, and perform processing based at least on the semantics of each attribute word in the target attribute word group to obtain a target attribute word for describing the target object.

[0030] Exemplarily, by combining the category to which the object belongs, a more refined semantic understanding of the attribute words can be achieved, so as to accurately semantically define and analyze the attribute words in different category scenarios, and improve the accuracy and fineness of classifying and summarizing the attribute words.

[0031] Among them, taking the target object as a brand as an example, the category to which the brand belongs can refer to the industry to which the brand belongs, such as industries like cosmetics, electronic products, food, automobiles, or smart homes, etc. The present disclosure places no restrictions on this.

[0032] In this embodiment, when the target object is a brand, the target attribute words of the target object can refer to the summary of various description contents of the brand or product from the brand's perspective by the brand side, such as description contents like the product functions of the brand, the origin of the product, the certification information of the brand, etc., or the summary of various description contents of the brand or product from the user's perspective, such as brand impression, product evaluation, etc. The target attribute words can also be called the mental words of brand mind. Among them, brand mind is used to represent the cognition and impression formed by consumers of the brand, and the mental words are the concrete expressions of brand mind, which are the key information that the brand hopes consumers will remember, covering aspects such as the personality characteristics of the brand, product advantages, and brand style. For example, the first attribute words can be the original attribute words directly extracted from the target content, such as "0 sucrose", "sugar-free", "zero sugar", "0 sugar", "no added sugar", "sugar-free", and the second attribute words can be obtained by classifying and summarizing multiple first attribute words. For example, "0 sucrose", "sugar-free", "zero sugar", "0 sugar" are summarized to get "zero sugar", and "no added sugar", "sugar-free" are summarized to get "sugar-free". The target attribute words can be "zero sugar" obtained by further grouping and summarizing the second attribute words, and can be specifically summarized according to the actual business scenario. The embodiments of the present disclosure place no restrictions on this.

[0033] Correspondingly, when the target object is an industry, the target attribute words of the target object can refer to the summary of various description contents of the brands under this industry for the brand or product, or the summary of various description contents of the brands under this industry for the brand or product by users, and can also be called the mental words of industry mind.

[0034] Subsequently, based on brand mind or industry mind, it can help the brand side more comprehensively and accurately understand the cognitive information of consumers about the brand or industry, and then optimize the promotion content and improve the content promotion efficiency.

[0035] By adopting the above method, it is possible to conduct a more detailed semantic understanding of the attribute words in combination with the category to which the object belongs, so as to obtain the complex semantics of the attribute words, which is convenient for accurate classification and semantic induction of the attribute words. And by performing semantic processing on the attribute words in the classified and inducted attribute word group, the target attribute words are obtained, which can further capture the complex semantics in the attribute words, and thus more accurate and representative attribute words can be obtained. In the object analysis scenario, based on the inducted target attribute words, the efficiency and accuracy of object analysis can be improved.

[0036] To facilitate the understanding of the content processing method provided by the present disclosure, the possible implementation manners in the present disclosure are described below.

[0037] In a possible manner, through a machine learning model, at least based on the category information of the target object, multiple first attribute words are classified and inducted to obtain multiple second attribute words for describing the target object, including: the machine learning model processes the multiple first attribute words as follows: respectively performing vector processing on the multiple first attribute words based on the category information of the target object to obtain multiple first word vectors; classifying the multiple first attribute words according to the multiple first word vectors to obtain multiple initial attribute word groups; and inducting the first attribute words in each initial attribute word group to obtain multiple second attribute words for describing the target object.

[0038] Exemplarily, in order to improve the accuracy of semantic understanding, the machine learning model can perform vector processing on the first attribute words in combination with the category information of the target object. The obtained word vectors can provide richer and more detailed semantic information compared with the word vectors obtained solely based on the attribute words, which is convenient for the model to more accurately capture the commonalities of multiple attribute words and more clearly divide the semantic boundaries.

[0039] Exemplarily, "smooth" can be classified as a texture attribute in the home furnishing industry, and then grouped into an attribute word group with other attribute words describing texture such as "shining" and "flat". In the beauty industry, it can be classified as an efficacy attribute and grouped into an attribute word group with other attribute words describing efficacy such as "delicate" and "silky", and so on. The present disclosure does not limit this. By combining the category information, the semantic understanding of the attribute words can be more accurate, and thus the classification and induction of the attribute words can be more accurate, obtaining more representative attribute words, such as attribute words that can cover the semantics of all attribute words in the attribute word group, or attribute words that can cover the semantics of the most attribute words in the attribute word group, or attribute words with the highest frequency of occurrence in the attribute word group, and so on. The present disclosure does not limit this.

[0040] In the embodiments of the present disclosure, a high-dimensional coding model can be adopted. Compared with a low-dimensional coding model, the high-dimensional coding model can convert high-dimensional and discrete text data into low-dimensional and continuous numerical vectors, enabling subsequent clustering models to directly process them, and being able to capture text semantic information, outputting vectors with closer distances for similar texts, reducing the computational complexity of the clustering model and improving the clustering effect.

[0041] In a possible way, the first attribute words in each initial attribute phrase are generalized to obtain multiple second attribute words for describing the target object, including: generalizing at least based on the semantics of the first attribute words in each initial attribute phrase to obtain multiple second attribute words for describing the target object. Determining the target attribute phrase based on the multiple second attribute words includes: performing vector processing on the multiple second attribute words respectively based on the category information of the category to which the target object belongs to obtain multiple second word vectors; classifying the multiple second attribute words at least according to the multiple second word vectors to obtain the target attribute phrase.

[0042] Exemplarily, the phenomenon of repetition of the finally output attribute words can be reduced by adding secondary clustering. Because after the first clustering and generalization are completed, some attribute words with similar semantics will appear together. By further performing a clustering operation, the attribute words with similar semantics are grouped into one category and generalized through a generalization model to obtain the final target attribute words, such as an attribute word that can cover all the semantics of the attribute words within the attribute phrase, or an attribute word that can cover the most semantics of the attribute words within the attribute phrase, etc., and the present disclosure does not limit this.

[0043] Exemplarily, as Figures 2-4 shown, in each clustering process, the attribute words can be vector-processed by a coding model in combination with the category information of the category to which the target object belongs to obtain word vectors with a fixed length, and then the distance between every two word vectors is calculated. For example, the Euclidean distance can be calculated using the Cartesian product, and the present disclosure does not limit this. The subsequent coding process and distance calculation process will not be elaborated further.

[0044] In a possible way, classifying the multiple first attribute words according to the multiple first word vectors to obtain multiple initial attribute phrases includes: clustering the multiple first attribute words based on the multiple first word vectors by a fifth clustering algorithm to obtain multiple initial attribute phrases, and the fifth clustering algorithm is used for clustering based on the modularity of the multiple first word vectors. Classifying the multiple second attribute words at least based on the multiple second word vectors to obtain the target attribute phrase includes: clustering the multiple second attribute words based on the multiple second word vectors by a sixth clustering algorithm to obtain the target attribute phrase, and the sixth clustering algorithm is used for clustering based on the connected components of the multiple second word vectors.

[0045] Exemplarily, as Figure 2As shown, edges can be constructed between two attribute words whose distances are less than a preset threshold, so as to perform clustering based on the constructed edges. The attribute words obtained from the first classification and induction are continued to be vector-processed to obtain corresponding word vectors, and then distance calculation, edge construction, and classification and induction are performed to obtain the final attribute words. The encoding processing, distance calculation, and induction processes of the two clusterings are similar, except that there are differences in the clustering algorithms and clustering processes.

[0046] It should be understood that, as Figure 2 shown, the first threshold in the first clustering and the second threshold in the second clustering can be set according to requirements, but the first threshold can be greater than the second threshold, indicating that in the second clustering process, the distances between the attribute words that can be clustered are closer, which means the semantics are closer, so that secondary clustering of the attribute words with closer semantics can be achieved.

[0047] Exemplarily, as Figure 2 shown, in the first clustering, an algorithm for clustering based on the modularity of multiple first word vectors can be used, such as the LP (Label Propagation) unsupervised clustering algorithm. The LP unsupervised clustering algorithm is suitable for clustering large-scale similar data. In the second clustering, an algorithm for clustering based on the connected components of multiple second word vectors can be used, such as the connected components unsupervised algorithm. The connected components unsupervised algorithm can further cluster high-similarity data based on the connected components, and can be specifically selected according to requirements. The embodiments of the present disclosure do not impose any restrictions on this.

[0048] It should be understood that first, clustering is performed based on the LP unsupervised clustering algorithm. The LP unsupervised clustering algorithm is good at capturing global semantic relationships and can quickly cluster large-scale and similar first attribute words. The data volume of the second clustering is smaller than that of the first clustering, but the similarity between the second attribute words is higher. If clustering continues based on the similarity between the attribute words, the effect is not good. Therefore, the connected components unsupervised clustering algorithm can be used for secondary clustering. The connected components unsupervised clustering algorithm is good at capturing local semantic relationships, so that the attribute words can be clustered and induced based on different dimensions, improving the accuracy of attribute word induction.

[0049] It should be noted that due to certain limitations in the semantic measurement of the encoding model, the clustering restrictions can be appropriately relaxed, and then compensated by the similarity model. As Figure 2As shown, in a possible way, in the step of classifying a plurality of second attribute words according to at least a plurality of second word vectors to obtain a target attribute phrase group, the plurality of second attribute words can be classified according to the plurality of second word vectors to obtain a candidate attribute phrase group, and then the candidate attribute phrase group is subjected to similarity verification based on a similarity model, and the candidate attribute phrase group that passes the similarity verification is determined as the target attribute phrase group.

[0050] Exemplarily, the similarity model is used to determine the semantic similarity between the input attribute words. By converting the attribute words into vector representations and calculating the similarity scores between them, such as cosine similarity, to determine whether they belong to the same attribute group, it can be implemented based on a trained large language model. For example, the model can be trained with attribute word samples marked with semantic similarity. The present disclosure does not limit this.

[0051] It should be understood that simply measuring the semantic distance between two attribute words based on the vector distance between them may not conform to the semantic distance in the business scenario. By adding similarity verification, it can be further verified whether the attribute words belong to the same attribute phrase group, improving the classification accuracy of the attribute words.

[0052] The above secondary clustering method based on the LP unsupervised clustering algorithm and the connected component unsupervised clustering algorithm can cluster and summarize attribute words based on different dimensions, improving the accuracy of attribute word summarization and reducing the repetition of the finally output attribute words.

[0053] In a possible way, classifying a plurality of first attribute words according to a plurality of first word vectors to obtain a plurality of initial attribute phrase groups includes: clustering the plurality of first attribute words based on a third clustering algorithm according to the plurality of first word vectors to obtain a plurality of initial attribute phrase groups. Classifying a plurality of second attribute words according to at least a plurality of second word vectors to obtain a target attribute phrase group includes: clustering the plurality of second attribute words based on a fourth clustering algorithm according to the plurality of second word vectors to obtain a target attribute phrase group; wherein, the third clustering algorithm and the fourth clustering algorithm are used to perform clustering based on the graph information between word vectors, and the walking scale of the third clustering algorithm is greater than the walking scale of the fourth clustering algorithm.

[0054] Exemplarily, the third clustering algorithm and the fourth clustering algorithm can be the Infomap unsupervised clustering algorithm (information graph algorithm). The Infomap unsupervised clustering algorithm is a community discovery algorithm based on information theory and random walk model, which can identify the optimal community partition by minimizing the coding length describing the network information flow. The specific parameters can be set according to requirements. The present disclosure does not limit this. Through the Infomap unsupervised clustering algorithm, it is possible to avoid assigning noise words to irrelevant clustering clusters when there are noises in the word vectors corresponding to the attribute words.

[0055] Exemplarily, as Figure 3 shown, during the two clustering processes, the Euclidean distance can be calculated through the Cartesian product to measure the distance between two word vectors, determine whether the words belong to the same set of similar words, construct a similarity network for the attribute words greater than a preset threshold. The third threshold is greater than the fourth threshold. Each node in the similarity network represents an attribute word. Run the Infomap unsupervised clustering algorithm on the similarity network, and divide the clustering clusters based on the information flow theory, clustering the closely connected nodes into the same cluster.

[0056] Exemplarily, the walk scale of the third clustering algorithm and the walk scale of the fourth clustering algorithm can be set according to requirements, and the walk scale of the third clustering algorithm is greater than the walk scale of the fourth clustering algorithm. In this way, the similarity network constructed during the first clustering is relaxed, and more similar attribute words can be divided into the same attribute word group to form a larger clustering cluster, which is convenient for subsequent cluster splitting. A smaller clustering cluster is formed during the second clustering to further refine the attribute word group.

[0057] In a possible way, at least generalize based on the semantics of the first attribute word in each initial attribute word group to obtain multiple second attribute words for describing the target object, including: generalizing based on the semantics of the first attribute word in each initial attribute word group to obtain multiple second candidate attribute words; for each second candidate attribute word, perform a third verification on the second candidate attribute word and the first attribute word in the corresponding initial attribute word group, and the third verification is used to verify whether the semantics of the second candidate attribute word and the first attribute word in the corresponding initial attribute word group are consistent; for each initial attribute word group, determine the first attribute word and the second candidate attribute word that pass the third verification in the initial attribute word group as the third attribute word group, and determine the second candidate attribute word as the second attribute word of the third attribute word group, determine the first attribute word that fails the third verification in the initial attribute word group as the fourth attribute word group, and use the third attribute word group as the new initial attribute word group, and repeat the step of generalizing based on the semantics of the first attribute word in each initial attribute word group to obtain multiple second candidate attribute words for describing the target object until all the attribute words in the new initial attribute word group pass the third verification.

[0058] Exemplarily, as Figure 3As shown, during the first clustering process, semantic induction can be performed on the initial attribute phrase group to obtain candidate attribute words corresponding to the initial attribute phrase group. Then, a consistency model is used to perform consistency verification on the candidate attribute words and the attribute words within the initial attribute phrase group. The candidate attribute words and the attribute words within the initial attribute phrase group that pass the consistency verification are used as an attribute phrase group. The attribute words within the initial attribute phrase group that do not pass the consistency verification continue to undergo semantic induction to obtain new candidate attribute words until all the attribute words within the initial attribute phrase group pass the consistency verification, resulting in one or more attribute phrase groups and the corresponding second attribute words for each attribute phrase group. By splitting the large clusters obtained from the first clustering, the classification fineness and accuracy of the attribute phrase groups can be improved, and the attribute phrase groups that cannot be induced into attribute words can be broken up into independent attribute phrase groups waiting for secondary clustering, thereby improving the accuracy of subsequent target attribute word induction.

[0059] Among them, the conditions for passing the consistency verification can be set according to requirements. For example, it can be that the candidate attribute word can contain all the attribute words within the initial attribute phrase group, or the inclusion degree of the candidate attribute word for the attribute words within the initial attribute phrase group is greater than a preset threshold. The present disclosure does not limit this.

[0060] In a possible manner, the target attribute phrase group is obtained by classifying the multiple second attribute words at least based on multiple second word vectors, including: performing a fourth verification on every two second attribute words, where the fourth verification is used to verify whether the semantic similarity between the two second attribute words is greater than a second similarity threshold; in response to the two second attribute words passing the fourth verification, setting the classification weight between the two second attribute words to 1, and in response to the two second attribute words not passing the fourth verification, setting the classification weight between the two second attribute words to the word vector distance between the two second attribute words; classifying the multiple second attribute words according to the classification weights between every two second attribute words to obtain the target attribute phrase group.

[0061] Exemplarily, as Figure 3 shown, during the second clustering process, a similarity model can be used to perform similarity verification on every two attribute words in the similarity network. For the attribute words that pass the similarity verification, the classification weight between the two can be set to 1, and for the attribute words that do not pass the similarity verification, the classification weight between the two can be set to the corresponding word vector distance. The word vector distance can be normalized to the 0 - 1 interval. In this way, during subsequent clustering, the probability of classifying the attribute words that pass the similarity verification into the same attribute phrase group can be increased, and the probability of classifying the attribute words that do not pass the similarity verification into the same attribute phrase group can be reduced, further improving the classification precision and accuracy of the target attribute phrase group.

[0062] It should be noted that since the encoding model is asymmetrically trained, the reliability of the similarity network based on clustering is not strong. It can only be compared relatively and cannot be compared absolutely. Simply relying on the word vector distance between attribute words is likely to result in the formation of miscellaneous clustering clusters. For example, attribute words with similarity are not clustered into the same attribute phrase group. The above-mentioned clustering method that combines the Infomap unsupervised clustering algorithm and the cluster splitting process can verify the semantic consistency of attribute words within the same clustering cluster through a consistency model, eliminate abnormal words, avoid the formation of miscellaneous attribute phrase groups, improve the accuracy and precision of attribute phrase groups, and thus improve the accuracy of target attribute word induction.

[0063] In a possible way, induction is performed at least based on the semantics of the first attribute words in each initial attribute phrase group to obtain multiple second attribute words for describing the target object, including: when the multiple first attribute words do not belong to the attribute words under the target attribute classification, induction is performed at least based on the semantics of the first attribute words in each initial attribute phrase group to obtain multiple second attribute words for describing the target object. The target attribute classification is a preset classification that cannot perform semantic induction in the preset attribute definition. The preset attribute definition is obtained by defining the attributes under each classification, and the attributes under each classification are obtained by classifying the attributes of the category to which the attribute description object belongs. Classify the multiple first attribute words according to the multiple first word vectors to obtain multiple initial attribute phrase groups, including: clustering the multiple first attribute words based on the multiple first word vectors through the first clustering algorithm to obtain multiple initial attribute phrase groups. Classify the multiple second attribute words at least based on the multiple second word vectors to obtain target attribute phrase groups, including: clustering the multiple second attribute words based on the multiple second word vectors through the second clustering algorithm to obtain target attribute phrase groups; wherein, the first clustering algorithm and the second clustering algorithm are used for clustering based on the graph information between word vectors, and the walking scale of the first clustering algorithm is smaller than the walking scale of the second clustering algorithm.

[0064] It should be noted that the target attribute classification can be an attribute classification such as the place of origin, product pronoun, etc. that cannot perform semantic induction, which can be specifically determined according to the actual business scenario, and the present disclosure does not limit this. The clustering granularity of such attributes needs to be very fine, and clustering based on semantics is likely to result in inaccurate induced attribute words. For example, when the product pronoun is "blue bottle shampoo", "small blue bottle shampoo", "green bottle shampoo", "small green bottle shampoo", it may be directly induced into "bottled shampoo". However, in the actual business scenario, the product pronouns representing shampoos of different colors may correspond to shampoos with different effects and do not require semantic induction.

[0065] Exemplarily, such as Figure 4As shown, when multiple first attribute words do not belong to the attribute words under the target attribute classification, multiple rounds of clustering can be performed on the multiple first attribute words. Among them, the first clustering algorithm and the second clustering algorithm can be the Infomap unsupervised clustering algorithm. During the first clustering and the second clustering, the Euclidean distance can be calculated through the Cartesian product to measure the distance between two word vectors, determine whether the words belong to the same group of similar words, and construct a similarity network for the attribute words greater than a preset threshold, where the fifth threshold is greater than the sixth threshold. Furthermore, run the Infomap unsupervised clustering algorithm on the similarity network, divide the clustering clusters based on the information flow theory, and cluster the tightly connected nodes into the same cluster.

[0066] Exemplarily, the walking scale of the first clustering algorithm and the walking scale of the second clustering algorithm can be set according to requirements, and the walking scale of the first clustering algorithm is less than the walking scale of the second clustering algorithm. This is because when clustering for the first time, a smaller walking scale can form smaller clustering clusters, making the similarity of the attribute words within the same clustering cluster higher. When clustering for the second time, a larger walking scale is convenient for clustering the attribute words obtained by summarizing different attribute word groups into the same attribute word group, which is convenient for subsequent summarization of the target attribute words.

[0067] It should be noted that during the second clustering process, the process of determining the classification weight based on the similarity model in the above-mentioned secondary clustering method can also be referred to. In this way, when clustering subsequently, the probability of classifying the attribute words that pass the similarity verification into the same attribute word group can be increased, and the probability of classifying the attribute words that do not pass the similarity verification into the same attribute word group can be reduced, further improving the classification accuracy and precision of the target attribute word group.

[0068] It should be understood that in this embodiment, preset attribute definitions can be obtained by pre-defining the attributes for different category scenarios. For example, different attribute classifications can be set for different industries such as beauty, food, electronic products, and automobiles. For example, the beauty industry can involve product functions and ingredients, the food industry can involve taste and texture, and automobiles can involve vehicle functions and performance, as Figure 5 shown as the preset attribute definition under the beauty industry.

[0069] In other possible implementation manners, it can also be the leiden algorithm, the LPA (Label Propagation Algorithm) semi-supervised clustering algorithm, and so on. The present disclosure does not limit this. However, through the Infomap unsupervised clustering algorithm, the silhouette coefficient of clustering can be effectively improved, and thus the accuracy of subsequent induction of target attribute words can be improved.

[0070] Among possible ways, the first attribute word in each initial attribute phrase is generalized to obtain multiple second attribute words for describing the target object, including: when multiple first attribute words belong to the attribute words under the target attribute classification, the first attribute word with the highest frequency of occurrence in each initial attribute phrase is used as the second attribute word corresponding to the initial attribute phrase, so as to obtain multiple second attribute words for describing the target object. The target attribute classification is a preset classification that cannot be semantically generalized in the preset attribute definition. The preset attribute definition is obtained by defining the attributes under each classification, and the attributes under each classification are obtained by classifying the attributes of the category to which the attribute description object belongs. Determining the target attribute phrase based on the multiple second attribute words includes: determining the initial attribute phrase where each second attribute word is located as the target attribute phrase.

[0071] It should be noted that for the attribute words under the target attribute classification, when extracting the first attribute words from the target content, the attribute words under the target attribute classification can be marked, without generalization during the first clustering, or without naming the clustering clusters for such attribute words, and then independent similarity verification is performed on the unnamed clustering clusters and then clustering is carried out. The present disclosure does not limit this.

[0072] Exemplarily, as Figure 4 shown, when multiple first attribute words belong to the attribute words under the target attribute classification, there is no need to perform a second clustering on the first attribute words, but the first attribute word with the highest frequency of occurrence is used as the second attribute word corresponding to the initial attribute phrase, and the initial attribute phrase where each second attribute word is located is directly determined as the target attribute phrase. By not performing semantic classification and generalization on the attribute words under the target attribute classification, the induction requirements of attribute words in business scenarios that do not require semantic induction can be met.

[0073] Among possible ways, using the first attribute word with the highest frequency of occurrence in each initial attribute phrase as the second attribute word corresponding to the initial attribute phrase includes: using the first attribute word with the highest frequency of occurrence in each initial attribute phrase as the first candidate attribute word corresponding to the initial attribute phrase; performing a first verification on the first candidate attribute word and the third attribute words other than the first candidate attribute word in the corresponding initial attribute phrase, where the first verification is used to verify whether the semantics of the first candidate attribute word and the third attribute words are consistent; determining the fourth attribute words that pass the first verification among the third attribute words and the first candidate attribute word as the first attribute phrase, and determining the first candidate attribute word as the second attribute word of the first attribute phrase; using the fifth attribute words that do not pass the first verification among the third attribute words as the second attribute phrase, and using the second attribute phrase as the new initial attribute phrase, and repeating the step of using the first attribute word with the highest frequency of occurrence in each initial attribute phrase as the first candidate attribute word corresponding to the initial attribute phrase until all the attribute words in the new initial attribute phrase pass the first verification.

[0074] Exemplarily, continuing with the above-mentioned "xx Shampoo" as an example, in an actual business scenario, the product pronouns representing shampoos of different colors may correspond to shampoos with different effects. If the most frequently occurring attribute word is directly used as the inductive word, it may lead to incomplete inductive attribute words. Therefore, clustering can be performed based on consistency verification to improve the integrity of the inductive attribute words.

[0075] Exemplarily, continuing to refer to Figure 4 , for example, the occurrence frequency of "Blue Bottle Shampoo" is 10%, the occurrence frequency of "Small Blue Bottle Shampoo" is 40%, the occurrence frequency of "Green Bottle Shampoo" is 20%, and the occurrence frequency of "Small Green Bottle Shampoo" is 30%. Taking the most frequently occurring "Small Blue Bottle Shampoo" as the inductive word, and then performing consistency verification. "Blue Bottle Shampoo" and "Small Blue Bottle Shampoo" have consistent semantics and pass the consistency verification. "Green Bottle Shampoo" and "Small Green Bottle Shampoo" have inconsistent semantics with "Small Blue Bottle Shampoo" and do not pass the consistency verification. Then, "Green Bottle Shampoo" and "Small Green Bottle Shampoo" are used as new attribute word groups, and continue to select the most frequently occurring "Small Green Bottle Shampoo". "Green Bottle Shampoo" and "Small Green Bottle Shampoo" have consistent semantics and pass the consistency verification. In the scenario of inducting brand mental words, it can improve the accuracy and integrity of brand mental word induction.

[0076] In a possible way, at least process based on the semantics of each attribute word in the target attribute word group to obtain a target attribute word for describing the target object, including: performing vector processing on each attribute word in the target attribute word group respectively based on the category information of the category to which the target object belongs to obtain multiple third word vectors; for each target word vector among the multiple third word vectors, determine a preset number of candidate word vectors closest to the target word vector among the multiple third word vectors, and perform a second verification on the sixth attribute word corresponding to the target word vector and the seventh attribute word corresponding to the candidate word vector. The second verification is used to verify whether the semantic similarity between the sixth attribute word and the seventh attribute word is greater than the first threshold; in response to the sixth attribute word and the seventh attribute word passing the second verification, perform induction on the sixth attribute word and the seventh attribute word to obtain a target attribute word for describing the target object; in response to the sixth attribute word and the seventh attribute word not passing the second verification, determine the sixth attribute word and the seventh attribute word as the target attribute word for describing the target object.

[0077] Exemplarily, continuing to refer to Figure 4 , in the multi-round clustering process based on the Infomap unsupervised clustering algorithm, for the attribute words belonging to the target attribute classification, directly use the attribute word group where the most frequently occurring attribute word is located as the target attribute word group. For the attribute words not belonging to the target attribute classification, the attribute word group obtained by performing secondary clustering on the attribute words can be used as the target attribute word group.

[0078] Furthermore, as Figure 4 shown, the vector processing can be performed on each attribute word in the target attribute phrase by combining the category information of the category to which the target object belongs through an encoding model to obtain word vectors of a fixed length, and then the distance between every two word vectors can be calculated. For example, the Euclidean distance can be calculated using the Cartesian product, and the present disclosure places no limitation thereon.

[0079] Then, continuing to refer to Figure 4 , a preset number of attribute words with the closest word vector distance to each attribute word are determined to form an attribute word list, and the similarity verification is performed on the attribute words in the attribute word list through a similarity model, and the semantic induction is performed on the attribute words that pass the similarity verification to obtain the target attribute words, and the semantic induction is no longer performed on the attribute words that do not pass the similarity verification, and they can all be used as the target attribute words.

[0080] Thus, the classification and induction of the attribute words in the attribute phrase can be performed again, and further verification can be performed based on the similarity model. Especially for the case where no attribute words can be induced, the integrity of the induced attribute words can be ensured, and the accuracy and fineness of the classification and induction of the attribute words can be further improved.

[0081] The above multi-round clustering method based on the Infomap unsupervised clustering algorithm can optimize the induction of the attribute words under the target attribute classification. By gradually optimizing the clustering results from the clustering algorithm, clustering process, and handling of special clustering cases, the difference between words in different clustering clusters and the similarity between words in the same clustering cluster are finally ensured, and the accuracy and fineness of the classification and induction of the attribute words are improved.

[0082] It should be noted that in addition to the three-round clustering as Figure 4 shown, the multi-round clustering method can also adopt more rounds of clustering, and the present disclosure places no limitation thereon. In addition, compared with the split-cluster clustering method, the similarity network construction of the multi-round clustering method is tightened, and multi-layer aggregation is performed through the InfoMap unsupervised clustering algorithm. Each time of clustering, an induction model is used for induction, and the InfoMap unsupervised clustering algorithm is run again on the result after induction.

[0083] It should be noted that in the related art, similarity models usually only focus on the geometric relationship between vectors and lack the understanding of context. However, in actual business scenarios, a polysemous word may have different meanings in different contexts, but may have similar vector representations in vector encoding. Then, the similarity model may wrongly consider them semantically similar. The recognition of synonyms also depends on semantic understanding. For example, "happy" and "glad" may have different emotional intensities and usages in some contexts, and the similarity model lacks the ability to understand and adapt to specific semantics. In addition, similarity models usually need to set thresholds to judge whether two objects are similar. However, setting the threshold too high may lead to missed judgments, and setting the threshold too low may lead to false judgments. Moreover, the results of similarity models are usually numerical scores, such as cosine similarity, lacking interpretability of the results. Users may not understand why the similarity score of two words is 0.8 instead of 0.9, lacking the reasoning process for the results.

[0084] In a possible way, the second verification is performed on the sixth attribute word corresponding to the target word vector and the seventh attribute word corresponding to the candidate word vector, including: the large language model performs the second verification on the sixth attribute word corresponding to the target word vector and the seventh attribute word corresponding to the candidate word vector based on a preset prompt word. The preset prompt word is used to instruct the large model to split each attribute word in the sixth attribute word and the seventh attribute word to obtain multiple eighth attribute words for describing the target object and modifiers for modifying the eighth attribute words, and perform the second verification on the sixth attribute word and the seventh attribute word based on the category information of the category to which the target object belongs, the multiple eighth attribute words, and the modifiers.

[0085] Exemplarily, the similarity model can be a large language model, and the Figure 5 logical chain of step-by-step comparison shown can be used as a prompt word and input into the large language model to guide the large language model to process the attribute words in a step-by-step reasoning manner, split the core words for describing the target object and the modifiers for modifying the core words in the attribute words, and combine the category information of the category to which the target object belongs. First, judge whether the core words are semantically consistent. If they are consistent, further compare whether the semantics of the modifiers are consistent. The finally output similarity verification result can include structured attribute words, whether the attribute words are similar, whether there is an inclusion relationship between the attribute words, etc.

[0086] Exemplarily, through the prompt word, the large language model will generate intermediate independent steps, consider context information when calculating similarity, have more detailed similarity judgments, and reduce the dependence on thresholds. Moreover, the large language model can generate a detailed reasoning process, making the reasoning process of the model more transparent and interpretable, capable of handling more complex verification tasks, no longer just a simple result score, facilitating subsequent error location and model optimization.

[0087] It should be noted that in the above-mentioned secondary clustering method, split-cluster clustering method, and multi-round clustering method, the processing process of the similarity model can be optimized based on the prompt words, and the present disclosure does not limit this.

[0088] In a possible way, at least generalize based on the semantics of the first attribute words in each initial attribute phrase to obtain multiple second attribute words for describing the target object, including: generalize based on the induction weight and semantics of the first attribute words in each initial attribute phrase to obtain multiple second attribute words for describing the target object, and the induction weight of the first attribute word is positively correlated with the occurrence frequency. At least process based on the semantics of each attribute word in the target attribute phrase to obtain the target attribute word for describing the target object, including: process based on the induction weight and semantics of each attribute word in the target attribute phrase to obtain the target attribute word for describing the target object, where the induction weight of each attribute word in the target attribute phrase is positively correlated with the occurrence frequency.

[0089] Exemplarily, when generalizing the semantics of the attribute words within the attribute phrase, the occurrence frequency of the attribute words can be considered. Therefore, when obtaining the first attribute word, it is necessary to count the occurrence frequency of the first attribute word. The induction weight of the attribute word is positively correlated with the occurrence frequency of the attribute word, so as to guide the induction model to generalize towards the semantics of the attribute words with high occurrence frequencies, improving the accuracy and precision of attribute word generalization.

[0090] It should be understood that the induction model usually combines natural language processing techniques and large language models to identify the similarities and correlations between words and generalize semantics based on this.

[0091] By adopting the above method, through the above optimizations of the clustering algorithm, similarity verification process, and induction process, more accurate and representative attribute words can be refined. For example, brand mind words can help brand owners clarify the core memory points of the brand in users' minds, optimize content creation to help brand owners conduct more accurate, efficient, and differentiated brand promotion on social media, understand user needs. For example, the attribute words extracted based on users' search content can directly reflect user needs, assist brand promotion by improving the content connection under popular search terms, or further analyze the preference and reputation of different products or different contents based on the attribute words, help the brand distinguish negative evaluations and optimize them in a timely manner, and at the same time adjust the product promotion direction and strengthen the advantages and highlights in positive evaluations, etc.

[0092] In the embodiments of the present disclosure, the above classification and induction process of attribute words can be adopted as Figure 7The architecture of Kafka + Flink is shown as follows. Among them, Kafka, as a high-throughput distributed message queue, can well solve the problem of peak and valley in data calls. Compared with the problem of unstable call traffic caused by the startup and data writing links of hourly tasks in the Spark architecture, Kafka can serve as a data buffer zone, and the attribute words generated by the upstream data source can be continuously sent to the Kafka topic. When the daily task starts, it does not directly pull a large number of attribute words from the data source, but controls the writing rate through a Flink task, so as to consume data from Kafka at a stable rate and avoid instant high-traffic calls.

[0093] In addition, Flink can process the data received from Kafka in real time, avoiding the data processing delay problem caused by hourly partitions in the Spark architecture, and then playing a role in peak shaving and valley filling. And Flink itself has a powerful fault recovery mechanism. Different from the breakpoint resumption mechanism based on external storage in the Spark architecture, Flink realizes task fault tolerance through checkpoints and state backends. When a task fails, Flink can quickly restore the task state from the nearest checkpoint and continue to process data without relying on an external storage system.

[0094] By optimizing the architecture, it is possible to solve the situation where a single task fails midway due to reasons such as hardware failures and network fluctuations, avoid high-traffic calls caused by repeated model calls, effectively improve the algorithm performance, and ensure the stability of operation.

[0095] Based on the same concept, an embodiment of the present disclosure further provides a content processing device, as Figure 8 shown. The content processing device 800 includes: An acquisition module 801, configured to acquire a plurality of first attribute words for describing a target object, where the plurality of first attribute words are obtained based on target content associated with the target object; A processing module 802, configured to classify and summarize the plurality of first attribute words at least based on category information of a category to which the target object belongs through a machine learning model, obtain a plurality of second attribute words for describing the target object, determine a target attribute word group based on the plurality of second attribute words, and perform processing at least based on the semantics of each attribute word in the target attribute word group to obtain a target attribute word for describing the target object.

[0096] Optionally, the processing module 802 includes: A first processing sub-module, configured to perform vector processing on the plurality of first attribute words respectively based on category information of a category to which the target object belongs to obtain a plurality of first word vectors; A classification module for classifying the multiple first attribute words according to the multiple first word vectors to obtain multiple initial attribute word groups; An induction module for inducing the first attribute words in each of the initial attribute word groups to obtain multiple second attribute words for describing the target object.

[0097] Optionally, the induction module includes: A first induction sub-module for inducing at least based on the semantics of the first attribute words in each of the initial attribute word groups to obtain multiple second attribute words for describing the target object; The processing module 802 includes: A second processing sub-module for performing vector processing on the multiple second attribute words respectively based on the category information of the category to which the target object belongs to obtain multiple second word vectors; A second classification module for classifying the multiple second attribute words at least according to the multiple second word vectors to obtain the target attribute word group.

[0098] Optionally, the first induction sub-module is used for: When the multiple first attribute words do not belong to the attribute words under the target attribute classification, inducing at least based on the semantics of the first attribute words in each of the initial attribute word groups to obtain multiple second attribute words for describing the target object, where the target attribute classification is a preset classification that cannot perform semantic induction in the preset attribute definition, the preset attribute definition is obtained by defining the attributes under each classification, and the attributes under each classification are obtained by classifying the attributes of the category to which the attribute description object belongs; The first classification module is used for: Clustering the multiple first attribute words based on the multiple first word vectors through a first clustering algorithm to obtain the multiple initial attribute word groups; The second classification module is used for: Clustering the multiple second attribute words based on the multiple second word vectors through a second clustering algorithm to obtain the target attribute word group; Wherein, the first clustering algorithm and the second clustering algorithm are used for clustering based on the graph information between the word vectors, and the walking scale of the first clustering algorithm is smaller than the walking scale of the second clustering algorithm.

[0099] Optionally, the induction module includes: A second induction sub-module, configured to, when the multiple first attribute words belong to the attribute words under a target attribute classification, use the first attribute word with the highest occurrence frequency in each initial attribute phrase as the second attribute word of the corresponding initial attribute phrase, to obtain multiple second attribute words for describing the target object, where the target attribute classification is a preset classification that cannot be semantically generalized in a preset attribute definition, the preset attribute definition is obtained by defining the attributes under each classification, and the attributes under each classification are obtained by classifying the attributes of the category to which the attribute description object belongs; The processing module 802 is configured to: Determine each initial attribute phrase where the second attribute word is located as the target attribute phrase.

[0100] Optionally, the second induction sub-module is configured to: Use the first attribute word with the highest occurrence frequency in each initial attribute phrase as the first candidate attribute word of the corresponding initial attribute phrase; Perform a first verification on the first candidate attribute word and the third attribute words other than the first candidate attribute word in the corresponding initial attribute phrase, where the first verification is used to verify whether the semantics of the first candidate attribute word and the third attribute words are consistent; Determine the fourth attribute words that pass the first verification among the third attribute words and the first candidate attribute word as the first attribute phrase, and determine the first candidate attribute word as the second attribute word of the first attribute phrase; Use the fifth attribute words that do not pass the first verification among the third attribute words as the second attribute phrase, and use the second attribute phrase as the new initial attribute phrase, and repeat the step of using the first attribute word with the highest occurrence frequency in each initial attribute phrase as the first candidate attribute word of the corresponding initial attribute phrase until all the attribute words in the new initial attribute phrase pass the first verification.

[0101] Optionally, the processing module 802 includes: A third processing sub-module, configured to perform vector processing on each attribute word in the target attribute phrase respectively based on the category information of the category to which the target object belongs, to obtain multiple third word vectors; A first verification module, configured to, for each target word vector among the multiple third word vectors, determine a preset number of candidate word vectors closest to the target word vector among the multiple third word vectors, and perform a second verification on the sixth attribute word corresponding to the target word vector and the seventh attribute word corresponding to the candidate word vector, where the second verification is used to verify whether the semantic similarity between the sixth attribute word and the seventh attribute word is greater than a first similarity threshold; The first response module is used to, in response to the sixth attribute word and the seventh attribute word passing the second verification, summarize the sixth attribute word and the seventh attribute word to obtain a target attribute word for describing the target object; The second response module is used to, in response to the sixth attribute word and the seventh attribute word not passing the second verification, determine the sixth attribute word and the seventh attribute word as the target attribute words for describing the target object.

[0102] Optionally, the second verification module is used to: Perform a second verification on the sixth attribute word corresponding to the target word vector and the seventh attribute word corresponding to the candidate word vector through a large language model based on a preset prompt word. The preset prompt word is used to instruct the large language model to split each attribute word in the sixth attribute word and the seventh attribute word to obtain multiple eighth attribute words for describing the target object and modifiers for modifying the eighth attribute words, and perform a second verification on the sixth attribute word and the seventh attribute word based on the category information of the category to which the target object belongs, the multiple eighth attribute words, and the modifiers.

[0103] Optionally, the first summarization sub-module is used to: Summarize based on the semantics of the first attribute words in each initial attribute word group to obtain multiple second candidate attribute words for describing the target object; For each second candidate attribute word, perform a third verification on the second candidate attribute word and the first attribute word in the corresponding initial attribute word group. The third verification is used to verify whether the semantics of the second candidate attribute word and the first attribute word in the corresponding initial attribute word group are consistent; For each initial attribute word group, determine the first attribute word and the second candidate attribute word that pass the third verification in the initial attribute word group as the third attribute word group, and determine the second candidate attribute word as the second attribute word of the third attribute word group. Determine the first attribute word that does not pass the third verification in the initial attribute word group as the fourth attribute word group, and use the third attribute word group as the new initial attribute word group. Repeat the step of summarizing based on the semantics of the first attribute words in each initial attribute word group to obtain multiple second candidate attribute words for describing the target object until all the attribute words in the new initial attribute word group pass the third verification.

[0104] Optionally, the second classification module is used to: Perform a fourth verification on every two second attribute words. The fourth verification is used to verify whether the semantic similarity between the two second attribute words is greater than a second similarity threshold; In response to the two second attribute words passing the fourth verification, set the classification weight between the two second attribute words to 1. In response to the two second attribute words not passing the fourth verification, set the classification weight between the two second attribute words to the word vector distance between the two second attribute words; Classify the multiple second attribute words according to the classification weight between each two second attribute words to obtain the target attribute phrase group.

[0105] Optionally, the first classification module is used for: Cluster the multiple first attribute words based on the multiple first word vectors through a third clustering algorithm to obtain the multiple initial attribute phrase groups; The second classification module is used for: Cluster the multiple second attribute words based on the multiple second word vectors through a fourth clustering algorithm to obtain the target attribute phrase group; Among them, the third clustering algorithm and the fourth clustering algorithm are used for clustering based on the graph information between word vectors, and the walking scale of the third clustering algorithm is greater than the walking scale of the fourth clustering algorithm.

[0106] Optionally, the second induction sub-module is used for: Induce based on the induction weight and semantics of the first attribute words in each initial attribute phrase group to obtain multiple second attribute words for describing the target object, and the induction weight of the first attribute words is positively correlated with the occurrence frequency; The processing module 802 is used for: Process based on the induction weight and semantics of each attribute word in the target attribute phrase group to obtain the target attribute word for describing the target object, where the induction weight of each attribute word in the target attribute phrase group is positively correlated with the occurrence frequency.

[0107] Optionally, the first classification module is used for: Cluster the multiple first attribute words based on the multiple first word vectors through a fifth clustering algorithm to obtain the multiple initial attribute phrase groups, and the fifth clustering algorithm is used for clustering based on the modularity of the multiple first word vectors; The second classification module is used for: Cluster the multiple second attribute words based on the multiple second word vectors through a sixth clustering algorithm to obtain the target attribute phrase group, and the sixth clustering algorithm is used for clustering based on the connected components of the multiple second word vectors.

[0108] Based on the same concept, embodiments of the present disclosure also provide a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of any of the above content processing methods are implemented.

[0109] Based on the same concept, embodiments of the present disclosure also provide an electronic device, which may include: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of any of the above content processing methods.

[0110] Based on the same concept, embodiments of the present disclosure also provide a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above content processing methods are implemented.

[0111] Next, refer to Figure 9 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing embodiments of the present disclosure. The terminal device in embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of embodiments of the present disclosure.

[0112] As Figure 9 shown, the electronic device 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage device 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0113] Generally, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 9An electronic device 900 is shown with various devices, but it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0114] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.

[0115] It should be noted that the above computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0116] In some embodiments, communication can be performed using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0117] The above computer-readable medium can be included in the above electronic device; it can also exist separately without being assembled into the electronic device.

[0118] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain a plurality of first attribute words for describing a target object, the plurality of first attribute words being obtained based on target content associated with the target object; classify and summarize the plurality of first attribute words at least based on category information of the category to which the target object belongs through a machine learning model to obtain a plurality of second attribute words for describing the target object, determine a target attribute phrase group based on the plurality of second attribute words, and perform processing at least based on the semantics of each attribute word in the target attribute phrase group to obtain a target attribute word for describing the target object.

[0119] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0121] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0122] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0123] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0125] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0126] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.

Claims

1. A content processing method, characterized in that, The content processing method includes: Obtaining a plurality of first attribute words for describing a target object, where the plurality of first attribute words are obtained based on target content associated with the target object; Classifying and summarizing the plurality of first attribute words at least based on category information of the category to which the target object belongs through a machine learning model, obtaining a plurality of second attribute words for describing the target object, determining a target attribute phrase group based on the plurality of second attribute words, and processing at least based on the semantics of each attribute word in the target attribute phrase group to obtain a target attribute word for describing the target object.

2. The content processing method according to claim 1, characterized in that The classifying and summarizing the plurality of first attribute words at least based on category information of the category to which the target object belongs through a machine learning model to obtain a plurality of second attribute words for describing the target object includes: Performing the following processing on the plurality of first attribute words through the machine learning model: Performing vector processing on the plurality of first attribute words respectively based on category information of the category to which the target object belongs to obtain a plurality of first word vectors; Classifying the plurality of first attribute words according to the plurality of first word vectors to obtain a plurality of initial attribute phrase groups; Summarizing the first attribute words in each of the initial attribute phrase groups to obtain a plurality of second attribute words for describing the target object.

3. The content processing method according to claim 2, wherein The summarizing the first attribute words in each of the initial attribute phrase groups to obtain a plurality of second attribute words for describing the target object includes: Summarizing at least based on the semantics of the first attribute words in each of the initial attribute phrase groups to obtain a plurality of second attribute words for describing the target object; The determining a target attribute phrase group based on the plurality of second attribute words includes: Performing vector processing on the plurality of second attribute words respectively based on category information of the category to which the target object belongs to obtain a plurality of second word vectors; Classifying the plurality of second attribute words at least according to the plurality of second word vectors to obtain the target attribute phrase group.

4. The content processing method according to claim 3, characterized in that The summarizing at least based on the semantics of the first attribute words in each of the initial attribute phrase groups to obtain a plurality of second attribute words for describing the target object includes: When the plurality of first attribute words do not belong to attribute words under a target attribute classification, summarizing at least based on the semantics of the first attribute words in each of the initial attribute phrase groups to obtain a plurality of second attribute words for describing the target object, where the target attribute classification is a preset classification that cannot perform semantic summarization in a preset attribute definition, the preset attribute definition is obtained by defining attributes under each classification, and the attributes under each classification are obtained by classifying attributes of the category to which the attribute description object belongs; The classifying the plurality of first attribute words according to the plurality of first word vectors to obtain a plurality of initial attribute phrase groups includes: Clustering the plurality of first attribute words based on the plurality of first word vectors through a first clustering algorithm to obtain the plurality of initial attribute phrase groups; The classifying the plurality of second attribute words at least based on the plurality of second word vectors to obtain a target attribute phrase group includes: Cluster the multiple second attribute words based on the multiple second word vectors through a second clustering algorithm to obtain the target attribute phrase group; Wherein, the first clustering algorithm and the second clustering algorithm are used to perform clustering based on the graph information between word vectors, and the walking scale of the first clustering algorithm is smaller than that of the second clustering algorithm.

5. The content processing method according to claim 2, characterized in that, The induction of the first attribute words in each of the initial attribute phrase groups to obtain multiple second attribute words for describing the target object includes: When the multiple first attribute words belong to the attribute words under the target attribute classification, the first attribute word with the highest occurrence frequency in each of the initial attribute phrase groups is used as the second attribute word corresponding to the initial attribute phrase group, so as to obtain multiple second attribute words for describing the target object. The target attribute classification is a preset classification that cannot perform semantic induction in the preset attribute definition. The preset attribute definition is obtained by defining the attributes under each classification, and the attributes under each classification are obtained by classifying the attributes of the category to which the attribute description object belongs; The determination of the target attribute phrase group based on the multiple second attribute words includes: Determine the initial attribute phrase group where each second attribute word is located as the target attribute phrase group.

6. The content processing method according to claim 5, wherein The use of the first attribute word with the highest occurrence frequency in each of the initial attribute phrase groups as the second attribute word corresponding to the initial attribute phrase group includes: Use the first attribute word with the highest occurrence frequency in each of the initial attribute phrase groups as the first candidate attribute word corresponding to the initial attribute phrase group; Perform a first verification on the first candidate attribute word and the third attribute words other than the first candidate attribute word in the corresponding initial attribute phrase group. The first verification is used to verify whether the semantics of the first candidate attribute word and the third attribute words are consistent; Determine the fourth attribute words that pass the first verification among the third attribute words and the first candidate attribute word as the first attribute phrase group, and determine the first candidate attribute word as the second attribute word of the first attribute phrase group; Use the fifth attribute words that do not pass the first verification among the third attribute words as the second attribute phrase group, and use the second attribute phrase group as the new initial attribute phrase group, and repeat the step of using the first attribute word with the highest occurrence frequency in each of the initial attribute phrase groups as the first candidate attribute word corresponding to the initial attribute phrase group until all the attribute words in the new initial attribute phrase group pass the first verification.

7. The content processing method according to any one of claims 4-6, characterized in that, The processing based on at least the semantics of each attribute word in the target attribute phrase group to obtain the target attribute word for describing the target object includes: Perform vector processing on each attribute word in the target attribute phrase group respectively based on the category information of the category to which the target object belongs to obtain multiple third word vectors; For each target word vector among the multiple third word vectors, determine a preset number of candidate word vectors closest to the target word vector among the multiple third word vectors, and perform a second verification on the sixth attribute word corresponding to the target word vector and the seventh attribute word corresponding to the candidate word vector. The second verification is used to verify whether the semantic similarity between the sixth attribute word and the seventh attribute word is greater than a first similarity threshold; In response to the sixth attribute word and the seventh attribute word passing the second verification, generalize the sixth attribute word and the seventh attribute word to obtain a target attribute word for describing the target object; In response to the sixth attribute word and the seventh attribute word failing to pass the second verification, determine the sixth attribute word and the seventh attribute word as the target attribute words for describing the target object.

8. The content processing method according to claim 7, wherein The performing a second verification on the sixth attribute word corresponding to the target word vector and the seventh attribute word corresponding to the candidate word vector includes: Performing a second verification on the sixth attribute word corresponding to the target word vector and the seventh attribute word corresponding to the candidate word vector through a large language model based on a preset prompt. The preset prompt is used to instruct the large language model to split each attribute word in the sixth attribute word and the seventh attribute word to obtain multiple eighth attribute words for describing the target object and modifiers for modifying the eighth attribute words, and perform a second verification on the sixth attribute word and the seventh attribute word based on the category information of the category to which the target object belongs, the multiple eighth attribute words, and the modifiers.

9. The content processing method according to claim 3, characterized in that The generalizing, at least based on the semantics of the first attribute words in each of the initial attribute word groups, to obtain multiple second attribute words for describing the target object includes: Generalizing based on the semantics of the first attribute words in each of the initial attribute word groups to obtain multiple second candidate attribute words for describing the target object; For each second candidate attribute word, perform a third verification on the second candidate attribute word and the first attribute word in the corresponding initial attribute word group. The third verification is used to verify whether the semantics of the second candidate attribute word and the first attribute word in the corresponding initial attribute word group are consistent; For each initial attribute word group, determine the first attribute word and the second candidate attribute word that pass the third verification in the initial attribute word group as a third attribute word group, and determine the second candidate attribute word as the second attribute word of the third attribute word group. Determine the first attribute word that fails to pass the third verification in the initial attribute word group as a fourth attribute word group, and use the third attribute word group as a new initial attribute word group, and repeat the step of generalizing based on the semantics of the first attribute words in each of the initial attribute word groups to obtain multiple second candidate attribute words for describing the target object until all the attribute words in the new initial attribute word group pass the third verification.

10. The content processing method according to claim 9, characterized in that, The classifying, at least according to the multiple second word vectors, the multiple second attribute words to obtain the target attribute word group includes: Perform a fourth verification on every two second attribute words, where the fourth verification is used to verify whether the semantic similarity between the two second attribute words is greater than a second similarity threshold; In response to the two second attribute words passing the fourth verification, set the classification weight between the two second attribute words to 1. In response to the two second attribute words failing the fourth verification, set the classification weight between the two second attribute words to the word vector distance between the two second attribute words; Classify the multiple second attribute words according to the classification weight between every two second attribute words to obtain the target attribute phrase group.

11. The content processing method according to claim 9 or 10, characterized in that The classifying the multiple first attribute words according to the multiple first word vectors to obtain multiple initial attribute phrase groups includes: Cluster the multiple first attribute words based on the multiple first word vectors through a third clustering algorithm to obtain the multiple initial attribute phrase groups; The at least classifying the multiple second attribute words based on the multiple second word vectors to obtain a target attribute phrase group includes: Cluster the multiple second attribute words based on the multiple second word vectors through a fourth clustering algorithm to obtain the target attribute phrase group; Wherein, the third clustering algorithm and the fourth clustering algorithm are used to perform clustering based on the graph information between word vectors, and the walking scale of the third clustering algorithm is greater than the walking scale of the fourth clustering algorithm.

12. The content processing method according to any one of claims 3-4, 9-10, characterized in that The at least generalizing based on the semantics of the first attribute words in each initial attribute phrase group to obtain multiple second attribute words for describing the target object includes: Generalize based on the induction weight and semantics of the first attribute words in each initial attribute phrase group to obtain multiple second attribute words for describing the target object, where the induction weight of the first attribute words is positively correlated with the occurrence frequency; The at least processing based on the semantics of each attribute word in the target attribute phrase group to obtain a target attribute word for describing the target object includes: Process based on the induction weight and semantics of each attribute word in the target attribute phrase group to obtain a target attribute word for describing the target object, where the induction weight of each attribute word in the target attribute phrase group is positively correlated with the occurrence frequency.

13. The content processing method according to claim 3, characterized in that, The classifying the multiple first attribute words according to the multiple first word vectors to obtain multiple initial attribute phrase groups includes: Cluster the multiple first attribute words based on the multiple first word vectors through a fifth clustering algorithm to obtain the multiple initial attribute phrase groups, where the fifth clustering algorithm is used to perform clustering based on the modularity of the multiple first word vectors; The at least classifying the multiple second attribute words based on the multiple second word vectors to obtain a target attribute phrase group includes: Cluster the multiple second attribute words based on the multiple second word vectors through a sixth clustering algorithm to obtain the target attribute phrase group, where the sixth clustering algorithm is used to perform clustering based on the connected components of the multiple second word vectors.

14. A content processing device, characterized in that, The content processing device includes: An acquisition module, configured to acquire a plurality of first attribute words for describing a target object, where the plurality of first attribute words are obtained based on target content associated with the target object; A processing module, configured to classify and summarize the plurality of first attribute words at least based on category information of a category to which the target object belongs through a machine learning model, obtain a plurality of second attribute words for describing the target object, determine a target attribute word group based on the plurality of second attribute words, and perform processing at least based on semantics of each attribute word in the target attribute word group to obtain a target attribute word for describing the target object.

15. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processing device, the steps of the method according to any one of claims 1-13 are implemented.

16. An electronic device, characterized in that, Comprising: A storage device on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1-13.

17. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1-13 are implemented.