Label labeling method and device, electronic equipment and storage medium

By tagging the information text of asset management products and combining the evaluation of concentration, offset and overlap, the tags are optimized using a knowledge graph model, which solves the problem of insufficient accuracy in traditional tagging methods and achieves higher tagging accuracy.

CN120892572APending Publication Date: 2025-11-04CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510996039.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Traditional manual or simple rule-based labeling methods cannot guarantee the accuracy of asset management product labeling, especially when faced with numerous unstructured features and complex, semantically ambiguous terms, resulting in insufficient labeling accuracy.

Method used

By acquiring target product information text, tagging it, and combining the evaluation of concentration, offset, and overlap, the knowledge graph model is used for optimization, ultimately generating accurate product tags.

Benefits of technology

It has improved the accuracy of asset management product labeling, ensuring the accuracy and consistency of labels, and optimized the labeling process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892572A_ABST
    Figure CN120892572A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a label labeling method and device, electronic equipment and a storage medium, belongs to the technical field of text processing, and is suitable for the field of finance. The method comprises the steps of obtaining a product information text of a target product; labeling the product information text to obtain an initial label; performing concentration degree evaluation on the initial label according to the product information text to obtain a label concentration degree; performing style drift evaluation on the initial label according to the product information text to obtain a label offset degree; obtaining a reference product, and performing overlap ratio evaluation on the initial label according to the reference product and the product information text to obtain a label overlap ratio; and performing label optimization on the initial label according to the label concentration ratio, the label offset degree and the label overlap ratio to obtain a product label of the target product. According to the embodiment of the invention, the label labeling accuracy of the asset management product can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of text processing technology, applicable to the financial sector, and particularly to a labeling method and apparatus, electronic device, and storage medium. Background Technology

[0002] A label is an identifier used to identify, classify, and describe things. In the fintech field, for example, labels can be used to identify and classify financial products such as funds. By setting labels, the characteristics or categories of a fund product can be described. For instance, in asset management, a fund product's label might include the industries it operates in. When labeling asset management products, the sheer volume of product information, its significant unstructured nature, and the existence of complex terminology and semantic ambiguity make it difficult to guarantee the accuracy of traditional manual or simple rule-based labeling methods. Therefore, how to effectively improve the accuracy of asset management product labeling has become an urgent technical problem to be solved. Summary of the Invention

[0003] The main objective of this application is to provide a labeling method, apparatus, electronic device, and storage medium, which aims to effectively improve the accuracy of labeling for asset management products.

[0004] To achieve the above objectives, a first aspect of this application proposes a labeling method, the method comprising:

[0005] Obtain the product information text of the target product;

[0006] The product information text is labeled to obtain initial tags;

[0007] The initial tags are evaluated for concentration based on the product information text to obtain the tag concentration.

[0008] The style drift of the initial label is evaluated based on the product information text to obtain the label offset.

[0009] Obtain a reference product, and evaluate the overlap of the initial label based on the reference product and the product information text to obtain the label overlap.

[0010] The initial labels are optimized based on the label concentration, label offset, and label overlap to obtain the product labels for the target product.

[0011] In some embodiments, the step of evaluating the overlap of the initial label based on the reference product and the product information text to obtain the label overlap includes:

[0012] Obtain the reference label of the reference product;

[0013] The initial label is calculated using the cosine similarity based on the reference label to obtain the label similarity.

[0014] If the tag similarity is greater than a preset first threshold, obtain the reference product text of the reference product, evaluate the product overlap of the initial tag based on the product information text, the reference product text and the tag similarity, and determine the tag overlap.

[0015] If the label similarity is less than or equal to a preset first threshold, the label similarity is taken as the label overlap.

[0016] In some embodiments, the step of evaluating the product overlap of the initial tag based on the product information text, the reference product text, and the tag similarity, and determining the tag overlap, includes:

[0017] The product information text is processed by graph embedding through a preset knowledge graph model to obtain a target information vector. The reference product text is also processed by graph embedding through the same knowledge graph model to obtain a reference information vector.

[0018] The semantic similarity is calculated based on the target information vector and the reference information vector to obtain the product similarity.

[0019] If the product similarity is greater than a preset second threshold, the preset overlap is used as the tag overlap; wherein the preset overlap is less than or equal to the preset first threshold.

[0020] If the product similarity is less than or equal to a preset second threshold, the label similarity is used as the label overlap.

[0021] In some embodiments, after optimizing the initial labels based on the label concentration, label offset, and label overlap to obtain the product label for the target product, the method further includes:

[0022] The knowledge graph model is updated by updating the nodes based on the product tags to obtain the updated knowledge graph model.

[0023] Obtain the nodes of the updated knowledge graph model to obtain a node set;

[0024] Based on the product information text, perform semantic relationship analysis on the node set to obtain a node relationship set;

[0025] The updated knowledge graph model is updated with node relationships based on the set of node relationships.

[0026] In some embodiments, the target product includes target sub-products; the product information text includes product proportion weights and expected equity growth rates of the target sub-products over multiple time periods, and includes cyclical style indicators of the target product over multiple time periods; the step of evaluating the style drift of the initial label based on the product information text to obtain the label offset includes:

[0027] For each time period, a weighted calculation is performed based on the product's proportion weight and the expected equity growth rate to determine the cyclical style factor of the target product.

[0028] Obtain the expected style factor of the initial label from the preset database;

[0029] Linear regression is performed based on the expected style factor, the periodic style index of each time period, and the periodic style factor of each time period to determine the style coefficient of each periodic style factor and the style coefficient of the expected style factor.

[0030] The label offset of the initial label is calculated based on the difference between the style coefficient of the expected style factor and the style coefficient of the periodic style factor.

[0031] In some embodiments, the target product includes at least one target sub-product; the product information text includes the product proportion weight of each target sub-product; and the step of evaluating the concentration of the initial tags based on the product information text to obtain the tag concentration includes:

[0032] The square of the product weight is calculated by squaring the weight of each product.

[0033] The sum of the squared weights of all the products is used to obtain the label concentration.

[0034] In some embodiments, the step of tagging the product information text to obtain initial tags includes:

[0035] The product information text is cleaned to obtain preprocessed text;

[0036] Text embedding is performed on the preprocessed text to obtain a text vector;

[0037] The text vector is labeled using a preset label classifier to obtain the initial label.

[0038] To achieve the above objectives, a second aspect of this application provides a labeling device, the device comprising:

[0039] The first acquisition module is used to acquire the product information text of the target product;

[0040] The labeling module is used to label the product information text to obtain initial labels;

[0041] The concentration assessment module is used to assess the concentration of the initial label based on the product information text to obtain the label concentration.

[0042] The offset evaluation module is used to perform style drift evaluation on the initial label based on the product information text to obtain the label offset.

[0043] The overlap assessment module is used to obtain a reference product and perform an overlap assessment on the initial label based on the reference product and the product information text to obtain the label overlap.

[0044] The label optimization module is used to optimize the initial label based on the label concentration, the label offset and the label overlap to obtain the product label of the target product.

[0045] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0046] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0047] This application proposes a labeling method, apparatus, electronic device, and storage medium. First, product information text of the target product is obtained; the product information text is labeled to obtain initial labels; the initial labels are evaluated for concentration based on the product information text to obtain label concentration; the initial labels are evaluated for style drift based on the product information text to obtain label offset; a reference product is obtained, and the initial labels are evaluated for overlap based on the reference product and product information text to obtain label overlap; the initial labels are optimized based on label concentration, label offset, and label overlap to obtain the product label of the target product. Thus, this application's embodiment obtains the product information text of the target product and labels it, then performs concentration, offset, and overlap evaluations on the initial labels, thereby optimizing the labels and ultimately obtaining accurate target product labels. This application's embodiment optimizes the labeling accuracy during the labeling process by comprehensively evaluating label concentration, offset, and overlap, solving the problem of low labeling accuracy for asset management products. Attached Figure Description

[0048] Figure 1 This is a flowchart of the labeling method provided in the embodiments of this application;

[0049] Figure 2 yes Figure 1 The flowchart of step S102 in the document;

[0050] Figure 3 yes Figure 1 The flowchart of step S103 in the process;

[0051] Figure 4 yes Figure 1 The flowchart of step S104 in the process;

[0052] Figure 5 yes Figure 1 The flowchart of step S105 in the process;

[0053] Figure 6 yes Figure 5 The flowchart of step S503 in the process;

[0054] Figure 7 This is a flowchart of a labeling method provided in another embodiment of this application;

[0055] Figure 8 This is a schematic diagram of the label marking device provided in the embodiments of this application;

[0056] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0058] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0060] A label is an identifier used to identify, classify, and describe things. In the fintech field, for example, labels can be used to identify and classify financial products such as funds. By setting labels, the characteristics or categories of a fund product can be described. For instance, in asset management, a fund product's label might include the industries it operates in. When labeling asset management products, the sheer volume of product information, its significant unstructured nature, and the existence of complex terminology and semantic ambiguity make it difficult to guarantee the accuracy of traditional manual or simple rule-based labeling methods. Therefore, how to effectively improve the accuracy of asset management product labeling has become an urgent technical problem to be solved.

[0061] Based on this, embodiments of this application provide a labeling method and apparatus, electronic device and storage medium, which aim to effectively improve the accuracy of labeling for asset management products.

[0062] This application provides a labeling method, apparatus, electronic device, and storage medium, which are specifically described through the following embodiments. First, the labeling method in this application is described.

[0063] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0064] Foundational artificial intelligence technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, large text processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0065] The labeling method provided in this application relates to the field of text processing technology and is applicable to the financial sector. The labeling method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the labeling method, but is not limited to the above forms.

[0066] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0067] Figure 1 This is an optional flowchart of the labeling method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.

[0068] Step S101: Obtain the product information text of the target product;

[0069] Step S102: Tag the product information text to obtain initial tags;

[0070] Step S103: Evaluate the concentration of the initial labels based on the product information text to obtain the label concentration.

[0071] Step S104: Evaluate the style drift of the initial label based on the product information text to obtain the label offset.

[0072] Step S105: Obtain a reference product, and evaluate the overlap of the initial labels based on the reference product and product information text to obtain the label overlap.

[0073] Step S106: Optimize the initial labels based on label concentration, label offset, and label overlap to obtain the product labels for the target product.

[0074] Steps S101 to S106 of this embodiment illustrate the following: First, product information text of the target product is obtained; the product information text is tagged to obtain initial tags; the initial tags are evaluated for concentration based on the product information text to obtain tag concentration; the initial tags are evaluated for style drift based on the product information text to obtain tag offset; a reference product is obtained, and the initial tags are evaluated for overlap based on the reference product and product information text to obtain tag overlap; the initial tags are optimized based on tag concentration, tag offset, and tag overlap to obtain the product tag of the target product. Thus, this embodiment obtains the product information text of the target product and tags it, then evaluates the concentration, offset, and overlap based on the initial tags, and optimizes the tags to obtain accurate target product tags. This embodiment optimizes the tag accuracy in the tagging process by comprehensively evaluating the concentration, offset, and overlap of tags, solving the problem of low tag accuracy in asset management products.

[0075] In step S101 of some embodiments, the target product refers to the asset management product that needs to be labeled. In some embodiments, the target product may be a fund product, a securities product, etc. In the asset management field, the target product usually refers to a specific fund product, such as an equity fund, a bond fund, or a mixed fund. Product information text refers to a detailed description of the various characteristics and attributes of the target product. For example: fund prospectus, research report, or announcement.

[0076] Please see Figure 2 In some embodiments, step S102 may include, but is not limited to, steps S201 to S205:

[0077] Step S201: Clean the product information text to obtain preprocessed text;

[0078] Step S202: Perform text embedding on the preprocessed text to obtain text vectors;

[0079] Step S203: Label the text vectors using a preset label classifier to obtain initial labels.

[0080] Steps S201 to S203 of this embodiment involve: first, text cleaning of the product information text to obtain preprocessed text; text embedding of the preprocessed text to obtain text vectors; and labeling the text vectors using a preset label classifier to obtain initial labels. Thus, this embodiment removes noise and irrelevant information from the product information text through text cleaning. Subsequently, text embedding converts the preprocessed text into text vectors. Finally, a preset label classifier is used to label the text vectors, generating preliminary labels. This embodiment effectively improves the quality of text data.

[0081] In step S201 of some embodiments, text cleaning refers to processing the original product information text to remove irrelevant data and ensure the cleanliness and high quality of the text data. Text cleaning can be implemented by removing special characters, numbers, meaningless whitespace characters, punctuation marks, etc., and processing duplicate information or typos in the text.

[0082] For example, a fund's product information text might contain something like, "Fund Manager: Zhang San. Risk Level: High. Historical Performance: 15% return in 2021." If the text contains irrelevant symbols, formatting errors, or duplicate content, the text cleaning process will remove this irrelevant information, retaining the useful content to obtain clean pre-processed text, such as "Fund Manager: Zhang San, Risk Level: High, Historical Performance: 15% return in 2021...".

[0083] In step S202 of some embodiments, text embedding is a method of converting text into a numerical representation, typically by mapping words or sentences in the text to a vector space. Text embedding is usually implemented using neural network models in deep learning, such as Word2Vec, GloVe, or BERT models. Through these neural network models, text can be transformed into a high-dimensional vector representation, where each dimension of the vector represents a semantic feature of the text.

[0084] For example, the BERT model is a pre-trained language model that effectively understands contextual information in text. By using the BERT model to embed pre-processed text, the text can be converted into a fixed-dimensional text vector. For example, "Fund Manager: Zhang San, Risk Level: High" will be transformed into a vector after processing by the BERT model.

[0085] In step S203 of some embodiments, the preset label classifier is a tool used to analyze and classify text vectors based on a trained classification model. It is typically implemented using a neural network model, such as a support vector machine (SVM), decision tree, or deep neural network (DNN). These classifiers learn from text vectors and can automatically generate matching labels based on the features of the text.

[0086] In one embodiment, the preset label classifier employs a deep neural network (DNN) model. Specifically, the classifier analyzes the features of the target product's product information text and automatically generates matching labels. The generated labels include three types:

[0087] Industry tags describe the industry sectors involved in the fund or asset management product. For example, if the product information text states that it invests in technology stocks or consumer goods, the default tag classifier will label it with "technology industry" or "consumer goods industry".

[0088] Fund strategy tags are used to describe a fund's investment strategy, such as "growth", "value", "balanced", etc.

[0089] The investment area label describes the investment area of ​​a fund or asset management product, referring to the geographical or market scope. For example, the product information text may mention that the fund mainly invests in the domestic stock market or the global market, in which case the preset label classifier will generate a "domestic market" or "international market" label.

[0090] Please see Figure 3 In some embodiments, the target product includes at least one target sub-product, and the product information text includes the product proportion weight of each target sub-product. Step S103 may include, but is not limited to, steps S301 to S302:

[0091] Step S301: Calculate the square of the weight percentage of each product to obtain the square of the product weight.

[0092] Step S302: Obtain the sum of the squared weights of all products to get the label concentration.

[0093] In steps S301 to S302 of this embodiment, the weight percentage of each product is squared to obtain the squared weight of the product; the sum of the squared weights of all products is then obtained to obtain the label concentration. Therefore, this embodiment determines the accuracy of the industry label of the target product through concentration calculation.

[0094] In steps S301 to S302 of some embodiments, the target sub-product refers to each component of the target product. For example, a fund may invest in stocks of different industries; these stocks of different industries are the target sub-products. Product weighting refers to the proportion of each target sub-product in the target product, usually expressed as a percentage, reflecting the impact of the target sub-product on the target product.

[0095] Label concentration is an indicator that measures the weight distribution of various target sub-products within a target product. It is calculated by squared the weight of each target sub-product and summing the squared weights of all target sub-products. A high label concentration indicates concentrated investment in the target product, meaning some target sub-products account for a large proportion; conversely, a low label concentration indicates a more even distribution of investment and a more dispersed labeling. Assessing label concentration helps determine whether the industry allocation of a target product meets expectations. If a target product is labeled "diversified investment" but has a very high label concentration, it indicates a significant bias in investment, potentially contradicting the expected "diversified investment" label.

[0096] Please see Figure 4 In some embodiments, the target product includes target sub-products; the product information text includes product proportion weights and expected equity growth rates of the target sub-products over multiple time periods, and includes cyclical style indicators of the target product over multiple time periods. Step S104 may include, but is not limited to, steps S401 to S404:

[0097] Step S401: For each time period, perform a weighted calculation based on the product proportion weight and the expected equity growth rate data to determine the cyclical style factor of the target product.

[0098] Step S402: Obtain the expected style factor of the initial label from the preset database;

[0099] Step S403: Perform linear regression based on the expected style factor, the periodic style index of each time period, and the periodic style factor of each time period to determine the style coefficient of each periodic style factor and the style coefficient of the expected style factor.

[0100] Step S404: Calculate the label offset of the initial label based on the difference between the style coefficient of the expected style factor and the style coefficient of the periodic style factor.

[0101] Steps S401 to S404 of this embodiment involve calculating a cyclical style factor based on the target product's product proportion weight and expected equity growth rate data, obtaining the expected style factor of the initial label from a preset database, and combining the cyclical style indicators and cyclical style factor for each time period. Through linear regression analysis, the style coefficients of the cyclical style factor and the expected style factor are calculated. By comparing the style coefficients of the expected style factor and the cyclical style factor, the label shift is calculated, ultimately determining whether the target product has experienced a style shift. This embodiment, by calculating the difference between the target product's cyclical style factor and the expected style factor, can effectively assess whether the target product has deviated from the investment style defined by the initial label, accurately assess whether the fund's style has shifted, and ensure the accuracy of labeling.

[0102] In step S401 of some embodiments, the time period refers to a predetermined time interval, such as a year or a month.

[0103] The expected growth rate of equity in the target sub-product is the expected growth rate for each time period. For example, the expected growth rate of product A is 10%. Similarly, the expected growth rate of production speed in a factory is 10%.

[0104] Cyclical style indicators represent the actual style performance of a fund in each time period, reflecting the investment style exhibited by the fund in actual operation. They are calculated from the fund's actual performance (such as returns, industry allocation, etc.). This application does not impose specific restrictions on the calculation of cyclical style indicators.

[0105] The cyclical style factor is a weighted calculation of the product share weight and expected equity growth rate of each sub-product of the target product in different time periods, reflecting the actual growth rate of the target product in a specific period.

[0106] In step S402 of some embodiments, the expected style factor is the defined style index of the initial label obtained from a preset database. It is usually set based on market expectations or historical data and reflects the expected growth rate of the target product under the initial label definition.

[0107] In step S403 of some embodiments, the linear regression is as shown in equation (1):

[0108] R=α+β1G1+B2G2+∈ (1),

[0109] Where R is the periodic style index, G1 is the periodic style factor, G2 is the expected style factor, α is the intercept term, ∈ is the random error term, β1 is the style coefficient of the periodic style factor, and β2 is the style coefficient of the expected style factor.

[0110] The style coefficients of each period style factor are the same; that is, during the regression process, only one value is calculated as the style coefficient of each period style factor.

[0111] Among them, the cyclical style index is known and serves as the dependent variable, the cyclical style factor is known and the expected style factor is known and serves as the independent variables. Through linear regression, such as the least squares method, the intercept term, the random error term, the style coefficient of the cyclical style factor, and the style coefficient of the expected style factor are determined.

[0112] For example: Actual style factor G1: [18,16,17,19], Expected style factor G2: [17,17,17,17], Periodic style index R1: [5.2,4.8,5.0,5.5]. Linear regression reveals an intercept of 0.4980, a style coefficient of 0.2300 for the periodic style factor, and a style coefficient of 0.0647 for the expected style factor.

[0113] In step S404 of some embodiments, the tag offset of the initial tag is calculated based on the difference between the style coefficient of the expected style factor and the style coefficient of the periodic style factor, specifically as |β1-β2| / β1. For example, if the style coefficient of the periodic style factor is 0.2300 and the style coefficient of the expected style factor is 0.0647, then the tag offset is 71.87%.

[0114] Label shift measures whether a fund's strategy label (such as "growth fund" or "value fund") deviates from its actual investment style during operation. For example, suppose a fund's initial label is defined as "growth fund," with an expected style factor (G2) set at 17%. However, the fund's actual investment style factor (G1) changes over different time periods. For instance, the fund's actual style factor might be 16% in the first quarter, 17% in the second quarter, and 18% in the third quarter. Based on this actual data and the expected style factor of the initial label, regression analysis can be used to calculate β1 and β2, and then the label shift can be further calculated.

[0115] If the label offset is less than a predetermined threshold, the fund's style performance can be considered consistent with the initial label. If the label offset is greater than the predetermined threshold, the fund's actual style differs significantly from the initial label, indicating that the fund's style has shifted and the label needs to be re-evaluated.

[0116] Please see Figure 5 In some embodiments, step S105 includes, but is not limited to, steps S501 to S504:

[0117] Step S501: Obtain the reference label of the reference product;

[0118] Step S502: Calculate the cosine similarity of the initial labels based on the reference labels to obtain the label similarity.

[0119] Step S503: If the tag similarity is greater than the preset first threshold, obtain the reference product text of the reference product, evaluate the product overlap of the initial tag based on the product information text, the reference product text and the tag similarity, and determine the tag overlap.

[0120] Step S504: If the label similarity is less than or equal to a preset first threshold, the label similarity is used as the label overlap.

[0121] Steps S501 to S504 of this embodiment involve calculating the cosine similarity between the initial tag and the reference tag to determine their similarity. When the tag similarity is greater than a preset first threshold, the overlap is further evaluated by obtaining reference product text. The overlap is calculated by combining the product information text, the reference product text, and the tag similarity. If the tag similarity is less than or equal to the preset first threshold, the tag similarity is directly used as the tag overlap. Therefore, this embodiment can accurately assess whether there is redundancy in the initial tag, thereby optimizing the simplicity and effectiveness of the tag.

[0122] In step S501 of some embodiments, the reference product is a historical product. A reference label refers to a label associated with the reference product, reflecting its investment style, objectives, industry classification, and other characteristics.

[0123] In step S502 of some embodiments, cosine similarity is a measure of the similarity between two vectors, with values ​​ranging from 0 to 1. The larger the value, the higher the similarity between the labels.

[0124] Please see Figure 6 In some embodiments, step S503 includes, but is not limited to, steps S601 to S604:

[0125] Step S601: The product information text is processed by graph embedding through a preset knowledge graph model to obtain the target information vector; the reference product text is processed by graph embedding through the knowledge graph model to obtain the reference information vector.

[0126] Step S602: Calculate semantic similarity based on the target information vector and the reference information vector to obtain product similarity;

[0127] Step S603: If the product similarity is greater than the preset second threshold, the preset overlap is used as the tag overlap; wherein the preset overlap is less than or equal to the preset first threshold.

[0128] Step S604: If the product similarity is less than or equal to the preset second threshold, the label similarity is used as the label overlap.

[0129] Steps S601 to S604 of this embodiment involve using a preset knowledge graph model to perform graph embedding processing on the text of the target product and the reference product, obtaining target information vectors and reference information vectors respectively. Next, semantic similarity is calculated between the target information vector and the reference information vector to obtain product similarity. Based on the magnitude of the product similarity, it is determined whether a preset overlap degree should be used as the label overlap degree. When the product similarity is greater than a preset second threshold, the preset overlap degree will be used as the label overlap degree; otherwise, the label overlap degree is determined based on the label similarity. This embodiment, by introducing a knowledge graph model for graph embedding processing and combining it with a semantic similarity calculation method, can accurately assess the similarity between the target product and the reference product, thereby determining whether they overlap, and, if product overlap is determined, whether the label overlap degree is reasonable.

[0130] In step S601 of some embodiments, graph embedding processing is the process of converting product information text into vector representation. First, based on a preset knowledge graph model, entities and their relationships are extracted from the product information text. Entities can be fund names, industry categories, investment strategies, risk levels, etc., while relationships represent various connections between entities, such as "belongs to" or "invests in". Next, the nodes (entities) and edges (relationships) in these graph structures are converted into high-dimensional vectors using a graph embedding algorithm.

[0131] Label overlap is used to indicate whether the labels of identical or similar products overlap. For identical / similar products, their labels should overlap; for dissimilar products, their labels should not overlap. Specifically, label overlap measures whether industry labels and investment area labels overlap.

[0132] It's important to note that knowledge graphs are introduced because, in asset management products, some products, while not entirely identical, share high similarities in industry or investment field labels. Knowledge graphs can help identify these similarities, ensuring accurate labeling. Through knowledge graphs, not only explicit relationships between products can be captured, but also potential, implicit connections can be uncovered, thus allowing for a more precise determination of whether labels should overlap.

[0133] For example, Fund A and Fund B invest in the "Internet industry" and "technology industry" respectively. Although their industry labels are different, the knowledge graph can identify that both belong to the "technology field" in a broad sense. Therefore, their labels should be considered overlapping. Similarly, Fund C invests in "green energy," which has a significantly different industry label from Fund A and Fund B. The knowledge graph can accurately determine that Fund C's label should not overlap with the labels of Fund A and Fund B.

[0134] It's important to note that graph data structures are used because they effectively represent complex relationships and semantic dependencies between entities. In the labeling process of asset management products, product information involves multiple entities (such as fund name, industry, investment strategy, etc.) and various relationships between them (such as "invests in" and "belongs to"). Traditional linear data structures struggle to express the complex interactions between these entities and relationships, while graph data structures can clearly represent these relationships through nodes and edges, thus better capturing the similarities and differences between products.

[0135] In step S602 of some embodiments, semantic similarity is used to measure the similarity between two texts or vectors. The degree of similarity is determined by calculating the distance between the two vectors in the vector space, such as cosine similarity or Mahalanobis distance.

[0136] In step S603 of some embodiments, if the product similarity is greater than a preset second threshold, it indicates that the target product and the reference product are semantically very similar, and therefore their labels can be considered to be highly consistent. Label overlap indicates a problem with the labels, and the labels should be adjusted to reduce the overlap. However, since similar / consistent product labels should be identical, assigning a preset value to reduce the label overlap indicates that the labeling is correct.

[0137] In step S604 of some embodiments, if the product similarity is less than or equal to a preset second threshold, it indicates that the target product and the reference product are semantically significantly different, and therefore their labels should not overlap. However, since the label similarity is greater than a preset first threshold, it indicates that there is a problem with the labeling.

[0138] In other words, products should be similar if the label similarity is greater than a preset first threshold.

[0139] If the products are indeed similar, the preset overlap rate is used as the label overlap rate, indicating that there is no labeling issue.

[0140] If the products are not similar, the label similarity will be used as the label overlap, indicating that there is a problem with the labeling.

[0141] In step S106 of some embodiments, the labels are optimized according to the concentration of the initial labels. Specifically, this includes: obtaining the concentration range corresponding to the initial labels, comparing the concentration range with the label concentration, and if the label concentration is not within the concentration range, then re-evaluating the industry labels of the initial labels.

[0142] If the label offset is greater than the preset offset threshold, the fund strategy label of the initial label will be re-evaluated.

[0143] If the overlap of tags exceeds a preset first threshold, the investment area tags and industry tags of the initial tags will be re-evaluated.

[0144] Please see Figure 7 In some embodiments, after step S106, steps S701 to S704 may also be included, but are not limited to:

[0145] Step S701: Update the nodes of the knowledge graph model according to the product tags to obtain the updated knowledge graph model;

[0146] Step S702: Obtain the nodes of the updated knowledge graph model to obtain the node set;

[0147] Step S703: Perform semantic relationship analysis on the node set based on the product information text to obtain the node relationship set;

[0148] Step S704: Update the node relationships in the updated knowledge graph model based on the set of node relationships.

[0149] Steps S701 to S704, as shown in the embodiments of this application, update the nodes of the knowledge graph model according to the product tags to ensure that the knowledge graph model can reflect new product characteristics in a timely manner as the product tags change.

[0150] In step S701 of some embodiments, node updating refers to the process of modifying nodes in the knowledge graph model or adding new nodes based on product tags. By analyzing and processing product tags, new entities or relationships related to nodes in the existing knowledge graph model are identified and integrated into the model, thereby ensuring that the knowledge graph model reflects the latest product information and tags.

[0151] In step S702 of some embodiments, the node set refers to a set containing all entities formed by acquiring all nodes in the updated knowledge graph model. These nodes represent different entities in the knowledge graph, such as fund names, investment areas, investment strategies, etc.

[0152] In step S703 of some embodiments, node semantic relationship analysis involves analyzing the nodes in the node set to identify the semantic connections and dependencies between them. This process involves in-depth analysis of the text descriptions, attributes, or tags between nodes to determine the correlations and similarities between them, forming a node relationship set that represents the semantic relationships between the nodes.

[0153] In step S704 of some embodiments, node relationship updating refers to adjusting or adding new relationships between nodes in the updated knowledge graph model based on the analyzed set of node relationships. By updating the connection methods between nodes, it ensures that the knowledge graph can accurately reflect the latest semantic relationships and interactions between nodes, making the structure of the knowledge graph more complete and consistent.

[0154] Please see Figure 8 This application also provides a labeling device that can implement the above-described labeling method. The device includes:

[0155] The first acquisition module 801 is used to acquire product information text of the target product;

[0156] The labeling module 802 is used to label product information text to obtain initial labels;

[0157] The concentration assessment module 803 is used to assess the concentration of the initial labels based on the product information text to obtain the label concentration.

[0158] The offset evaluation module 804 is used to evaluate the style drift of the initial label based on the product information text to obtain the label offset.

[0159] The overlap assessment module 805 is used to obtain reference products and assess the overlap of the initial labels based on the reference products and product information text to obtain the label overlap.

[0160] The label optimization module 806 is used to optimize the initial labels based on label concentration, label offset and label overlap to obtain the product label of the target product.

[0161] The specific implementation of this label labeling device is basically the same as the specific implementation of the label labeling method described above, and will not be repeated here.

[0162] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described labeling method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0163] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0164] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0165] The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the labeling method of the embodiments of this application.

[0166] The input / output interface 903 is used to implement information input and output;

[0167] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0168] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0169] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0170] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described labeling method.

[0171] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0172] The labeling method, labeling device, electronic device, and storage medium provided in this application first obtain product information text of the target product; then, label the product information text to obtain initial labels; next, perform a concentration assessment on the initial labels based on the product information text to obtain label concentration; finally, perform a style drift assessment on the initial labels based on the product information text to obtain label offset; then, obtain a reference product, and perform an overlap assessment on the initial labels based on the reference product and the product information text to obtain label overlap; finally, optimize the initial labels based on label concentration, label offset, and label overlap to obtain the product label of the target product. Thus, this application embodiment obtains product information text of the target product and performs labeling, then performs concentration, offset, and overlap assessments on the initial labels, and optimizes the labels to ultimately obtain accurate target product labels. This application embodiment optimizes the labeling accuracy in the labeling process by comprehensively evaluating the concentration, offset, and overlap of labels, solving the problem of low labeling accuracy for asset management products.

[0173] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0174] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0175] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0176] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0177] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0178] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0179] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0180] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0181] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0182] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0183] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A labeling method, characterized in that, The method includes: Obtain the product information text of the target product; The product information text is labeled to obtain initial tags; The initial tags are evaluated for concentration based on the product information text to obtain the tag concentration. The style drift of the initial label is evaluated based on the product information text to obtain the label offset. Obtain a reference product, and evaluate the overlap of the initial label based on the reference product and the product information text to obtain the label overlap. The initial labels are optimized based on the label concentration, label offset, and label overlap to obtain the product labels for the target product.

2. The method according to claim 1, characterized in that, The step of evaluating the overlap of the initial label based on the reference product and the product information text to obtain the label overlap includes: Obtain the reference label of the reference product; The initial label is calculated using the cosine similarity based on the reference label to obtain the label similarity. If the tag similarity is greater than a preset first threshold, obtain the reference product text of the reference product, evaluate the product overlap of the initial tag based on the product information text, the reference product text and the tag similarity, and determine the tag overlap. If the label similarity is less than or equal to a preset first threshold, the label similarity is taken as the label overlap.

3. The method according to claim 2, characterized in that, The step of evaluating the product overlap of the initial tag based on the product information text, the reference product text, and the tag similarity, and determining the tag overlap, includes: The product information text is processed by graph embedding through a preset knowledge graph model to obtain a target information vector. The reference product text is also processed by graph embedding through the same knowledge graph model to obtain a reference information vector. The semantic similarity is calculated based on the target information vector and the reference information vector to obtain the product similarity. If the product similarity is greater than a preset second threshold, the preset overlap is used as the tag overlap; wherein the preset overlap is less than or equal to the preset first threshold. If the product similarity is less than or equal to a preset second threshold, the label similarity is used as the label overlap.

4. The method according to claim 3, characterized in that, After optimizing the initial labels based on the label concentration, label offset, and label overlap to obtain the product label for the target product, the method further includes: The knowledge graph model is updated by updating the nodes based on the product tags to obtain the updated knowledge graph model. Obtain the nodes of the updated knowledge graph model to obtain a node set; Based on the product information text, perform semantic relationship analysis on the node set to obtain a node relationship set; The updated knowledge graph model is updated with node relationships based on the set of node relationships.

5. The method according to claim 1, characterized in that, The target product includes target sub-products; the product information text includes product proportion weights and expected equity growth rates of the target sub-products over multiple time periods, and includes cyclical style indicators of the target product over multiple time periods; the step of evaluating the style drift of the initial tags based on the product information text to obtain tag offset includes: For each time period, a weighted calculation is performed based on the product's proportion weight and the expected equity growth rate to determine the cyclical style factor of the target product. Obtain the expected style factor of the initial label from the preset database; Linear regression is performed based on the expected style factor, the periodic style index of each time period, and the periodic style factor of each time period to determine the style coefficient of each periodic style factor and the style coefficient of the expected style factor. The label offset of the initial label is calculated based on the difference between the style coefficient of the expected style factor and the style coefficient of the periodic style factor.

6. The method according to claim 1, characterized in that, The target product includes at least one target sub-product; the product information text includes the product proportion weight of each target sub-product; the step of evaluating the concentration of the initial tags based on the product information text to obtain the tag concentration includes: The square of the product weight is calculated by squaring the weight of each product. The sum of the squared weights of all the products is used to obtain the label concentration.

7. The method according to any one of claims 1 to 6, characterized in that, The step of tagging the product information text to obtain initial tags includes: The product information text is cleaned to obtain preprocessed text; Text embedding is performed on the preprocessed text to obtain a text vector; The text vector is labeled using a preset label classifier to obtain the initial label.

8. A label marking device, characterized in that, The device includes: The first acquisition module is used to acquire product information text of the target product; The labeling module is used to label the product information text to obtain initial labels; The concentration assessment module is used to assess the concentration of the initial label based on the product information text to obtain the label concentration. The offset evaluation module is used to perform style drift evaluation on the initial label based on the product information text to obtain the label offset. The overlap assessment module is used to obtain a reference product and perform an overlap assessment on the initial label based on the reference product and the product information text to obtain the label overlap. The label optimization module is used to optimize the initial label based on the label concentration, the label offset and the label overlap to obtain the product label of the target product.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the labeling method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the labeling method according to any one of claims 1 to 7.