A label acquisition method and device based on an LLM model and a medium

By using an LLM-based label acquisition method that combines rules and a large model, labels are automatically generated, solving the problems of low efficiency and poor consistency in manual label allocation in existing technologies, and achieving efficient and accurate label generation.

CN121303112BActive Publication Date: 2026-06-26BEIJING SHOUFA INTELLIGENT TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SHOUFA INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2025-09-28
Publication Date
2026-06-26

Smart Images

  • Figure CN121303112B_ABST
    Figure CN121303112B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, in particular to a label acquisition method and device based on an LLM model and a medium. The method comprises the following steps: processing an original text set provided by a target server to obtain an original text cluster set; acquiring a first label set corresponding to the original text cluster set and a second label set corresponding to the original text cluster set according to the original text cluster set corresponding to the original text set; and processing the first label set and the second label set to obtain an original target label set. It is known that the first label set is a label processed by a large model, and the second label set is a label processed by a rule. On the one hand, the labels can be quickly generated through the rule and the large model, manual label allocation is avoided, and the processing efficiency is improved. On the other hand, the labels can be generated through two different ways, the lack or inaccuracy of the labels is avoided, and the inconsistent label allocation caused by the subjective judgment difference of the same operator is also avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a tag acquisition method, device and medium based on an LLM model. Background Technology

[0002] With the advent of the big data era, data processing is widely used by some users, especially enterprise users. For example, for some documents in enterprises, such as patent documents, achievement documents, and talent data tables, data processing methods are used to process historical or updated documents. Generally, enterprises will tag historical or updated documents when processing data. However, in current technology, tagging is done manually, which leads to subsequent classification of enterprise documents or organization of data. This tagging method relies entirely on manual tag allocation, which is extremely inefficient when dealing with massive amounts of data. Differences in subjective judgment among operators also lead to inconsistent tag allocation. Summary of the Invention

[0003] The purpose of this invention is to provide a label acquisition method, device, and medium based on an LLM model, so as to label by rules and large models, avoid relying entirely on manual label allocation, and improve processing efficiency.

[0004] This invention protects a label acquisition method based on an LLM model, the method comprising the following steps:

[0005] S100, Obtain the original text set A = {A1, ..., A2} provided by the target server. i , ..., A m}, A i It refers to the i-th original text provided by the i-th target server, where i ranges from 1 to m, and m is the number of original texts provided by the target server and m is a positive integer greater than 1.

[0006] S200, process A to obtain the original text cluster set B = {B1, ..., B1} corresponding to A. j , ..., B n}, B j It is the j-th original text cluster, where j ranges from 1 to n, and n is the number of original text clusters and is a positive integer greater than 1.

[0007] S300, Based on B, obtain the first tag set corresponding to B and the second tag set corresponding to B.

[0008] S400, process the first tag set corresponding to B and the second tag set corresponding to B to obtain the target tag set corresponding to B.

[0009] Furthermore, step S200 includes the following steps:

[0010] S201, for A i Keyword extraction was performed on the abstract text to obtain A. i The corresponding original keyword set A 0 i ={A 0 i1 , ..., A 0 ix , ..., A 0 ip}, A 0 ix It is A i The x-th original keyword in the abstract text, where x ranges from 1 to p, and p is A i The number of original keywords in the abstract text, where p is a positive integer greater than 1.

[0011] S202, according to A i The corresponding original tag set C i ={C i1 , ..., C ix , ..., C ip}, generate A i The corresponding intermediate tag subset D i Among them, A 0 ix The corresponding original tag subset C ix ={C 1 ix , ..., C r ix , ..., C s ix}, C r ix It is A 0 ix The r-th original label, where r ranges from 1 to s, and s is A 0 ix The original number of tags is s, where s is a positive integer greater than 1, and the original tags are obtained by inputting the original keywords into the preset LLM model.

[0012] S203, Based on the intermediate tag set D = {D1, ..., D2} of the original text i , ..., D m}, thus obtaining the initial similarity matrix F of the original text, where F satisfies the following condition:

[0013] Among them, F 11 It is the similarity between D1 and D1, F 1i It is D1 and D i The similarity between them, F1m It is D1 and D m The similarity between them, F i1 It is D i Similarity with D1, F ii It is D i With D i The similarity between them, F im It is D i With D m The similarity between them, F m1 It is D m Similarity with D1, F mi It is D m With D i The similarity between them, F mm It is D m With D m The similarity between them, where F 11 , ..., F ii , ..., F mm All are equal to 1.

[0014] S204, Generate B based on F.

[0015] Furthermore, in step S203, A i The corresponding intermediate tag subset D i =(D i1 , ..., D ia , ..., D ib ), where D ia It is A i The corresponding a-th intermediate label, where the value of a ranges from 1 to b and b is a positive integer greater than 1, is an original label whose priority is not less than a preset first priority threshold.

[0016] Furthermore, regarding C i After deduplication, we get C. i The corresponding specified tag subset E i ={E i1 , ..., E ic , ..., E id}, E ic It is A i The corresponding c-th specified tag, where c ranges from 1 to d and d is a positive integer greater than 1, is a tag obtained by deduplicating a subset of the original tags.

[0017] Furthermore, E i The corresponding specified tag priority set E 0 i ={E 0 i1 , ..., E0 ic , ..., E 0 id}, E 0 ic It is E ic The corresponding specified tag priority, E 0 ic Meets the following conditions:

[0018] Where, N ic This refers to A i China E ic The number of original keywords corresponding to the specified tag, N 0 ic This refers to E in A. ic The number of original keywords corresponding to the specified tag, p 0 It represents the total number of original keywords corresponding to A.

[0019] Furthermore, in step S203, F 1i Meets the following conditions:

[0020] Among them, K 0 1i It is D1 and D i The number of identical intermediate labels, K 1i It is D1 and D i The number of intermediate tags after deduplication.

[0021] Furthermore, the S300 steps include the following steps:

[0022] S301, obtain B j ={B j1 , ..., B jy , ..., B jq}, B jy It refers to the y-th original text in the j-th original text cluster, where the value of y ranges from 1 to q, and q is the number of original texts in the original text cluster and q is a positive integer greater than 1.

[0023] S302, B jy The abstract text is input into the preset LLM model to obtain B. jy The corresponding first initial tag set, wherein the first initial tag set includes several first initial tags, the first initial tags being obtained through B jy The summary text is input into the preset LLM model to obtain the result.

[0024] S303, extract B jy The original keyword set is input into the designed LLM model to obtain B.jy The corresponding second initial label set, wherein the second initial label set includes several second initial labels, the second initial labels being extracted from B jy The original keyword set is input into the LLM model to obtain the result.

[0025] S304, B jy The abstract text is combined with the first preset tag library to obtain B. jy The corresponding third initial tag set includes several third initial tags, and the first preset tag library includes several first preset tags and several preset texts corresponding to each preset tag. When the preset text matches B... jy When the similarity between the summary texts is greater than the preset similarity threshold, the first preset tag corresponding to the preset text is used as the third initial tag.

[0026] S305, extract B jy The original keyword set is combined with the second preset tag library to obtain B. jy The corresponding fourth initial tag set includes several fourth initial tags. The second preset tag library includes several second preset tags and several preset keywords corresponding to each second preset tag. When a preset keyword matches an original keyword, the preset tag corresponding to the preset keyword is used as the fourth initial tag.

[0027] S306, according to all B jy The corresponding first initial tag set and B jy The corresponding second initial tag set is used to obtain the first tag set corresponding to B; wherein, after merging the first initial tag set and the second initial tag set, the first tag set is constructed by selecting initial tags from the merged initial tag set whose occurrence frequency is greater than a preset occurrence frequency threshold.

[0028] S307, according to all B jy The corresponding third initial tag set and B jy The corresponding fourth initial tag set is used to obtain the second tag set corresponding to B; wherein, after merging the third initial tag set and the fourth initial tag set, the merged initial tag set is used as the second tag set.

[0029] Furthermore, the target label set corresponding to B includes several target labels, wherein step S400 includes the following steps to determine the target labels:

[0030] S401, process the first tag set corresponding to B and the second tag set corresponding to B to obtain the key priority set H = {H1, ..., H2} corresponding to B. g H z}, H g It is the priority of the g-th specified tag, where the value of g ranges from 1 to z, and z is the number of key tags and z is a positive integer greater than 1; wherein, the key tags are tags in the tag set generated after the first tag set and the second tag set.

[0031] S402, when H g When the value is greater than or equal to ΔH, the key label is determined as the target label, where ΔH is a preset priority threshold.

[0032] S403, when H g When < △H, the key label is determined to be a non-target label.

[0033] The present invention also protects an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the described LLM-based tag acquisition method.

[0034] The present invention also protects a computer-readable storage medium storing a computer program, characterized in that the computer program, when executed by a processor, implements the aforementioned label acquisition method based on the LLM model.

[0035] Compared with the prior art, the present invention has at least the following beneficial effects:

[0036] This invention discloses a tag acquisition method based on an LLM model. The method includes the following steps: acquiring an original text set provided by a target server; processing the original text set provided by the target server to obtain an original text cluster set corresponding to the original text set provided by the target server; acquiring a first tag set and a second tag set corresponding to the original text cluster set based on the original text cluster set; processing the first tag set and the second tag set corresponding to the original text cluster set to obtain a target tag set corresponding to the original text cluster set. It can be seen that the first tag set corresponding to the original text cluster set is a tag processed by a large model, and the second tag set corresponding to the original text cluster set is a tag processed by rules. On the one hand, tags can be quickly generated through rules and a large model, avoiding complete reliance on manual tag allocation and improving processing efficiency. On the other hand, tags can be generated in two different ways, avoiding the situation of missing or inaccurate tags, and also avoiding the inconsistency in tag allocation caused by differences in subjective judgment between operators. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart of a label acquisition method based on an LLM model provided in Embodiment 1 of the present invention. Detailed Implementation

[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] Example 1

[0041] like Figure 1 As shown in the figure, this embodiment provides a label acquisition method based on an LLM model, the method including the following steps:

[0042] S100, Obtain the original text set A = {A1, ..., A2} provided by the target server. i , ..., A m}, A i It refers to the i-th original text provided by the i-th target server, where i ranges from 1 to m, and m is the number of original texts provided by the target server and m is a positive integer greater than 1.

[0043] Specifically, the target server is the server of the target object, for example, the target object is a certain enterprise.

[0044] Specifically, the original text is text provided by the target object, such as patent text, achievement text, talent recruitment text, etc. provided by the target object.

[0045] S200, process A to obtain the original text cluster set B = {B1, ..., B1} corresponding to A. j , ..., B n}, B j It is the j-th original text cluster, where j ranges from 1 to n, and n is the number of original text clusters and is a positive integer greater than 1.

[0046] In one specific embodiment, step S200 includes the following steps:

[0047] S201, for Ai Keyword extraction was performed on the abstract text to obtain A. i The corresponding original keyword set A 0 i ={A 0 i1 , ..., A 0 ix , ..., A 0 ip}, A 0 ix It is A i The x-th original keyword in the abstract text, where x ranges from 1 to p, and p is A i The number of original keywords in the abstract text, where p is a positive integer greater than 1; those skilled in the art are familiar with any existing method for extracting keywords, and will not elaborate further here.

[0048] S202, according to A i The corresponding original tag set C i ={C i1 , ..., C ix , ..., C ip}, generate A i The corresponding intermediate tag subset D i Among them, A 0 ix The corresponding original tag subset C ix ={C 1 ix , ..., C r ix , ..., C s ix}, C r ix It is A 0 ix The r-th original label, where r ranges from 1 to s, and s is A 0 ix The original number of tags is s, where s is a positive integer greater than 1, and the original tags are obtained by inputting the original keywords into the preset LLM model.

[0049] Specifically, A i The corresponding intermediate tag subset D i =(D i1 , ..., D ia , ..., D ib ), where D ia It is A i The corresponding a-th intermediate label, where a ranges from 1 to b and b is a positive integer greater than 1, is an original label whose priority is not less than a preset first priority threshold; further understood as: for Ci After deduplication, we get C. i The corresponding specified tag subset E i ={E i1 , ..., E ic , ..., E id}, E ic It is A i The corresponding c-th specified tag, where c ranges from 1 to d and d is a positive integer greater than 1, is a tag obtained by deduplicating a subset of the original tags; according to E i , obtain E i The corresponding specified tag priority set E 0 i ={E 0 i1 , ..., E 0 ic , ..., E 0 id}, E 0 ic It is E ic The corresponding specified tag priority, where E 0 ic Meets the following conditions:

[0050] Where, N ic This refers to A i China E ic The number of original keywords corresponding to the specified tag, N 0 ic This refers to E in A. ic The number of original keywords corresponding to the specified tag, p 0 It is the total number of original keywords corresponding to A; when E 0 ic >△E 1 , when E 0 ic >△E 1 , △E 1 It is the preset first priority threshold, which determines E. 0 ic The corresponding indicator label is the intermediate label.

[0051] Those skilled in the art can set a preset first priority threshold according to actual needs, which will not be elaborated here.

[0052] The above-mentioned method can generate the priority of a specified tag by comparing the ratio of the number of original keywords in the corresponding original text to the ratio of the number of original keywords in the entire original text. This shows that the influence of tags between different texts is taken into account, and suitable intermediate tags that can reflect the characteristics of the target object are selected, so as to facilitate the subsequent generation of an accurate target object tag system.

[0053] S203, Based on the intermediate tag set D = {D1, ..., D2} of the original text i , ..., D m}, thus obtaining the initial similarity matrix F of the original text, where F satisfies the following condition:

[0054] Among them, F 11 It is the similarity between D1 and D1, F 1i It is D1 and D i The similarity between them, F 1m It is D1 and D m The similarity between them, F i1 It is D i Similarity with D1, F ii It is D i With D i The similarity between them, F im It is D i With D m The similarity between them, F m1 It is D m Similarity with D1, F mi It is D m With D i The similarity between them, F mm It is D m With D m The similarity between them, where F 11 , ..., F ii , ..., F mm All equal to 1; further understanding: F 1i Meets the following conditions:

[0055] Among them, K 0 1i It is D1 and D i The number of identical intermediate labels, K 1i It is D1 and D i The number of intermediate tags after deduplication.

[0056] S204, Generate B based on F.

[0057] Specifically, step S204 also includes the following steps:

[0058] S2041, process F to obtain the intermediate similarity matrix F of the original text. 0 , of which F 0 Meets the following conditions:

[0059]

[0060] S2042, according to F 0 , obtain F 0 The corresponding key feature vector λ F =[λ 1 F , ..., λ i F , ..., λ m F ], λ i F It is F 0 The corresponding i-th key feature value, λ F It is sorted from smallest to largest, λ i F Meets the following conditions:

[0061] S2043, according to λ F , obtain λ F The corresponding first characteristic matrix F λ F λ Meets the following conditions:

[0062] Among them, F λ It is an m×m matrix;

[0063] S2044, according to λ F , obtain λ F The corresponding second characteristic matrix F 0 λ F 0 λ Meets the following conditions:

[0064] Among them, F 0 λ It is an m×m matrix with all elements being 0 except for the diagonal elements;

[0065] S2045, according to F λ and F 0 λ , thus obtaining λ F The corresponding exp(F) 0 ), exp(F 0 It meets the following conditions:

[0066] Where exp() is an exponential function with base e, and T is the transpose sign;

[0067] S2046, according to exp(F 0 ), obtain F 0 The corresponding target feature vector λ0 = [λ 1 0, ..., λ r 0, ..., λ s 0], λ r 0 is the r-th target feature value, where r ranges from 1 to s, and s is the number of target feature values, which is a positive integer greater than 1. Further, s can be a value of k pre-defined in the k-means clustering model; where λ r 0 meets the following conditions:

[0068] Generate B;

[0069] S2047, based on λ0 and F 0 Generate B, which can be further understood as: based on λ0 and F 0 A new matrix B was generated. 0 Then, based on the new matrix B 0 The k-means clustering model is used to generate B. Those skilled in the art are familiar with the methods of implementing clustering using the k-means clustering model in the prior art, and will not elaborate further here.

[0070] Furthermore, B 0 It is an m×s matrix, B 0 ir It is B 0 The i-th value in the r-th list, according to F 0 *V0=λ r 0*V0, calculate V0, and then calculate B. 0 ir .

[0071] The above-mentioned method allows the matrix to be input into the k-means clustering model, thereby clustering the original text. This facilitates the processing of the clustered texts and the determination of accurate target labels by identifying the key relationships between the original texts within the cluster.

[0072] S300, Based on B, obtain the first tag set corresponding to B and the second tag set corresponding to B.

[0073] In one specific embodiment, step S300 includes the following steps:

[0074] S301, Obtain B j ={B j1 , ..., B jy , ..., Bjq}, B jy It refers to the y-th original text in the j-th original text cluster, where the value of y ranges from 1 to q, and q is the number of original texts in the original text cluster and q is a positive integer greater than 1.

[0075] S302, B jy The abstract text is input into the preset LLM model to obtain B. jy The corresponding first initial tag set; further understood as: the first initial tag set includes several first initial tags, which are obtained through B jy The abstract text is input into a preset LLM model to obtain the text tags. Those skilled in the art are familiar with any existing method for obtaining text tags through a preset LLM model, which will not be described in detail here.

[0076] S303, extract B jy The original keyword set is input into the designed LLM model to obtain B. jy The corresponding second initial label set; further understood as: the second initial label set includes several second initial labels, which are extracted from B jy The original keyword set is input into the preset LLM model to obtain the keyword tags. Those skilled in the art know any existing method of obtaining keyword tags through a preset LLM model, which will not be described in detail here.

[0077] S304, B jy The abstract text is combined with the first preset tag library to obtain B. jy The corresponding third initial tag set; further understood as: the third initial tag set includes several third initial tags, the first preset tag library includes several first preset tags and several preset texts corresponding to each preset tag, when the preset text is related to B jy When the similarity between the summary texts is greater than the preset similarity threshold, the first preset tag corresponding to the preset text is used as the third initial tag.

[0078] S305, extract B jy The original keyword set is combined with the second preset tag library to obtain B. jy The corresponding fourth initial tag set; further understood as: the fourth initial tag set includes several fourth initial tags, the second preset tag library includes several second preset tags and several preset keywords corresponding to each second preset tag, when the preset keyword is consistent with the original keyword, the preset tag corresponding to the preset keyword is used as the fourth initial tag.

[0079] S306, according to all B jy The corresponding first initial tag set and B jyThe corresponding second initial tag set is used to obtain the first tag set corresponding to B; further understood as: after merging the first initial tag set and the second initial tag set, the first tag set is constructed by selecting initial tags from the merged initial tag set whose occurrence frequency is greater than a preset occurrence frequency threshold; those skilled in the art set the preset occurrence frequency threshold according to actual needs, which will not be elaborated here.

[0080] S307, according to all B jy The corresponding third initial tag set and B jy The corresponding fourth initial tag set is used to obtain the second tag set corresponding to B; this can be further understood as: after merging the third initial tag set and the fourth initial tag set, the merged initial tag set is used as the second tag set.

[0081] The above describes how the first and second tags of the original text are generated using two different methods: a preset LLM model and a preset tag library, based on the original text's summary and keywords. This avoids the inaccuracy of tags generated solely by the preset LLM model and the incompleteness of tags generated solely by the preset tag library. It also avoids inconsistencies in tag allocation caused by differences in the operator's subjective judgment.

[0082] S400, process the first tag set corresponding to B and the second tag set corresponding to B to obtain the target tag set corresponding to B.

[0083] In a specific embodiment, the target tag set corresponding to B includes several target tags, and step S400 includes the following steps to determine the target tags:

[0084] S401, process the first tag set corresponding to B and the second tag set corresponding to B to obtain the key priority set H = {H1, ..., H2} corresponding to B. g H z}, H g It is the priority of the g-th specified tag, where the value of g ranges from 1 to z, and z is the number of key tags and z is a positive integer greater than 1; further understood as: the key tags are the tags in the tag set generated after the first tag set and the second tag set.

[0085] Furthermore, H g The following conditions must be met:

[0086] H 0 g It is H g The frequency of the corresponding key tags; further understood as: H g The number of times the corresponding key tag appears in the first tag set corresponding to B and the second tag set corresponding to B.

[0087] S402, when H g When the threshold is ≥△H, the key label is determined as the target label, where △H is a preset priority threshold. Those skilled in the art can set the preset priority threshold according to actual needs, which will not be elaborated here.

[0088] S403, when H g When < △H, the key label is determined to be a non-target label.

[0089] The above-mentioned method can calculate the priority of key tags based on their frequency, and then determine the target tags corresponding to users based on the priority of key tags. This allows for the generation of tags in two different ways, avoiding the situation of missing or inaccurate tags, and also avoiding the inconsistency in tag allocation caused by differences in subjective judgment among operators.

[0090] In summary, this embodiment provides a label acquisition method based on an LLM model. The method includes the following steps: obtaining an original text set provided by a target server; processing the original text set provided by the target server to obtain an original text cluster set corresponding to the original text set provided by the target server; obtaining a first label set and a second label set corresponding to the original text cluster set based on the original text cluster set; processing the first label set and the second label set corresponding to the original text cluster set to obtain a target label set corresponding to the original text cluster set. It can be seen that the first label set corresponding to the original text cluster set is a label processed by a large model, and the second label set corresponding to the original text cluster set is a label processed by rules. On the one hand, it can quickly generate labels through rules and a large model, avoiding complete reliance on manual label allocation and improving processing efficiency. On the other hand, it can generate labels in two different ways, avoiding the situation of missing or inaccurate labels, and also avoiding the inconsistency in label allocation caused by differences in subjective judgment between operators.

[0091] Example 2:

[0092] This second embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it performs the following steps:

[0093] S100, Obtain the original text set A = {A1, ..., A2} provided by the target server. i , ..., A m}, A i This refers to the i-th original text provided by the i-th target server, where i ranges from 1 to m, and m is the number of original texts provided by the target server, and m is a positive integer greater than 1.

[0094] S200, process A to obtain the original text cluster set B = {B1, ..., B1} corresponding to A. j , ..., B n}, B j It is the j-th original text cluster, where j ranges from 1 to n, and n is the number of original text clusters and is a positive integer greater than 1;

[0095] S300, Based on B, obtain the first tag set corresponding to B and the second tag set corresponding to B;

[0096] S400, process the first tag set corresponding to B and the second tag set corresponding to B to obtain the target tag set corresponding to B.

[0097] Example 3:

[0098] This third embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it performs the following steps:

[0099] S100, Obtain the original text set A = {A1, ..., A2} provided by the target server. i , ..., A m}, A i This refers to the i-th original text provided by the i-th target server, where i ranges from 1 to m, and m is the number of original texts provided by the target server, and m is a positive integer greater than 1.

[0100] S200, process A to obtain the original text cluster set B = {B1, ..., B1} corresponding to A. j , ..., B n}, B j It is the j-th original text cluster, where j ranges from 1 to n, and n is the number of original text clusters and is a positive integer greater than 1;

[0101] S300, Based on B, obtain the first tag set corresponding to B and the second tag set corresponding to B;

[0102] S400, process the first tag set corresponding to B and the second tag set corresponding to B to obtain the target tag set corresponding to B.

[0103] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0104] While specific embodiments of the invention have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of the invention. It should also be understood that various modifications can be made to the embodiments without departing from the scope and spirit of the invention. The scope of the invention is defined by the appended claims.

Claims

1. A label acquisition method based on an LLM model, characterized in that, The method includes the following steps: S100, Obtain the original text set A = {A1, ..., A2} provided by the target server. i , ..., A m }, A i It refers to the i-th original text provided by the i-th target server, where m is the number of original texts provided by the target server and m is a positive integer greater than 1; S200, process A to obtain the original text cluster set B = {B1, ..., B1} corresponding to A. j , ..., B n }, B j It is the j-th original text cluster, and n is the number of original text clusters, where n is a positive integer greater than 1; S300, based on B, obtain the first tag set corresponding to B and the second tag set corresponding to B, including: S301, Obtain B j ={B j1 , ..., B jy , ..., B jq }, B jy It refers to the y-th original text in the j-th original text cluster, where q is the number of original texts in the original text cluster and q is a positive integer greater than 1; S302, B jy The abstract text is input into the preset LLM model to obtain B. jy The corresponding first initial tag set, which includes several first initial tags; S303, extract B jy The original keyword set is input into the preset LLM model to obtain B. jy The corresponding second initial tag set, which includes several second initial tags; S304, B jy The abstract text is combined with the first preset tag library to obtain B. jy The corresponding third initial tag set includes several third initial tags. The first preset tag library includes several first preset tags and several preset texts corresponding to each first preset tag. When the preset text matches B... jy When the similarity of the summary text is greater than the preset similarity threshold, the first preset tag corresponding to the preset text is used as the third initial tag; S305, extract B jy The original keyword set is combined with the second preset tag library to obtain B. jy The corresponding fourth initial tag set includes several fourth initial tags. The second preset tag library includes several second preset tags and several preset keywords corresponding to each second preset tag. When the preset keywords are consistent with the original keywords, the second preset tag corresponding to the preset keywords is used as the fourth initial tag. S306, according to all B jy The corresponding first initial tag set and B jy The corresponding second initial tag set is used to obtain the first tag set corresponding to B; wherein, the first initial tag set and the second initial tag set are merged, and the initial tags with a frequency greater than a preset frequency threshold are selected from the merged initial tag set to construct the first tag set; S307, according to all B jy The corresponding third initial tag set and B jy Given the corresponding fourth initial tag set, obtain the second tag set corresponding to B; wherein, the initial tag set obtained by merging the third initial tag set and the fourth initial tag set is used as the second tag set; S400, process the first tag set corresponding to B and the second tag set corresponding to B to obtain the target tag set corresponding to B.

2. The label acquisition method based on the LLM model according to claim 1, characterized in that, The S200 procedure includes the following steps: S201, for A i Keyword extraction was performed on the abstract text to obtain A. i The corresponding original keyword set A 0 i ={A 0 i1 , ..., A 0 ix , ..., A 0 ip }, A 0 ix It is A i The x-th original keyword in the abstract text, where x ranges from 1 to p, and p is A i The number of original keywords in the abstract text, where p is a positive integer greater than 1; S202, according to A i The corresponding original tag set C i Generate A i The corresponding intermediate tag subset D i The original labels are obtained by inputting the original keywords into the preset LLM model. S203, Based on the intermediate tag set D = {D1, ..., D2} of the original text i , ..., D m }, thus obtaining the initial similarity matrix F of the original text, where F satisfies the following condition: Among them, F 11 It is the similarity between D1 and D1, F 1i It is D1 and D i The similarity between them, F 1m It is D1 and D m The similarity between them, F i1 It is D i Similarity with D1, F ii It is D i With D i The similarity between them, F im It is D i With D m The similarity between them, F m1 It is D m Similarity with D1, F mi It is D m With D i The similarity between them, F mm It is D m With D m The similarity between them, where F 11 , ..., F ii , ..., F mm All equal to 1; S204, Generate B based on F.

3. The label acquisition method based on the LLM model according to claim 2, characterized in that, In step S203, A i The corresponding intermediate tag subset D i =(D i1 , ..., D ia , ..., D ib ), where D ia It is A i The corresponding a-th intermediate label, where the value of a ranges from 1 to b and b is a positive integer greater than 1, is an original label whose priority is not less than a preset first priority threshold.

4. The label acquisition method based on the LLM model according to claim 3, characterized in that, For C i After deduplication, we get C. i The corresponding specified tag subset E i ={E i1 , ..., E ic , ..., E id }, E ic It is A i The corresponding c-th specified tag, where c ranges from 1 to d and d is a positive integer greater than 1.

5. The label acquisition method based on the LLM model according to claim 4, characterized in that, E i The corresponding specified tag priority set E 0 i ={E 0 i1 , ..., E 0 ic , ..., E 0 id }, E 0 ic It is E ic The corresponding specified tag priority, E 0 ic Meets the following conditions: Where, N ic This refers to A i China E ic The number of original keywords corresponding to the specified tag, N 0 ic This refers to E in A. ic The number of original keywords corresponding to the specified tag, p 0 It represents the total number of original keywords corresponding to A.

6. The label acquisition method based on the LLM model according to claim 2, characterized in that, In step S203, F 1i Meets the following conditions: Among them, K 0 1i It is D1 and D i The number of identical intermediate labels, K 1i It is D1 and D i The number of intermediate tags after deduplication.

7. The label acquisition method based on the LLM model according to claim 1, characterized in that, The target label set corresponding to B includes several target labels, wherein step S400 includes the following steps to determine the target labels: S401, process the first tag set and the second tag set corresponding to B to obtain the key priority set H = {H1, ..., H2} corresponding to B. g H z }, H g It is the priority of the g-th key tag, where the value of g ranges from 1 to z, and z is the number of key tags and z is a positive integer greater than 1; wherein, the key tags are the tags in the tag set generated after the first tag set and the second tag set; S402, when H g When the value is greater than or equal to △H, the key label is determined as the target label, where △H is a preset priority threshold. S403, when H g When < △H, the key label is determined to be a non-target label.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the label acquisition method based on the LLM model as described in any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the label acquisition method based on the LLM model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for fusing different types of labels and storage medium

    CN118861967A