Information identification optimization method and system based on large language model

By extracting the initial target information category tags on the large language model, generating derivative and semantic topic tags, and performing label aggregation, the problem of deviation and inaccuracy of classification results in complex text classification tasks is solved, and the accurate identification and efficient classification of harmful information is achieved.

CN120067904APending Publication Date: 2025-05-30XIAMEN MEIYABAIKE INFORMATION SECURITY RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411937811.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When facing complex text classification tasks, the classification results deviate from expectations, the classification accuracy does not meet the requirements, and the excessive labeling system increases the difficulty of large language models to analyze text and affects the analysis speed, and may even exceed the upper limit of text length that it can process.

Method used

By capturing complex implicit relationships, the initial target information category tags are extracted using a large language model, derivative topic tags and semantic topic tags are generated, and the final harmful information category tags are obtained through tag aggregation.

Benefits of technology

It improves the accurate recognition ability of harmful information in complex texts, enhances classification accuracy, and reduces the difficulty and delay of large language models when processing texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067904A_ABST
    Figure CN120067904A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information recognition, in particular to an information recognition optimization method and system based on a large language model, and the method comprises the steps: inputting a to-be-detected text X into a constructed large language model for information extraction, and obtaining an initial target information category label C1, a target information explanation EX, a harmful topic HarmfulTopic and a topic keyword HarmfulKW; generating a derivative topic label C2 by utilizing the target information explanation EX; a corresponding semantic topic label C3 is generated in combination with the obtained harmful topic HarmfulTopic and the topic keyword HarmfulKW; the C1, the C2 and the C3 form a label set C {C1, C2 and C3}, then voting is carried out on the set, and a label aggregation result is obtained and serves as a final harmful information category label Li; and outputting the obtained label aggregation result as a recognition result of the to-be-detected text X. According to the method, the strong comprehensive analysis capability, the flexible text generation skill and the efficient topic summarization capability of the large language model are utilized, and the recognition accuracy of the target information is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information recognition, and particularly relates to an information recognition optimization method and system based on a large language model. Background Art

[0002] With the popularization of social media, various social media platforms have become important platforms for people's daily communication. Although they carry a huge amount of information exchange, they inevitably accompany the spread of various harmful contents. These harmful data and information pose a serious threat to social stability, network security and personal rights. Moreover, in most cases, these harmful data and information will enter the user's field of vision in the form of text on social media.

[0003] In order to eliminate these harmful data and information that appear on various social media platforms, the key is to accurately classify them. When performing text classification tasks to achieve the recognition of target information, large language models have demonstrated excellent capabilities in this technical field. Especially when facing challenging tasks such as cross-domain migration and few-shot learning, they have significant advantages compared with traditional supervised classification algorithms: there is no need to separately adjust the model structure or loss function for each specific task, which greatly enhances flexibility and generalization ability. However, the application of large language models is also accompanied by a series of challenges, especially the high training cost, usage cost and computational latency caused by their huge parameter scale. These problems are particularly prominent when only dealing with a single or a small number of tasks.

[0004] To overcome these problems, researchers have explored how to effectively reduce large language models while maintaining their powerful capabilities to meet the needs of specific tasks: one is through knowledge distillation technology, using a more compact "student" model to learn and simulate the intermediate feature representations and final outputs of a large-scale "teacher" model, so as to reduce the model size and computational complexity while maintaining performance; the other is to use the knowledge and generalization ability learned by large language models on extensive data to generate or enhance training data, and then train a small model with excellent performance. This method not only reduces the training cost, but also promotes the effective transfer and reuse of knowledge. However, for the recognition and classification of predefined information category labels by large language models, even on the basis of designing professional prompt engineering, there are still the following problems: (1) Large language models cannot well understand custom target information categories and will generate "nearby labels" that are not in the predefined categories, resulting in classification results deviating from expectations; (2) The label system is too small and the classification accuracy cannot be achieved; (3) The label system is too large, and it is necessary to increase the representation of prompt words, which will increase the difficulty of the large language model in analyzing and processing text and affect the analysis speed, and even exceed the upper limit of the text length that the large language model can process.

[0005] In view of the above deficiencies, this solution utilizes the powerful generalization ability and in-depth understanding ability of large language models to propose an optimized method for information recognition based on large language models, which can accurately identify harmful information hidden in complex texts by capturing complex implicit relationships. Summary of the Invention

[0006] The object of the present invention is to provide an optimized method and system for information recognition based on large language models, so as to solve the problems that the classification results deviate from expectations, the classification accuracy fails to meet the requirements, and the excessive label system increases the difficulty of analyzing and processing texts by large language models and affects the analysis speed when facing complex text classification tasks, and even exceeds the upper limit of the text length that it can process.

[0007] An optimized method for information recognition based on large language models provided by the present invention includes the following steps:

[0008] S1: Text acquisition: Input the text and use it as the text to be detected X;

[0009] S2: Information extraction: Input the text to be detected X obtained in step S1 above into a pre-constructed large language model for information extraction, and obtain the initial target information category label C 1 , target information explanation E X , harmful topic HarmfulTopic and topic keywords HarmfulKW; where HarmfulTopic = {T 1 ,...T m}, including multiple topics involved in the text to be detected X; HarmfulKW = {W 1 ,...W k} includes topic keywords or topic phrases of the above multiple topics;

[0010] S3: Generation of derivative topic labels: Use the target information explanation E obtained in step S2 above X to generate derivative topic labels C 2 ; where the generation method of the derivative topic label C 2 is as follows:

[0011] C 2 = Det(E x )

[0012] where Det is a large language model detector that converts the explanatory information into the corresponding derivative topic label, and C 2 ∈ L, and L is a predefined set of target information categories;

[0013] S4: Semantic Topic Tag Generation: Generate the corresponding semantic topic tag C by combining the harmful topic HarmfulTopic and topic keywords HarmfulKW obtained in the above step S2 3 ;

[0014] S5: Tag Aggregation: Aggregate the initial target information category tags C obtained in step S2 1 , the derivative topic tags C generated in step S3 2 and the semantic topic tags C generated in step S4 3 to form a tag set C{C 1 , C 2 , C 3}, and then vote on this set to obtain the tag aggregation result as the final harmful information category tag L i , L i = Vote(C), where L i ∈L;

[0015] S6: Recognition Result Output: Output the tag aggregation result obtained in step S5 as the recognition result of the text X to be detected

[0016] An information recognition optimization method based on a large language model as described above is further preferably that step S4 specifically includes

[0017] S4.1: Convert the harmful topic HarmfulTopic and topic keywords HarmfulKW into key vectors KV through the semantic encoder Encode. The calculation formula of KV is

[0018] KV = Encode(HarmfulTopic, HarmfulKW)

[0019] where Encode is the Bert encoder of the Chinese-based language model; after being encoded by the Encode encoder, the key vector KV = {HT, HKW}, where HT is the semantic representation of the harmful topic HarmfulTopic after semantic encoding, HT = {HT 1 , HT 2 ,..., HT m}; HKW is the semantic representation of the topic keyword HarmfulKW after semantic encoding, HKW = {HKW 1 , HKW 2 ,..., HKW k};

[0020] S4.2: Semantically encode the topic tags W in the knowledge database through the semantic encoder Encode to obtain the corresponding knowledge data vector WV, where W = {W ij};

[0021] S4.3: Calculate the similarity between the key vector KV and the knowledge data vector WV to obtain the similarity matrix CS. The calculation formula is as follows:

[0022] CS = f(KV, WV)

[0023] where f is the cosine similarity, KV = [KV 1 , KV 2 , …, KV T ,

[0024] CS is expressed as:

[0025]

[0026] where n is the total number of predefined target information label categories; T is the number of harmful topics after semantic encoding;

[0027] S4.4: Perform label mapping on the similarity matrix CS and the labels in the label database through the label mapping T to obtain the semantic topic label C 3 , C 3 = T(CS).

[0028] An information recognition optimization method based on a large language model as described above is further preferably that step S4.4 specifically includes: First, convert the above similarity matrix CS into a sub-label matrix The specific conversion rule is: If KV 1 WV 11 is greater than ε, it is denoted as the corresponding label W 11 , otherwise it is denoted as 0, where ε is a predefined threshold; the converted sub-label matrix is as follows:

[0029]

[0030] Second, use the mapping relationship in the label database to convert the sub-label matrix into a predefined label matrix as follows:

[0031]

[0032] Finally, determine the label L with the highest frequency of occurrence in i as the final semantic label C 3 , where the calculation formula of C 3 is as follows:

[0033]

[0034] An information recognition optimization method based on a large language model as described above is further preferably that the construction steps of the label database specifically include: obtaining topic labels W through three methods: expert construction, user customization, and automatic update ij , and using the data interface, establishing a mapping relationship with the predefined target information label categories in combination with the knowledge database to obtain the label database

[0035] An information recognition optimization method based on a large language model as described above is further preferably that the label database further includes an automatic update step, which specifically includes: first, starting the intelligent agent Agent corresponding to the label database through the identified target information category label to generate new topic labels W ij , where different target information category labels correspond to different Agents; then, using the newly generated W ij and the corresponding target information category labels to update the label database

[0036] The present invention also discloses an information recognition optimization system based on a large language model, including:

[0037] Text acquisition module: used to take the input text as the text to be detected X

[0038] Information extraction module: used to input the text to be detected X obtained by the text acquisition module into the pre-constructed large language model for information extraction, and obtain the initial target information category label C 1 , target information explanation E X , harmful topic HarmfulTopic and topic keyword HarmfulKW; where HarmfulTopic = {T 1 ,...T m}, including multiple topics involved in the text to be detected X; HarmfulKW = {W 1 ,...W k} includes the topic keywords or topic phrases of the above multiple topics

[0039] Derivative topic label generation module: used to generate derivative topic labels C X using the target information explanation E 2 obtained by the above information extraction module; where the generation method of the derivative topic label C 2 is as follows:

[0040] C 2 = Det(E x )

[0041] Among them, Det is a large language model detector that converts explanatory information into corresponding derivative topic tags, where C 2 ∈L, and L is a predefined set of target information categories;

[0042] Semantic topic tag generation module: used to generate corresponding semantic topic tags C by combining the harmful topic HarmfulTopic and topic keywords HarmfulKW obtained by the above information extraction module 3 ;

[0043] Tag aggregation module: used to aggregate the initial target information category tags C obtained by the above information extraction module 1 , the derivative topic tags C generated by the derivative topic tag generation module 2 , and the semantic topic tags C generated by the semantic topic tag generation module 3 to form a tag set C{C 1 , C 2 , C 3}, and then vote on this set to obtain the tag aggregation result as the final harmful information category label L i , L i =Vote(C), where L i ∈L;

[0044] Recognition result output module: used to output the tag aggregation result obtained by the tag aggregation module as the recognition result of the text X to be detected.

[0045] An information recognition optimization system based on a large language model as described above is further preferably that the semantic topic tag generation module further includes:

[0046] Key vector generation module: used to convert the harmful topic HarmfulTopic and topic keywords HarmfulKW into key vectors KV through the semantic encoder Encode, and the calculation formula of KV is:

[0047] KV = Encode(HarmfulTopic, HarmfulKW)

[0048] Among them, Encode is a Chinese-based language model Bert encoder; after being encoded by the Encode encoder, the key vector KV = {HT, HKW}, where HT is the semantic representation of the harmful topic HarmfulTopic after semantic encoding, HT = {HT 1 , HT 2 ,..., HT m}; HKW is the semantic representation of the topic keyword HarmfulKW after semantic encoding, HKW = {HKW 1 , HKW2 ,...,HKW k};

[0049] Knowledge data vector generation module: used to semantically encode the topic tags W in the knowledge database through the semantic encoder Encode to obtain the corresponding knowledge data vector WV, where W = {W ij};

[0050] Similarity calculation module: used to calculate the similarity between the key vector KV and the knowledge data vector WV to obtain the similarity matrix CS. The calculation formula is as follows:

[0051] CS = f(KV, WV)

[0052] where f is the cosine similarity, KV = [KV 1 , KV 2 , …, KV T ,

[0053] CS is expressed as:

[0054]

[0055] where n is the total number of predefined target information label categories; T is the number of harmful topics after semantic encoding;

[0056] Mapping module: used to map the similarity matrix CS with the labels in the label database through the label mapping T to obtain the semantic topic label C 3 , C 3 = T(CS).

[0057] An information recognition optimization system based on a large language model as described above is further preferably that the mapping module includes:

[0058] Similarity matrix conversion module: used to convert the above similarity matrix CS into a sub-label matrix The specific conversion rule is: if KV 1 W 11 is greater than ε, it is denoted as the corresponding label W 11 , otherwise it is denoted as 0, where ε is a predefined threshold; the converted sub-label matrix is as follows:

[0059]

[0060] Sub-label matrix conversion module: used to convert the sub-label matrix into a predefined label matrix As follows:

[0061]

[0062] Determination module: Determine the tag L with the highest occurrence frequency from as the final semantic tag C i , where the calculation formula of C 3 is as follows: 3 For an information recognition optimization system based on a large language model as described above, further preferably, the tag database module further includes a construction module: used to obtain topic tags W

[0063]

[0064] through three methods: expert construction, user customization, and automatic update, and use the data interface to establish a mapping relationship with the predefined target information tag categories in combination with the knowledge database to obtain the tag database. ij For an information recognition optimization system based on a large language model as described above, further preferably, the tag database module further includes an automatic update module, which includes: a new topic tag generation module: starting the intelligent agent Agent corresponding to the tag database through the identified target information category tags to generate new topic tags W

[0065] , where different target information category tags correspond to different Agents; an update sub-module: using the newly generated W ij and the corresponding target information category tags to update the tag database. ij For an information recognition optimization system based on a large language model as described above, further preferably, the tag database module further includes an automatic update module, which includes: a new topic tag generation module: starting the intelligent agent Agent corresponding to the tag database through the identified target information category tags to generate new topic tags W

[0066] The beneficial effects of the present invention are: Text acquisition: Input the text as the text to be detected X; Information extraction: Input the text to be detected X obtained according to the above step S1 into the pre-constructed large language model for information extraction to obtain the initial target information category tag C 1 , the target information explanation E X , the harmful topic HarmfulTopic, and the topic keyword HarmfulKW; where HarmfulTopic = {T 1 ,...T m}, including multiple topics involved in the text to be detected X; HarmfulKW = {W 1 ,...W k} includes the topic keywords or topic phrases of the above multiple topics; Derived topic tag generation: Use the target information explanation E X obtained in the above step S2 to generate derived topic tags C 2; Semantic topic label generation: Combine the harmful topic HarmfulTopic and topic keywords HarmfulKW obtained in the above step S2 to generate the corresponding semantic topic label C 3 ; Label aggregation: Aggregate the obtained initial target information category labels C 1 , the generated derivative topic labels C 2 and the semantic topic labels C 3 to form a label set C{C 1 , C 2 , C 3}, then vote on this set to obtain the label aggregation result, which is used as the final harmful information category label L i , L i = Vote(C), where L i ∈ L; Recognition result output: According to the obtained label aggregation result, output it as the recognition result of the text X to be detected. The present invention makes full use of the powerful comprehensive analysis ability, flexible text generation skills and efficient topic summarization ability of the large language model, aims at generating detailed text explanation materials and in-depth topic insight information, and through the constructed mapping mechanism, closely connects these features with the target information category, enhancing the accuracy of the model's recognition of the target information. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0068] Figure 1 FIG. is a flowchart of the steps of the information recognition optimization method based on a large language model of the present invention.

[0069] Figure 2 FIG. is a flowchart of the steps of the semantic topic label generation of the present invention.

[0070] Figure 3 FIG. is a flowchart of the steps of the label database construction of the present invention.

[0071] Figure 4 FIG. is a flowchart of the steps of the automatic update of the label database of the present invention.

[0072] Figure 5 FIG. is a schematic diagram of the composition structure of the information recognition optimization system based on a large language model of the present invention.

[0073] Figure 6 FIG. is a schematic diagram of the composition structure of the semantic topic label generation module of the present invention.

[0074] Figure 7 This is a schematic diagram of the composition structure of the mapping module of the present invention.

[0075] Figure 8 This is a schematic diagram of the composition structure of the label database module of the present invention. Detailed implementation manners

[0076] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features, and effects of the information recognition optimization method and system based on the large language model proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0077] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0078] The following specifically describes the specific solution of the information recognition optimization method based on the large language model provided by the present invention with reference to the accompanying drawings.

[0079] Please refer to Figure 1 , which shows a flowchart of the steps of the information recognition optimization method based on the large language model provided by an embodiment of the present invention. As Figure 1 shown, the information recognition optimization method provided by the embodiments of the present application includes the following steps S1 to S6:

[0080] S1: Text acquisition: Input the text and use it as the text to be detected X;

[0081] It should be noted that: The text in this step S1 can be the web content in the Internet platform crawled by the web crawler, or the text input by the user on the Internet web page or application.

[0082] S2: Information extraction: Input the text to be detected X obtained according to the above step S1 into the pre-constructed large language model for information extraction, and obtain the initial target information category label C 1 , target information explanation E X , harmful topic HarmfulTopic, and topic keyword HarmfulKW; where HarmfulTopic = {T 1 ,...T m}, which contains multiple topics involved in the text to be detected X; HarmfulKW = {W 1 ,...Wk Topic keywords or topic phrases that include the above multiple topics;

[0083] It should be noted that: The large language model used in this step S2 is pre-built and is used for text information extraction. The information obtained after being detected by the large language model includes at least the following four items: the initial target information category label C 1 , the target information explanation E X , the harmful topic HarmfulTopic and the topic keyword HarmfulKW; among them, the harmful topic HarmfulTopic and the topic keyword HarmfulKW can appear in pairs or not; and the topic keyword HarmfulKW can be not only the keyword of the relevant topic but also the phrase corresponding to these topics.

[0084] S3: Generation of derivative topic labels: Use the target information explanation E obtained in the above step S2 X to generate the derivative topic label C 2 ; among them, the generation method of the derivative topic label C 2 is as follows:

[0085] C 2 = Det(E x )

[0086] Det is a large language model detector that converts the explanatory information into the corresponding derivative topic label, where C 2 ∈L, and L is a predefined set of target information categories;

[0087] It should be noted that: Here, the powerful comprehensive analysis ability of the large language model is fully utilized to extract rich explanatory information, and the obtained explanatory information is used to generate the corresponding labels, so as to provide a guarantee for obtaining accurate recognition results in subsequent information processing. The large language model detector Det used here is a well-known information detector and will not be elaborated here; and the obtained derivative topic label C 2 belongs to the range of the set of target information categories predefined by the user.

[0088] S4: Generation of semantic topic labels: Combine the harmful topic HarmfulTopic and the topic keyword HarmfulKW obtained in the above step S2 to generate the corresponding semantic topic label C 3 ;

[0089] It should be noted that: here, by leveraging the powerful comprehensive analysis and topic summarization capabilities of the large language model, deeper relevant topic information hidden in the input text is mined (here mainly targeting harmful topics); the mined topic information includes harmful topics HarmfulTopic and topic keywords HarmfulKW, and the corresponding topic tags are generated based on the above information, thus providing a guarantee for subsequent obtaining accurate recognition results. Compared with the initial target information category tags, this deeper relevant topic information is more concealed and less noticeable to users; therefore, by utilizing the characteristics of the powerful large language model, semantic topic tags are generated for the mined deep topic information, and combining with the initial target information category tags can be more conducive to the accuracy of information recognition.

[0090] S5: Label aggregation: Combine the initial target information category label C obtained in step S2 1 with the derived topic label C generated in step S3 2 and the semantic topic label C generated in step S4 3 to form a label set C{C 1 , C 2 , C 3}, then vote on this set to obtain the label aggregation result, which is used as the final harmful information category label L i , L i =Vote(C), where L i ∈L;

[0091] It should be noted that: here, by voting on the label set composed of the initial target information category label C 1 , the derived topic label C 2 and the semantic topic label C 3 , the label aggregation result is obtained; the voting strategy usually is that the one with the highest hit rate of the preferred label category is the final label; if the hit rates of the three are equal, output C 3 as the final label.

[0092] S6: Output of recognition result: According to the label aggregation result obtained in step S5, it is output as the recognition result of the text X to be detected.

[0093] It should be noted that: the output methods here can include popping up a dialog box, voice reminder, etc.

[0094] Please refer to Figure 2 , which shows the step flow chart of semantic topic label generation provided by an embodiment of the present invention. As Figure 2 shown, the semantic topic label generation steps provided by the embodiments of the present application include the following steps S4.1 to S4.4:

[0095] The specific steps for generating the semantic topic tags are as follows:

[0096] S4.1: Use the semantic encoder Encode to convert the harmful topic HarmfulTopic and the topic keyword HarmfulKW into a key vector KV. The calculation formula for KV is:

[0097] KV = Encode(HarmfulTopic, HarmfulKW)

[0098] Where Encode is the Bert encoder, a Chinese-based language model; after encoding by the Encode encoder, the key vector KV = {HT, HKW}, where HT is the semantic representation of the harmful topic HarmfulTopic after semantic encoding, HT = {HT 1 , HT 2 ,..., HT m}; HKW is the semantic representation of the topic keyword HarmfulKW after semantic encoding, HKW = {HKW 1 , HKW 2 ,..., HKW k};

[0099] It should be noted that here, the harmful topic HarmfulTopic and the topic keyword HarmfulKW are semantically encoded using the semantic encoder Encode to generate the corresponding key vectors. The key vectors generated in this way contain the relevant information of the harmful topic and the topic keyword. Performing semantic encoding to obtain the corresponding vector information is for calculating the similarity between information. Here, it is for realizing the similarity between the harmful topic and the topic tags in the knowledge database.

[0100] S4.2: Use the semantic encoder Encode to semantically encode the topic tag W in the knowledge database to obtain the corresponding knowledge data vector WV, where W = {W ij};

[0101] It should be noted that as explained in step S4.1, using the semantic encoder to encode the topic tag W to obtain the corresponding knowledge data vector is also for calculating the similarity between the two.

[0102] S4.3: Calculate the similarity between the key vector KV and the knowledge data vector WV to obtain the similarity matrix CS. The calculation formula is as follows:

[0103] CS = f(KV, WV)

[0104] Where f is the cosine similarity, KV = [KV 1 , KV2 , …, KV T ,

[0105] CS is expressed as:

[0106]

[0107] Where n is the total number of predefined target information label categories; T is the number of harmful topics;

[0108] It should be noted that: the key vector KV here is a T-dimensional vector group, and the obtained knowledge data vector WV is an n*n vector matrix; when solving the similarity matrix between the two, each element KV of the vector group i needs to be calculated separately with each element WV in the vector matrix ij to obtain a similarity matrix with T rows and n*n columns.

[0109] S4.4: Map the similarity matrix CS with the labels in the label database through the label mapping T to obtain the semantic topic label C 3 , C 3 = T(CS).

[0110] It should be noted that: here, the similarity matrix needs to be first converted into a corresponding sub-label matrix, and the conversion rule here is by means of threshold comparison; then each element in the sub-label matrix is corresponded to the predefined label L with reference to the label database i to obtain the predefined label matrix. Finally, the semantic topic label is determined from the predefined label matrix by voting.

[0111] The specific steps of the label mapping T include:

[0112] First, convert the above similarity matrix CS into a sub-label matrix The specific conversion rule is: if KV 1 WV 11 is greater than ε, it is denoted as the corresponding label W 11 , otherwise it is denoted as 0, where ε is a predefined threshold; the converted sub-label matrix is as follows:

[0113]

[0114] Secondly, use the mapping relationship in the label database to convert the sub-label matrix into a predefined label matrix as follows:

[0115]

[0116] Finally, determine the tag L with the highest occurrence frequency from as the final semantic tag C i , where the calculation formula of C 3 is as follows: 3 It should be noted that: the voting strategy is that the one with the highest hit rate of the preferred tag category is the final tag; if the hit rates are equal, output C

[0117]

[0118] as the final tag. 3

[0119] Please refer to Figure 3 , which shows the step flowchart of building the tag database provided by an embodiment of the present invention, as Figure 3 shown, the specific steps of building the tag database are as follows:

[0120] Obtain the topic tag W through three methods: expert construction, user-defined, and automatic update ij , and use the data interface to establish a mapping relationship with the predefined target information tag category in combination with the knowledge database, so as to obtain the tag database.

[0121] It should be noted that: obtaining the topic tag through three methods here is to ensure the flexibility and dynamics of the tags stored in the tag database, so that the final recognition result is more accurate. And this mapping relationship is a one-to-many mapping relationship.

[0122] Figure 4 Please refer to , which shows the step flowchart of automatically updating the tag database provided by an embodiment of the present invention, as Figure 4 shown, the specific steps of automatically updating the tag database are as follows:

[0123] First, start the intelligent agent Agent corresponding to the tag database through the identified target information category tag to generate a new topic tag W ij , where different target information category tags correspond to different Agents;

[0124] Then, use the newly generated W ij and the corresponding target information category tag to update the tag database.

[0125] It should be noted that: there are multiple different intelligent agents Agent here, which is for the diversity when generating new topic tags, so as to be more conducive to enriching the tag database.

[0126] Figure 5 Please refer to Figure 5, which shows the schematic composition structure of the information recognition optimization system based on a large language model provided by an embodiment of the present invention, as Figure 5 shown, the information recognition optimization system based on a large language model includes:

[0127] Text acquisition module: used to take the input text as the text to be detected X according to the input text;

[0128] It should be noted that: the text obtained by the text acquisition module can be the web content in the Internet platform crawled by the web crawler, or the text input by the user in the web page or application program on the Internet.

[0129] Information extraction module: used to input the text to be detected X obtained by the text acquisition module into a pre-constructed large language model for information extraction, and obtain the initial target information category label C 1 , target information explanation E X , harmful topic HarmfulTopic and topic keyword HarmfulKW; where HarmfulTopic = {T 1 ,...T m}, including multiple topics involved in the text to be detected X; HarmfulKW = {W 1 ,...W k} includes the topic keywords or topic phrases of the above multiple topics;

[0130] It should be noted that: the large language model used in the information extraction module is pre-constructed and is used for text information extraction. The information obtained after being detected by the large language model includes at least the following four items: the initial target information category label C 1 , target information explanation E X , harmful topic HarmfulTopic and topic keyword HarmfulKW; among them, the harmful topic HarmfulTopic and topic keyword HarmfulKW can appear in pairs or not; and the topic keyword HarmfulKW can be not only the keyword of the relevant topic but also the phrase corresponding to these topics.

[0131] Derivative topic label generation module: used to generate derivative topic label C X using the target information explanation E 2 obtained by the above information extraction module; where the generation method of the derivative topic label C 2 is as follows:

[0132] C 2 = Det(E x )

[0133] Among them, Det is a large language model detector that converts explanatory information into corresponding derivative topic tags, where C 2 ∈L, and L is a predefined set of target information categories;

[0134] It should be noted that: here, the powerful comprehensive analysis ability of the large language model is fully utilized to extract rich explanatory information, and the obtained explanatory information is used to generate corresponding tags, so as to provide guarantee for subsequent information processing to obtain accurate recognition results. The large language model detector Det used here is a well-known information detector and will not be elaborated here; while the obtained derivative topic tag C 2 belongs to the range of the set of target information categories predefined by the user.

[0135] Semantic topic tag generation module: used to generate corresponding semantic topic tags C 3 ;

[0136] It should be noted that: here, the powerful comprehensive analysis and topic summarization ability of the large language model is used to mine the deeper relevant topic information hidden in the input text (here mainly targeting harmful topics); the topic information mined here includes the harmful topic HarmfulTopic and the topic keyword HarmfulKW, and the above information is used to generate corresponding topic tags, so as to provide guarantee for subsequent auxiliary to obtain accurate recognition results. Compared with the initial target information category label, this deeper relevant topic information is more hidden and less noticeable to users; therefore, using the characteristics of the powerful large language model, the mined deep topic information is generated into semantic topic tags, and combining with the initial target information category label can be more conducive to the accuracy of information recognition.

[0137] Label aggregation module: used to combine the initial target information category label C 1 obtained by the above information extraction module, the derivative topic tag C 2 generated by the derivative topic tag generation module, and the semantic topic tag C 3

[0138] generated by the semantic topic tag generation module to form a label set C{C 1 , C 2 , C 3}, and then vote on this set to obtain the label aggregation result as the final harmful information category label L i , L i =Vote(C), where L i ∈L;

[0139] It should be noted that: here, by voting on the label set composed of the initial target information category label C 1 , the derived topic label C 2 , and the semantic topic label C 3 , the label aggregation result is obtained; the voting strategy usually is that the one with the highest hit rate of the preferred label category is the final label; if the hit rates of the three are equal, output C 3 as the final label.

[0140] Recognition result output module: used to output the label aggregation result obtained by the label aggregation module as the recognition result of the text X to be detected.

[0141] It should be noted that: the output methods here may include dialog box pop-up, voice reminder, etc.

[0142] Please refer to Figure 6 , which shows the schematic composition structure diagram of the semantic topic label generation module provided by an embodiment of the present invention. As Figure 6 shown, the semantic topic label generation module further includes:

[0143] Key vector generation module: used to convert the harmful topic HarmfulTopic and the topic keyword HarmfulKW into a key vector KV through the semantic encoder Encode. The calculation formula of KV is:

[0144] KV = Encode(HarmfulTopic, HarmfulKW)

[0145] where Encode is the Bert encoder based on the Chinese language model; after being encoded by the Encode encoder, the key vector KV = {HT, HKW}, where HT is the semantic representation of the harmful topic HarmfulTopic after semantic encoding, HT = {HT 1 , HT 2 ,..., HT m}; HKW is the semantic representation of the topic keyword HarmfulKW after semantic encoding, HKW = {HKW 1 , HKW 2 ,..., HKW k};

[0146] It should be noted that here, the harmful topic "HarmfulTopic" and the topic keyword "HarmfulKW" are semantically encoded using the semantic encoder "Encode" to generate the corresponding key vectors for both. The key vectors generated in this way contain the relevant information of the harmful topic and the topic keyword. Performing semantic encoding to obtain the corresponding vector information is for solving the similarity between information. Here, it is to achieve the similarity between the harmful topic and the topic labels in the knowledge database.

[0147] Knowledge data vector generation module: It is used to semantically encode the topic label "W" in the knowledge database through the semantic encoder "Encode" to obtain the corresponding knowledge data vector "WV", where W = {W ij};

[0148] It should be noted that: as explained in the part of the key vector generation module, encoding the topic label "W" through the semantic encoder to obtain the corresponding knowledge data vector is also for solving the similarity between the two.

[0149] Similarity calculation module: It is used to calculate the similarity between the key vector "KV" and the knowledge data vector "WV" to obtain the similarity matrix "CS", and the calculation formula is as follows:

[0150] CS = f(KV, WV)

[0151] where f is the cosine similarity, KV = [KV 1 , KV 2 , …, KV T ,

[0152] CS is expressed as:

[0153]

[0154] where n is the total number of predefined target information label categories; T is the number of harmful topics after semantic encoding;

[0155] It should be noted that: the key vector "KV" here is a T-dimensional vector group, and the obtained knowledge data vector "WV" is an n*n vector matrix; when solving the similarity matrix between the two, each element KV i of the vector group needs to be calculated separately with each element WV ij of the vector matrix, thus obtaining a similarity matrix with T rows and n*n columns.

[0156] Mapping module: It is used to perform label mapping on the similarity matrix "CS" and the labels in the label database through the label mapping "T" to obtain the semantic topic label "C 3 , C 3= T(CS).

[0157] It should be noted that: here, the similarity matrix needs to be first converted into a corresponding sub-label matrix, and the conversion rule here is by means of threshold comparison; then each element in the sub-label matrix is corresponded to a predefined label L with reference to the label database i , so as to obtain a predefined label matrix. Finally, the semantic topic label is determined from the predefined label matrix by means of voting.

[0158] Please refer to Figure 7 , which shows a schematic diagram of the composition structure of the mapping module provided by an embodiment of the present invention, as Figure 7 shown, the mapping module further includes:

[0159] Similarity matrix conversion module: used to convert the above similarity matrix CS into a sub-label matrix The specific conversion rule is: if KV 1 W 11 is greater than ε, it is denoted as the corresponding label W 11 , otherwise it is denoted as 0, where ε is a predefined threshold; the converted sub-label matrix is as follows:

[0160]

[0161] Sub-label matrix conversion module: used to convert the sub-label matrix into a predefined label matrix as follows:

[0162]

[0163] Determination module: determine from the label L with the highest occurrence frequency i , as the final semantic label C 3 , where the calculation formula of C 3 is as follows:

[0164]

[0165] It should be noted that: the voting strategy is that the one with the highest hit rate of the preferred label category is the final label; if the hit rates are equal, output C 3 as the final label.

[0166] Please refer to Figure 8 , which shows a schematic diagram of the composition structure of the label database module provided by an embodiment of the present invention.

[0167] In this embodiment, the tag database module includes a construction module: which is used to obtain topic tags W through three methods: expert construction, user customization, and automatic update ij , and by using the data interface and combining with the knowledge database, establish a mapping relationship with the predefined target information tag categories, so as to obtain the tag database.

[0168] It should be noted that: obtaining topic tags through three methods here is to ensure the flexibility and dynamics of the tags stored in the tag database, so that the final recognition result is more accurate. And this mapping relationship is a one-to-many mapping relationship.

[0169] For example Figure 8 As shown, in this embodiment, the tag database module further includes an automatic update module, and the automatic update module includes a new topic tag generation module and an update sub-module:

[0170] New topic tag generation module: Through the identified target information category tags, start the intelligent agent Agent corresponding to the tag database to generate new topic tags W ij , where different target information category tags correspond to different Agents;

[0171] Update sub-module: Use the newly generated W ij and the corresponding target information category tags to update the tag database.

[0172] It should be noted that: there are multiple different intelligent agents Agent here to ensure the diversity when generating new topic tags, which is more conducive to enriching the tag database.

[0173] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An information recognition optimization method based on a large language model, characterized in that: The method comprises the following steps: S1: Text acquisition: input text and use it as the text to be detected X; S2: Information extraction: Input the to-be-detected text X obtained in step S1 above into the pre-built large language model for information extraction, and obtain the initial target information category label C1 and target information explanation E X , harmful topic HarmfulTopic and topic keyword HarmfulKW; where HarmfulTopic={T1,...T m }, including multiple topics involved in the text to be detected X; HarmfulKW = {W1,...W k }Topic keywords or topic phrases containing multiple topics mentioned above; S3: Derived topic tag generation: Use the target information obtained in step S2 to explain E X Generate a derived topic label C2; wherein the derived topic label C2 is generated as follows: C2=It(E x ) Among them, Det is a large language model detector that converts explanation information into corresponding derived topic labels, where C2∈L, L is a predefined set of target information categories; S4: Semantic topic tag generation: Combine the harmful topic HarmfulTopic and topic keyword HarmfulKW obtained in step S2 above to generate the corresponding semantic topic tag C3; S5: Tag aggregation: The initial target information category label C1 obtained in step S2, the derived topic label C2 generated in step S3, and the semantic topic label C3 generated in step S4 are combined into a label set C{C1, C2, C3}, and then the set is voted to obtain the label aggregation result as the final harmful information category label L i , L i =Vote(C), where L i ∈L; S6: Recognition result output: According to the tag aggregation result obtained in step S5, it is output as the recognition result of the text to be detected X.

2. The information recognition optimization method based on a large language model according to claim 1, characterized in that: The specific steps of generating the semantic topic tags include: S4.1: The harmful topic HarmfulTopic and topic keyword HarmfulKW are converted into a key vector KV through the semantic encoder Encode. The calculation formula of KV is: KV=Encode(HarmfulTopic,HarmfulKW) Among them, Encode is a language model Bert encoder based on Chinese; after being encoded by Encode encoder, the key vector KV = {HT, HKW}, where HT is the semantic representation of the harmful topic HarmfulTopic after semantic encoding, HT = {HT1, HT2, ..., HT m }; HKW is the semantic representation of the topic keyword HarmfulKW after semantic encoding, HKW = {HKW1, HKW2, ..., HKW k }; S4.2: The topic tag W in the knowledge database is semantically encoded by the semantic encoder Encode to obtain the corresponding knowledge data vector WV, where W = {W ij }; S4.3: Calculate the similarity between the key vector KV and the knowledge data vector WV to obtain the similarity matrix CS. The calculation formula is as follows: CS=f(KV,WV) Where f is cosine similarity, KV = [KV1, KV2, …, KV T ], CS is expressed as: Where n is the total number of predefined target information tag categories; T is the number of harmful topics after semantic encoding; S4.4: Through label mapping T, the similarity matrix CS is labeled with the labels in the label database to obtain the semantic topic label C3, C3 = T (CS).

3. The information recognition optimization method based on a large language model according to claim 2 is characterized in that: The specific steps of the label mapping T include: First, convert the above similarity matrix CS into a sub-label matrix The specific conversion rules are: if KV1WV 11 is greater than ε, then it is recorded as the corresponding label W 11 , otherwise it is recorded as 0, where ε is the predetermined threshold; the transformed sub-label matrix as follows: Secondly, the sub-label matrix is ​​transformed into Convert to a predefined label matrix as follows: Finally, from Determine the label L that appears most frequently i , as the final semantic label C3, where the calculation formula of C3 is as follows:

4. The information recognition optimization method based on a large language model according to claim 2 is characterized in that: The steps for constructing the label database are specifically as follows: Obtain topic tags through three methods: expert construction, user customization and automatic update ij ,Using the data interface, combined with the knowledge database, a mapping relationship with the predefined target information label categories is established, thus obtaining the label database.

5. The information recognition optimization method based on a large language model according to any one of claims 2 to 4, characterized in that: The tag database also includes an automatic update step, which is as follows: First, by identifying the target information category label, the agent corresponding to the label database is started to generate a new topic label W ij , where different target information category labels correspond to different Agents; Then, use the newly generated W ij And the corresponding target information category label, update the label database.

6. An information recognition optimization system based on a large language model, characterized in that: The information recognition optimization method based on a large language model according to any one of claims 1 to 5 comprises: Text acquisition module: used to take the input text as the text to be detected X; Information extraction module: used to input the to-be-detected text X obtained by the text acquisition module into the pre-built large language model for information extraction, and obtain the initial target information category label C1 and the target information explanation E X , harmful topic HarmfulTopic and topic keyword HarmfulKW; where HarmfulTopic={T1,...T m }, including multiple topics involved in the text to be detected X; HarmfulKW = {W1,...W k }Topic keywords or topic phrases containing multiple topics mentioned above; Derived topic tag generation module: used to explain the target information obtained by the above information extraction module X Generate a derived topic label C2; wherein the derived topic label C2 is generated as follows: C2=It(E x ) Among them, Det is a large language model detector that converts explanation information into corresponding derived topic labels. C2∈L, L is a set of predefined target information categories; Semantic topic tag generation module: used to combine the harmful topic HarmfulTopic obtained by the above information extraction module and the topic keyword HarmfulKW to generate the corresponding semantic topic tag C3; Label aggregation module: used to combine the initial target information category label C1 obtained by the above information extraction module, the derived topic label C2 generated by the derived topic label generation module, and the semantic topic label C3 generated by the semantic topic label generation module into a label set C{C1, C2, C3}, and then vote on the set to obtain the label aggregation result as the final harmful information category label L i , L i =Vote(C), where L i ∈L; Recognition result output module: used to output the tag aggregation result obtained by the tag aggregation module as the recognition result of the text X to be detected.

7. The information recognition optimization system based on a large language model according to claim 6 is characterized in that: The semantic topic tag generation module also includes: Key vector generation module: used to convert harmful topics HarmfulTopic and topic keywords HarmfulKW into key vectors KV through semantic encoder Encode. The calculation formula of KV is: KV=Encode(HarmfulTopic,HarmfulKW) Among them, Encode is a language model Bert encoder based on Chinese; after being encoded by Encode encoder, the key vector KV = {HT, HKW}, where HT is the semantic representation of the harmful topic HarmfulTopic after semantic encoding, HT = {HT1, HT2, ..., HT m }; HKW is the semantic representation of the topic keyword HarmfulKW after semantic encoding, HKW = {HKW1, HKW2, ..., HKW k }; Knowledge data vector generation module: used to encode the topic tags W in the knowledge database through the semantic encoder Perform semantic encoding to obtain the corresponding knowledge data vector WV, where W = {W ij }; Similarity calculation module: used to calculate the similarity between the key vector KV and the knowledge data vector WV to obtain the similarity matrix CS. The calculation formula is as follows: CS=f(KV,WV) Where f is cosine similarity, KV = [KV1, KV2, …, KV T ], CS is expressed as: Where n is the total number of predefined target information tag categories; T is the number of harmful topics after semantic encoding; Mapping module: used to map the similarity matrix CS with the labels in the label database through label mapping T to obtain the semantic topic label C3, C3 = T (CS).

8. The information recognition optimization system based on a large language model according to claim 7 is characterized in that: The mapping module comprises: Similarity matrix conversion module: used to convert the above similarity matrix CS into a sub-label matrix The specific conversion rules are: if KV1W 11 is greater than ε, then it is recorded as the corresponding label W 11 , otherwise it is recorded as 0, where ε is the predetermined threshold; the transformed sub-label matrix as follows: Sub-label matrix conversion module: used to convert the sub-label matrix into Convert to a predefined label matrix as follows: Determine the module: From Determine the label L that appears most frequently i , as the final semantic label C3, where the calculation formula of C3 is as follows:

9. The information recognition optimization system based on a large language model according to claim 7, characterized in that: The tag database module includes a construction module for obtaining topic tags W through expert construction, user customization and automatic update. ij ,Using the data interface, combined with the knowledge database, a mapping relationship with the predefined target information label categories is established, thus obtaining the label database.

10. The information recognition optimization system based on a large language model according to any one of claims 7 to 9, characterized in that: The label database module also includes an automatic update module, which includes: New topic tag generation module: By identifying the target information category label, start the agent corresponding to the label database to generate a new topic tag W ij , where different target information category labels correspond to different Agents; Update submodule: Use the newly generated W ij And the corresponding target information category label, update the label database.