Content extraction and analysis method, device and system for large models

By combining word segmentation vector clustering and sensitive thesaurus for text generated by the big model, the problem of ignoring structured features when direct word segmentation comparison of big model generation content is solved, and the accuracy and safety of detection results are improved.

CN119514540BActive Publication Date: 2025-06-06INSTITUTE OF NETWORK TECHNOLOGY (YANTAI)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510088849.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-06
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Content generated by large models is prone to ignore structured features when directly comparing word segmentation, resulting in mis-checking or missed detection, affecting user experience and security.

Method used

By preprocessing the text generated by the big model, the word segmentation vector is obtained, and clustering is performed according to its importance, and the word segmentation cluster center is obtained. Combined with sensitive thesaurus, determine whether the text is compliant.

Benefits of technology

Reduce missed or missed inspections, improve the accuracy of detection results, and ensure the compliance and security of generated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119514540B_ABST
    Figure CN119514540B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data processing technology, and specifically to a content extraction and analysis method, device and system for large models. The method includes: preprocessing the current text generated by the large model to obtain the word segmentation vector of the current text; clustering the word segmentation vector according to the importance of the word segmentation vector to obtain multiple word segmentation cluster centers; determining whether the current text is compliant based on the multiple word segmentation cluster centers and in combination with a sensitive word library. The embodiment of the present application clusters the word segmentation vectors according to the importance of the word segmentation vectors, and then combines the clustered word segmentation cluster centers with the sensitive word library to determine whether the current text is compliant, thereby reducing the occurrence of false detection or missed detection and improving the accuracy of the detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of data processing technology, and specifically relates to a content extraction and analysis method, device and system for large models. Background Art

[0002] Large models refer to machine learning models with large-scale parameters and complex computational structures. They are usually built with deep neural networks and have billions or even hundreds of billions of parameters. Large models are widely used in various fields, including natural language processing, computer vision, speech recognition, and recommendation systems. They learn complex patterns and features by training massive data, have strong generalization capabilities, and can make accurate predictions on unseen data.

[0003] As the training data and parameters of text models continue to expand, especially when the training data is not fully analyzed, the model is more likely to show some unpredictable and more complex capabilities, which may generate inappropriate information content, affect the user experience, and cause unnecessary risks. Therefore, extracting and analyzing the content generated by large models is an important way to ensure that the content generated by the model is accurate and does not lead to bad guidance. However, by directly comparing the word segmentation of the model-generated content, it is easy to ignore the structural features of the content generated by the large model, which may cause false detection or missed detection. Summary of the invention

[0004] In order to solve the above problems, the embodiments of the present application provide a content extraction and analysis method, device and system for large models.

[0005] According to a first aspect of an embodiment of the present application, a content extraction and analysis method for a large model is provided, the method comprising:

[0006] Preprocessing the current text generated by the large model to obtain a word segmentation vector of the current text;

[0007] Clustering the word segmentation vectors according to the importance of the word segmentation vectors to obtain multiple word segmentation cluster centers;

[0008] Whether the current text complies with the regulations is determined based on the multiple word segmentation cluster centers and in combination with a sensitive word library.

[0009] Optionally, obtaining a plurality of word segmentation cluster centers includes:

[0010] According to the position relationship of the sentences and segments of the current text, setting a segment label for the word segmentation vector;

[0011] According to the paragraph label, obtaining the description density of the word segmentation vector;

[0012] According to the description density, obtaining a structural distribution parameter of the word segmentation vector;

[0013] Obtaining a text positioning coefficient of the word segmentation vector according to the description density and the structural distribution parameter;

[0014] According to the text positioning coefficient, obtain the similarity of any two word segmentation vectors;

[0015] The current text is clustered according to the similarity to obtain the multiple word segmentation cluster centers.

[0016] Optionally, setting a paragraph label for the word segment vector according to the position relationship of the sentence segments of the current text includes:

[0017] Get the first line indent of each paragraph of the current text;

[0018] According to the indent amount, obtaining the description class of each paragraph;

[0019] According to the description class, obtaining the text structure of the current text;

[0020] According to the text structure, a paragraph label of the word segmentation vector is obtained.

[0021] Optionally, obtaining the description density of the word segmentation vector includes:

[0022] Obtaining the summary importance of the word segmentation vector;

[0023] Obtaining information entropy of the word segmentation vector;

[0024] According to the summary importance and the information entropy, the description density of the word segmentation vector is obtained.

[0025] Optionally, obtaining the structural distribution parameter of the word segmentation vector includes:

[0026] Obtaining the adjusted Euclidean norm of the word segmentation vector;

[0027] Obtaining the text distance of the word segmentation vector;

[0028] According to the adjusted Euclidean norm and the text distance, a structural distribution parameter of the word segmentation vector is obtained.

[0029] Optionally, obtaining the text positioning coefficient of the word segmentation vector includes:

[0030] Obtaining a first ranking position of the word segmentation vector, where the first ranking position is the ranking position of the word segmentation vector in a structural distribution parameter sequence, where the structural distribution parameter sequence is obtained by arranging the structural distribution parameters of all word segmentation vectors in ascending order;

[0031] Obtaining a second ranking position of the word segmentation vector, where the second ranking position is the ranking position of the word segmentation vector in a description density sequence, where the description density sequence is obtained by arranging the description densities of all word segmentation vectors in ascending order;

[0032] A text positioning coefficient of the word segmentation vector is obtained according to the first ranking position, the second ranking position, the description density and the structural distribution parameter.

[0033] Optionally, obtaining the similarity between any two word segmentation vectors includes:

[0034] Get the Euclidean norm of any two word segmentation vectors;

[0035] According to the text positioning coefficient and the Euclidean norm of any two word segmentation vectors, the similarity between any two word segmentation vectors is obtained.

[0036] Optionally, determining whether the current text is compliant includes:

[0037] Converting the sensitive words in the sensitive word library into sensitive word vectors;

[0038] Obtaining the modified cosine similarity between the sensitive word vector and the centers of the multiple word clusters;

[0039] When the modified cosine similarity is greater than a preset threshold, the current text is determined to be non-compliant and not displayed, and the large model is instructed to regenerate the current text.

[0040] According to a second aspect of an embodiment of the present application, a content extraction and analysis device for a large model is provided, the device comprising:

[0041] A first acquisition module is used to preprocess the current text generated by the large model to obtain a word segmentation vector of the current text;

[0042] A second acquisition module is used to cluster the word segmentation vectors according to the importance of the word segmentation vectors to obtain multiple word segmentation cluster centers;

[0043] A determination module is used to determine whether the current text is compliant based on the multiple word segmentation cluster centers and in combination with a sensitive word library.

[0044] According to a third aspect of an embodiment of the present application, a content extraction and analysis system for a large model is provided, the system comprising a server, the server comprising:

[0045] a memory having a computer program stored thereon;

[0046] A processor is used to execute the computer program in the memory to implement the steps of any method described in the first aspect.

[0047] In summary, the embodiments of the present application provide a content extraction and analysis method, device and system for a large model, the method comprising: preprocessing the current text generated by the large model to obtain the word segmentation vector of the current text; clustering the word segmentation vector according to the importance of the word segmentation vector to obtain multiple word segmentation cluster centers; determining whether the current text is compliant based on the multiple word segmentation cluster centers and in combination with a sensitive word library. The embodiments of the present application cluster the word segmentation vectors according to the importance of the word segmentation vectors, and then combines the clustered word segmentation cluster centers with the sensitive word library to determine whether the current text is compliant, thereby reducing the occurrence of false detection or missed detection and improving the accuracy of the detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the implementation scheme of the present application, the drawings required for use in the implementation scheme will be briefly introduced below. It should be understood that the drawings only show certain implementation schemes of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on the drawings without paying creative work.

[0049] Figure 1 is a flow chart of a content extraction and analysis method for a large model according to an exemplary embodiment;

[0050] Figure 2 is a flow chart of a method for obtaining multiple word segmentation cluster centers according to an exemplary embodiment;

[0051] Figure 3 is a flowchart of a method for setting paragraph labels for a word segmentation vector according to an exemplary embodiment;

[0052] Figure 4 is a schematic diagram of a paragraph label according to an exemplary embodiment;

[0053] Figure 5 is a flowchart of a method for obtaining description density of a word segmentation vector according to an exemplary embodiment;

[0054] Figure 6 is a flowchart of a method for obtaining structural distribution parameters of a word segmentation vector according to an exemplary embodiment;

[0055] Figure 7 is a flow chart showing a method for obtaining a text positioning coefficient of a word segmentation vector according to an exemplary embodiment;

[0056] Figure 8 is a flowchart of a method for obtaining the similarity of any two word segmentation vectors according to an exemplary embodiment;

[0057] Fig. 9 is a flow chart showing a method for determining whether a current text is compliant according to an exemplary embodiment;

[0058] Fig.10 is a block diagram of a content extraction and analysis device for a large model according to an exemplary embodiment;

[0059] Fig.11 is a block diagram of a content extraction and analysis system for a large model according to an exemplary embodiment;

[0060] Fig.12 It is a block diagram of a server according to an exemplary embodiment. DETAILED DESCRIPTION

[0061] In order to clearly illustrate the technical features of the present solution, the present application is described in detail below through specific implementation methods and in conjunction with the accompanying drawings.

[0062] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not intended to limit the scope of protection of the present application.

[0063] It should be understood that the various steps described in the method implementation of the present application can be performed in different orders and / or performed in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0064] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0065] It should be noted that the concepts such as "first" and "second" mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0066] It should be noted that the modifications of "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more". In the description of this application, unless otherwise specified, "multiple" means two or more than two, and other quantifiers are similar; "at least one item", "one or more items" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one item a can represent any number of a; for another example, one or more items of a, b and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple; "and / or" is a kind of association relationship that describes the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural.

[0067] Although the operations or steps are described in a specific order in the drawings in the embodiments of the present application, it should not be understood that it is required to perform these operations or steps in the specific order shown or in a serial order, or to perform all the operations or steps shown to obtain the desired results. In the embodiments of the present application, these operations or steps can be performed in series; these operations or steps can also be performed in parallel; or some of these operations or steps can be performed.

[0068] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0069] First, the application scenario of this application is explained. The generative large language model receives the user's input or context and converts it into the model's input format, then encodes the input unit into a vector representation, and then uses self-attention mechanism, multi-layer encoding, position encoding and other methods to process this information, and recursively generates the final text based on context modeling through different strategies (such as greedy decoding, beam search, temperature control, etc.), and post-processes the generated text to ensure the quality, accuracy and suitability of the content. The scenario optimized by this method is mainly for the processing stage after the text is generated.

[0070] During the operation of the large model, after obtaining and analyzing the user input text, the large model outputs the corresponding response text. Since the large model is trained with ultra-large-scale data and can generate realistic text information, it may sometimes generate some bad information based on the user's input text content. In order to reduce the impact of the above situation, it is necessary to extract and analyze the generated content before sending it to determine its compliance. However, by directly performing word segmentation comparison on the model-generated content, it is easy to ignore the structural features of the content generated by the large model, which may cause false detection or missed detection. Therefore, a new content extraction and analysis method for large models is urgently needed. The present application is described below in conjunction with specific embodiments.

[0071] Figure 1 FIG. 1 is a flowchart of a content extraction and analysis method for a large model according to an exemplary embodiment. Figure 1 As shown, the embodiment of the present application provides a content extraction and analysis method for a large model, and the method may include the following steps:

[0072] In step S10, the current text generated by the large model is preprocessed to obtain a word segmentation vector of the current text.

[0073] In this step, the current text generated by the large model is preprocessed to obtain the word segmentation vector of the current text. For example, the text can be segmented using a word segmentation algorithm (such as the jieba module); the word segmentation result is converted into a vector form using a word vector conversion method (such as the word2vec model), thereby obtaining the word segmentation vector of the current text.

[0074] In step S20, the word segmentation vectors are clustered according to the importance of the word segmentation vectors to obtain a plurality of word segmentation cluster centers.

[0075] In this step, the word segmentation vectors are clustered according to their importance to obtain multiple word segmentation cluster centers. Exemplarily, the word segmentation vectors can be set with paragraph labels according to the position relationship of the current text, and then the description density of the word segmentation vectors can be obtained according to the paragraph labels, and then the structural distribution parameters of the word segmentation vectors can be obtained according to the description density, and then the text positioning coefficient of the word segmentation vectors can be obtained according to the description density and the structural distribution parameters, and then the similarity of any two word segmentation vectors can be obtained according to the text positioning coefficient, and finally the current text can be clustered according to the similarity to obtain multiple word segmentation cluster centers.

[0076] In step S30, whether the current text complies with the regulations is determined based on the multiple word segmentation cluster centers and in combination with a sensitive word library.

[0077] In this step, the current text is determined to be compliant based on multiple word segmentation cluster centers and in combination with the sensitive word library. For example, the sensitive words in the sensitive word library can be converted into sensitive word vectors, and then the corrected cosine similarity between the sensitive word vectors and the multiple word segmentation cluster centers is obtained. Then, when the corrected cosine similarity is greater than a preset threshold, the current text is determined to be non-compliant and will not be displayed, and the large model is instructed to regenerate the current text.

[0078] In summary, the embodiment of the present application provides a content extraction and analysis method for a large model, the method comprising: preprocessing the current text generated by the large model to obtain the word segmentation vector of the current text; clustering the word segmentation vector according to the importance of the word segmentation vector to obtain multiple word segmentation cluster centers; determining whether the current text is compliant based on the multiple word segmentation cluster centers and in combination with a sensitive word library. The embodiment of the present application determines whether the current text is compliant by clustering the word segmentation vector according to the importance of the word segmentation vector, and then combining the clustered word segmentation cluster centers with the sensitive word library, thereby reducing the occurrence of false detection or missed detection and improving the accuracy of the detection results.

[0079] Figure 2 FIG. 1 is a flow chart showing a method for obtaining multiple word segmentation cluster centers according to an exemplary embodiment. Figure 2 As shown, the acquisition of multiple word segmentation cluster centers may include the following steps:

[0080] In step S201, a paragraph label is set for the word segment vector according to the position relationship between the sentences and segments of the current text.

[0081] In this step, a segmentation label is set for the word segmentation vector according to the position relationship of the sentence segments of the current text. For example, the first line indentation of each segment of the current text can be obtained first, and then the description class of each segment can be obtained according to the indentation, and then the text structure of the current text can be obtained according to the description class, and then the segmentation label of the word segmentation vector can be obtained according to the text structure.

[0082] In step S202, the description density of the word segmentation vector is obtained according to the paragraph label.

[0083] In this step, the description density of the word segmentation vector is obtained according to the paragraph label. For example, the summary importance of the word segmentation vector can be obtained first, and then the information entropy of the word segmentation vector can be obtained, and then the description density of the word segmentation vector can be obtained according to the summary importance and the information entropy.

[0084] In step S203, the structural distribution parameters of the word segmentation vector are obtained according to the description density.

[0085] In this step, the structural distribution parameters of the word segmentation vector are obtained according to the description density. For example, the adjusted Euclidean norm of the word segmentation vector can be obtained first, and then the text distance of the word segmentation vector can be obtained, and then the structural distribution parameters of the word segmentation vector can be obtained according to the adjusted Euclidean norm and the text distance.

[0086] In step S204, the text localization coefficient of the word segmentation vector is obtained according to the description density and the structural distribution parameter.

[0087] In this step, the text positioning coefficient of the word segmentation vector is obtained according to the description density and the structural distribution parameter. Exemplarily, the first ranking position of the word segmentation vector can be obtained first, and the first ranking position is the ranking position of the word segmentation vector in the structural distribution parameter sequence, and the structural distribution parameter sequence is obtained by arranging the structural distribution parameters of all word segmentation vectors in order from small to large, and then the second ranking position of the word segmentation vector is obtained, and the second ranking position is the ranking position of the word segmentation vector in the description density sequence, and the description density sequence is obtained by arranging the description density of all word segmentation vectors in order from small to large, and then the text positioning coefficient of the word segmentation vector is obtained according to the first ranking position, the second ranking position, the description density and the structural distribution parameter.

[0088] In step S205, the similarity between any two word segmentation vectors is obtained according to the text positioning coefficient.

[0089] In this step, the similarity between any two word segmentation vectors is obtained according to the text positioning coefficient. For example, the Euclidean norm of any two word segmentation vectors can be obtained first, and then the similarity between any two word segmentation vectors can be obtained according to the text positioning coefficient and the Euclidean norm of any two word segmentation vectors.

[0090] In step S206, the current text is clustered according to the similarity to obtain the multiple word segmentation cluster centers.

[0091] In this step, the current text is clustered according to the similarity between any two word segmentation vectors to obtain multiple word segmentation cluster centers. Exemplarily, the method for clustering the current text is a prior art, and the K-means clustering algorithm can be used to cluster the word segmentation vectors in the current text to obtain multiple word segmentation cluster centers.

[0092] Figure 3 FIG. 1 is a flow chart showing a method for setting paragraph labels for a word segmentation vector according to an exemplary embodiment. Figure 3 As shown, setting a paragraph label for the word segment vector according to the sentence segment position relationship of the current text may include the following steps:

[0093] In step S2011, the first line indentation of each paragraph of the current text is obtained.

[0094] In this step, the first line indent of each paragraph of the current text is obtained. Exemplarily, the line break in the current text can be identified first, and the current text can be segmented according to the position of the line break, and then the first line indent of each paragraph can be calculated.

[0095] In step S2012, the description class of each segment is obtained according to the indentation amount.

[0096] In this step, the description class of each paragraph is obtained according to the indentation. For example, the text paragraphs can be divided into multiple description classes according to the indentation size, wherein paragraphs with the same indentation are in the same description class, and each description class is called the first description class, the second description class, etc. in the order of the indentation from small to large, and so on.

[0097] In step S2013, the text structure of the current text is obtained according to the description class.

[0098] In this step, the text structure of the current text is obtained according to the description class. For example, for a text with a "general-specific-general" structural feature, a paragraph with a large indentation is usually a detailed description of the level to which it belongs. Therefore, the structural information representing the current text can be extracted in a manner similar to a forest structure.

[0099] Take each paragraph in the description class with the smallest indent as the root node, and divide the current text into multiple blocks according to the position of the root node in the text description. For the text block corresponding to each root node, the second description class, the third description class (and so on) of this block are respectively used as the second-level child nodes and the third-level child nodes (and so on), etc., where the parent node of the nth-level child node should be the n-1th description class closest to the paragraph corresponding to the child node in the text; the child nodes of the same parent node are arranged in order from top to bottom according to the text. Arrange the nodes of each layer in order from top to bottom according to the text to obtain the text structure of the current text with a tree forest structure.

[0100] In step S2014, the paragraph label of the word segmentation vector is obtained according to the text structure.

[0101] In this step, the paragraph label of the word segmentation vector is obtained according to the text structure. Exemplarily, each child node of the same parent node in the tree forest structure can be numbered, and the numbers can be set in sequence as 1, 2, 3... and so on; for the root node, the number can be set as 1, 2, 3... and so on according to the tree sorting order; for any node, the traversal number is set to the root node of the tree to which it belongs and the digits of the node are traversed (the number of digits is the number of nodes traversed), and then the maximum digit number is selected as the digit of the paragraph label. For nodes whose traversal number digits are less than the number of digits of the label, 0 is added after the number until the number of digits is equal to the number of digits of the label, and then the number is used as the paragraph label of the current node, and the paragraph labels of the word segmentation in the paragraph corresponding to the node are all the paragraph labels of the node.

[0102] Figure 4 FIG. 1 is a schematic diagram showing a paragraph label according to an exemplary embodiment. Figure 4 As shown, for the convenience of demonstration, a binary tree is selected. The actual construction process may be a multi-branch tree. Taking the second deepest leftmost node Q as an example, its segment label is 1110.

[0103] Figure 5 FIG. 1 is a flowchart of a method for obtaining the description density of a word segmentation vector according to an exemplary embodiment. Figure 5 As shown, obtaining the description density of the word segmentation vector may include the following steps:

[0104] In step S2021, the summary importance of the word segmentation vector is obtained.

[0105] In this step, we obtain the summary importance of the word vector Generally speaking, the more frequently a word segmentation vector appears in a text, the more relevant the text description is to the word segmentation. However, for highly structured texts generated by large models, the importance of word segmentations to the text description varies depending on their position. In order to accurately obtain the general importance of a word segmentation to the text, it is necessary to analyze it in combination with structural information.

[0106] For example, the number of zeros in the paragraph tag of each word in the text can be calculated first, and the number can be normalized to obtain ; Then the participle The accumulated sum is divided by the number of occurrences of the current word segmentation vector in the current text , get the summary importance of the current word vector .

[0107] In step S2022, the information entropy of the word segmentation vector is obtained.

[0108] In this step, the information entropy of the current word segmentation vector is obtained The information entropy It is the information entropy of the current word segmentation vector and the word segmentation vector at the adjacent position in the current text.

[0109] In step S2023, the description density of the word segmentation vector is obtained according to the summary importance and the information entropy.

[0110] In this step, the summary importance of the current word segmentation vector is and information entropy , get the description density of the current word segmentation vector For example, the description density of the current word segmentation vector is It can be obtained by the following formula:

[0111] Formula 1

[0112] in, It represents the ratio of the number of occurrences of the current word segmentation vector in the current text to the total number of word segmentation vectors in the text. The addition of 0.01 is to avoid the special case where the information entropy is zero.

[0113] For structured text, the importance of text representations at different positions is different: in the paragraph forest, the smaller the number of layers of the paragraph node, the greater the probability that it is a general description. When a segmentation word appears in such a paragraph, it is more important than other segmentations. The number of layers of each participle is expressed by accumulating them and then calculating the ratio with the number to get The larger the value is, the stronger the generalization ability of the structural features of the word in the text is.

[0114] In addition, in order to reduce the description density of meaningless modifiers The influence of is calculated by calculating the information entropy of the current word vector and the word vector of the adjacent word segmentation. Analyze the meaningless probability of the current word segmentation vector: The adjacent context of a meaningless modifier is usually difficult to predict in a text representing specific information. Therefore, the uncertainty with respect to its adjacent words is expressed by calculating its information entropy. The larger the value, the higher the probability that the current word segmentation is a meaningless modifier.

[0115] pass right After correction, the description density of the current word segmentation vector in the current text is obtained .

[0116] Figure 6 FIG. 1 is a flow chart showing a method for obtaining structural distribution parameters of a word segmentation vector according to an exemplary embodiment. Figure 6As shown, the obtaining of the structural distribution parameters of the word segmentation vector may include the following steps:

[0117] In step S2031, the adjusted Euclidean norm of the word segmentation vector is obtained.

[0118] In this step, get the adjusted Euclidean norm of the current word segmentation vector i For example, the adjusted Euclidean norm of the current word segmentation vector i can be obtained by normalizing the sum of the number of all descendant nodes of the nodes corresponding to any two adjacent word segmentation paragraph labels of the current word segmentation vector. The larger the value, the higher the probability that the current word segmentation vector i is an important generalized topic word in the current text.

[0119] In step S2032, the text distance of the word segmentation vector is obtained.

[0120] In this step, the text distance of the current word segmentation vector i is obtained . The text distance of the current word segmentation vector i Indicates the difference in paragraph numbers between any two adjacent paragraph positions of the current word segmentation vector i. The larger the value, the more sparsely distributed the current word segmentation vector i is in the current text, and the higher the probability that the word segmentation vector is a topic word throughout the current text.

[0121] In step S2033, the structural distribution parameters of the word segmentation vector are obtained according to the adjusted Euclidean norm and the text distance.

[0122] In this step, the Euclidean norm is adjusted according to Distance from text , get the structural distribution parameters of word segmentation vector i . Exemplarily, the structural distribution parameter of the current word segmentation vector i is It can be obtained by the following formula:

[0123] Formula 2

[0124] in, Indicates the number of occurrences of the current word segmentation vector i in the current text.

[0125] Structural distribution parameters of word segmentation vector i The larger it is, the more important and widespread the current word vector i is in the current text, and the higher the probability that the word vector is an important topic word throughout the current text.

[0126] Figure 7 FIG. 1 is a flow chart showing a method for obtaining a text positioning coefficient of a word segmentation vector according to an exemplary embodiment. Figure 7As shown, the obtaining of the text positioning coefficient of the word segmentation vector may include the following steps:

[0127] In step S2041, a first ranking position of the word segmentation vector is obtained, where the first ranking position is the ranking position of the word segmentation vector in a structural distribution parameter sequence, and the structural distribution parameter sequence is obtained by arranging the structural distribution parameters of all word segmentation vectors in ascending order.

[0128] In this step, get the first ranking of the word segmentation vector c , first ranking is the ranking position of the word segmentation vector c in the structural distribution parameter sequence A. The structural distribution parameter sequence A is obtained by arranging the structural distribution parameters of all word segmentation vectors in order from small to large.

[0129] In step S2042, a second ranking position of the word segmentation vector is obtained, where the second ranking position is the ranking position of the word segmentation vector in a description density sequence, where the description density sequence is obtained by arranging the description densities of all word segmentation vectors in ascending order.

[0130] In this step, the second ranking of the word segmentation vector c is obtained , the second ranking is the ranking position of the word segmentation vector c in the description density sequence B. The description density sequence B is obtained by arranging the description densities of all word segmentation vectors in ascending order.

[0131] In step S2043, a text positioning coefficient of the word segmentation vector is obtained according to the first ranking rank, the second ranking rank, the description density and the structural distribution parameter.

[0132] In this step, according to the first ranking , second ranking , describing density and structural distribution parameters , get the text positioning coefficient of the word segmentation vector c For example, the text localization coefficient of the word segmentation vector c is It can be obtained by the following formula:

[0133] Formula 3

[0134] The larger the value, the higher the probability that the current word represents important topic content in the current text.

[0135] For some word segmentation vectors with strong interpretability, their description density may be low, and further judgment is needed in combination with the penetration of structural distribution parameters. Characterizes the relative size of the structural distribution parameter and description density of the word segmentation vector c in the current text: when the description density of the word segmentation vector c is low and the structural distribution parameter is large, The larger the value, the wider the distribution range of the word segmentation vector c is when the distribution amount (number of occurrences) is small, and the higher the probability that the current word segmentation can represent important topic content. The disadvantage in the text positioning calculation process is that After adjustment, the text positioning coefficient of the current word segmentation vector is obtained . Text positioning factor The larger it is, the higher the probability that the current word is an important topic.

[0136] Figure 8 FIG. 1 is a flowchart of a method for obtaining the similarity of any two word segmentation vectors according to an exemplary embodiment. Figure 8 As shown, the method of obtaining the similarity of any two word segmentation vectors may include the following steps:

[0137] In step S2051, the Euclidean norm of any two word segmentation vectors is obtained.

[0138] In this step, we get the Euclidean norm of any two word segmentation vectors a and b. . Euclidean norm Characterizes the similarity between any two word segmentation vectors a and b.

[0139] In step S2052, the similarity between any two word segmentation vectors is obtained based on the text positioning coefficient and the Euclidean norm of any two word segmentation vectors.

[0140] In this step, the text localization coefficients of any two word segmentation vectors a and b are and , and the Euclidean norm , get the similarity of any two word segmentation vectors a and b For example, the similarity between any two word segmentation vectors a and b is It can be obtained by the following formula:

[0141] Formula 4

[0142] The similarity between any two word segmentation vectors a and b It can also be understood as the distance between any two word segmentation vectors a and b.

[0143] Fig. 9 FIG. 1 is a flow chart showing a method for determining whether a current text is compliant according to an exemplary embodiment. Fig. 9As shown, the determination of whether the current text is compliant may include the following steps:

[0144] In step S301, the sensitive words in the sensitive word library are converted into sensitive word vectors.

[0145] In this step, the sensitive words in the sensitive word library are converted into sensitive word vectors. For example, the method of converting sensitive words into sensitive word vectors can refer to the embodiment in step S10, and this application will not repeat it here.

[0146] In step S302, the modified cosine similarity between the sensitive word vector and the centers of the multiple word clusters is obtained.

[0147] In this step, the modified cosine similarity between the sensitive word vector and the centers of multiple word clusters is obtained. Exemplarily, the calculation process of the modified cosine similarity is prior art and will not be described in detail here.

[0148] In step S303, when the modified cosine similarity is greater than a preset threshold, the current text is determined to be non-compliant and not displayed, and the large model is instructed to regenerate the current text.

[0149] In this step, when the corrected cosine similarity is greater than the preset threshold, the current text is determined to be non-compliant and not displayed, and the large model is instructed to regenerate the current text. Exemplarily, the preset threshold may be 0.65, and the implementer may set the value of the preset threshold according to the actual situation.

[0150] In summary, the embodiment of the present application provides a content extraction and analysis method for a large model, the method comprising: preprocessing the current text generated by the large model to obtain the word segmentation vector of the current text; clustering the word segmentation vector according to the importance of the word segmentation vector to obtain multiple word segmentation cluster centers; determining whether the current text is compliant based on the multiple word segmentation cluster centers and in combination with a sensitive word library. The embodiment of the present application clusters the word segmentation vectors according to the importance of the word segmentation vectors, and then combines the clustered word segmentation cluster centers with the sensitive word library to determine whether the current text is compliant, thereby reducing the occurrence of false detection or missed detection and improving the accuracy of the detection results.

[0151] Fig.10 FIG. 1 is a block diagram of a content extraction and analysis device for a large model according to an exemplary embodiment. Fig.10 As shown, the embodiment of the present application provides a content extraction and analysis device for a large model, and the device may include the following modules:

[0152] The first acquisition module 910 is used to preprocess the current text generated by the large model to obtain a word segmentation vector of the current text.

[0153] The second acquisition module 920 is used to cluster the word segmentation vectors according to the importance of the word segmentation vectors to obtain multiple word segmentation cluster centers.

[0154] The determination module 930 is used to determine whether the current text complies with the regulations based on the multiple word segmentation cluster centers and in combination with the sensitive word library.

[0155] Optionally, the second acquisition module 920 is further configured to:

[0156] According to the position relationship of the sentences and segments of the current text, setting a segment label for the word segmentation vector;

[0157] According to the paragraph label, obtaining the description density of the word segmentation vector;

[0158] According to the description density, obtaining a structural distribution parameter of the word segmentation vector;

[0159] Obtaining a text positioning coefficient of the word segmentation vector according to the description density and the structural distribution parameter;

[0160] According to the text positioning coefficient, obtain the similarity of any two word segmentation vectors;

[0161] The current text is clustered according to the similarity to obtain the multiple word segmentation cluster centers.

[0162] Optionally, the second acquisition module 920 is further configured to:

[0163] Get the first line indent of each paragraph of the current text;

[0164] According to the indent amount, obtaining the description class of each paragraph;

[0165] According to the description class, obtaining the text structure of the current text;

[0166] According to the text structure, a paragraph label of the word segmentation vector is obtained.

[0167] Optionally, the second acquisition module 920 is further configured to:

[0168] Obtaining the summary importance of the word segmentation vector;

[0169] Obtaining information entropy of the word segmentation vector;

[0170] According to the summary importance and the information entropy, the description density of the word segmentation vector is obtained.

[0171] Optionally, the second acquisition module 920 is further configured to:

[0172] Obtaining the adjusted Euclidean norm of the word segmentation vector;

[0173] Obtaining the text distance of the word segmentation vector;

[0174] According to the adjusted Euclidean norm and the text distance, a structural distribution parameter of the word segmentation vector is obtained.

[0175] Optionally, the second acquisition module 920 is further configured to:

[0176] Obtaining a first ranking position of the word segmentation vector, where the first ranking position is the ranking position of the word segmentation vector in a structural distribution parameter sequence, where the structural distribution parameter sequence is obtained by arranging the structural distribution parameters of all word segmentation vectors in ascending order;

[0177] Obtaining a second ranking position of the word segmentation vector, where the second ranking position is the ranking position of the word segmentation vector in a description density sequence, where the description density sequence is obtained by arranging the description densities of all word segmentation vectors in ascending order;

[0178] A text positioning coefficient of the word segmentation vector is obtained according to the first ranking position, the second ranking position, the description density and the structural distribution parameter.

[0179] Optionally, the second acquisition module 920 is further configured to:

[0180] Get the Euclidean norm of any two word segmentation vectors;

[0181] According to the text positioning coefficient and the Euclidean norm of any two word segmentation vectors, the similarity between any two word segmentation vectors is obtained.

[0182] Optionally, the determining module 930 is further configured to:

[0183] Converting the sensitive words in the sensitive word library into sensitive word vectors;

[0184] Obtaining the modified cosine similarity between the sensitive word vector and the centers of the multiple word clusters;

[0185] When the modified cosine similarity is greater than a preset threshold, the current text is determined to be non-compliant and not displayed, and the large model is instructed to regenerate the current text.

[0186] In summary, the embodiment of the present application provides a content extraction and analysis device for a large model, the device comprising: a first acquisition module, used to pre-process the current text generated by the large model, and obtain the word segmentation vector of the current text; a second acquisition module, used to cluster the word segmentation vector according to the importance of the word segmentation vector, and obtain multiple word segmentation cluster centers; a determination module, used to determine whether the current text is compliant based on the multiple word segmentation cluster centers and in combination with a sensitive word library. The embodiment of the present application clusters the word segmentation vectors according to the importance of the word segmentation vectors, and then combines the clustered word segmentation cluster centers with the sensitive word library to determine whether the current text is compliant, thereby reducing the occurrence of false detection or missed detection and improving the accuracy of the detection results.

[0187] The present application also provides a computer-readable storage medium on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the large-model oriented content extraction and analysis method provided by the present application are implemented.

[0188] Fig.11 FIG. 1 is a block diagram of a content extraction and analysis system for a large model according to an exemplary embodiment. Fig.11 As shown, an embodiment of the present application provides a content extraction and analysis system 1000 for a large model, including a server 1100.

[0189] Fig.12 is a block diagram of a server according to an exemplary embodiment. Fig.12 The server 1100 includes a processor 1122, which further includes one or more processors, and a memory resource represented by a memory 1132 for storing instructions executable by the processor 1122, such as an application. The application stored in the memory 1132 may include one or more modules, each corresponding to a set of instructions. In addition, the processor 1122 is configured to execute instructions to perform the above-mentioned large model-oriented content extraction and analysis method.

[0190] The server 1100 may also include a power component 1126 configured to perform power management of the server 1100, a communication component 1150 configured to connect the server 1100 to a network, and an input / output interface 1158. The server 1100 may operate based on an operating system stored in the memory 1132.

[0191] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable electronic device, and the computer program has a code portion for executing the above-mentioned large model-oriented content extraction and analysis method when executed by the programmable electronic device.

[0192] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application.

Claims

1. A content extraction and analysis method for a large model, characterized in that: The method comprises: Preprocessing the current text generated by the large model to obtain a word segmentation vector of the current text; Clustering the word segmentation vectors according to the importance of the word segmentation vectors to obtain multiple word segmentation cluster centers; Determining whether the current text complies with the regulations based on the multiple word segmentation cluster centers and in combination with a sensitive word library; The step of obtaining multiple word segmentation cluster centers includes: According to the position relationship of the sentences and segments of the current text, setting a segment label for the word segmentation vector; According to the paragraph label, obtaining the description density of the word segmentation vector; According to the description density, obtaining a structural distribution parameter of the word segmentation vector; Obtaining a text positioning coefficient of the word segmentation vector according to the description density and the structural distribution parameter; According to the text positioning coefficient, obtain the similarity of any two word segmentation vectors; The current text is clustered according to the similarity to obtain the multiple word segmentation cluster centers.

2. The large model-oriented content extraction and analysis method according to claim 1 is characterized in that: The step of setting a segment label for the word segment vector according to the position relationship of the segments of the current text includes: Get the first line indent of each paragraph of the current text; According to the indent amount, obtaining the description class of each paragraph; According to the description class, obtaining the text structure of the current text; According to the text structure, a paragraph label of the word segmentation vector is obtained.

3. The content extraction and analysis method for large models according to claim 1 is characterized in that: The obtaining the description density of the word segmentation vector includes: Obtaining the summary importance of the word segmentation vector; Obtaining information entropy of the word segmentation vector; According to the summary importance and the information entropy, the description density of the word segmentation vector is obtained.

4. The large model-oriented content extraction and analysis method according to claim 1 is characterized in that: The obtaining of the structural distribution parameters of the word segmentation vector includes: Obtaining the adjusted Euclidean norm of the word segmentation vector; Obtaining the text distance of the word segmentation vector; According to the adjusted Euclidean norm and the text distance, a structural distribution parameter of the word segmentation vector is obtained.

5. The large model-oriented content extraction and analysis method according to claim 1 is characterized in that: The obtaining of the text positioning coefficient of the word segmentation vector includes: Obtaining a first ranking position of the word segmentation vector, where the first ranking position is the ranking position of the word segmentation vector in a structural distribution parameter sequence, where the structural distribution parameter sequence is obtained by arranging the structural distribution parameters of all word segmentation vectors in ascending order; Obtaining a second ranking position of the word segmentation vector, where the second ranking position is the ranking position of the word segmentation vector in a description density sequence, where the description density sequence is obtained by arranging the description densities of all word segmentation vectors in ascending order; A text positioning coefficient of the word segmentation vector is obtained according to the first ranking position, the second ranking position, the description density and the structural distribution parameter.

6. The large model-oriented content extraction and analysis method according to claim 1 is characterized in that: The obtaining of the similarity between any two word segmentation vectors includes: Get the Euclidean norm of any two word segmentation vectors; According to the text positioning coefficient and the Euclidean norm of any two word segmentation vectors, the similarity between any two word segmentation vectors is obtained.

7. The large model-oriented content extraction and analysis method according to claim 1 is characterized in that: Determining whether the current text is compliant includes: Converting the sensitive words in the sensitive word library into sensitive word vectors; Obtaining the modified cosine similarity between the sensitive word vector and the centers of the multiple word clusters; When the modified cosine similarity is greater than a preset threshold, the current text is determined to be non-compliant and not displayed, and the large model is instructed to regenerate the current text.

8. A content extraction and analysis device for a large model, the device being used to implement the steps of the method according to any one of claims 1 to 7, characterized in that: The device comprises: A first acquisition module is used to preprocess the current text generated by the large model to obtain a word segmentation vector of the current text; A second acquisition module is used to cluster the word segmentation vectors according to the importance of the word segmentation vectors to obtain multiple word segmentation cluster centers; A determination module is used to determine whether the current text is compliant based on the multiple word segmentation cluster centers and in combination with a sensitive word library.

9. A content extraction and analysis system for large models, characterized in that: The system comprises a server, wherein the server comprises: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Product evaluation analysis method and device, computer equipment and storage medium

    CN109800307A

  • Data product security compliance inspection method and device, and server

    CN116150349A