Generative large model-oriented text homology analysis method
Through high-dimensional chaotic semantic embedding algorithm and multi-iteration mapping mechanism, combined with multi-modal fusion and graph structure analysis, the complex semantic feature capture difficulty in generating text in generative large-scale models and low recognition accuracy of homologous text clusters is solved, and more efficient and robust text homology analysis is achieved.
Patent Information
- Application Number
- CN202510008588.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to accurately capture the complex semantic features of generated texts of generative large models, and the homologous text cluster recognition accuracy and robustness are low.
High-dimensional chaotic semantic embedding algorithm is used to convert text data into high-dimensional semantic embedding vectors, and mixed distance metrics are introduced for similarity analysis. The identification of homologous text clusters is optimized through dynamic clustering algorithms based on density peaks, multiple iterative mapping and dynamic gradient perturbation mechanisms. Multimodal fusion is carried out in combination with text context information and metadata, and deep analysis is carried out using graph structure and time series analysis methods.
It significantly improves the complexity and expression ability of text semantic embedding, enhances the robustness and accuracy of text representation, and can effectively identify and optimize complex semantic transformations in homologous text clusters.
Smart Images

Figure CN119940368A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text data analysis, and in particular to a text homology analysis method oriented to a generative large model. Background Art
[0002] In the field of natural language processing (NLP), generative large models, such as GPT and BERT, have become the core technology for solving various text generation and understanding tasks. These models can generate texts with rich semantics and contextual relevance through training with a large amount of corpus. However, with the widespread application of these models, text homology analysis has become increasingly important, that is, how to identify the semantic similarity between different texts, especially when these texts may be generated by the same generative model or there are complex semantic transformations. Although traditional text similarity analysis methods can measure the similarity between texts to a certain extent, they often seem to be powerless when faced with high-dimensional and complex texts generated by generative large models. In addition, existing text clustering technologies, such as K-means and hierarchical clustering, can cluster similar texts, but when faced with texts generated by generative large models, because these texts may contain complex semantic transformations and contextual dependencies, traditional clustering algorithms often find it difficult to achieve ideal results when identifying homologous text clusters, especially for texts with highly nonlinear and multimodal information, the performance of traditional clustering methods is usually limited.
[0003] The Chinese invention patent with the authorization announcement number CN114595306B discloses a text similarity calculation system and method based on distance-aware self-attention mechanism and multi-angle modeling, including the following steps: Step S1: perform word segmentation and stop word removal preprocessing on the original texts P and Q respectively, and convert the preprocessed texts into<P,Q> Use Word2vec model pre-training to get text pairs<P,Q> The word vectors are randomly initialized and fed into GRU learning.<P,Q> Character embedding representation and exact match mark between text P and Q; Step S2: embed the text pair after word embedding representation<P,Q> Use double-layer BiLSTM to encode and get text pair<P,Q> Step S3: Use the proposed distance-aware self-attention mechanism to encode text P and Q respectively, capture the deep features of text P and Q, and then obtain the text pair<P,Q> Enhanced semantic vector representation; Step S4: Use the interactive attention mechanism Co-attention to model texts P and Q to capture the interactive information between texts P and Q; Step S5: Use a multi-angle similarity modeling algorithm, and use element-by-element similarity, bilinear distance, and cosine similarity to calculate the similarity between enhanced feature vectors from multiple different angles to obtain a multi-angle similarity aggregation vector; Step S6: Perform maximum pooling and average pooling on the multi-angle similarity aggregation vector, extract key features, and then send them to the fully connected layer and softmax to calculate the final similarity score, which is finally converted into a specific similarity score and output. The above-mentioned public scheme can effectively enhance the representation of text and improve the accuracy of similarity calculation. The present application proposes a technical scheme different from the above-mentioned public scheme, which aims to solve the technical problems that in the process of processing text data generated by a generative large model, the complex semantic features of text similarity cannot be accurately captured and the recognition accuracy and robustness of homologous text clusters are low. Summary of the invention
[0004] Technical purpose: In order to overcome the deficiencies in the prior art, the present invention provides a text homology analysis method for a generative large model.
[0005] Technical solution: To achieve the above purpose, the present invention discloses a text homology analysis method for a generative large model, comprising the following steps:
[0006] S1: After preprocessing the text data for the generative large model, the high-dimensional chaotic semantic embedding algorithm is used to convert the preprocessed text data into a high-dimensional semantic embedding vector. Based on the high-dimensional semantic embedding vector, a hybrid distance metric is introduced to perform similarity analysis to identify the preliminary semantic similarity between texts.
[0007] S2: A dynamic clustering algorithm based on density peak is used to perform dynamic clustering analysis based on preliminary semantic similarity to generate preliminary homologous text clusters. Multiple iterative mapping and dynamic gradient perturbation mechanisms are then introduced to further analyze the preliminary homologous text clusters, identify complex semantic transformations, and obtain optimized homologous text clusters.
[0008] S3: Combine the optimized homologous text clusters with the contextual information and metadata of the text to perform multimodal fusion to obtain the fused multimodal homologous text clusters. Use the graph structure to analyze the fused multimodal homologous text clusters, obtain the node representation and homology relationship diagram based on the graph structure, and apply the time series analysis method to dynamically track the generated text to obtain the homology analysis and source tracking results of the text.
[0009] Furthermore, the high-dimensional chaotic semantic embedding algorithm in step S1 includes the following steps:
[0010] S101: Decompose the preprocessed text data into independent word sets, and generate each word w i The initial word vector representation v i , and introduces a hybrid semantic kernel function;
[0011] S102: After constructing the hybrid semantic kernel function, a chaotic mapping mechanism is introduced to enhance the nonlinear characteristics of the hybrid semantic kernel function;
[0012] S103: Chaotic mapping mechanism will be introduced after the word w i The mixed semantic kernel function value generates a multi-dimensional semantic embedding vector, and the multi-dimensional semantic embedding vectors are combined to obtain a high-dimensional semantic embedding vector representation S of the entire text;
[0013] S104: The high-dimensional semantic embedding vector representation S is compressed into a low-dimensional space through compression mapping to reduce computational complexity, and then restored to the high-dimensional space through reconstruction mapping to restore its original semantic information. A hybrid distance metric is introduced based on the high-dimensional semantic embedding vector, and the hybrid distance metric value is converted into a similarity index. The preliminary semantic similarity is determined according to the set similarity threshold.
[0014] Furthermore, the calculation formula of the mixed semantic kernel function in step S101 is:
[0015]
[0016] Among them, K(w i ) is the word w i The semantic kernel function value, α j For word w i The dynamic weight of v is λ, which is a constant that controls the change of weight. i ,v j The word wi and word w j The initial word vector representation of , μ is a constant that controls the translation of the kernel function, σ is a constant that controls the scaling of the kernel function, δ controls the amplitude and phase offset of the nonlinear part of the semantic kernel function, θ ij For word w i and word w j The angle between the vectors,
[0017]
[0018] Among them, ∥v i ∥,∥v j ∥ are the initial word vectors v i ,v j The norm of .
[0019] Furthermore, the step S102 uses the Logistic map as the chaotic system and perturbs the mixed semantic kernel function value by using a perturbation calculation formula, and the perturbation calculation formula is:
[0020] K′(w i )=K(w i )·ρ t+1 ,
[0021] ρ t+1 = r·ρ t ·(1-ρ t ),
[0022] Among them, K′(w i ) is the value of the mixed semantic kernel function after the chaotic mapping mechanism is introduced, ρ t is the state variable of the chaotic system, indicating the state of the chaotic system at time t, ρ t+1 is the state variable of the chaotic system, indicating the state of the chaotic system at time t+1, and r is the control parameter of the chaotic system.
[0023] Furthermore, in step S103, the word w i The calculation formula for generating a multi-dimensional semantic embedding vector using the mixed semantic kernel function value is:
[0024]
[0025] Among them, E i For word w i The multi-dimensional semantic embedding vector of k is the weight of the kth dimension, k and j are the dimension indices in the multidimensional space, is the average value of the dimension index, τ is the parameter controlling the weight distribution, d is the number of dimensions of the multidimensional semantic space, ω kis the frequency parameter of the sine transform, γ k is the amplitude parameter of the sine transform, ∈ i For the sound item.
[0026] Furthermore, the compression mapping process formula in step S104 is:
[0027] C=σ(W c ·S+b c ),
[0028] Among them, C is the compressed low-dimensional representation, W c is the weight matrix of the compressed mapping, controlling the mapping relationship from high dimension to low dimension, b c is the bias vector, σ(·) is the activation function;
[0029] The formula for reconstructing the mapping to restore the low-dimensional representation C to the high-dimensional space is:
[0030]
[0031] Among them, R is the high-dimensional representation after reconstruction mapping, w r To reconstruct the mapping weight matrix, control the mapping relationship from low dimension to high dimension, b r is the bias vector, is the nonlinear adjustment coefficient;
[0032] The formula for the hybrid distance metric is:
[0033]
[0034] Where D(T1,T2) is the hybrid distance measure between text T1 and text T2. and are the values of the reconstructed representation vectors of text T1 and text T2 in the high-dimensional space in the i-th dimension, p is a parameter for adjusting the contribution of each dimension in the Euclidean distance calculation, and q is a parameter for adjusting the balance relationship between the two parts in the overall distance measurement formula;
[0035] The formula for the similarity index Sim(T1,T2) is:
[0036]
[0037] Furthermore, the multiple iterative mapping mechanism in step S2 projects the semantic features of the text into a deeper semantic space by performing multiple nonlinear mappings in a high-dimensional space, and includes the following steps:
[0038] S201: The text T s The high-dimensional semantic embedding vector R s Perform nonlinear mapping to an intermediate semantic space. The nonlinear mapping formula is:
[0039]
[0040] in, For text T s The representation vector in the intermediate semantic space after the initial mapping, W1 is the initial mapping matrix, and the high-dimensional semantic embedding vector R s Mapped to the intermediate semantic space, b1 is the bias vector, tanh(·) is the hyperbolic tangent activation function, providing nonlinear mapping capability;
[0041] S202: In order to capture the deep semantic transformation in the text, the intermediate semantic space is mapped multiple times iteratively. The iterative mapping is defined as:
[0042]
[0043] in, For text T s In the The semantic representation vector after the iterative mapping, For the The weight matrix of the iterative mapping maps the semantic representation vector of the previous iteration to the new semantic space. For text T s In the The semantic representation vector after the iterative mapping, For the The bias vector of the iterative mapping, For the The nonlinear perturbation coefficient of the iterative mapping, For the The frequency parameter in the sine function in the iterative mapping;
[0044] Furthermore, in order to identify and optimize the complex semantic transformations in homologous text clusters, a dynamic gradient perturbation mechanism is introduced, and its formula is:
[0045]
[0046] in, For text T s After the Kth iterative mapping, the semantic representation vector after dynamic gradient perturbation is is the gradient of the semantic representation vector after K-th iteration mapping, is the dynamic perturbation coefficient, which adjusts the intensity of gradient perturbation, n is the total number of texts,
[0047] In order to dynamically adjust the global information of all texts, accumulation represents the cumulative sum of all texts.
[0048] Furthermore, in order to optimize the semantic representation, a feature decoupling and recombination mechanism is introduced to perform feature decoupling on the semantic representation vector after the dynamic gradient perturbation mechanism to generate a decoupled vector. The feature decoupling formula is:
[0049]
[0050] Among them, D s is the decoupled text T s The semantic representation vector of is the decoupling matrix, which is used to linearly transform the semantic representation vector, b d is the decoupling bias vector, is the decoupling strength parameter, which controls the smoothness of the decoupling process, m is the number of texts in the preliminary homologous text cluster, For text T j After the Kth iterative mapping, the semantic representation vector after dynamic gradient perturbation;
[0051] The decoupled semantic representation vectors are recombined through the recombination mechanism formula to generate the optimized semantic representation O s , the recombination mechanism formula is:
[0052]
[0053] Where T is the transpose.
[0054] Furthermore, step S3 includes the following steps:
[0055] S301: extracting context information and metadata of the text, and converting the context information and metadata into embedded vectors using a word vector model to form an embedded vector set of context information and a metadata embedded vector; the context information is the context and background information related to the text data generated by the generative large model, and the metadata is additional data related to the training and generation process of the generative large model;
[0056] S302: performing multimodal fusion on the optimized homologous text clusters, the embedding vector set of context information and the metadata embedding vectors through a fusion algorithm to obtain fused multimodal homologous text clusters;
[0057] S303: Further analyzing the fused multimodal homologous text clusters using a graph structure, including the following steps:
[0058] A1: Each text is regarded as a node in the graph. The edges between nodes represent the similarity relationship between texts. The weight of the edge is determined by calculating the Euclidean distance between multimodal representation vectors. The weight is calculated by the Gaussian kernel function to ensure that the edges between similar texts are closer.
[0059] A2: After obtaining the graph structure, perform graph convolution on each node and generate a new node representation by aggregating the information of its neighboring nodes;
[0060] A3: Based on the node representation of the graph structure, we further construct a homology graph and identify the deep connections between texts by analyzing the nodes and edges in the homology graph.
[0061] S304: After constructing the homology graph, a time series analysis method is introduced to dynamically track the generated text to identify the source of the text and its evolution path, generating a comprehensive text homology analysis result.
[0062] The beneficial effects of the present invention are:
[0063] 1. By introducing a high-dimensional chaotic semantic embedding algorithm, the complexity and expressiveness of text semantic embedding are significantly improved; through multiple nonlinear mappings and chaotic perturbation mechanisms, the generated high-dimensional semantic embedding vector can capture more subtle and complex semantic features in the text. This method can effectively cope with the diversity and complexity of semantic expression when processing texts generated by generative large models, and enhance the robustness and accuracy of text representation.
[0064] 2. The introduction of hybrid distance metric makes text similarity analysis more accurate. The hybrid distance metric combines the advantages of Euclidean distance and cosine similarity. By adjusting parameters, it flexibly controls the calculation method of similarity, so as to more accurately reflect the semantic similarity between texts. This method can not only identify the preliminary similarity between texts, but also effectively distinguish texts with complex semantic changes.
[0065] 3. In cluster analysis, a dynamic clustering algorithm based on density peak is combined with multiple iterative mapping and dynamic gradient perturbation mechanism to further improve the recognition accuracy of homologous text clusters; the dynamic clustering algorithm based on density peak automatically identifies the clustering center by analyzing the local density and distance in the similarity matrix, making text clustering more intelligent and dynamic; multiple iterative mapping and dynamic gradient perturbation mechanism further optimize the clustering results, and through multiple nonlinear mappings and dynamic adjustment of global information, it identifies and captures the complex semantic changes in the text, so that the finally generated homologous text clusters have higher semantic consistency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a flow chart of the overall method in the present invention;
[0067] Figure 2 This is a flow chart of the high-dimensional chaotic semantic embedding algorithm in the present invention;
[0068] Figure 3 It is a flow chart of the multiple iterative mapping mechanism in the present invention;
[0069] Figure 4 This is a flow chart of the method of step S3 in the present invention. DETAILED DESCRIPTION
[0070] The following is combined with Figure 1 —Attachment Figure 4 The principles and features of the present invention are described, and the examples given are only used to explain the present invention but not to limit the scope of the present invention.
[0071] A text homology analysis method for generative large models, such as Figure 1 As shown, the following steps are included:
[0072] S1: After preprocessing the text data for the generative large model, a high-dimensional chaotic semantic embedding algorithm is used to convert the preprocessed text data into a high-dimensional semantic embedding vector, and a hybrid distance metric is introduced on the basis of the high-dimensional semantic embedding vector to perform similarity analysis to identify the preliminary semantic similarities between texts; in this embodiment, the text data for the generative large model is preprocessed, and the preprocessing technology includes but is not limited to noise removal, text normalization, word segmentation, part-of-speech tagging, stop word removal, and word form restoration, so as to obtain the preprocessed text data. The preprocessing technology is well known to those skilled in the art and will not be described in detail.
[0073] like Figure 2 As shown, further, a high-dimensional chaotic semantic embedding algorithm is used to convert the preprocessed text data into a high-dimensional semantic embedding vector. The high-dimensional chaotic semantic embedding algorithm includes the following steps:
[0074] S101: Decompose the preprocessed text data into independent word sets, and generate each word w i The initial word vector representation v i , the initial word vector representation v i It can be obtained through existing word vector models such as Word2Vec or GloVe, and a mixed semantic kernel function is introduced.
[0075] Furthermore, the calculation formula of the mixed semantic kernel function in step S101 is:
[0076]
[0077] Among them, K(w i ) is the word wi The semantic kernel function value of word w i The semantic complexity of j For word w i The dynamic weight of word w in the context j Word w i The influence of semantics, λ is a constant that controls the change of weight and determines the influence of distance on weight; v i ,v j The word w i and word w j The initial word vector representation of , μ is a constant, which controls the translation of the kernel function, σ is a constant, which controls the scaling of the kernel function and affects the sensitivity of the similarity calculation. δ controls the amplitude and phase offset of the nonlinear part of the semantic kernel function, θ ij For word w i and word w j The angle between vectors reflects the similarity of their directions.
[0078]
[0079] Among them, ∥v i ∥,∥v j ∥ are the initial word vectors v i ,v j The norm of .
[0080] S102: After constructing the hybrid semantic kernel function, a chaotic mapping mechanism is introduced to enhance the nonlinear characteristics of the hybrid semantic kernel function. Logistic mapping is used as a chaotic system, and the hybrid semantic kernel function value is perturbed by a perturbation calculation formula. The perturbation calculation formula is:
[0081] K′(w i )=K(w i )·ρ t+1 ,
[0082] ρ t+1 =γ·ρ t ·(1-ρ t ),
[0083] Among them, K′(w i ) is the value of the mixed semantic kernel function after the chaotic mapping mechanism is introduced, which represents the semantic complexity after disturbance, ρ t is the state variable of the chaotic system, indicating the state of the chaotic system at time t, ρ t+1 is the state variable of the chaotic system, indicating the state of the chaotic system at time t+1, and r is the control parameter of the chaotic system, which adjusts the chaos of the system.
[0084] The introduction of the chaotic mapping mechanism increases the dynamic complexity, making the mixed semantic kernel function generated in the same context have higher variability, thereby improving the robustness of the algorithm when processing different texts.
[0085] S103: After completing the chaotic perturbation of the hybrid semantic kernel function, in order to generate a more expressive multidimensional semantic embedding vector and more comprehensively represent the semantic features of the word, the chaotic mapping mechanism is introduced. i The mixed semantic kernel function value is mapped to the multi-dimensional semantic space to generate a multi-dimensional semantic embedding vector.
[0086] Word i The mixed semantic kernel function value generates a multi-dimensional semantic embedding vector through nonlinear mapping. The calculation formula is as follows:
[0087]
[0088] Among them, E i For word w i The multidimensional semantic embedding vector of represents the semantic features of words in multidimensional space, β k is the weight of the kth dimension, which adjusts the semantic features of different dimensions. k and j are the dimension indexes in the multidimensional space. is the average value of the dimension index, indicating the central dimension, τ is the parameter that controls the weight distribution, affecting the balance between different dimensions, d is the number of dimensions of the multidimensional semantic space, indicating the number of dimensions of the entire semantic space, ω k is the frequency parameter of the sine transform, which controls the vibration frequency of the semantic kernel function in the kth dimension, γ k is the amplitude parameter of the sine transform, which controls the amplitude of the semantic kernel function in the kth dimension, ∈ i It is a sound term that represents the randomness and uncertainty in the semantic embedding of words to prevent overfitting. The introduction of random noise is used to enhance the robustness of the model and make the embedding vector more universal.
[0089] Furthermore, the multi-dimensional semantic embedding vectors are combined to obtain a high-dimensional semantic embedding vector representation S of the entire text.
[0090] S104: The high-dimensional semantic embedding vector representation S is compressed into a low-dimensional space through compression mapping, thereby reducing the computational complexity, and then restored to the high-dimensional space through reconstruction mapping to restore its original semantic information. A hybrid distance metric is introduced based on the high-dimensional semantic embedding vector, and the hybrid distance metric value is converted into a similarity index. The preliminary semantic similarity is determined according to the set similarity threshold.
[0091] The compression mapping process formula is:
[0092] C=σ(W c ·S+bc ),
[0093] Among them, C is the compressed low-dimensional representation, which represents the compressed text semantic information, and W c is the weight matrix of the compressed mapping, controlling the mapping relationship from high dimension to low dimension, b c is the bias vector, which adjusts the center point of the mapping result. σ(·) is the activation function, and you can choose tanh(·) or ReLU.
[0094] The compression mapping compresses the high-dimensional embedding vector representation S into a low-dimensional space through matrix operations, which reduces the computational burden while maintaining the semantic information.
[0095] Subsequently, the compressed low-dimensional representation C is restored to the high-dimensional space through reconstruction mapping, and the formula is as follows:
[0096]
[0097] Among them, R is the high-dimensional representation after reconstruction mapping, that is, the high-dimensional semantic embedding vector, which contains the restored text semantic information, and W r To reconstruct the mapping weight matrix, control the mapping relationship from low dimension to high dimension, b r is the bias vector, which adjusts the center point of the reconstruction result. is the nonlinear adjustment coefficient, which is used to control the degree of nonlinearity after reconstruction.
[0098] Based on the high-dimensional semantic embedding vector, a hybrid distance metric is introduced to perform similarity analysis and identify the preliminary semantic similarity between texts. The hybrid distance metric formula is:
[0099]
[0100] Among them, D(T1, T2) is the mixed distance measure between text T1 and text T2, which indicates the degree of similarity between the two in the high-dimensional semantic space. The smaller the distance, the closer the semantics of the two are, and the larger the distance, the greater the semantic difference between the two. and are the values of the reconstructed representation vectors of text T1 and text T2 in the high-dimensional space on the i-th dimension, obtained through the reconstruction mapping process, which reflects the specific representation of the semantic features of the text in the high-dimensional space. p is a parameter for adjusting the contribution of each dimension in the Euclidean distance calculation. Different values will affect the sensitivity of the overall distance. q is a parameter for adjusting the balance relationship between the two parts in the overall distance measurement formula. Different values can adjust the nonlinearity of the final distance.
[0101] Specifically, the hybrid distance metric is converted into a similarity index Sim(T1, T2), and a similarity threshold is set according to the expert experience method. The similarity threshold can be adjusted according to the specific application scenario. When the similarity index of the two texts is greater than or equal to the similarity threshold, the two texts are judged to have preliminary semantic similarity. Otherwise, it is considered that the semantic difference between the two texts is large, and the preliminary semantic similarity analysis is achieved through the above steps.
[0102] The formula for the similarity index Sim(T1,T2) is:
[0103]
[0104] S2: A dynamic clustering algorithm based on density peak is used to perform dynamic clustering analysis according to preliminary semantic similarity to generate preliminary homologous text clusters. Multiple iterative mapping and dynamic gradient perturbation mechanisms are then introduced to further analyze the preliminary homologous text clusters, identify complex semantic transformations, and obtain optimized homologous text clusters.
[0105] In the preliminary semantic similarity analysis, the similarity between all texts is calculated through a mixed distance metric, and the similarity values are stored in a similarity matrix. Then, the existing dynamic clustering algorithm based on density peak is used for dynamic clustering analysis to generate preliminary homologous text clusters. The main idea of the dynamic clustering algorithm based on density peak is to automatically identify the clustering center by analyzing the local density and distance in the similarity matrix, thereby dynamically clustering the text and generating preliminary homologous text clusters.
[0106] like Figure 3 As shown in the figure, further, based on the preliminary homologous text clusters, a more complex semantic mapping space is constructed. In order to capture the complex semantic transformation in the text, a multiple iterative mapping mechanism is introduced. The multiple iterative mapping mechanism projects the semantic features of the text into a deeper semantic space through multiple nonlinear mappings in a high-dimensional space. It includes the following steps:
[0107] S201: The text T s The high-dimensional semantic embedding vector R s Perform nonlinear mapping to an intermediate semantic space. The nonlinear mapping formula is:
[0108]
[0109] in, For text T s The representation vector in the intermediate semantic space after the initial mapping, W1 is the initial mapping matrix, and the high-dimensional semantic embedding vector R sMapped to the intermediate semantic space, b1 is the bias vector used to adjust the mapped result to avoid the situation of all-zero input, and tanh(·) is the hyperbolic tangent activation function, which provides nonlinear mapping capability.
[0110] S202: In order to capture the deep semantic transformation in the text, the intermediate semantic space is mapped multiple times iteratively. The iterative mapping is defined as:
[0111]
[0112] in, For text T s In the The semantic representation vector after the iterative mapping, For the The weight matrix of the iterative mapping maps the semantic representation vector of the previous iteration to the new semantic space. For text T s In the The semantic representation vector after the iteration mapping is used as the input of the current iteration. For the The bias vector of the iterative mapping is used to adjust the mapping result. For the The nonlinear perturbation coefficient of the iterative mapping controls the influence of the sine function. For the The frequency parameter in the sine function in the iterative mapping adjusts the frequency of the nonlinear perturbation.
[0113] Furthermore, in order to identify and optimize the complex semantic transformations in homologous text clusters, a dynamic gradient perturbation mechanism is introduced, and its formula is:
[0114]
[0115] in, For text T s After the Kth iterative mapping, the semantic representation vector after dynamic gradient perturbation is is the gradient of the semantic representation vector after the K-th iterative mapping, which is used to capture the direction of change of text features. is the dynamic perturbation coefficient, which adjusts the intensity of gradient perturbation, n is the total number of texts,
[0116] In order to dynamically adjust the global information of all texts, accumulation represents the cumulative sum of all texts.
[0117] Furthermore, in order to optimize the semantic representation, a feature decoupling and recombination mechanism is introduced to perform feature decoupling on the semantic representation vector after the dynamic gradient perturbation mechanism to generate a decoupled vector. The feature decoupling formula is:
[0118]
[0119] Among them, D s is the decoupled text T s The semantic representation vector of is the decoupling matrix, which is used to linearly transform the semantic representation vector, b d is the decoupling bias vector, used to adjust the decoupling result. is the decoupling strength parameter, which controls the smoothness of the decoupling process, m is the number of texts in the preliminary homologous text cluster, For text T j After the K-th iterative mapping, the semantic representation vector is dynamically perturbed by the gradient.
[0120] Then, the decoupled semantic representation vectors are recombined through the recombination mechanism formula to generate the optimized semantic representation O s , the recombination mechanism formula is:
[0121]
[0122] Where T is the transpose.
[0123] Based on the optimized semantic representation, the hybrid distance metric is calculated and converted into a similarity index. Then, based on the similarity index, the existing dynamic clustering algorithm based on density peak is used to perform dynamic clustering analysis to generate optimized homologous text clusters.
[0124] S3: Combine the optimized homologous text clusters with the contextual information and metadata of the text to perform multimodal fusion to obtain the fused multimodal homologous text clusters. Use the graph structure to analyze the fused multimodal homologous text clusters, obtain the node representation and homology relationship diagram based on the graph structure, and apply the time series analysis method to dynamically track the generated text to obtain the homology analysis and source tracking results of the text.
[0125] like Figure 4 As shown, further, step S3 includes the following steps:
[0126] S301: Extract the context information and metadata of the text, and use the word vector model to convert the context information and metadata into an embedding vector to form an embedding vector set of context information and a metadata embedding vector; the context information is the context and background information related to the text data generated by the generative large model, including the relevant content before and after the text is generated, the conversation history, the paragraph structure, the theme, etc. The metadata is additional data related to the training and generation process of the generative large model, including model parameters, the source of the training corpus, the model configuration (such as the number of layers, the number of hidden units), the hyperparameters in the training process (such as the learning rate, the optimizer), etc.; after extracting the context information and metadata, convert them into an embedding vector by using a word vector model such as Word2Vec or BERT to form an embedding vector set of context information and a metadata embedding vector.
[0127] S302: Perform multimodal fusion on the optimized homologous text clusters, the embedding vector set of context information and the metadata embedding vectors through a fusion algorithm such as weighted summation to obtain fused multimodal homologous text clusters.
[0128] S303: After the multimodal fusion is completed, the fused multimodal homologous text clusters are further analyzed using the graph structure, including the following steps:
[0129] A1: Each text is regarded as a node in the graph. The edges between nodes represent the similarity relationship between texts. The weight of the edge is determined by calculating the Euclidean distance between multimodal representation vectors. The weight is calculated by Gaussian kernel function or other nonlinear functions to ensure that the edges between similar texts are closer.
[0130] A2: After obtaining the graph structure, a graph convolution operation is performed on each node to generate a new node representation by aggregating the information of its neighboring nodes. In this process, the deep semantic associations between texts can be effectively captured by aggregating the similarity relationships between texts layer by layer.
[0131] A3: Based on the node representation of the graph structure, a homology graph is further constructed. The homology graph can reflect the similarity between texts and reveal the internal structure of homology text clusters. By analyzing the nodes and edges in the homology graph, the deep connection between texts can be identified. Especially when dealing with complex semantic transformations, graph structure analysis can provide more intuitive and accurate results.
[0132] S304: After constructing the homology relationship graph, the time series analysis method is introduced to dynamically track the generated text to identify the source of the text and its evolution path, and generate a comprehensive text homology analysis result. Specifically, the timestamp information of the text is used to model the changes of each text in the time dimension by establishing an autoregressive model. The autoregressive model predicts the evolution trend of the text representation based on the time series data, thereby inferring the possible source of the text and its change trajectory over time. The results of dynamic tracking can understand the generation and dissemination process of the text and identify which texts evolved from the same source, which is of great significance for the homology analysis of the text.
[0133] Through the above steps, the optimized homologous text clusters are multimodally fused with contextual information and metadata, and a comprehensive text homology analysis result is generated through graph structure analysis and time series modeling. The text homology analysis result can accurately identify the complex semantic relationship between texts, trace the source of texts, and provide a reliable basis for subsequent text analysis. This method significantly improves the complexity and expression ability of text semantic embedding by introducing a high-dimensional chaotic semantic embedding algorithm; through multiple nonlinear mappings and chaotic perturbation mechanisms, the generated high-dimensional semantic embedding vector can capture more subtle and complex semantic features in the text. This method can effectively cope with the diversity and complexity of its semantic expression when processing texts generated by generative large models, and enhance the robustness and accuracy of text representation.
[0134] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A text homology analysis method for a generative large model, characterized by: The following steps are involved: S1: After preprocessing the text data for the generative large model, the high-dimensional chaotic semantic embedding algorithm is used to convert the preprocessed text data into a high-dimensional semantic embedding vector. Based on the high-dimensional semantic embedding vector, a hybrid distance metric is introduced to perform similarity analysis to identify the preliminary semantic similarity between texts. S2: A dynamic clustering algorithm based on density peak is used to perform dynamic clustering analysis based on preliminary semantic similarity to generate preliminary homologous text clusters. Multiple iterative mapping and dynamic gradient perturbation mechanisms are then introduced to further analyze the preliminary homologous text clusters, identify complex semantic transformations, and obtain optimized homologous text clusters. S3: Combine the optimized homologous text clusters with the contextual information and metadata of the text to perform multimodal fusion to obtain the fused multimodal homologous text clusters. Use the graph structure to analyze the fused multimodal homologous text clusters, obtain the node representation and homology relationship diagram based on the graph structure, and apply the time series analysis method to dynamically track the generated text to obtain the homology analysis and source tracking results of the text.
2. The text homology analysis method for generative large models according to claim 1 is characterized in that: The high-dimensional chaotic semantic embedding algorithm in step S1 comprises the following steps: S101: Decompose the preprocessed text data into independent word sets, and generate each word w i The initial word vector representation v i , and introduces a hybrid semantic kernel function; S102: After constructing the hybrid semantic kernel function, a chaotic mapping mechanism is introduced to enhance the nonlinear characteristics of the hybrid semantic kernel function; S103: Chaotic mapping mechanism will be introduced after the word w i The mixed semantic kernel function value generates a multi-dimensional semantic embedding vector, and the multi-dimensional semantic embedding vectors are combined to obtain a high-dimensional semantic embedding vector representation s of the entire text; S104: compress the high-dimensional semantic embedding vector representation s into a low-dimensional space through compression mapping to reduce computational complexity, and then restore it to the high-dimensional space through reconstruction mapping to restore its original semantic information. A hybrid distance metric is introduced based on the high-dimensional semantic embedding vector, and the hybrid distance metric value is converted into a similarity index. According to the set similarity threshold, the preliminary semantic similarity is determined.
3. The text homology analysis method for generative large models according to claim 2 is characterized in that: The calculation formula of the mixed semantic kernel function in step S101 is: Among them, K(w i ) is the word w i The semantic kernel function value, α j For word w i The dynamic weight of v is λ, which is a constant that controls the change of weight. i , v j The word w i and word w j The initial word vector representation of , μ is a constant that controls the translation of the kernel function, σ is a constant that controls the scaling of the kernel function, δ controls the amplitude and phase offset of the nonlinear part of the semantic kernel function, θ ij For word w i and word w j The angle between the vectors, Among them, ||v i ||,||v j || are the initial word vectors v i , v j The norm of .
4. The text homology analysis method for generative large models according to claim 3 is characterized in that: The step S102 uses the Logistic map as a chaotic system and perturbs the mixed semantic kernel function value by using a perturbation calculation formula, wherein the perturbation calculation formula is: K′(in i )=K(in i )·ρ t+1 , r t+1 =r·ρ t ·(1-p t ), Among them, K′(w i ) is the value of the mixed semantic kernel function after the chaotic mapping mechanism is introduced, ρ t is the state variable of the chaotic system, indicating the state of the chaotic system at time t, ρ t+1 is the state variable of the chaotic system, indicating the state of the chaotic system at time t+1, and r is the control parameter of the chaotic system.
5. The text homology analysis method for generative large models according to claim 4 is characterized in that: In step S103, the word w i The calculation formula for generating a multi-dimensional semantic embedding vector using the mixed semantic kernel function value is: Among them, E i For word w i The multi-dimensional semantic embedding vector of k is the weight of the kth dimension, k and j are the dimension indices in the multidimensional space, is the average value of the dimension index, τ is the parameter controlling the weight distribution, d is the number of dimensions of the multidimensional semantic space, ω k is the frequency parameter of the sine transform, γ k is the amplitude parameter of the sine transform, ∈ i For the sound item.
6. The text homology analysis method for generative large models according to claim 5 is characterized in that: The compression mapping process formula in step S104 is: C=σ(W c ·S+b c ), Among them, C is the compressed low-dimensional representation, W c is the weight matrix of the compressed mapping, controlling the mapping relationship from high dimension to low dimension, b c is the bias vector, σ(·) is the activation function; The formula for reconstructing the mapping to restore the low-dimensional representation C to the high-dimensional space is: Among them, R is the high-dimensional representation after reconstruction mapping, W r To reconstruct the mapping weight matrix, control the mapping relationship from low dimension to high dimension, b r is the bias vector, is the nonlinear adjustment coefficient; The formula for the hybrid distance metric is: Where D(T1, T2) is the hybrid distance measure between text T1 and text T2. and are the values of the reconstructed representation vectors of text T1 and text T2 in the high-dimensional space in the i-th dimension, p is the parameter for adjusting the contribution of each dimension in the Euclidean distance calculation, and q is the parameter for adjusting the balance relationship between the two parts in the overall distance measurement formula; The formula of similarity index Sim(T1, T2) is:
7. The text homology analysis method for generative large models according to claim 1 is characterized in that: The multiple iterative mapping mechanism in step S2 is to project the semantic features of the text into a deeper semantic space by performing multiple nonlinear mappings in a high-dimensional space, which includes the following steps: S201: The text T s The high-dimensional semantic embedding vector R s Perform nonlinear mapping to an intermediate semantic space. The nonlinear mapping formula is: in, For text T s The representation vector in the intermediate semantic space after the initial mapping, W1 is the initial mapping matrix, and the high-dimensional semantic embedding vector R s Mapped to the intermediate semantic space, b1 is the bias vector, tanh(·) is the hyperbolic tangent activation function, providing nonlinear mapping capability; S202: In order to capture the deep semantic transformation in the text, the intermediate semantic space is mapped multiple times iteratively. The iterative mapping is defined as: in, For text T s In the The semantic representation vector after the iterative mapping, For the The weight matrix of the iterative mapping maps the semantic representation vector of the previous iteration to the new semantic space. For text T s In the The semantic representation vector after the iterative mapping, For the The bias vector of the iterative mapping, For the The nonlinear perturbation coefficient of the iterative mapping, For the The frequency parameter in the sine function in the iteration map.
8. The text homology analysis method for generative large models according to claim 7 is characterized in that: In order to identify and optimize the complex semantic transformation in homologous text clusters, a dynamic gradient perturbation mechanism is introduced, and its formula is: in, For text T s After the Kth iterative mapping, the semantic representation vector after dynamic gradient perturbation is is the gradient of the semantic representation vector after K-th iteration mapping, is the dynamic perturbation coefficient, which adjusts the intensity of gradient perturbation, n is the total number of texts, In order to dynamically adjust the global information of all texts, accumulation represents the cumulative sum of all texts.
9. The text homology analysis method for generative large models according to claim 8 is characterized in that: In order to optimize the semantic representation, a feature decoupling and recombination mechanism is introduced to perform feature decoupling on the semantic representation vector after the dynamic gradient perturbation mechanism to generate a decoupled vector. The feature decoupling formula is: Among them, D s is the decoupled text T s The semantic representation vector of is the decoupling matrix, which is used to linearly transform the semantic representation vector, b d is the decoupling bias vector, is the decoupling strength parameter, which controls the smoothness of the decoupling process, m is the number of texts in the preliminary homologous text cluster, For text T j After the Kth iterative mapping, the semantic representation vector after dynamic gradient perturbation; The decoupled semantic representation vectors are recombined through the recombination mechanism formula to generate the optimized semantic representation O s , the recombination mechanism formula is: Where T is the transpose.
10. The text homology analysis method for generative large models according to claim 1 is characterized in that: The step S3 comprises the following steps: S301: extracting context information and metadata of the text, and converting the context information and metadata into embedded vectors using a word vector model to form an embedded vector set of context information and a metadata embedded vector; the context information is the context and background information related to the text data generated by the generative large model, and the metadata is additional data related to the training and generation process of the generative large model; S302: performing multimodal fusion on the optimized homologous text clusters, the embedding vector set of context information and the metadata embedding vectors through a fusion algorithm to obtain fused multimodal homologous text clusters; S303: Further analyzing the fused multimodal homologous text clusters using a graph structure, including the following steps: A1: Each text is regarded as a node in the graph. The edges between nodes represent the similarity relationship between texts. The weight of the edge is determined by calculating the Euclidean distance between multimodal representation vectors. The weight is calculated by the Gaussian kernel function to ensure that the edges between similar texts are closer. A2: After obtaining the graph structure, perform graph convolution on each node and generate a new node representation by aggregating the information of its neighboring nodes; A3: Based on the node representation of the graph structure, we further construct a homology graph and identify the deep connections between texts by analyzing the nodes and edges in the homology graph. S304: After constructing the homology graph, a time series analysis method is introduced to dynamically track the generated text to identify the source of the text and its evolution path, generating a comprehensive text homology analysis result.
Citation Information
Patent Citations
Text similarity calculation system and method based on distance-aware self-attention mechanism and multi-angle modeling
CN114595306B