Text semantic partitioning method and device, equipment, medium and product
By using dynamic similarity thresholds and weighted average embedding vectors in text chunking, the problem of poor text chunking in the prior art is solved, and more accurate and reasonable text chunking is achieved.
Patent Information
- Application Number
- CN202510517731.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
In the prior art, when text blocking, the use of fixed similarity thresholds leads to poor distinction, and may divide text blocks that are too long or too short, and cannot adapt to the dynamic division of document content of different topics.
By obtaining the weighted average embed vector of the sliding window of the pending document, compute the similarity array, and adjust it based on the target coefficient of variation, determine the dynamic similarity threshold, and dynamically divide the text blocks.
Improve the accuracy and rationality of text chunking, adapt to the structure and content of different documents, and avoid excessively long or too short segmentation.
Smart Images

Figure CN120031046A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a text semantic segmentation method, device, equipment, medium and product. Background Art
[0002] Text segmentation technology is an important part of natural language processing. Its purpose is to segment long text into shorter text blocks with certain semantic integrity to facilitate subsequent processing (such as retrieval, generation, classification, etc.).
[0003] In the related art, a fixed similarity threshold is used to divide the text of all documents. However, this method has poor discrimination when dividing documents of different types and topics, and may divide text blocks that are too long or too short, and cannot adapt to the dynamic division of document content of different topics.
[0004] Therefore, how to improve the accuracy and rationality of text segmentation is an urgent problem that needs to be solved. Summary of the invention
[0005] The present application provides a text semantic segmentation method, apparatus, device, medium and product to at least solve the problem of low accuracy and rationality of text segmentation in the related art.
[0006] The present application provides a text semantic segmentation method, the method comprising: Obtaining the weighted average embedding vector of each sliding window corresponding to the document to be processed; the weighted average embedding vector of any sliding window is obtained by weighted summing the embedding vectors of each clause of the document to be processed in the window; Obtaining a similarity array of the document to be processed; the similarity array includes: similarities between weighted average embedding vectors of adjacent sliding windows; Adjust the similarity array based on the target variation coefficient to obtain the target similarity array; Determining a dynamic similarity threshold according to the target coefficient of variation and segmentation information, wherein the segmentation information is used to indicate the range of the number of clauses contained in the segment; Dividing the target similarity array according to the dynamic similarity threshold to obtain a plurality of sub-similarity arrays; The clauses corresponding to the respective sub-similarity arrays are respectively determined as segments of the document to be processed.
[0007] The present application also provides a text semantic segmentation device, the device comprising: An acquisition module is used to acquire the weighted average embedding vectors of each sliding window corresponding to the document to be processed; the weighted average embedding vector of any sliding window is obtained by weighted summing the embedding vectors of each clause of the document to be processed in the window; A calculation module, used to obtain a similarity array of the document to be processed; the similarity array includes: similarities between weighted average embedding vectors of adjacent sliding windows; An adjustment module, used for adjusting the similarity array based on the target variation coefficient to obtain a target similarity array; An analysis module, configured to determine a dynamic similarity threshold according to the target coefficient of variation and segmentation information; the segmentation information is used to indicate the range of the number of clauses contained in the segment; A division module, used for dividing the target similarity array according to the dynamic similarity threshold to obtain a plurality of sub-similarity arrays; The determination module is used to determine the clauses corresponding to each sub-similarity array as the segments of the document to be processed.
[0008] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned text semantic segmentation methods when executing the computer program.
[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned text semantic segmentation methods are implemented.
[0010] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned text semantic segmentation methods when executed by a processor.
[0011] Through the present application, the weighted average embedding vector of each sliding window corresponding to the document to be processed is obtained, wherein the weighted average embedding vector of any sliding window is obtained by weighted summing the embedding vectors of each clause of the document to be processed in the window, and the similarity array of the document to be processed is obtained, and the similarity array includes: the similarity between the weighted average embedding vectors of each adjacent sliding window, adjusting the similarity array based on the target variation coefficient, obtaining the target similarity array, and determining the dynamic similarity threshold according to the target variation coefficient and segmentation information, wherein the segmentation information is used to indicate the number range of clauses contained in the segment, and dividing the target similarity array according to the dynamic similarity threshold to obtain multiple sub-similarity arrays, and determining the clauses corresponding to each sub-similarity array as the segment of the document to be processed. By weighted summing the embedding vectors of each clause in the sliding window, the overall semantic features of the clauses in the window can be reflected, the influence of noise can be reduced, and the similarity between the weighted average embedding vectors of adjacent sliding windows can be calculated, which can more accurately identify the semantically continuous parts of the document. The similarity array is adjusted based on the target variation coefficient, so that the segmentation effect of different documents remains relatively stable. The similarity threshold is dynamically adjusted according to the target variation coefficient and segmentation information, so that the segmentation process can adapt to the structure and content of different documents. The segmentation information indicates the range of the number of clauses contained in the segment, so that the segmentation results are more in line with actual needs and avoid overly long or short segments. The target similarity array is divided according to the dynamic similarity threshold, and the corresponding clauses are determined as the segments of the document to be processed, which improves the accuracy and rationality of text segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0013] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 A flowchart of a text semantic segmentation method provided in an embodiment of the present application; Figure 2 A structural schematic diagram of a text semantic segmentation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0015] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0016] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0017] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0018] Terminology explanation: Chunk: also known as chunking, refers to the process of breaking down large chunks of text into smaller fragments. This is particularly important in retrieval-enhanced generation (RAG) systems, which can help optimize the accuracy of content recalled from vector databases.
[0019] RAG: Retrieval-Augmented Generation, is a technology that combines information retrieval (IR) and natural language generation (NLG). It enhances the model's ability to generate text by retrieving relevant information from an external knowledge base and feeding it as context to a large language model (LLM).
[0020] Vector: Mathematically, it is usually represented by v, which means vector in Chinese. In this application, it means using a machine learning encoder model to convert the target text block into a vector form. Vector is related to the meaning of Embedding and is generally used as the output of Embedding.
[0021] Embedding: Embedding refers to the process of mapping original, usually discrete text or image data (such as phrases, sentences, documents, pictures, etc., which are the "chunks" mentioned above) into a continuous, low-dimensional vector space. These generated vectors are called Embedding vectors. Each Embedding vector is regarded as a semantic representation of its corresponding sound, shape, graphic and text element, which can capture the semantic relationship, grammatical structure and contextual dependency between words. For example, similar or related words will be closer in the Embedding space, reflecting their semantic commonality or relevance.
[0022] L2 norm: Also known as the Euclidean norm, the L2 norm of a vector is a commonly used measure of the length of a vector. It is particularly important in mathematics and engineering, especially in signal processing, linear algebra, and machine learning. The L2 norm is based on the concept of distance in Euclidean space and can be understood as the straight-line distance from the origin to the location pointed to by the vector.
[0023] L2 regularization: Also known as Ridge regularization or weight decay, it is a commonly used regularization technique that limits the size of model parameters by adding a penalty term proportional to the sum of squares of weight parameters to the loss function. The main function of L2 regularization is to reduce the complexity of the model by penalizing excessive weight values, thereby preventing overfitting.
[0024] Common text segmentation methods include: segmentation by fixed length, but this method is easy to cut off the text with the same semantics. Based on this, there is also a fixed-length segmentation method that retains a certain number of overlapping redundant words. Furthermore, there is a sentence-based segmentation method, such as segmentation based on punctuation marks, which is also easy to separate sentences with the same semantics. This method is essentially the same as the fixed-length segmentation method. In addition, there are some special segmentation methods for document formats, which perform specific processing for different formats; the most ideal method at present is the semantic segmentation method, which divides text blocks according to semantically similar programs. However, the existing technology uses a fixed similarity threshold to segment the text of all documents. When segmenting documents of different types and topics, the distinction is poor, and text blocks that are too long or too short may be segmented. When changing parameters such as sentence length, the Embedding model needs to be used again to calculate the embedding vectors of all document-related sentence clusters. Repeated calculations also waste a lot of time and computing power, reducing the availability of business services.
[0025] Based on the above problems, an embodiment of the present application provides a text semantic segmentation method, and describes the method in detail in conjunction with the execution flow of the text semantic segmentation method.
[0026] Reference Figure 1 As shown, the text semantic segmentation method provided by the embodiment of the present invention includes the following steps: S11. Obtain the weighted average embedding vector of each sliding window corresponding to the document to be processed.
[0027] The weighted average embedding vector of any sliding window is obtained by weighted summing of the embedding vectors of each clause of the document to be processed in the window.
[0028] Specifically, the weighted average embedding vector of each sliding window corresponding to the document to be processed is obtained by calculation.
[0029] In some embodiments, the above step S11 (obtaining the weighted average embedding vector of each sliding window corresponding to the document to be processed) can be implemented by the following steps: (1) Obtain a document to be processed, and split the document to be processed into multiple clauses.
[0030] Specifically, the document to be processed is obtained, all contents of the document to be processed are read, and the document to be processed is split into independent sentences according to punctuation marks, that is, the document to be processed is split into multiple clauses.
[0031] (2) Inputting the multiple clauses into the target embedding model for embedding processing to obtain embedding vectors of the multiple clauses.
[0032] Specifically, all independent clauses are embedded to obtain independent embedding vectors of all clauses of the document to be processed. Embedding is the process of mapping the independent clauses of the document to be processed into a continuous vector space.
[0033] In order to address the shortcoming that each time the window size changes, the weighted vector encoding method with variable clause length based on the independent clause embedding vector provided in the embodiment of the present disclosure only needs to calculate the independent clause embedding vector once, so as to realize dynamic modification of the window size without recalculating the independent clause embedding vector, and the computational coding efficiency is significantly improved compared with the method of recalculating the window sentence embedding vector in the related art.
[0034] (3) Obtain the target weight combination.
[0035] The target weight combination includes: weights corresponding to the embedding vectors of multiple clauses of the sliding window.
[0036] The size of the sliding window can be set according to the actual application scenario, and no specific restrictions are imposed here. For example, when the size of the sliding window is 3, each sliding window contains 3 consecutive clauses.
[0037] In some embodiments, the above step (3) can be implemented as follows: 1) Determine whether the target database includes the target embedding model and the weight combination corresponding to the target window size.
[0038] The target database includes: a plurality of embedding model identifiers and a plurality of weight combinations corresponding to a plurality of sliding window sizes; the target window size is the size of the sliding window. The embedding model identifier includes but is not limited to the name of the embedding model, the number of the embedding model, etc.
[0039] Specifically, since the target database includes: multiple embedding model identifiers and multiple weight combinations corresponding to multiple sliding window sizes, the query condition is the target embedding model identifier and the target sliding window size, so that a query is performed in the target database according to the target embedding model and the target window size to determine whether there is a weight combination corresponding to the target embedding model identifier and the target sliding window size in the target database.
[0040] 2) If the target database includes a combination weight corresponding to the target embedding model and the target window size, the combination weight is used as the target combination weight.
[0041] Specifically, if the target database includes the target embedding model and the combination weight corresponding to the target window size, the combination weight is used as the target combination weight. In this way, when calculating the embedding vector of the sentence in the sliding window, there is no need to reuse the embedding model for embedding processing. The existing target combination weight can be directly read from the target database, and then a simple mathematical operation can be performed with the independent embedding vector. This process is significantly faster than using the embedding model for calculation.
[0042] By calculating and saving the combined weight matrix, repeated calculations are avoided each time the window is adjusted, which greatly reduces the consumption of computing resources and improves the efficiency of text processing.
[0043] 3) If the target database does not include a target embedding model and a combination weight corresponding to the target window size, a target weight combination is obtained based on the target window size and the embedding vectors of the multiple clauses.
[0044] Specifically, initially, there is no combination weight information in the target database of the system. The combination weight information is calculated and obtained by the following calculation method.
[0045] In some embodiments, the above step 3) can be implemented by the following steps: a. Obtain a target embedding vector and an initial weight combination of the multiple clauses.
[0046] The target embedding vector is a combined embedding vector obtained by combining multiple clauses in the target window into a target clause and then performing embedding processing.
[0047] Optionally, the above step a) can be performed by the following steps: Combining multiple clauses within the target window into a target clause; Performing embedding processing on the target clause according to the target embedding model to obtain a target embedding vector of the target clause; Initialize the weights of multiple clauses in the target window to obtain an initial weight combination of the multiple clauses.
[0048] Among them, based on the weight constraint condition, the weights of multiple clauses in the target window are initialized to obtain the initial weight combination of the multiple clauses. The weight constraint condition includes equality constraint and boundary constraint. The equality constraint means that the sum of the weights of each clause in the target window is 1; the boundary constraint means that the weights of each clause in the target window are between 0 and 1, that is, all the weights satisfy .
[0049] Specifically, multiple clauses in the target window are combined into a target clause, and the target clause is embedded according to the target embedding model to obtain a target embedding vector of the target clause; the weights of multiple clauses in the target window are initialized to obtain an initial weight combination of the multiple clauses.
[0050] For example, it is assumed that the initial weights are set to be uniformly distributed, that is, , Indicates the size of the sliding window. In this embodiment, The value of can be 3. Therefore, the weights of multiple clauses in the target window are initialized, and the initial weight combination of multiple clauses is obtained as .
[0051] b. Optimizing the initial weight combination according to the target embedding vector and the embedding vectors of the multiple clauses to obtain a target weight combination.
[0052] In some embodiments, the above step b) can be performed by the following steps: According to the initial weight combination of the multiple clauses, weighted averaging processing is performed on the embedding vectors of the multiple clauses in the target window to obtain an initial weighted average embedding vector; According to the initial weighted average embedding vector and the target embedding vector, adjusting the initial weight combination to obtain a current weight combination of the multiple clauses; According to the current weight combination of the multiple clauses, weighted averaging processing is performed on the embedding vectors of the multiple clauses in the target window to obtain a current weighted average embedding vector; When the current weighted average embedding vector and the target embedding vector meet a preset convergence condition, a target weight combination is obtained.
[0053] Specifically, first, according to the initial weight combination of multiple clauses, the embedding vectors of multiple clauses in the target window are weighted averaged to obtain the initial weighted average embedding vector, and then according to the initial weighted average embedding vector and the target embedding vector, the initial weight combination is adjusted to obtain the current weight combination of multiple clauses in the target window, and then according to the current weight combination of multiple clauses in the target window, the embedding vectors of multiple clauses in the target window are weighted averaged to obtain the current weighted average embedding vector, until the current weighted average embedding vector and the target embedding vector meet the preset convergence condition, the target weight combination is obtained.
[0054] Assume that the size of the sliding window is 3, and the independent embedding vectors of all clauses in a set of windows are and a target embedding vector , where each set of embedding vectors v in E is a multidimensional vector, usually 512 or 768 dimensions, and a set of weights is found by the following method , so that the weighted average embedding vector As close as possible to the target embedding vector T. Among them, the weighted average embedding vector is: . Where W·E is a matrix multiplication, which means applying the weight vector W to each column of the embedding matrix E and summing the results.
[0055] Optionally, adjusting the initial weight combination according to the initial weighted average embedding vector and the target embedding vector to obtain the current weight combination of the multiple clauses includes:
[0056] in, represents the value of the objective function, represents the Euclidean distance between the weighted average embedding vector and the target embedding vector, is a constant, Indicates The current weight of the clause. represents the weighted average embedding vector, represents the target embedding vector. The objective function is used to measure the difference between the weighted average embedding vector and the target embedding vector.
[0057] Specifically, the first term (before the plus sign) in the above formula is the Euclidean distance between the weighted average embedding vector and the target embedding vector. The second term (after the plus sign) is the L2 regularization term, which is used to prevent overfitting. By limiting the size of the weight, the model is smoother, thereby improving the generalization ability. The parameter is a constant used to control the strength of regularization. In this embodiment The value can be 0.01.
[0058] By minimizing the objective function , find the optimal weight vector , that is, find the target weight combination. The preset optimization function is used to optimize the objective function. Among them, the preset optimization function includes but is not limited to the Trust Region Method or SLSQP (Sequential Least Squares Programming). Specifically, first select an initial point (initial solution), calculate the search direction and update the current point based on the information of the current point, and check whether the preset convergence conditions are met (such as the gradient is small enough, the objective function value changes very little, etc.). If the convergence conditions are met, stop the iteration and return the optimal solution; otherwise continue to iterate. Use numerical optimization methods to find the optimal weight vector and save it in the target database for subsequent use. This process is a one-time job, and the saved weight vector file can be used continuously.
[0059] (4) Based on the target weight combination and the embedding vectors of the multiple clauses, obtain a weighted average embedding vector of each sliding window of the document to be processed.
[0060] Optionally, the above step (4) can be implemented as follows: Get an embedding vector matrix consisting of the embedding vectors of multiple clauses of any sliding window; A dot product operation is performed on the target weight combination and the embedding vector matrix to obtain a weighted average embedding vector of any sliding window.
[0061] Specifically, each clause in the sliding window has a corresponding embedding vector, and these vectors can form a matrix. By performing a dot product operation on the target weight combination and the embedding vector matrix, a new vector is obtained. This new vector is the weighted average of all word embedding vectors in the sliding window, which represents the semantic information of the entire sliding window. After the user selects the size of the sliding window, the target weight combination is dot-producted with the embedding vector matrix corresponding to the clause in the sliding window to obtain the weighted average embedding vector of the current sliding window. The window is then slid sequentially to obtain the weighted average embedding vector of all sliding windows of the entire document to be processed.
[0062] S12: Obtain a similarity array of the documents to be processed.
[0063] The similarity array includes: the similarity between the weighted average embedding vectors of each adjacent sliding window.
[0064] Specifically, the similarity array of the document to be processed is obtained by calculating the similarity between the weighted average embedding vectors of each adjacent sliding window. Exemplarily, the cosine similarity method can be used to calculate the similarity between the weighted average embedding vectors of each adjacent sliding window. The similarity between two vectors is measured by calculating the cosine value of the angle between them. The value range of cosine similarity is [-1,1]. The closer the value is to 1, the more similar the vectors are. In addition, the straight-line distance between two vectors can be measured by Euclidean distance, or other reasonable similarity calculation methods can be used to measure the similarity between two vectors.
[0065] S13. Adjust the similarity array based on the target variation coefficient to obtain a target similarity array.
[0066] The target coefficient of variation is specified by the user so that the similarity fluctuation level of the paragraph division basis of the same batch of documents reaches the specified level. For example, in this embodiment, the target coefficient of variation can be set to 0.35.
[0067] Specifically, if there is one document, the similarity array of the document is adjusted based on the target variation coefficient to obtain a set of target similarity arrays. If there are multiple documents, the similarity arrays of multiple documents are adjusted based on the target variation coefficient to obtain multiple sets of target similarity arrays.
[0068] In some embodiments, before executing the above step S13 (adjusting the similarity array based on the target variation coefficient to obtain the target similarity array), the following steps may also be executed: Based on the coefficient of variation of the similarity array, a target coefficient of variation is determined; the value of the target coefficient of variation includes: greater than the coefficient of variation of the similarity array; or; less than the coefficient of variation of the similarity array.
[0069] Specifically, the target coefficient of variation is determined based on the coefficient of variation of the similarity array. When there is at least one similarity array, the target coefficient of variation can be determined based on the coefficient of variation of the at least one similarity array, and the value of the target coefficient of variation includes: greater than the coefficient of variation of the at least one similarity array; or; less than the coefficient of variation of the at least one similarity array.
[0070] Exemplarily, which value is specifically selected depends on whether the degree of discreteness of the similarity array is expected to increase or decrease. For example, if the similarity values are expected to be more concentrated, the target coefficient of variation can be selected to be smaller than the coefficient of variation of at least one similarity array. Conversely, if the similarity values are expected to be more dispersed, the target coefficient of variation can be selected to be larger than the coefficient of variation of at least one similarity array. Based on the determined target coefficient of variation, the similarity calculation method or data processing method can be adjusted to achieve the desired degree of discreteness. For example, this can be achieved by adjusting the parameters of the similarity calculation, using different similarity measurement methods, or preprocessing the data (such as standardization or normalization).
[0071] By calculating the coefficient of variation of the similarity array and determining the target coefficient of variation according to actual needs, the discrete degree of the similarity value can be better controlled, thereby improving the accuracy and effectiveness of the similarity calculation.
[0072] In some embodiments, before executing the above step S13 (adjusting the similarity array based on the target variation coefficient to obtain the target similarity array), the following steps may also be executed: Calculate the standard deviation and mean of the similarity array; The coefficient of variation of the similarity array is calculated according to the ratio of the standard deviation to the mean of the similarity array.
[0073] Among them, CV, Coefficient of Variation, is a statistic used to measure the relative volatility of a data set, which represents the ratio of the standard deviation to the mean. Since the coefficient of variation is the ratio of the standard deviation to the mean, it is a dimensionless quantity and is suitable for comparing data of different dimensions or scales. The coefficient of variation reflects the degree of fluctuation of the data relative to its mean, and is suitable for comparing the stability or dispersion of different data sets.
[0074] Specifically, the coefficient of variation is obtained by statistically analyzing the target similarity array. That is, first, all data values in the similarity array are added together, and then divided by the number of data to calculate the average value of the similarity array. Then, the square of the difference between each data value and the average value is calculated, and then all square values are added together, divided by the number of data, and the square root is taken to obtain the standard deviation of the similarity array. Finally, the standard deviation is divided by the average value and multiplied by 100% to obtain the coefficient of variation of the similarity array.
[0075] For example, in the initial case, it is assumed that two sets of similarity arrays of two documents are obtained, and the data are as follows: The first set of data: data1 = [0.912, 0.865, 0.413, 0.675, 0.755, 0.322, 0.876, 0.821, 0.367, 0.688]; The second set of data: data2 = [0.612, 0.765, 0.513, 0.775, 0.715, 0.522, 0.776, 0.621, 0.867, 0.668].
[0076] Through the above calculation steps, the coefficient of variation of the first set of data is about 31.62%, and the coefficient of variation of the second set of data is about 16.20%.
[0077] Optionally, the above step S13 can be implemented in the following manner: Based on the mean and standard deviation of the similarity array and the target coefficient of variation, each value in the similarity array is calculated to obtain a target similarity array.
[0078] Optionally, the above steps can be implemented as follows: Calculate each value in the target similarity array corresponding to the similarity array according to the adjustment formula:
[0079] in, Represents the adjusted first numerical values, Represents the first numerical values, represents the target coefficient of variation, represents the mean of the similarity array, Represents the standard deviation of the similarity array.
[0080] Specifically, assuming that there are multiple documents, the similarity arrays of the multiple documents are adjusted based on the target variation coefficient to obtain multiple sets of target similarity arrays. Based on the mean and standard deviation of each similarity array and the target variation coefficient, each value in each similarity array is calculated to obtain the target similarity array.
[0081] Exemplarily, the formula is adjusted according to the above formula, and the original data is transformed in sequence with the entire array as the unit. After data1 and data2 are converted, they are respectively: Target similarity array 1: [0.9379, 0.8859, 0.3856, 0.6756, 0.7642, 0.2849, 0.8981, 0.8372, 0.3347, 0.6900]; Target similarity array 2: [0.5291, 0.8597, 0.3152, 0.8813, 0.7517, 0.3347, 0.8835, 0.5486, 1.0801, 0.6501].
[0082] Through the above coefficient of variation calculation steps, the coefficient of variation of the first set of data is approximately 35.01%, and the coefficient of variation of the second set of data is approximately 35.01%.
[0083] S14. Determine a dynamic similarity threshold according to the target variation coefficient and segment information.
[0084] The segmentation information is used to indicate the range of the number of clauses contained in the segmentation. The segmentation information can be understood as a user segmentation preference value, which reflects the degree of preference for different document segmentations. For example, the user segmentation preference value can be a large paragraph segmentation, a medium paragraph segmentation, and a small paragraph segmentation, wherein the large paragraph segmentation, the medium paragraph segmentation, and the small paragraph segmentation can be represented by 0, 1, and 2, respectively. The specific division of the large paragraph segmentation, the medium paragraph segmentation, and the small paragraph segmentation can be determined by setting the range of the number of clauses contained in the segmentation.
[0085] Specifically, a dynamic similarity threshold is determined based on the target variation coefficient and segmentation information. The dynamic similarity threshold is generated based on the target variation coefficient and the user segmentation preference value, and is used to adapt to the semantic distribution of different documents.
[0086] Exemplarily, the dynamic similarity threshold may be set to 0.5, or other reasonable values.
[0087] S15. Divide the target similarity array according to the dynamic similarity threshold to obtain a plurality of sub-similarity arrays.
[0088] Specifically, the target similarity array is divided according to the dynamic similarity threshold to obtain multiple sub-similarity arrays.
[0089] Exemplarily, taking the adjusted target similarity array 1 [0.9379, 0.8859, 0.3856, 0.6756, 0.7642, 0.2849, 0.8981, 0.8372, 0.3347, 0.6900] as an example for division, and dividing according to the dynamic similarity threshold of 0.5, we get: Group 1: [0.9379, 0.8859]; Group 2: [0.3856, 0.6756, 0.7642]; Group 3: [0.2849, 0.8981, 0.8372]; Group 4: [0.3347, 0.6900].
[0090] S16. Determine the clauses corresponding to the respective sub-similarity arrays as segments of the document to be processed.
[0091] Specifically, the text sentences corresponding to the grouped data can be segmented in sequence. That is, the clauses corresponding to each sub-similarity array are respectively determined as the segments of the document to be processed.
[0092] Combining dynamic similarity threshold and dynamic clause length vector encoding, it solves the shortcomings of fixed threshold and fixed window size in traditional methods and significantly improves the flexibility and accuracy of semantic segmentation.
[0093] The text semantic segmentation method provided by the embodiment of the present disclosure obtains the weighted average embedding vector of each sliding window corresponding to the document to be processed, wherein the weighted average embedding vector of any sliding window is obtained by weighted summing the embedding vectors of each clause of the document to be processed in the window, obtains the similarity array of the document to be processed, and the similarity array includes: the similarity between the weighted average embedding vectors of each adjacent sliding window, adjusts the similarity array based on the target variation coefficient, obtains the target similarity array, determines the dynamic similarity threshold according to the target variation coefficient and segmentation information, wherein the segmentation information is used to indicate the number range of clauses contained in the segment, divides the target similarity array according to the dynamic similarity threshold to obtain multiple sub-similarity arrays, and determines the clauses corresponding to each sub-similarity array as the segment of the document to be processed. By weighted summing the embedding vectors of each clause in the sliding window, the overall semantic features of the clauses in the window can be reflected, the influence of noise can be reduced, and the similarity between the weighted average embedding vectors of adjacent sliding windows can be calculated, which can more accurately identify the semantically continuous parts of the document. The similarity array is adjusted based on the target variation coefficient, so that the segmentation effect of different documents remains relatively stable. The similarity threshold is dynamically adjusted according to the target variation coefficient and segmentation information, so that the segmentation process can adapt to the structure and content of different documents. The segmentation information indicates the range of the number of clauses contained in the segment, so that the segmentation results are more in line with actual needs and avoid overly long or short segments. The target similarity array is divided according to the dynamic similarity threshold, and the corresponding clauses are determined as the segments of the document to be processed, which improves the accuracy and rationality of text segmentation.
[0094] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0095] Figure 2 A structural diagram of a text semantic segmentation device 200 provided by the present disclosure is shown in FIG. Figure 2 As shown, the device of this embodiment includes: an acquisition module 210, a calculation module 220, an adjustment module 230, an analysis module 240, a division module 250 and a determination module 260, wherein: The acquisition module 210 is used to acquire the weighted average embedding vector of each sliding window corresponding to the document to be processed; the weighted average embedding vector of any sliding window is obtained by weighted summing the embedding vectors of each clause of the document to be processed in the window; The calculation module 220 is used to obtain the similarity array of the document to be processed; the similarity array includes: the similarity between the weighted average embedding vectors of each adjacent sliding window; An adjustment module 230, configured to adjust the similarity array based on a target coefficient of variation to obtain a target similarity array; An analysis module 240 is used to determine a dynamic similarity threshold according to the target coefficient of variation and segmentation information; the segmentation information is used to indicate the range of the number of clauses contained in the segment; A division module 250, configured to divide the target similarity array according to the dynamic similarity threshold to obtain a plurality of sub-similarity arrays; The determination module 260 is used to determine the clauses corresponding to each sub-similarity array as the segments of the document to be processed.
[0096] As an optional implementation of the embodiment of the present disclosure, the acquisition module 210 includes: A splitting unit, used for acquiring a document to be processed and splitting the document to be processed into multiple clauses; A processing unit, configured to input the plurality of clauses into a target embedding model for embedding processing, and obtain embedding vectors of the plurality of clauses; An acquisition unit, used to acquire a target weight combination; the target weight combination includes: weights corresponding to the embedding vectors of the plurality of clauses of the sliding window; A calculation unit is used to obtain a weighted average embedding vector of each sliding window of the document to be processed based on the target weight combination and the embedding vectors of the multiple clauses.
[0097] As an optional implementation of the embodiment of the present disclosure, the acquisition unit is specifically used to: Determine whether the target database includes the target embedding model and the weight combination corresponding to the target window size; the target database includes: multiple embedding model identifiers and multiple weight combinations corresponding to multiple sliding window sizes; the target window size is the size of the sliding window; If the target database includes a combination weight corresponding to the target embedding model and the target window size, the combination weight is used as the target combination weight; If the target database does not include a target embedding model and a combination weight corresponding to the target window size, a target weight combination is obtained based on the target window size and the embedding vectors of the multiple clauses.
[0098] As an optional implementation of the embodiment of the present disclosure, the acquisition unit is further specifically configured to: Obtaining a target embedding vector and an initial weight combination of the plurality of clauses; the target embedding vector is: a combined embedding vector obtained by combining the plurality of clauses in the target window into a target clause and then performing embedding processing; The initial weight combination is optimized according to the target embedding vector and the embedding vectors of the multiple clauses to obtain a target weight combination.
[0099] As an optional implementation of the embodiment of the present disclosure, the acquisition unit is further specifically configured to: Combining multiple clauses within the target window into a target clause; Performing embedding processing on the target clause according to the target embedding model to obtain a target embedding vector of the target clause; Initialize the weights of multiple clauses in the target window to obtain an initial weight combination of the multiple clauses.
[0100] As an optional implementation of the embodiment of the present disclosure, the acquisition unit is further specifically configured to: According to the initial weight combination of the multiple clauses, weighted averaging processing is performed on the embedding vectors of the multiple clauses in the target window to obtain an initial weighted average embedding vector; According to the initial weighted average embedding vector and the target embedding vector, adjusting the initial weight combination to obtain a current weight combination of the multiple clauses; According to the current weight combination of the multiple clauses, weighted averaging processing is performed on the embedding vectors of the multiple clauses in the target window to obtain a current weighted average embedding vector; When the current weighted average embedding vector and the target embedding vector meet a preset convergence condition, a target weight combination is obtained.
[0101] As an optional implementation of the embodiment of the present disclosure, adjusting the initial weight combination to obtain the current weight combination of the multiple clauses according to the initial weighted average embedding vector and the target embedding vector includes:
[0102] in, represents the value of the objective function, represents the Euclidean distance between the weighted average embedding vector and the target embedding vector, is a constant, Indicates The current weight of the clause, represents the weighted average embedding vector, represents the target embedding vector, and the objective function is used to measure the difference between the weighted average embedding vector and the target embedding vector.
[0103] As an optional implementation of the embodiment of the present disclosure, the computing unit is specifically configured to: Get an embedding vector matrix consisting of the embedding vectors of multiple clauses of any sliding window; A dot product operation is performed on the target weight combination and the embedding vector matrix to obtain a weighted average embedding vector of any sliding window.
[0104] As an optional implementation of the embodiment of the present disclosure, the device further includes a coefficient calculation module, and the coefficient calculation module is used to: Calculate the standard deviation and mean of the similarity array; The coefficient of variation of the similarity array is calculated according to the ratio of the standard deviation to the mean of the similarity array.
[0105] As an optional implementation of the embodiment of the present disclosure, the adjustment module is specifically used to: Based on the mean and standard deviation of the similarity array and the target coefficient of variation, each value in the similarity array is calculated to obtain a target similarity array.
[0106] As an optional implementation of the embodiment of the present disclosure, the adjustment module is specifically used to: Calculate each value in the target similarity array corresponding to the similarity array according to the adjustment formula:
[0107] in, Represents the adjusted first numerical values, Represents the first numerical values, represents the target coefficient of variation, represents the mean of the similarity array, Represents the standard deviation of the similarity array.
[0108] The description of the features in the embodiment corresponding to the text semantic segmentation device 200 can be found in the relevant description of the embodiment corresponding to the text semantic segmentation method, which will not be repeated here.
[0109] The text semantic segmentation device provided by the embodiment of the present disclosure obtains the weighted average embedding vector of each sliding window corresponding to the document to be processed, wherein the weighted average embedding vector of any sliding window is obtained by weighted summing the embedding vectors of each clause of the document to be processed in the window, obtains the similarity array of the document to be processed, and the similarity array includes: the similarity between the weighted average embedding vectors of each adjacent sliding window, adjusts the similarity array based on the target variation coefficient, obtains the target similarity array, determines the dynamic similarity threshold according to the target variation coefficient and segmentation information, wherein the segmentation information is used to indicate the number range of clauses contained in the segment, divides the target similarity array according to the dynamic similarity threshold to obtain multiple sub-similarity arrays, and determines the clauses corresponding to each sub-similarity array as the segment of the document to be processed. By weighted summing the embedding vectors of each clause in the sliding window, the overall semantic features of the clauses in the window can be reflected, the influence of noise can be reduced, and the similarity between the weighted average embedding vectors of adjacent sliding windows can be calculated, which can more accurately identify the semantically continuous parts of the document. The similarity array is adjusted based on the target variation coefficient, so that the segmentation effect of different documents remains relatively stable. The similarity threshold is dynamically adjusted according to the target variation coefficient and segmentation information, so that the segmentation process can adapt to the structure and content of different documents. The segmentation information indicates the range of the number of clauses contained in the segment, so that the segmentation results are more in line with actual needs and avoid overly long or short segments. The target similarity array is divided according to the dynamic similarity threshold, and the corresponding clauses are determined as the segments of the document to be processed, which improves the accuracy and rationality of text segmentation.
[0110] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned text semantic segmentation method embodiments.
[0111] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned text semantic segmentation method embodiments when running.
[0112] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0113] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned text semantic segmentation method embodiments are implemented.
[0114] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned text semantic segmentation method embodiments are implemented.
[0115] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0116] The above is a detailed introduction to a text semantic segmentation method provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A text semantic segmentation method, characterized in that: The method comprises: Obtaining the weighted average embedding vectors of each sliding window corresponding to the document to be processed; the weighted average embedding vector of any sliding window is obtained by weighted summing the embedding vectors of each clause of the document to be processed in the window; Obtaining a similarity array of the document to be processed; the similarity array includes: similarities between weighted average embedding vectors of adjacent sliding windows; Adjust the similarity array based on the target variation coefficient to obtain the target similarity array; Determining a dynamic similarity threshold according to the target coefficient of variation and segmentation information, wherein the segmentation information is used to indicate the range of the number of clauses contained in the segment; Dividing the target similarity array according to the dynamic similarity threshold to obtain a plurality of sub-similarity arrays; The clauses corresponding to the respective sub-similarity arrays are respectively determined as segments of the document to be processed.
2. The text semantic segmentation method according to claim 1, characterized in that: The step of obtaining the weighted average embedding vector of each sliding window corresponding to the document to be processed includes: Obtain a document to be processed, and split the document to be processed into multiple clauses; Inputting the multiple clauses into a target embedding model for embedding processing to obtain embedding vectors of the multiple clauses; Obtain a target weight combination; the target weight combination includes: weights corresponding to the embedding vectors of the plurality of clauses of the sliding window; Based on the target weight combination and the embedding vectors of the multiple clauses, a weighted average embedding vector of each sliding window of the document to be processed is obtained.
3. The text semantic segmentation method according to claim 2, characterized in that: The obtaining of the target weight combination comprises: Determine whether the target database includes the target embedding model and the weight combination corresponding to the target window size; the target database includes: multiple embedding model identifiers and multiple weight combinations corresponding to multiple sliding window sizes; the target window size is the size of the sliding window; If the target database includes a combination weight corresponding to the target embedding model and the target window size, the combination weight is used as the target combination weight; If the target database does not include a target embedding model and a combination weight corresponding to the target window size, a target weight combination is obtained based on the target window size and the embedding vectors of the multiple clauses.
4. The text semantic segmentation method according to claim 3, characterized in that: The acquiring a target weight combination based on the target window size and the embedding vectors of the plurality of clauses comprises: Obtaining a target embedding vector and an initial weight combination of the plurality of clauses; the target embedding vector is: a combined embedding vector obtained by combining the plurality of clauses in the target window into a target clause and then performing embedding processing; The initial weight combination is optimized according to the target embedding vector and the embedding vectors of the multiple clauses to obtain a target weight combination.
5. The text semantic segmentation method according to claim 4, characterized in that: The obtaining of the target embedding vector and the initial weight combination of the plurality of clauses comprises: Combining multiple clauses within the target window into a target clause; Performing embedding processing on the target clause according to the target embedding model to obtain a target embedding vector of the target clause; Initialize the weights of multiple clauses in the target window to obtain an initial weight combination of the multiple clauses.
6. The text semantic segmentation method according to claim 4, characterized in that: The step of optimizing the initial weight combination according to the target embedding vector and the embedding vectors of the multiple clauses to obtain a target weight combination includes: According to the initial weight combination of the multiple clauses, weighted averaging processing is performed on the embedding vectors of the multiple clauses in the target window to obtain an initial weighted average embedding vector; According to the initial weighted average embedding vector and the target embedding vector, adjusting the initial weight combination to obtain a current weight combination of the multiple clauses; According to the current weight combination of the multiple clauses, weighted averaging processing is performed on the embedding vectors of the multiple clauses in the target window to obtain a current weighted average embedding vector; When the current weighted average embedding vector and the target embedding vector meet a preset convergence condition, a target weight combination is obtained.
7. The text semantic segmentation method according to claim 6, characterized in that: The adjusting the initial weight combination according to the initial weighted average embedding vector and the target embedding vector to obtain the current weight combination of the multiple clauses includes: in, represents the value of the objective function, represents the Euclidean distance between the weighted average embedding vector and the target embedding vector, is a constant, Indicates The current weight of the clause, represents the weighted average embedding vector, represents the target embedding vector, and the objective function is used to measure the difference between the weighted average embedding vector and the target embedding vector.
8. The text semantic segmentation method according to claim 2, characterized in that: The step of obtaining a weighted average embedding vector of each sliding window of the document to be processed based on the target weight combination and the embedding vectors of the multiple clauses includes: Get an embedding vector matrix consisting of the embedding vectors of multiple clauses of any sliding window; A dot product operation is performed on the target weight combination and the embedding vector matrix to obtain a weighted average embedding vector of any sliding window.
9. The text semantic segmentation method according to claim 1, characterized in that: The similarity array is adjusted based on the target variation coefficient, and before obtaining the target similarity array, the method further includes: Calculate the standard deviation and mean of the similarity array; The coefficient of variation of the similarity array is calculated according to the ratio of the standard deviation to the mean of the similarity array.
10. The text semantic segmentation method according to claim 9, characterized in that: The step of adjusting the similarity array based on the target variation coefficient to obtain the target similarity array includes: Based on the mean and standard deviation of the similarity array and the target coefficient of variation, each value in the similarity array is calculated to obtain a target similarity array.
11. The text semantic segmentation method according to claim 10, characterized in that: The step of calculating each value in the similarity array based on the mean and standard deviation of the similarity array and the target coefficient of variation to obtain a target similarity array includes: Calculate each value in the target similarity array corresponding to the similarity array according to the adjustment formula: in, Represents the adjusted first numerical values, Represents the first numerical values, represents the target coefficient of variation, represents the mean of the similarity array, Represents the standard deviation of the similarity array.
12. A text semantic segmentation device, characterized in that: The device comprises: An acquisition module is used to acquire the weighted average embedding vectors of each sliding window corresponding to the document to be processed; the weighted average embedding vector of any sliding window is obtained by weighted summing the embedding vectors of each clause of the document to be processed in the window; A calculation module, used to obtain a similarity array of the document to be processed; the similarity array includes: similarities between weighted average embedding vectors of each adjacent sliding window; An adjustment module, used for adjusting the similarity array based on the target variation coefficient to obtain a target similarity array; An analysis module, used to determine a dynamic similarity threshold according to the target coefficient of variation and segmentation information; the segmentation information is used to indicate the range of the number of clauses contained in the segment; A division module, used for dividing the target similarity array according to the dynamic similarity threshold to obtain a plurality of sub-similarity arrays; The determination module is used to determine the clauses corresponding to each sub-similarity array as the segments of the document to be processed.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the text semantic segmentation method as claimed in any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the text semantic segmentation method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the text semantic segmentation method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
A document subject term automatic extraction method
CN109726402A
Automatic classification and threshold optimization method and system for specific content of long text
CN118227796A
Intelligent search method based on retrieval enhancement generation
CN118733712A
Intellectual property operation management system based on big data
CN119741156A
Patent data analysis method based on dynamic context window
CN119760117A