Text segmentation method, computer equipment, readable storage medium and program product
By identifying and prioritizing the spaced symbol segmentation of text, and combining semantic recognition and merge, the problem of inaccurate large text segmentation in the knowledge base is solved, and the accuracy and semantic correlation of text segmentation are achieved.
Patent Information
- Application Number
- CN202510831062.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
In the prior art, the text segmentation method of the knowledge base cannot accurately process large text, resulting in model input limitations and poor user reading experience.
By identifying the interval symbols in the target text, selecting the highest priority interval symbols for paragraph segmentation, and semantic recognition and merging are performed after segmentation to ensure that the segmented text complies with the preset length threshold and semantic association.
Improves the accuracy and user experience of text segmentation, ensuring that the segmented text complies with model input limitations and has semantic relevance.
Smart Images

Figure CN120354843A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and particularly to a text segmentation method, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] A knowledge base, also known as an enterprise knowledge base or an intelligent knowledge base, is a system specifically designed to store, manage, and retrieve knowledge within an organization or in a specific domain. It helps users efficiently find the information they need by collecting and organizing data from multiple sources. With the rapid rise of artificial intelligence, especially the breakthroughs in deep learning, the form of knowledge bases is undergoing a transformation. Models (such as large models) are gradually becoming the core technology of the new generation of knowledge bases. However, the length of text that a model can input is limited. Therefore, there is an urgent need for a way to accurately segment documents. Summary of the Invention
[0003] Based on this, in view of the above technical problems, it is necessary to provide a text segmentation method, device, computer device, computer-readable storage medium, and computer program product that can improve the accuracy of text segmentation.
[0004] In a first aspect, the present application provides a text segmentation method, including:
[0005] Obtain a target text and identify the delimiter symbols in the target text;
[0006] Symbol selection step: Select a target delimiter symbol with the highest symbol priority from the delimiter symbols in the target text;
[0007] Perform paragraph segmentation on the target text according to the target delimiter symbol to obtain a plurality of first paragraph segmentation texts;
[0008] If there is a target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold corresponding to the target text, update the target paragraph segmentation text to the target text and return to execute the symbol selection step until there is no target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold, to obtain paragraph segmentation text information;
[0009] Perform semantic recognition on each paragraph text in the paragraph segmentation text information respectively to obtain semantic recognition information, and merge each paragraph text respectively according to the semantic recognition information to obtain a target segmentation text.
[0010] In a second aspect, the present application further provides a text segmentation device, including:
[0011] An acquisition module, configured to acquire a target text and identify interval symbols in the target text;
[0012] A selection module, configured to perform a symbol selection step: select a target interval symbol with the highest symbol priority from the interval symbols in the target text;
[0013] A segmentation module, configured to perform paragraph segmentation on the target text according to the target interval symbol to obtain a plurality of first paragraph segmentation texts;
[0014] An update module, configured to, if there is a target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than a preset text length threshold corresponding to the target text, update the target paragraph segmentation text to the target text and return to execute the symbol selection step until there is no target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold, to obtain paragraph segmentation text information;
[0015] A merging module, configured to respectively perform semantic recognition on each paragraph text in the paragraph segmentation text information to obtain semantic recognition information, and merge each paragraph text according to the semantic recognition information to obtain a target segmentation text.
[0016] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0017] Acquire a target text and identify interval symbols in the target text;
[0018] Symbol selection step: select a target interval symbol with the highest symbol priority from the interval symbols in the target text;
[0019] Perform paragraph segmentation on the target text according to the target interval symbol to obtain a plurality of first paragraph segmentation texts;
[0020] If there is a target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than a preset text length threshold corresponding to the target text, update the target paragraph segmentation text to the target text and return to execute the symbol selection step until there is no target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold, to obtain paragraph segmentation text information;
[0021] Respectively perform semantic recognition on each paragraph text in the paragraph segmentation text information to obtain semantic recognition information, and merge each paragraph text according to the semantic recognition information to obtain a target segmentation text.
[0022] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0023] Obtain a target text and identify the interval symbols in the target text;
[0024] Symbol selection step: Select a target interval symbol with the highest symbol priority from the interval symbols in the target text;
[0025] Perform paragraph segmentation on the target text according to the target interval symbol to obtain a plurality of first paragraph segmentation texts;
[0026] If there is a target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold corresponding to the target text, update the target paragraph segmentation text to the target text and return to execute the symbol selection step until there is no target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold, to obtain paragraph segmentation text information;
[0027] Perform semantic recognition on each paragraph text in the paragraph segmentation text information respectively to obtain semantic recognition information, and merge each paragraph text respectively according to the semantic recognition information to obtain a target segmentation text.
[0028] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0029] Obtain a target text and identify the interval symbols in the target text;
[0030] Symbol selection step: Select a target interval symbol with the highest symbol priority from the interval symbols in the target text;
[0031] Perform paragraph segmentation on the target text according to the target interval symbol to obtain a plurality of first paragraph segmentation texts;
[0032] If there is a target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold corresponding to the target text, update the target paragraph segmentation text to the target text and return to execute the symbol selection step until there is no target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold, to obtain paragraph segmentation text information;
[0033] Semantically recognize each paragraph text in the paragraph segmentation text information respectively to obtain semantic recognition information, and merge each paragraph text respectively according to the semantic recognition information to obtain the target segmentation text.
[0034] The above text segmentation method, device, computer device, computer-readable storage medium and computer program product obtain a target text and identify the interval symbols in the target text; symbol selection step: select a target interval symbol with the highest symbol priority from the interval symbols in the target text; perform paragraph segmentation on the target text according to the target interval symbol to obtain a plurality of first paragraph segmentation texts; if there is a target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold corresponding to the target text, update the target paragraph segmentation text to the target text and return to execute the symbol selection step until there is no target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold, to obtain paragraph segmentation text information; semantically recognize each paragraph text in the paragraph segmentation text information respectively to obtain semantic recognition information, and merge each paragraph text respectively according to the semantic recognition information to obtain the target segmentation text.
[0035] In this way, first, based on the priority of the interval symbols, a preliminary division of the paragraphs of the target text is performed. Considering that the interval symbols are determined by the will of the author of the target text and can reflect the text layout of the target text to a certain extent, the paragraph segmentation text information has a certain segmentation effect; after the preliminary segmentation, semantic recognition is performed on each paragraph text obtained by the segmentation, and based on the semantic recognition information, each paragraph text is merged to ensure that there is a certain semantic association between the merged paragraph texts. Therefore, the accuracy of text segmentation is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for describing the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 It is an application environment diagram of the text segmentation method in an embodiment;
[0038] Figure 2 It is a flowchart of the text segmentation method in an embodiment;
[0039] Figure 3A flowchart showing the steps of merging each paragraph of text according to semantic recognition information in an embodiment to obtain a target segmented text;
[0040] Figure 4 A structural block diagram of a text segmentation device in an embodiment;
[0041] Figure 5 An internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0042] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0043] It should be noted that the information and data involved in the present application (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by users or fully authorized by all parties, and the acquisition, transmission, storage, use and processing of relevant data all comply with the relevant provisions of national laws and regulations. For the content pushed to users (for example, the first paragraph segmented text, semantic recognition information, target segmented text, etc.), users can refuse or can conveniently refuse content push, etc. In the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned, and they should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.
[0044] The text segmentation method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers. Obtain the target text through the server 104, and identify the interval symbols in the target text; Symbol selection step: Select the target interval symbol with the highest symbol priority from the interval symbols in the target text; According to the target interval symbol, perform paragraph segmentation on the target text to obtain multiple first paragraph segmentation texts; If there is a target paragraph segmentation text in the multiple first paragraph segmentation texts whose text length is greater than the preset text length threshold corresponding to the target text, update the target paragraph segmentation text to the target text, and return to execute the symbol selection step until there is no target paragraph segmentation text in the multiple first paragraph segmentation texts whose text length is greater than the preset text length threshold, to obtain paragraph segmentation text information; Perform semantic recognition on each paragraph text in the paragraph segmentation text information respectively to obtain semantic recognition information, and according to the semantic recognition information, merge each paragraph text respectively to obtain the target segmentation text. The server 104 can push at least one of the first paragraph segmentation text, semantic recognition information, and target segmentation text to the terminal 102. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0045] In an exemplary embodiment, as Figure 2 shown, a text segmentation method is provided. Taking the server 104 in Figure 1 as an example and described in an abbreviated form of the main body, it includes the following steps 202 to step 210. Among them:
[0046] Step 202, obtain the target text and identify the interval symbols in the target text.
[0047] Among them, the target text in step 202 is the text waiting for text segmentation. The target text can be the text waiting to be input into the model, or the text waiting to be typeset to a preset position. The preset position can be a newspaper column block, or other positions such as web page layout positions that require restricting the text length. The interval symbol is a connection symbol used between paragraphs or sentences in the text. The interval symbols include the connection symbols between paragraphs and the connection symbols between sentences. The connection symbols between paragraphs include, but are not limited to, line break symbols and consecutive line break symbols, etc. The connection symbols between sentences include, but are not limited to, spaces, full stops, semicolons, commas, and pauses, etc.
[0048] In this way, considering that the text processing of the model usually has a limit on the input text length, therefore, taking the text waiting to be input into the model as the target text and performing subsequent text segmentation, so that the obtained target segmented text will not exceed the input text length limit of the model, thus ensuring that the input text received by the model each time is complete; considering preset positions such as newspaper column blocks and web page layout positions, in order to ensure the user's reading experience, it is necessary to limit the text length of each preset position. Therefore, taking the text waiting to be typeset to the preset position as the target text and performing subsequent text segmentation, so that the obtained target segmented text can conform to the conventional reading method after typesetting.
[0049] As an embodiment, identifying the interval symbols in the target text includes: identifying the symbols in the target text and screening the interval symbols from the symbols in the target text.
[0050] Step 204, select the target interval symbol with the highest symbol priority from the interval symbols in the target text.
[0051] Among them, the symbol priority of the connection symbols between paragraphs in the interval symbols in step 204 is higher than the symbol priority of the connection symbols between sentences; the symbol priorities of the connection symbols between paragraphs and the connection symbols between sentences can be set by the user as needed, or in the following way: the higher the text relevance of the two parts of the text connected by the connection symbol, the lower the symbol priority of the connection symbol. The determination methods of text relevance include, but are not limited to, Pearson correlation coefficient, Spearman rank correlation coefficient, cosine similarity, and Jaccard similarity, etc.
[0052] Exemplarily, before step 204, the method further includes: obtaining the symbol priorities of the interval symbols in the target text and executing step 204.
[0053] Furthermore, as an embodiment, obtaining the symbol priorities of the interval symbols in the target text includes: obtaining the symbol priorities of the interval symbols in the target text set by the user.
[0054] As another embodiment, obtaining the symbol priorities of each interval symbol in the target text includes: for each interval symbol, obtaining the text relevance between the two parts of text connected by the interval symbol, and generating the symbol priority of the interval symbol according to the text relevance between the two parts of text connected by the interval symbol.
[0055] In this way, through shallow text relevance analysis, the lower the symbol priority of the interval symbol with higher text relevance between the two parts of text connected correspondingly is set, ensuring that two parts of text with a certain relevance are not split as much as possible.
[0056] Step 206: Perform paragraph segmentation on the target text according to the target interval symbol to obtain multiple first paragraph segmentation texts.
[0057] Exemplarily, step 206 includes: performing paragraph segmentation on the target text according to the position where the target interval symbol is located to obtain multiple first paragraph segmentation texts.
[0058] If there is a target paragraph segmentation text with a text length greater than the preset text length threshold corresponding to the target text among the multiple first paragraph segmentation texts, then execute step 208 to update the target paragraph segmentation text to the target text.
[0059] Among them, the preset text length threshold is adapted to the object corresponding to the input of the target text (including but not limited to the above-mentioned model and preset position). Specifically, the longer the input limit text length of the object corresponding to the input of the target text is, the longer the preset text length threshold is, and the preset text length threshold is less than the input limit text length of the object corresponding to the input of the target text.
[0060] In this way, since the preset text length threshold is used to limit the length of the pre-segmented paragraph text, and the longer the input limit text length of the object corresponding to the input of the target text is, it indicates that the object corresponding to the input of the target text has a stronger accommodation ability. Then, the preset text length threshold can also be set longer to ensure that the information that can be displayed matches the accommodation ability of the object corresponding to the input of the target text; and since text merging is required later, the preset text length threshold needs to be set less than the input limit text length of the object corresponding to the input of the target text, so as to ensure that the finally obtained target segmentation text does not exceed the accommodation ability of the object corresponding to the input of the target text.
[0061] Return to execute step 204 until there is no target paragraph segmentation text with a text length greater than the preset text length threshold among the multiple first paragraph segmentation texts, and obtain the paragraph segmentation text information.
[0062] In this way, it is ensured that the target text is segmented into paragraphs in sequence according to the symbol priority, and then, with the preset text length threshold as the start and end marks of the loop, each first paragraph segmentation text in the segmented paragraph segmentation text information not only meets the length requirement but also ensures the segmentation accuracy.
[0063] Step 210: Semantically recognize each paragraph text in the paragraph segmentation text information to obtain semantic recognition information, and merge each paragraph text according to the semantic recognition information to obtain the target segmentation text.
[0064] Among them, the semantic recognition methods in Step 210 include but are not limited to rule-based semantic recognition methods, such as dictionary matching method, pattern matching method, and syntax analysis tree method, etc., and machine learning-based methods, such as word embedding model, language model, and graph neural network, etc.
[0065] As an embodiment, merging each paragraph text according to the semantic recognition information to obtain the target segmentation text includes: evaluating the association degree between each paragraph text according to the semantic recognition information to obtain association information; merging each paragraph text according to each association information to obtain the target segmentation text.
[0066] As an embodiment, after merging each paragraph text according to the semantic recognition information to obtain the target segmentation text, the above method further includes: determining the target paragraph segmentation information according to the target segmentation text; constructing target metadata according to the target paragraph segmentation information and the target text.
[0067] As an embodiment, constructing target metadata according to the target paragraph segmentation information and the target text includes: constructing initial metadata according to the target text; adding the target paragraph segmentation information to the initial metadata to obtain the target metadata.
[0068] In the above text segmentation method, first, the target text is pre-segmented into paragraphs based on the priority of the delimiter symbols. Considering that the delimiter symbols are determined by the will of the writer of the target text and can reflect the text layout of the target text to a certain extent, the paragraph segmentation text information has a certain segmentation effect. After pre-segmentation, each segmented paragraph text is semantically recognized, and based on the semantic recognition information, each paragraph text is merged, ensuring that there is a certain semantic association between the merged paragraph texts. Therefore, the accuracy of text segmentation is improved.
[0069] In an exemplary embodiment, as Figure 3As shown, a method for accurately merging paragraph texts is provided. According to semantic recognition information, each paragraph text is merged respectively to obtain the target segmented text, including steps 302 to 310. Among them:
[0070] Step 302, divide each paragraph text into multiple paragraph text groups.
[0071] Exemplarily, dividing each paragraph text into multiple paragraph text groups includes: determining the text information of each paragraph text, where the text information is used to characterize at least one of the segmentation status and text complexity of each paragraph text, and dividing each paragraph text into multiple paragraph text groups according to the text information of each paragraph text.
[0072] Further, determining the text information of each paragraph text includes: the text information includes text complexity, obtaining the text length and text character information of the paragraph text, and determining the text character repetition rate of the paragraph text according to the text character information, where the text character repetition rate is the average value of the character repetition rates of the characters in the paragraph text; determining the text complexity of the paragraph text according to the text character repetition rate and text length of the paragraph text, where the lower the text character repetition rate of the paragraph text, the higher the determined text complexity of the paragraph text, and the longer the text length of the paragraph text, the higher the determined text complexity of the paragraph text.
[0073] As an embodiment, the text information includes the total number of paragraphs; dividing each paragraph text into multiple paragraph text groups according to the text information of each paragraph text includes: determining any divisor of the total number of paragraphs of each paragraph text as the number of paragraphs in the text group corresponding to each paragraph text, or determining the largest divisor or median divisor of the total number of paragraphs of each paragraph text as the number of paragraphs in the text group corresponding to each paragraph text, where the divisor of the total number of paragraphs means that when the total number of paragraphs is divided by the divisor of the total number of paragraphs, it can be divided evenly without a remainder; dividing each paragraph text into multiple paragraph text groups according to the number of paragraphs in the text group.
[0074] In this way, it can be ensured that for all paragraph texts, the number of paragraphs in each group divided based on the number of paragraphs in the text group is consistent.
[0075] It can be understood that only ensuring that the number of paragraphs in each group divided based on the number of paragraphs in the text group is consistent may result in the following situation: when the text length distributions of each paragraph text are uneven, but when dividing with a consistent number of paragraphs in the text group, it is easy to have a situation where the total text length of each paragraph text included in one paragraph text group is relatively long, or the total text length of each paragraph text included in another paragraph text group is relatively short, resulting in the technical defect of uneven paragraph text division.
[0076] To overcome the technical defect of uneven division of the above paragraph text, as another embodiment, the text information includes text complexity; the complexity difference between the total text complexities of the paragraph texts included in each divided paragraph text group is within a preset difference range, and the preset difference range can be set by the user according to needs or can be an empirical value.
[0077] Specifically, according to the text information of each paragraph text, each paragraph text is divided into multiple paragraph text groups, including: obtaining the pre-grouping quantity, where the pre-grouping quantity can be set by the user according to needs or can be determined by the total text complexity corresponding to all paragraph texts. Specifically, the higher the total text complexity corresponding to all paragraph texts, the larger the pre-grouping quantity; determining the ratio between the total text complexity corresponding to all paragraph texts and the pre-grouping quantity as the total text complexity of each paragraph text included in each group; for the first paragraph text, dividing the first paragraph text into the first paragraph text group; for each non-first paragraph text, if the sum of the text complexity of the non-first paragraph text and the text complexity of the previous paragraph text group is not greater than the total text complexity of each paragraph text included in each group, then divide the non-first paragraph text into the previous paragraph text group, and if the sum of the text complexity of the non-first paragraph text and the text complexity of the previous paragraph text group is not greater than the total text complexity of each paragraph text included in each group, then divide the non-first paragraph text into the current paragraph text group.
[0078] For example, the paragraph texts include paragraphs a, b, c, d, and e in sequence. Paragraph a is the first paragraph text and is divided into the first paragraph text group. The sum of the text complexity of paragraph b and paragraph a is not greater than the total text complexity of each paragraph text included in each group, and paragraph b is divided into the first paragraph text group; the sum of the text complexity of paragraph c, paragraph a, and paragraph b is greater than the total text complexity of each paragraph text included in each group, and paragraph c is used as the second paragraph text group; the sum of the text complexity of paragraph d and paragraph c is not greater than the total text complexity of each paragraph text included in each group, and paragraph d is divided into the second paragraph text group; the sum of the text complexity of paragraph e, paragraph d, and paragraph c is not greater than the total text complexity of each paragraph text included in each group, and paragraph e is divided into the second paragraph text group.
[0079] In this way, it can be ensured that the total text complexities of the divided paragraph text groups are relatively uniform.
[0080] It can be understood that when dividing the paragraph text groups based only on the text information, after obtaining the paragraph text groups, it may be necessary to merge the paragraph texts in the paragraph text groups. As a result, it is possible that all the paragraph texts in the paragraph text groups need to be merged. Without considering the text length limit of the divided text of the paragraph text groups, it is easy to have the situation that the final merged target segmented text does not meet the input text length limit corresponding to the target text, resulting in the technical defect of inaccurate text segmentation.
[0081] To overcome the above technical defect of inaccurate text segmentation caused by the situation that the final merged target segmented text does not meet the input text length limit corresponding to the target text, as an embodiment, each paragraph text is divided into multiple paragraph text groups, including: obtaining the restricted text length corresponding to the target text, and determining the text information of each paragraph text, where the text information is used to represent at least one of the segmentation status and text complexity of each paragraph text; dividing each paragraph text into multiple paragraph text groups according to the text information of each paragraph text and the restricted text length.
[0082] In this way, considering the situation that all the paragraph texts in the paragraph text groups may need to be merged, using both the text information of the paragraph texts and the restricted text length as the basis for dividing the paragraph text groups can ensure that the multiple divided paragraph text groups do not exceed the input text length limit corresponding to the target text. Furthermore, it can ensure that the final merged target segmented text also does not exceed the input text length limit corresponding to the target text, improving the accuracy of text segmentation.
[0083] Among them, the restricted text length is adapted to the object corresponding to the input of the target text. Specifically, the longer the input restricted text length of the object corresponding to the input of the target text, the longer the restricted text length, and the restricted text length is greater than the preset text length threshold.
[0084] Furthermore, as an embodiment, the text information includes the text complexity of each paragraph text; dividing each paragraph text into multiple paragraph text groups according to the text information of each paragraph text and the restricted text length includes: determining the number of pre-grouped paragraphs according to the text complexity of each paragraph text; pre-division step: pre-dividing each paragraph text according to the number of pre-grouped paragraphs; determining the pre-grouped text length of the multiple pre-divided text groups obtained by pre-division according to the text length of each paragraph text; if the pre-grouped text lengths are not all less than or equal to the restricted text length, then reduce the number of pre-grouped paragraphs and return to execute the pre-division step until the pre-grouped text lengths are all less than or equal to the restricted text length, and determine the multiple pre-divided text groups as multiple paragraph text groups.
[0085] Among them, the number of pre-grouped paragraphs is the total number of paragraph texts included in the pre-divided text groups.
[0086] Optionally, according to the text complexity of each paragraph text, the specific implementation manner of determining the number of pre-grouped paragraphs may refer to the above-mentioned specific implementation content of dividing each paragraph text into multiple paragraph text groups according to the text information of each paragraph text, which will not be elaborated here.
[0087] In this way, it is ensured that the multiple paragraph text groups obtained by division not only match the text complexity of each paragraph text, but also do not exceed the input text length limit corresponding to the target text.
[0088] As an embodiment, reducing the number of pre-grouped paragraphs includes: when the number of pre-grouped paragraphs includes multiple ones, at least one of the number of pre-grouped paragraphs can be reduced.
[0089] For each paragraph text group, step 304 is executed. According to the semantic recognition information, the correlation degree between each pair of paragraph texts included in the paragraph text group is evaluated respectively to obtain correlation information.
[0090] Among them, the correlation information can be used to represent at least one of the correlation degree between each pair of adjacent paragraph texts included in the paragraph text group and the correlation degree between each pair of non-adjacent paragraph texts.
[0091] Exemplarily, step 304 includes: for each paragraph text group, according to the semantic recognition information, the text semantic features of each paragraph text included in the paragraph text group are extracted respectively, where the text semantic features are used to represent at least one of the theme feature, view feature, term feature, event feature, and concept feature of the paragraph text; according to the text semantic features of each paragraph text included in the paragraph text group, the correlation degree between each pair of paragraph texts included in the paragraph text group is evaluated respectively to obtain correlation information.
[0092] As an embodiment, according to the text semantic features of each paragraph text in the paragraph text group, the correlation degree between each pair of paragraph texts included in the paragraph text group is evaluated respectively to obtain correlation information, including: for each paragraph text pair composed of two paragraph texts, according to the text semantic features of the two paragraph texts in the paragraph text pair, the correlation between the paragraph text pair is evaluated respectively from each semantic feature dimension to obtain the semantic dimension correlation information corresponding to each semantic feature dimension, where the semantic feature dimensions include but are not limited to the above-mentioned theme feature, view feature, term feature, event feature, and concept feature; according to the semantic dimension weights corresponding to each semantic feature dimension, the semantic dimension correlation information corresponding to each semantic feature dimension is weighted respectively to obtain the semantic weighted correlation information corresponding to each semantic feature dimension, where the semantic dimension weights can be set by the user as needed, and the semantic dimension weights include the weights of at least one of the theme feature, view feature, term feature, event feature, and concept feature of the paragraph text on the evaluation of the inter-paragraph correlation; the semantic weighted correlation information corresponding to each semantic feature dimension is fused to obtain the correlation information.
[0093] If there is a target paragraph text pair in the paragraph text group whose correlation information meets the preset correlation condition, then step 306 is executed to merge the target paragraph text pair in the paragraph text group.
[0094] Among them, if the correlation degree represented by the correlation information is greater than the preset degree threshold, it is determined that the correlation information meets the preset correlation condition. The preset degree threshold can be set by the user as needed or can be an empirical value, which is not limited here.
[0095] It can be understood that the above-mentioned correlation information may represent the correlation degree between two non-adjacent paragraph texts. Therefore, there may be a situation where the two paragraph texts included in the target paragraph text pair are not adjacent paragraphs. At this time, if the target paragraph text pair is directly merged, there may be a technical defect that the merged text is poorly described or logically incorrect, resulting in a low accuracy of paragraph merging.
[0096] To overcome the above technical defect that the accuracy of paragraph merging is low due to the possible situation that the merged text is poorly described or logically incorrect, before merging the target paragraph text pair in the paragraph text group, the above method further includes: if the two paragraph texts included in the target paragraph text pair are adjacent paragraphs, then execute the step of merging the target paragraph text pair in the paragraph text group; if the two paragraph texts included in the target paragraph text pair are not adjacent paragraphs, then determine the text fluency between the two paragraph texts included in the target paragraph text pair, and execute the step of merging the target paragraph text pair in the paragraph text group when the text fluency meets the preset fluency condition.
[0097] Furthermore, determine the text fluency between the two passage texts included in the target passage text pair, including: extracting the last text part of the earlier passage text in the target passage text pair, and the first text part of the later passage text in the target passage text pair; evaluating the connection fluency and logical coherence between the last text part and the first text part; determining the text fluency between the two passage texts included in the target passage text pair according to the connection fluency and logical coherence between the last text part and the first text part, where the higher the connection fluency, the higher the text fluency; the higher the logical coherence, the higher the text fluency.
[0098] Among them, the connection fluency and logical coherence between the last text part and the first text part can be realized by natural language processing (NLP, Natural Language Processing) technology.
[0099] If there is no target passage text pair in the passage text group whose relevance information meets the preset relevance condition, then execute step 308 to keep the passage text group unchanged.
[0100] When all passage text groups have been processed, execute step 310 to combine each passage text group to obtain the target segmented text.
[0101] In this embodiment, each passage text is divided into multiple passage text groups, so as to conduct targeted relevance degree evaluation on each passage text group, and then realize the merging of text pairs in the passage text group. Considering that related passages usually do not have a large interval, therefore, it is not necessary to evaluate the relevance degree between all passage texts. By grouping and separately evaluating the relevance degree, the efficiency of passage text merging is ensured.
[0102] As a detailed embodiment, obtain the target text and identify the interval symbols in the target text; symbol selection step: select the target interval symbol with the highest symbol priority from the interval symbols in the target text; according to the target interval symbol, perform paragraph segmentation on the target text to obtain multiple first paragraph segmentation texts; if there is a target paragraph segmentation text with a text length greater than the preset text length threshold corresponding to the target text among the multiple first paragraph segmentation texts, update the target paragraph segmentation text to the target text and return to execute the symbol selection step until there is no target paragraph segmentation text with a text length greater than the preset text length threshold among the multiple first paragraph segmentation texts, and obtain the paragraph segmentation text information; respectively perform semantic recognition on each paragraph text in the paragraph segmentation text information to obtain semantic recognition information, and obtain the restricted text length corresponding to the target text, and determine the text information of each paragraph text, where the text information is used to represent at least one of the segmentation status and text complexity of each paragraph text; determine the number of pre-grouped paragraphs according to the text complexity of each paragraph text; pre-partitioning step: pre-partition each paragraph text according to the number of pre-grouped paragraphs; determine the pre-grouped text length of the multiple pre-partitioned text groups obtained by pre-partitioning according to the text length of each paragraph text; if the pre-grouped text lengths are not all less than or equal to the restricted text length, reduce the number of pre-grouped paragraphs and return to execute the pre-partitioning step until the pre-grouped text lengths are all less than or equal to the restricted text length, and determine the multiple pre-partitioned text groups as multiple paragraph text groups.
[0103] Further, for each group of paragraph texts, according to the semantic recognition information, the text semantic features of each paragraph text included in the group of paragraph texts are respectively extracted, where the text semantic features are used to represent at least one of the theme features, view features, term features, event features, and concept features of the paragraph text; according to the text semantic features of each paragraph text included in the group of paragraph texts, the degree of association between each pair of paragraph texts included in the group of paragraph texts is respectively evaluated to obtain association information; if there is a target paragraph text pair in the group of paragraph texts whose association information meets the preset association condition, then when the two paragraph texts included in the target paragraph text pair are adjacent paragraphs, perform the step of merging the target paragraph text pair in the group of paragraph texts; when the two paragraph texts included in the target paragraph text pair are not adjacent paragraphs, determine the text fluency between the two paragraph texts included in the target paragraph text pair, and when the text fluency meets the preset fluency condition, perform the step of merging the target paragraph text pair in the group of paragraph texts; if there is no target paragraph text pair in the group of paragraph texts whose association information meets the preset association condition, keep the group of paragraph texts unchanged; when all groups of paragraph texts are processed, combine the groups of paragraph texts to obtain the target segmented text; determine the target paragraph segmentation information according to the target segmented text; construct the target metadata according to the target paragraph segmentation information and the target text.
[0104] In this way, first, based on the priority of the delimiter symbols, the target text is pre-divided into paragraphs. Considering that the delimiter symbols are determined by the will of the author of the target text and can reflect the text layout of the target text to a certain extent, the paragraph segmentation text information has a certain segmentation effect; after the pre-segmentation, semantic recognition is performed on each segmented paragraph text, and based on the semantic recognition information, each paragraph text is merged, ensuring that there is a certain semantic association between the merged paragraph texts. Therefore, the accuracy of text segmentation is improved.
[0105] Further, each paragraph text is divided into multiple groups of paragraph texts, so as to conduct targeted evaluation of the degree of association for each group of paragraph texts, and then realize the merging of text pairs in the group of paragraph texts. Considering that paragraphs with associations usually are not far apart, it is not necessary to evaluate the degree of association between all paragraph texts. By grouping and evaluating the degree of association separately, the efficiency of paragraph text merging is ensured.
[0106] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0107] Based on the same inventive concept, an embodiment of the present application further provides a text segmentation device for implementing the above-mentioned text segmentation method. The implementation solutions for solving problems provided by this device are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more embodiments of the following text segmentation devices can refer to the limitations on the text segmentation method in the above text and will not be repeated here.
[0108] In an exemplary embodiment, as Figure 4 shown, a text segmentation device 400 is provided, including: an acquisition module 402, a selection module 404, a segmentation module 406, an update module 408, and a merging module 410, where:
[0109] The acquisition module 402 is configured to acquire a target text and identify the delimiter symbols in the target text;
[0110] The selection module 404 is configured to perform a symbol selection step: select a target delimiter symbol with the highest symbol priority from the delimiter symbols in the target text;
[0111] The segmentation module 406 is configured to perform paragraph segmentation on the target text according to the target delimiter symbol to obtain a plurality of first paragraph segmentation texts;
[0112] The update module 408 is configured to, if there is a target paragraph segmentation text with a text length greater than the preset text length threshold corresponding to the target text among the plurality of first paragraph segmentation texts, update the target paragraph segmentation text to the target text and return to execute the symbol selection step until there is no target paragraph segmentation text with a text length greater than the preset text length threshold among the plurality of first paragraph segmentation texts, to obtain paragraph segmentation text information;
[0113] The merging module 410 is configured to perform semantic recognition on each paragraph text in the paragraph segmentation text information to obtain semantic recognition information, and merge each paragraph text according to the semantic recognition information to obtain a target segmentation text.
[0114] In one embodiment, the merging module 410 is further configured to divide each paragraph text into a plurality of paragraph text groups; for each paragraph text group, according to the semantic recognition information, evaluate the association degree between each pair of paragraph texts included in the paragraph text group respectively to obtain association information; if there is a target paragraph text pair in the paragraph text group whose association information meets the preset association condition, then merge the target paragraph text pair in the paragraph text group; if there is no target paragraph text pair in the paragraph text group whose association information meets the preset association condition, then keep the paragraph text group unchanged; when all paragraph text groups are processed, combine each paragraph text group to obtain the target segmented text.
[0115] In one embodiment, the merging module 410 is further configured to obtain the restricted text length corresponding to the target text, and determine the text information of each paragraph text, where the text information is used to characterize at least one of the segmentation status and text complexity of each paragraph text; divide each paragraph text into a plurality of paragraph text groups according to the text information of each paragraph text and the restricted text length.
[0116] In one embodiment, the text information includes the text complexity of each paragraph text; the merging module 410 is further configured to determine the number of pre-grouped paragraphs according to the text complexity of each paragraph text; pre-partitioning step: pre-partition each paragraph text according to the number of pre-grouped paragraphs; determine the pre-grouped text length of the plurality of pre-partitioned text groups obtained by pre-partitioning according to the text length of each paragraph text; if the pre-grouped text lengths are not all less than or equal to the restricted text length, then reduce the number of pre-grouped paragraphs and return to execute the pre-partitioning step until the pre-grouped text lengths are all less than or equal to the restricted text length, and determine the plurality of pre-partitioned text groups as a plurality of paragraph text groups.
[0117] In one embodiment, before merging the target paragraph text pair in the paragraph text group, the merging module 410 is further configured to, if the two paragraph texts included in the target paragraph text pair are adjacent paragraphs, execute the step of merging the target paragraph text pair in the paragraph text group; if the two paragraph texts included in the target paragraph text pair are not adjacent paragraphs, then determine the text fluency between the two paragraph texts included in the target paragraph text pair, and execute the step of merging the target paragraph text pair in the paragraph text group when the text fluency meets the preset fluency condition.
[0118] In one embodiment, the merging module 410 is further configured to, for each paragraph text group, respectively extract text semantic features of each paragraph text included in the paragraph text group according to the semantic recognition information, where the text semantic features are used to represent at least one of the topic feature, view feature, term feature, event feature, and concept feature of the paragraph text; and respectively evaluate the association degree between each pair of paragraph texts included in the paragraph text group according to the text semantic features of each paragraph text included in the paragraph text group, so as to obtain association information.
[0119] In one embodiment, after respectively merging each paragraph text according to the semantic recognition information to obtain the target segmented text, the apparatus further includes: a construction module, configured to determine target paragraph segmentation information according to the target segmented text; and construct target metadata according to the target paragraph segmentation information and the target text.
[0120] Each module in the above text segmentation apparatus can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or be independent of the processor, or can be stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0121] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 5As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements a text segmentation method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0122] Those skilled in the art can understand that Figure 5 the structure shown in
[0123] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0124] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0125] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0126] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0127] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0128] The above embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A text segmentation method, characterized in that The method includes: Obtain a target text and identify the interval symbols in the target text; Symbol selection step: Select a target interval symbol with the highest symbol priority from the interval symbols in the target text; Perform paragraph segmentation on the target text according to the target interval symbol to obtain a plurality of first paragraph segmentation texts; If there is a target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold corresponding to the target text, update the target paragraph segmentation text to the target text and return to execute the symbol selection step until there is no target paragraph segmentation text in the plurality of first paragraph segmentation texts whose text length is greater than the preset text length threshold, to obtain paragraph segmentation text information; Perform semantic recognition on each paragraph text in the paragraph segmentation text information respectively to obtain semantic recognition information, and merge each paragraph text respectively according to the semantic recognition information to obtain a target segmentation text.
2. The method according to claim 1, wherein The merging each paragraph text respectively according to the semantic recognition information to obtain a target segmentation text includes: Divide each paragraph text into a plurality of paragraph text groups; For each paragraph text group, evaluate the degree of association between each pair of paragraph texts included in the paragraph text group according to the semantic recognition information to obtain association information; If there is a target paragraph text pair in the paragraph text group whose association information meets the preset association condition, merge the target paragraph text pair in the paragraph text group; If there is no target paragraph text pair in the paragraph text group whose association information meets the preset association condition, keep the paragraph text group unchanged; When all the paragraph text groups are processed, combine each paragraph text group to obtain a target segmentation text.
3. The method according to claim 2, wherein The dividing each paragraph text into a plurality of paragraph text groups includes: Obtain the restricted text length corresponding to the target text, and determine the text information of each paragraph text, where the text information is used to represent at least one of the segmentation status and text complexity of each paragraph text; Divide each paragraph text into a plurality of paragraph text groups according to the text information of each paragraph text and the restricted text length.
4. The method according to claim 3, characterized in that The text information includes the text complexity of each paragraph text; the dividing each paragraph text into a plurality of paragraph text groups according to the text information of each paragraph text and the restricted text length includes: Determine the number of pre-grouped paragraphs according to the text complexity of each paragraph text; Pre-division step: Pre-divide each paragraph text according to the number of pre-grouped paragraphs; Determine the pre-grouped text length of a plurality of pre-grouped text groups obtained by pre-division according to the text length of each paragraph text; If the pre-grouped text lengths are not all less than or equal to the restricted text length, reduce the number of pre-grouped paragraphs and return to execute the pre-division step until the pre-grouped text lengths are all less than or equal to the restricted text length, and determine the plurality of pre-grouped text groups as a plurality of paragraph text groups.
5. The method according to claim 3, wherein Before merging the target paragraph text pairs in the paragraph text group, the method further includes: If the two paragraph texts included in the target paragraph text pair are adjacent paragraphs, perform the step of merging the target paragraph text pairs in the paragraph text group; If the two paragraph texts included in the target paragraph text pair are not adjacent paragraphs, determine the text fluency between the two paragraph texts included in the target paragraph text pair, and perform the step of merging the target paragraph text pairs in the paragraph text group when the text fluency meets the preset fluency condition.
6. The method according to any one of claims 3 to 5, characterized in that The method of evaluating the association degree between each pair of paragraph texts included in the paragraph text group according to the semantic recognition information to obtain association information includes: For each paragraph text group, according to the semantic recognition information, extract the text semantic features of each paragraph text included in the paragraph text group, where the text semantic features are used to characterize at least one of the theme feature, view feature, term feature, event feature, and concept feature of the paragraph text; According to the text semantic features of each paragraph text included in the paragraph text group, evaluate the association degree between each pair of paragraph texts included in the paragraph text group to obtain association information.
7. The method according to claim 1, wherein After merging each paragraph text according to the semantic recognition information to obtain the target segmented text, the method further includes: Determine the target paragraph segmentation information according to the target segmented text; Construct the target metadata according to the target paragraph segmentation information and the target text.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Document segmentation method, electronic equipment and computer readable storage medium
CN117688926A
Text segmentation method and device, computer readable storage medium and equipment
CN118607531A
Long text analysis method and device, storage medium and terminal
CN119150880A
Text paragraph recognition method and device, computer equipment, readable storage medium and program product
CN119272769A
Electronic document segmentation method, requires assigning weighting factor to each cell indicating agreement of cell content with each key-word
DE10339467A1