Method and device for segmenting unstructured file

By classifying unstructured files and combining the corresponding segmentation strategies for file type matching, natural language processing, compression segmentation and dynamic planning strategies are adopted to solve the problems of low efficiency and insufficient accuracy of unstructured file segmentation, and efficient and accurate file segmentation are achieved, reducing management costs.

CN120337913AActive Publication Date: 2025-07-18JIANGSU TIANYUAN TENERING CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510400828.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

In the prior art, when segmenting unstructured files, the segmentation efficiency is low and the accuracy is difficult to guarantee. Especially when the training data set is very different from the target files, segmentation errors are prone to occur.

Method used

By classifying unstructured files, the segmentation strategy is used to match the corresponding segmentation strategy according to the file size and type, natural language processing, compression segmentation and dynamic programming strategies are used for segmentation, including natural language processing strategy, compression segmentation and dynamic programming strategies.

Benefits of technology

Improves the storage and retrieval efficiency of unstructured files, reduces management costs, and reduces segmentation errors and computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337913A_ABST
    Figure CN120337913A_ABST
Patent Text Reader

Abstract

The invention discloses an unstructured file segmentation method and device. The method comprises the following steps: acquiring a plurality of unstructured files; according to the file volume of each unstructured file, classifying the unstructured files to obtain a classification result; and based on the classification result, combining the file type of each unstructured file, and adopting different segmentation strategies to segment the unstructured files to obtain a file segmentation result. According to the method, the non-structured files are classified, and different segmentation strategies are adopted by combining the classification result with the file types, so that the segmentation of the non-structured files is realized, the effect of improving the storage and retrieval efficiency of the non-structured files is achieved, and a basis is provided for subsequent analysis and value mining of the non-structured files; and meanwhile, the unstructured file management cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of file segmentation, and particularly relates to a method and device for segmenting unstructured files. Background Art

[0002] Unstructured data is data with irregular or incomplete data structures, without a predefined data model, and is not convenient to be represented by a two-dimensional logical table of a database. It includes all formats of office documents, texts, pictures, HTML, various reports, images, audio, and video information, etc. The formats of unstructured data are very diverse, and the standards are also diverse. Moreover, unstructured information is more difficult to standardize and understand than structured information in terms of technology. Therefore, the storage, retrieval, publication, and utilization of unstructured data require more intelligent technologies to achieve. For example, the massive storage, intelligent retrieval, knowledge mining, content protection, and value-added development and utilization of unstructured data.

[0003] In the current related technologies, when segmenting large unstructured files, they do not segment according to the file types in the large unstructured files, resulting in low segmentation efficiency. It may also divide sub-files of different file types into the same segmented file, increasing the probability of errors in the final combined file. Patent CN119513057A provides a method and system for parallel synchronization of unstructured files, including intelligently segmenting a target file using a pre-trained file chunking model, and generating a two-level cascaded index number and a data hash value. Then, according to the data transmission priority model, the data sub-chunks are transmitted to the target-side database through an adaptive parallel transmission channel and are asynchronously verified. Finally, using the pre-trained file recombination model, the data sub-chunks are recombined based on the two-level cascaded index number and the data block association map to generate a synchronized target file, and the confirmation information is submitted to the blockchain network. Among them, the data sub-chunks are obtained by inputting the initial chunking scheme into a dynamic programming model, constructing a cost function with semantic integrity features, data balance features, and structural coherence features as optimization objectives, inputting the cost function into the dynamic programming model, iteratively calculating the optimal state transition sequence through the dynamic programming model to obtain the final chunking scheme, determining the final chunking boundary position according to the final chunking scheme, and splitting the target unstructured file according to the final chunking boundary position to generate multiple data sub-chunks.

[0004] Although a dynamic programming model is used in the related technologies during the file segmentation process, the segmentation accuracy of the dynamic programming model is based on the training data set and training effect of the model. If there are significant differences between the unstructured files in the training data set and the target file, the segmentation effect of the final obtained data sub-chunks cannot be guaranteed.

[0005] How to improve the efficiency of unstructured file segmentation while ensuring the accuracy of unstructured file segmentation is a problem that needs to be solved currently. Summary of the Invention

[0006] In view of the defects existing in the above-mentioned prior art, the present invention provides a method and device for segmenting unstructured files. The method includes: obtaining a plurality of unstructured files; classifying the unstructured files according to the file volume of each unstructured file to obtain a classification result; based on the classification result, combining the file types of each unstructured file, matching a corresponding segmentation strategy, and segmenting the unstructured files to obtain a file segmentation result. By classifying unstructured files and matching corresponding segmentation strategies according to the classification result in combination with the file types, the segmentation of unstructured files is realized, the effect of improving the storage and retrieval efficiency of unstructured files is achieved, a basis is provided for subsequent analysis and value mining of unstructured files, and at the same time, the management cost of unstructured files is reduced.

[0007] In a first aspect, the present invention provides a method for segmenting unstructured files, specifically including the following steps:

[0008] Obtaining a plurality of unstructured files;

[0009] Classifying the unstructured files according to the file volume of each unstructured file to obtain a classification result;

[0010] Based on the classification result, combining the file types of each unstructured file, matching a corresponding segmentation strategy, and segmenting the unstructured files to obtain a file segmentation result.

[0011] Further, the classification result includes at least one of a first file set, a second file set, and a third file set;

[0012] Classifying the unstructured files according to the file volume of each unstructured file to obtain a classification result, specifically including:

[0013] If the file volume of the unstructured file is less than the first volume threshold, add the unstructured file to the first file set;

[0014] If the file volume of the unstructured file is greater than or equal to the first volume threshold and less than the second volume threshold, add the unstructured file to the second file set;

[0015] If the file volume of the unstructured file is greater than or equal to the second volume threshold, add the unstructured file to the third file set.

[0016] Further, based on the classification result, combining the file types of each unstructured file, matching a corresponding segmentation strategy, and segmenting the unstructured files to obtain a file segmentation result, specifically including:

[0017] For the unstructured files in the first file set, a natural language processing strategy is adopted for segmentation to obtain a file segmentation result;

[0018] For the unstructured files in the second file set, a compression segmentation strategy is adopted for segmentation to obtain a file segmentation result;

[0019] If the file type of the unstructured file in the third file set is a text file, traverse the unstructured file, split the unstructured file based on the punctuation marks in the unstructured file, combine to obtain multiple sub - segments, and adopt a dynamic programming strategy for segmentation to obtain a file segmentation result;

[0020] If the file type of the unstructured file in the third file set is a non - text file, determine the number of segmentation points of the unstructured file, and segment the unstructured file based on the number of segmentation points of the unstructured file to obtain a file segmentation result.

[0021] Furthermore, adopting a natural language processing strategy for segmentation to obtain a file segmentation result specifically includes:

[0022] Based on a pre - obtained natural language processing model, perform semantic analysis on the unstructured files in the first file set to obtain segmentation identifiers corresponding to the unstructured files;

[0023] According to the segmentation identifiers corresponding to the unstructured files, segment the unstructured files in the first file set to obtain a file segmentation result.

[0024] Furthermore, splitting the unstructured file based on the punctuation marks in the unstructured file, combining to obtain multiple sub - segments, and adopting a dynamic programming strategy for segmentation to obtain a file segmentation result specifically includes:

[0025] Split the unstructured file according to the punctuation marks in the unstructured file to obtain multiple sub - segments;

[0026] Based on the order of the sub - segments, combine consecutive multiple sub - segments to form candidate sub - texts. Among them, each candidate sub - text includes consecutive n sub - segments, the text volume of each candidate sub - text is less than a preset sub - text threshold, and the text volume of the candidate sub - text and the (n + 1) - th sub - segment is greater than the preset sub - text threshold, n ∈ N + ;

[0027] According to the order between the candidate sub - texts, construct a directed acyclic graph. Among them, each node in the directed acyclic graph includes at least one candidate sub - text, and the edge in the directed acyclic graph is the cross - degree of two candidate sub - texts;

[0028] Based on the directed acyclic graph, analyze the forward shortest path from the source point to the sink point in the directed acyclic graph, and determine the forward splitting point corresponding to the forward shortest path;

[0029] Based on the directed acyclic graph, analyze the reverse shortest path from the sink point to the source point in the directed acyclic graph, and determine the reverse splitting point corresponding to the reverse shortest path;

[0030] Compare the forward shortest path and the reverse shortest path, determine the target splitting point from the forward splitting point and the reverse splitting point, and obtain the file splitting result.

[0031] Furthermore, the cross-degree is specifically expressed as:

[0032] JCD(i - 1, i) = JCW(i - 1, i) 2

[0033] Wherein, JCD(i - 1, i) is the cross-degree between the (i - 1)-th candidate sub-text and the i-th candidate sub-text in the directed acyclic graph, and JCW(i - 1, i) is the total number of characters of the cross-text corresponding to the (i - 1)-th candidate sub-text and the i-th candidate sub-text in the directed acyclic graph.

[0034] Furthermore, based on the directed acyclic graph, analyze the forward shortest path from the source point to the sink point in the directed acyclic graph, and determine the forward splitting point corresponding to the forward shortest path, specifically including:

[0035] Based on the source point of the directed acyclic graph, according to the cross-degree between the source point and other nodes, sequentially determine the next node until reaching the sink point of the directed acyclic graph to determine the forward shortest path, wherein each node in the forward shortest path corresponds to a target sub-text;

[0036] Take the end of the cross-text of any two target sub-texts in the forward shortest path as the forward splitting point.

[0037] Furthermore, based on the directed acyclic graph, analyze the reverse shortest path from the sink point to the source point in the directed acyclic graph, and determine the reverse splitting point corresponding to the reverse shortest path, specifically including:

[0038] Based on the sink point of the directed acyclic graph, according to the cross-degree between the sink point and other nodes, sequentially determine the next node until reaching the source point of the directed acyclic graph to determine the reverse shortest path, wherein each node in the reverse shortest path corresponds to a target sub-text;

[0039] Take the end of the cross-text of any two target sub-texts in the reverse shortest path as the reverse splitting point.

[0040] Furthermore, compare the forward shortest path and the reverse shortest path, determine the target splitting point from the forward splitting point and the reverse splitting point, and obtain the file splitting result, specifically including:

[0041] Compare the sum of the forward crossing degrees corresponding to the forward shortest path and the sum of the reverse crossing degrees corresponding to the reverse shortest path;

[0042] If the sum of the forward crossing degrees is less than the sum of the reverse crossing degrees, use the forward segmentation point as the target segmentation point to segment the unstructured file to obtain the file segmentation result;

[0043] If the sum of the forward crossing degrees is greater than the sum of the reverse crossing degrees, use the reverse segmentation point as the target segmentation point to segment the unstructured file to obtain the file segmentation result;

[0044] If the sum of the forward crossing degrees is equal to the sum of the reverse crossing degrees, determine the target segmentation point based on the number of the forward segmentation point and the reverse segmentation point, and segment the unstructured file to obtain the file segmentation result.

[0045] In a second aspect, the present invention also provides a device for segmenting an unstructured file, which adopts the method for segmenting an unstructured file as described in any one of the above, and includes:

[0046] A data acquisition module, configured to acquire a plurality of unstructured files;

[0047] A file classification module, configured to classify the unstructured files according to the file volume of each unstructured file to obtain a classification result;

[0048] A file segmentation module, configured to match a corresponding segmentation strategy based on the classification result and in combination with the file type of each unstructured file, and segment the unstructured file to obtain the file segmentation result.

[0049] The method and device for segmenting an unstructured file provided by the present invention at least include the following beneficial effects:

[0050] (1) By classifying the unstructured files and matching the corresponding segmentation strategies according to the classification result in combination with the file type, the segmentation of the unstructured files is realized, the effect of improving the storage and retrieval efficiency of the unstructured files is achieved, a basis is provided for subsequent analysis and value mining of the unstructured files, and at the same time, the management cost of the unstructured files is reduced.

[0051] (2) By adopting a compression and segmentation strategy of first compressing and then segmenting the unstructured files, the file volume of the unstructured files can be significantly reduced, the storage, transmission and processing efficiency of the unstructured files can be improved, and at the same time, segmentation errors and computing resource consumption can be reduced. Description of the Drawings

[0052] Figure 1 It is a flowchart of the method for segmenting an unstructured file provided by an embodiment of the present invention;

[0053] Figure 2 Flow chart for determining classification results provided by an embodiment of the present invention;

[0054] Figure 3 Flow chart for matching corresponding segmentation strategies provided by an embodiment of the present invention;

[0055] Figure 4 Flow chart for segmentation using the dynamic programming strategy provided by an embodiment of the present invention;

[0056] Figure 5 Schematic diagram for determining candidate sub - texts provided by an embodiment of the present invention;

[0057] Figure 6 Schematic diagram for constructing a directed acyclic graph provided by an embodiment of the present invention;

[0058] Figure 7 Flow chart for determining the file segmentation result provided by an embodiment of the present invention;

[0059] Figure 8 Structural block diagram of a segmentation device for unstructured files provided by an embodiment of the present invention.

[0060] Among them, 201 is a data acquisition module; 202 is a file classification module; 203 is a file segmentation module. Detailed implementation manners

[0061] In order to better understand the above - mentioned technical solutions, the following will describe the above - mentioned technical solutions in detail in conjunction with the accompanying drawings of the specification and specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0062] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Multiple" generally includes at least two.

[0063] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non - exclusive inclusion, so that a commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a commodity or device. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the commodity or device including the said element.

[0064] Unstructured files have rich data types, diverse formats, and no fixed format. At the same time, unstructured files are generated and changed rapidly, with complex, unpredictable, and huge amounts of content, making them difficult to manage.

[0065] In the practical application of unstructured files, due to the scattered storage of unstructured file data, it is difficult to conduct unified classification and management. Currently, the common practice of the related technology for splitting unstructured files is to split the unstructured files to be split according to a certain fixed value, such as splitting according to a fixed file size or a fixed number of file lines. This method has low splitting efficiency and low splitting accuracy.

[0066] With the acceleration of the digitalization process, the proportion of unstructured data in enterprise data assets is getting higher and higher, usually accounting for more than 80% of the total data volume. Therefore, how to efficiently and accurately split and manage unstructured data has become a key issue in data governance.

[0067] The present invention provides a method for splitting unstructured files. The method includes: obtaining a plurality of unstructured files; classifying the unstructured files according to the file volume of each unstructured file to obtain a classification result; based on the classification result, combining the file type of each unstructured file, matching a corresponding splitting strategy, and splitting the unstructured files to obtain a file splitting result. By classifying unstructured files and matching corresponding splitting strategies according to the classification result in combination with the file type, the splitting of unstructured files is realized, the effect of improving the storage and retrieval efficiency of unstructured files is achieved, a basis is provided for subsequent analysis and value mining of unstructured files, and at the same time, the management cost of unstructured files is reduced.

[0068] As Figure 1 shown, an embodiment of the present invention provides a method for splitting unstructured files, and the specific steps are as follows:

[0069] S101: Obtain a plurality of unstructured files.

[0070] S102: Classify the unstructured files according to the file volume of each unstructured file to obtain a classification result.

[0071] Specifically, the file volume refers to the size of the unstructured file, that is, a measure of the storage space occupied by the unstructured file. Based on the judgment of the file volume of the unstructured file, unstructured files with different file volumes are divided into a classification result including a first file set, a second file set, and a third file set.

[0072] Further, referring to Figure 2 , if the file volume of the unstructured file is less than the first volume threshold, the unstructured file is added to the first file set;

[0073] If the file size of the unstructured file is greater than or equal to the first size threshold and less than the second size threshold, add the unstructured file to the second file set;

[0074] If the file size of the unstructured file is greater than or equal to the second size threshold, add the unstructured file to the third file set.

[0075] In a specific example, the first size threshold is 2GB and the second size threshold is 500GB. In other embodiments, the first size threshold and the second size threshold can be set according to actual partitioning requirements, and are not limited thereto.

[0076] By setting the first size threshold and the second size threshold, the unstructured files are classified and added to different file sets respectively. For the unstructured files in different file sets, corresponding splitting strategies are matched to improve the splitting accuracy of the unstructured files.

[0077] It can be understood that multiple unstructured files may all be smaller than the first size threshold, then the first file set is a non-empty set, and the second file set and the third file set are empty sets. Similarly, there may only be the second file set or the third file set. Therefore, the first file set, the second file set, and the third file set can exist alone or all exist. The unstructured files in the first file set, the second file set, and the third file set can be processed simultaneously, or the unstructured files in a single file set can be processed, and are not limited thereto.

[0078] S103: Based on the classification result, combined with the file types of each unstructured file, match the corresponding splitting strategy, and split the unstructured file to obtain a file splitting result.

[0079] Specifically, referring to Figure 3 , for the unstructured files in the first file set, adopt a natural language processing strategy to split and obtain a file splitting result;

[0080] For the unstructured files in the second file set, adopt a compression splitting strategy to split and obtain a file splitting result;

[0081] If the file type of the unstructured file in the third file set is a non-text file, determine the number of splitting points of the unstructured file, and split the unstructured file based on the number of splitting points of the unstructured file to obtain a file splitting result;

[0082] If the file type of the unstructured files in the third file set is a text file, traverse the unstructured files, split the unstructured files based on the punctuation marks in the unstructured files, combine to obtain multiple sub-fragments, and use the dynamic programming strategy for splitting to obtain the file splitting result.

[0083] Based on the foregoing descriptions of the first file set, the second file set, and the third file set, the first file set, the second file set, and the third file set can exist independently or all exist. Regardless of whether the first file set, the second file set, and the third file set all exist, the processing of the unstructured files in each file set can be carried out simultaneously or sequentially, or the unstructured files in a single file set can be processed, and there is no sequence of steps.

[0084] In the implementation manner provided by the present invention, for the unstructured files in the first file set, a natural language processing strategy is adopted for splitting, that is, natural language processing technology is used for model training to obtain the splitting identifiers of the unstructured files in the first file set, and then the splitting of the unstructured files in the first file set is completed. During the model training process, first, the training texts in the training data set need to be cleaned to remove the noise data in the training texts, such as advertisements, headers and footers, etc., and the encoding format is unified to provide clean data for subsequent processing. Then use a natural language processing library (such as SpaCy, NLTK) to perform sentence splitting on the text and decompose the text into individual sentences. After that, each sentence is converted into an embedding vector so that the model can understand the meaning of the sentence. By calculating the semantic similarity between sentences, it is judged which sentences are semantically related to provide a basis for subsequent chunking. The semantically related sentences are grouped together to obtain text chunks, and according to specific requirements and model limitations, the size of the split chunks is adjusted and controlled to ensure that the size of the chunks is suitable for subsequent processing and analysis. Through the analysis of the text chunks or each sentence, the splitting identifiers of each training text can be obtained. The splitting of the training text is completed through the splitting identifiers. In a specific example, the unstructured files in the first file set can be understood and analyzed through natural language processing (NLP) to obtain the splitting identifiers, and then the splitting identifiers corresponding to the file content of the unstructured files in the first file set are identified and the splitting is completed. For example, the splitting identifiers can be "Thank you for watching", "Best regards", etc., or can be the delimiters of the file content.

[0085] Furthermore, a compression splitting strategy is adopted for splitting to obtain the file splitting result, which specifically includes:

[0086] First, compress the unstructured files in the second file set to obtain compressed files. Then, based on a preset compressed file volume threshold, divide the file volume of the compressed files by the compressed file volume threshold to determine the number of compression segmentation points for the unstructured files. After determining the number of compression segmentation points, divide the compressed files into equal volumes to obtain a file segmentation result. For example, if the volume of the compressed file is 10 GB and the preset compressed file volume threshold is 2 GB, then the number of compression segmentation points is 10 / 2 = 5. The compressed file is then evenly divided into sub-files of 2 GB to obtain a file segmentation result.

[0087] By adopting a compression and segmentation strategy of first compressing and then segmenting unstructured files, the file volume of unstructured files can be significantly reduced, and the storage, transmission, and processing efficiency of unstructured files can be improved. At the same time, segmentation errors and computational resource consumption can be reduced.

[0088] Furthermore, determining the number of segmentation points for unstructured files and segmenting the unstructured files based on the number of segmentation points for unstructured files to obtain a file segmentation result specifically includes:

[0089] First, obtain the file volume of the non-text files in the third file set. Then, based on a preset non-text file volume threshold, divide the file volume of the non-text files by the non-text file volume threshold to determine the number of segmentation points for the unstructured files. After determining the number of segmentation points, divide the non-text files into equal volumes to obtain a file segmentation result. For example, if the volume of the non-text file is 200 GB and the preset non-text file volume threshold is 20 GB, then the number of segmentation points is 200 / 20 = 10. The non-text file is then evenly divided into sub-files of 20 GB to obtain a file segmentation result. In the embodiment provided by the present invention, the segmentation of non-text files can be completed by a media splitter. In other embodiments, for the segmentation of non-text files, the above natural language processing strategy can also be adopted, that is, by analyzing the non-text files to obtain corresponding segmentation identifiers, and based on the segmentation identifiers, completing the segmentation of the non-text files to obtain a file segmentation result.

[0090] Furthermore, referring to Figure 4 , adopt a dynamic programming strategy for segmentation to obtain a file segmentation result, specifically including:

[0091] Split the unstructured file according to the punctuation marks in the unstructured file to obtain multiple sub-fragments;

[0092] Based on the order of the sub-fragments, combine multiple consecutive sub-fragments to form candidate sub-texts, where each candidate sub-text includes n consecutive sub-fragments, the text volume of each candidate sub-text is less than a preset sub-text threshold, and the text volume of the candidate sub-text and the (n + 1)-th sub-fragment is greater than or equal to the preset sub-text threshold, n ∈ N+ ;

[0093] Construct a directed acyclic graph according to the order among candidate sub-texts, where each node in the directed acyclic graph includes at least one candidate sub-text, and the edge in the directed acyclic graph is the cross-degree of two candidate sub-texts;

[0094] Based on the directed acyclic graph, analyze the forward shortest path from the source point to the sink point in the directed acyclic graph, and determine the forward splitting point corresponding to the forward shortest path;

[0095] Based on the directed acyclic graph, analyze the reverse shortest path from the sink point to the source point in the directed acyclic graph, and determine the reverse splitting point corresponding to the reverse shortest path;

[0096] Compare the forward shortest path and the reverse shortest path, determine the target splitting point from the forward splitting point and the reverse splitting point, and obtain the file splitting result.

[0097] In a specific implementation manner, refer to Figure 5 , split the unstructured file according to the punctuation marks in the unstructured file to obtain multiple sub-fragments, number the sub-fragments in order, and the numbers of the sub-fragments are a1, a2, a3, a4, a5, ……, am. Then, based on a preset sub-text threshold, reorganize the sub-fragments to obtain multiple candidate sub-texts, and mark the candidate sub-texts in order as b1, b2, b3, b4, b5, ……, bq. In a specific example, for the candidate sub-text b1, b1 is composed of a1, a2, and a3, and the text volumes of a1, a2, and a3 are less than the preset sub-text threshold, and the text volumes of a1, a2, a3, and a4 are greater than or equal to the preset sub-text threshold, that is, for the candidate sub-text b1, n is 3. There is only this one combination method for the candidate sub-text b1. For other candidate sub-texts, there may be multiple combination methods. Taking the candidate sub-text b2 as an example, the candidate sub-text b2 can be composed of a2, a3, and a4, or can be composed of a3 and a4, or can be composed of a4 and a5. It can be understood that in the above example, the combination methods of the candidate sub-text b2 all meet the limitations on the sub-text threshold. Through the combination methods of the candidate sub-text b2, it can be understood that all other candidate sub-texts may include multiple combination methods. After completing the combination of all candidate sub-texts, construct a directed acyclic graph based on all the candidate sub-texts. Refer to Figure 6, the source point in the directed acyclic graph is the candidate sub-text b1. The source point is the node with an in-degree of zero. The nodes connected to the candidate sub-text b1 are different candidate sub-texts b2. The edge between b1 and b2 is the cross-degree of b1 and b2, that is, the square of the total number of cross-text characters between b1 and b2. For example, if b1 consists of a1, a2, and a3, and b2 consists of a2, a3, and a4, then the cross-text of b1 and b2 is a2 and a3, and the square of the total number of characters of a2 and a3 is used as the cross-degree of b1 and the corresponding b2. Similarly, the construction of other candidate sub-texts is completed, and then the construction of the directed acyclic graph is completed. The sink point in the directed acyclic graph is the candidate sub-text bq, and the sink point is the node with an out-degree of zero.

[0098] Furthermore, the cross-degree is specifically expressed as:

[0099] JCD(i - 1, i) = JCW(i - 1, i) 2

[0100] where JCD(i - 1, i) is the cross-degree between the (i - 1)-th candidate sub-text and the i-th candidate sub-text in the directed acyclic graph, and JCW(i - 1, i) is the total number of characters corresponding to the cross-text between the (i - 1)-th candidate sub-text and the i-th candidate sub-text in the directed acyclic graph.

[0101] Furthermore, determining the forward splitting point specifically includes:

[0102] Based on the source point of the directed acyclic graph, according to the cross-degree between the source point and other nodes, the next node is determined in turn until the sink point of the directed acyclic graph is reached, and the forward shortest path is determined. Among them, each node in the forward shortest path corresponds to a target sub-text;

[0103] Take the end of the cross-text of any two target sub-texts in the forward shortest path as the forward splitting point.

[0104] In a specific implementation manner, each layer in the directed acyclic graph contains multiple nodes, and the nodes in each layer represent various possibilities corresponding to a certain candidate sub-text. By recursively calculating the shortest path between the candidate node u and the candidate node v, the calculation formula is specifically as follows:

[0105] F u,v = F u,v-1 + minJCD(v - 1, v), u < v - 1

[0106] where u and v are the numbers of the candidate sub-texts, F u,v is the shortest path between the target sub-text u and the target sub-text v, F u,v-1 is the shortest path between the target sub-text u and the target sub-text v - 1, and JCD(v - 1, v) is the minimum value of the cross-degree between the target sub-text v - 1 and the candidate sub-text v.

[0107] Similarly, calculate the shortest paths between candidate sub-texts b1 to bq, that is, the forward shortest path, and use the end of the cross-text of any two target sub-texts in the forward shortest path as the forward segmentation points.

[0108] In a specific example, to calculate F b1,bq , it is necessary to calculate F b1,b(q-1) . Based on recursive calculation, first, it is necessary to calculate F b1,b2 . It can be understood that there are multiple possibilities for candidate sub-text b2. According to minJCD(b1, b2), the target sub-text b2 corresponding to the minimum cross-degree can be determined from multiple candidate sub-texts b2, and then the target sub-text b3,..., the target sub-text bq are determined in turn to obtain the forward shortest path. Based on the cross-text situation of adjacent target sub-texts in the forward shortest path, mark the end of each cross-text as the forward segmentation point.

[0109] Furthermore, determine the reverse segmentation points, specifically including:

[0110] Based on the sink point of the directed acyclic graph, according to the cross-degree between the sink point and other nodes, determine the next node in turn until reaching the source point of the directed acyclic graph to determine the reverse shortest path, where each node in the reverse shortest path corresponds to a target sub-text;

[0111] Use the end of the cross-text of any two target sub-texts in the reverse shortest path as the reverse segmentation point.

[0112] In a specific implementation, construct a virtual node N after the node corresponding to candidate sub-text bq, calculate the distance between virtual node N and the last candidate sub-text bq, and use the candidate sub-text bq with the shortest distance to virtual node N as the target sub-text. Then calculate the shortest path between virtual node N and candidate sub-text b1 and use it as the reverse shortest path, and at the same time use the end of the cross-text of any two target sub-texts in the reverse shortest path as the reverse segmentation point.

[0113] It can be understood that the methods for determining the reverse shortest path and the forward shortest path are the same, and the difference lies in the starting and ending points of the two candidate text paths.

[0114] In a specific example, to calculate F N,b1 , it is necessary to calculate F N,b2 . Based on recursive calculation, first, it is necessary to calculate F N,bqIt can be understood that there are multiple possibilities for the candidate subtext bq. According to minJCD(N,bq), the target subtext bq corresponding to the minimum intersection degree can be determined from multiple candidate subtexts bq, and the target subtexts bq-1, ..., and target subtext b1 are determined in turn to obtain the reverse shortest path. Based on the cross texts in each adjacent target subtext in the reverse shortest path, the end of each cross text is marked as a reverse segmentation point.

[0115] Further, refer to Figure 7 , determine the file segmentation results, including:

[0116] Compare the sum of the forward cross-degree corresponding to the forward shortest path with the sum of the reverse cross-degree corresponding to the reverse shortest path;

[0117] If the sum of the forward cross-degree is less than the sum of the reverse cross-degree, the forward segmentation point is used as the target segmentation point to segment the unstructured file and obtain the file segmentation result;

[0118] If the sum of the forward cross-degree is greater than the sum of the reverse cross-degree, the reverse segmentation point is used as the target segmentation point to segment the unstructured file and obtain the file segmentation result;

[0119] If the sum of the forward cross-degrees is equal to the sum of the reverse cross-degrees, the target segmentation points are determined based on the number of forward segmentation points and reverse segmentation points, and the unstructured file is segmented to obtain a file segmentation result.

[0120] In a possible implementation, the intersections on the forward shortest path are summed to obtain the forward intersection sum, and the intersections on the reverse shortest path are summed to obtain the reverse intersection sum. The forward intersection sum is compared with the reverse intersection sum. If the forward intersection sum is less than the reverse intersection sum, the forward segmentation point is used as the target segmentation point to segment the unstructured file to obtain the file segmentation result. If the forward intersection sum is greater than the reverse intersection sum, the reverse segmentation point is used as the target segmentation point to segment the unstructured file to obtain the file segmentation result. The smaller the intersection sum is, the fewer the intersection texts of each target subtext in the corresponding shortest path are, and the smaller the degree of repetition of the file segmentation result is. If the forward intersection sum is the same as the reverse intersection sum, the number of forward segmentation points and reverse segmentation points is further compared, and the smaller number is used as the target segmentation point to segment the unstructured file to obtain the file segmentation result. While ensuring the segmentation effect, selecting fewer segmentation points can improve the segmentation efficiency of unstructured text. If the number of forward segmentation points and reverse segmentation points is the same, you can select any one of them as the target segmentation point to complete the segmentation of the unstructured text.

[0121] In a specific example, first create an empty dictionary DG to represent a directed graph. Determine the number of candidate sub-texts, that is, obtain the length N0 of the candidate sub-text list. For the index k of each candidate sub-text, generate directed edges from k to other candidate sub-texts. Specifically, for each k, create an edge from k to all indices between k + 1 and min(alls[k][-1]+1, N0). If there are no indices in this range, only create an edge from k to k + 1. Among them, alls[k][-1] is the index of the last segment of the k-th candidate sub-text, alls[k][-1]+1 is the index of the next possible segment of the k-th candidate sub-text, and min(alls[k][-1]+1, N0) is to select the smaller value from alls[k][-1]+1 and N0 to ensure that alls[k][-1]+1 does not exceed the total number N0 of candidate sub-texts.

[0122] Then, take N as a virtual node to represent the end point. Initialize the path dictionary routes, and set routes[N] to (0, -1), indicating that the path length from the virtual node N to itself is 0 and there is no predecessor node. Start from the candidate sub-text index N - 1 and traverse in reverse order to index 0. Traverse the neighbors of each node: for the current node i, traverse all its neighbor nodes j in the directed graph DG, and calculate the path weight from node i to node j. JCD(i, j) is the cross-degree between the i-th candidate sub-text and the j-th candidate sub-text in the directed acyclic graph. Calculate the total path length and record the shortest path: for each node i, among all possible neighbor nodes j, select the j with the shortest path length, and record this path length and the corresponding predecessor node j into routes[i]. Finally, obtain the dynamic programming table routes. Each key-value pair in the routes dictionary represents the shortest path length from this node to the end point and the corresponding predecessor node. Then, by backtracking the path, starting from the starting node and backtracking according to the predecessor nodes recorded in routes, the shortest path from the starting point to the end point can be obtained.

[0123] The forward splitting point can be obtained through the routes dictionary. By backtracking the path, the reverse splitting point can be obtained. Then, compare the forward splitting point and the reverse splitting point to determine the target splitting point and complete the splitting.

[0124] The method for splitting unstructured files further includes:

[0125] Obtain the total splitting duration of the unstructured file and the file size of the unstructured file;

[0126] Divide the file size of the unstructured file by the total splitting duration to obtain the splitting rate;

[0127] The average segmentation rate is obtained by adding up the segmentation rates of all unstructured texts and taking the mean value.

[0128] Referring to Figure 8 , an embodiment of the present invention provides a device for segmenting unstructured files, including:

[0129] A data acquisition module 201, configured to acquire a plurality of unstructured files;

[0130] A file classification module 202, configured to classify the unstructured files according to the file sizes of the respective unstructured files to obtain a classification result;

[0131] A file segmentation module 203, configured to, based on the classification result and in combination with the file types of the respective unstructured files, match corresponding segmentation strategies to segment the unstructured files to obtain a file segmentation result.

[0132] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0133] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and variations.

Claims

1. A method for splitting unstructured files, characterized in that Including: Obtain multiple unstructured files; Classify the unstructured files according to the file size of each unstructured file to obtain a classification result; Based on the classification result, combine the file types of each unstructured file, match the corresponding splitting strategy, and split the unstructured files to obtain a file splitting result.

2. The method for splitting unstructured documents according to claim 1, wherein The classification result includes at least one of a first file set, a second file set, and a third file set; Classify the unstructured files according to the file size of each unstructured file to obtain a classification result, specifically including: If the file size of the unstructured file is less than the first size threshold, add the unstructured file to the first file set; If the file size of the unstructured file is greater than or equal to the first size threshold and less than the second size threshold, add the unstructured file to the second file set; If the file size of the unstructured file is greater than or equal to the second size threshold, add the unstructured file to the third file set.

3. The method for splitting unstructured documents according to claim 2, characterized in that, Based on the classification result, combine the file types of each unstructured file, match the corresponding splitting strategy, and split the unstructured files to obtain a file splitting result, specifically including: For the unstructured files in the first file set, adopt a natural language processing strategy for splitting to obtain a file splitting result; For the unstructured files in the second file set, adopt a compression splitting strategy for splitting to obtain a file splitting result; If the file type of the unstructured file in the third file set is a text file, traverse the unstructured file, split the unstructured file based on the punctuation marks in the unstructured file, combine to obtain multiple sub-fragments, and adopt a dynamic programming strategy for splitting to obtain a file splitting result; If the file type of the unstructured file in the third file set is a non-text file, determine the number of splitting points of the unstructured file, and split the unstructured file based on the number of splitting points of the unstructured file to obtain a file splitting result.

4. The method for segmenting an unstructured document according to claim 3, wherein Adopt a natural language processing strategy for splitting to obtain a file splitting result, specifically including: Based on a pre-obtained natural language processing model, perform semantic analysis on the unstructured files in the first file set to obtain splitting identifiers corresponding to the unstructured files; According to the splitting identifiers corresponding to the unstructured files, split the unstructured files in the first file set to obtain a file splitting result.

5. The method for splitting an unstructured document according to claim 3, wherein Split the unstructured file based on the punctuation marks in the unstructured file, combine to obtain multiple sub-fragments, and adopt a dynamic programming strategy for splitting to obtain a file splitting result, specifically including: Split the unstructured file according to the punctuation marks in the unstructured file to obtain multiple sub-fragments; Based on the order of sub - segments, combine multiple consecutive sub - segments to form candidate sub - texts. Among them, each candidate sub - text includes n consecutive sub - segments, the text volume of each candidate sub - text is less than a preset sub - text threshold, and the text volume of the candidate sub - text and the (n + 1)th sub - segment is greater than the preset sub - text threshold, where n ∈ N + ; Construct a directed acyclic graph according to the order between candidate sub-texts, where each node in the directed acyclic graph includes at least one candidate sub-text, and the edges in the directed acyclic graph are the cross-degrees of two candidate sub-texts; Based on the directed acyclic graph, analyze the forward shortest path from the source point to the sink point in the directed acyclic graph, and determine the forward splitting point corresponding to the forward shortest path; Based on the directed acyclic graph, analyze the reverse shortest path from the sink point to the source point in the directed acyclic graph, and determine the reverse splitting point corresponding to the reverse shortest path; Compare the forward shortest path and the reverse shortest path, determine the target segmentation point from the forward segmentation point and the reverse segmentation point, and obtain the file segmentation result.

6. The method for splitting an unstructured document according to claim 5, wherein The cross-degree is specifically expressed as: JCD(i-1,i) = JCW(i-1,i) 2 Among them, JCD(i - 1, i) is the cross-degree between the (i - 1)-th candidate sub-text and the i-th candidate sub-text in the directed acyclic graph, and JCW(i - 1, i) is the total number of characters corresponding to the cross-text between the (i - 1)-th candidate sub-text and the i-th candidate sub-text in the directed acyclic graph.

7. The method for splitting an unstructured document according to claim 5, wherein Based on the directed acyclic graph, analyze the forward shortest path from the source point to the sink point in the directed acyclic graph, and determine the forward segmentation point corresponding to the forward shortest path, specifically including: Based on the source point of the directed acyclic graph, according to the cross-degree between the source point and other nodes, successively determine the next node until reaching the sink point of the directed acyclic graph to determine the forward shortest path, where each node in the forward shortest path corresponds to a target sub-text; Take the end of the cross-text of any two target sub-texts in the forward shortest path as the forward segmentation point.

8. The method for splitting an unstructured document according to claim 5, wherein, Based on the directed acyclic graph, analyze the reverse shortest path from the sink point to the source point in the directed acyclic graph, and determine the reverse segmentation point corresponding to the reverse shortest path, specifically including: Based on the sink point of the directed acyclic graph, according to the cross-degree between the sink point and other nodes, successively determine the next node until reaching the source point of the directed acyclic graph to determine the reverse shortest path, where each node in the reverse shortest path corresponds to a target sub-text; Take the end of the cross-text of any two target sub-texts in the reverse shortest path as the reverse segmentation point.

9. The method for splitting unstructured documents according to claim 5, characterized in that Compare the forward shortest path and the reverse shortest path, determine the target segmentation point from the forward segmentation point and the reverse segmentation point, and obtain the file segmentation result, specifically including: Compare the forward cross-degree sum corresponding to the forward shortest path and the reverse cross-degree sum corresponding to the reverse shortest path; If the forward cross-degree sum is less than the reverse cross-degree sum, take the forward segmentation point as the target segmentation point and segment the unstructured file to obtain the file segmentation result; If the forward cross-degree sum is greater than the reverse cross-degree sum, take the reverse segmentation point as the target segmentation point and segment the unstructured file to obtain the file segmentation result; If the forward cross-degree sum is equal to the reverse cross-degree sum, determine the target segmentation point based on the number of forward segmentation points and reverse segmentation points, and segment the unstructured file to obtain the file segmentation result.

10. A splitting device for unstructured files, characterized in that, Adopt the segmentation method of the unstructured file as described in any one of claims 1 - 9, including: A data acquisition module for acquiring multiple unstructured files; A file classification module for classifying the unstructured files according to the file size of each unstructured file to obtain a classification result; A file segmentation module for segmenting the unstructured files based on the classification result, combining the file types of each unstructured file, matching the corresponding segmentation strategy, and obtaining the file segmentation result.

Citation Information

Patent Citations

  • Parallel synchronization method and system for unstructured files

    CN119513057A

  • Medical record text data segmentation method and device, readable storage medium and electronic equipment

    CN112786132A

  • Unstructured data marking method and device, equipment and storage medium

    CN113190680A

  • Classification method suitable for unstructured data

    CN118210917A

  • File sorting device, its method and storage medium

    JP1999328198A