A method, device, medium, product and equipment for segmenting content of letter of credit field

By determining the segmentation symbols in the letter of credit column and combining machine learning technology, the problem of low segmentation accuracy of the content of the letter of credit column in the existing technology is solved, and more accurate clause splitting and analysis are achieved.

CN114140224BActive Publication Date: 2025-05-16CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111464010.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-05-16
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

The prior art has low accuracy when segmenting the content of the letter of credit column, and it is impossible to effectively split multiple clauses in complex clauses.

Method used

By determining the segmentation start and continuity in the specified column of the letter of credit, combined with machine learning and an improved language model, accurate segmentation of the content of the letter of credit column is achieved.

Benefits of technology

It improves the accuracy of segmentation of the content of the letter of credit column, can more effectively split complex terms, and improves the accuracy of letter of credit analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140224B_ABST
    Figure CN114140224B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device, medium, product and equipment for segmenting the content of a letter of credit field. Among them, it can be judged whether the starting position of the first line of the content of the designated field includes a designated segmentation start character. If it is included, the corresponding segmentation continuation character will be searched from the second line, and the sum of the number of occurrences of the segmentation start character and the segmentation continuation character will be determined. If it is determined that a segmentation start character of a single designated symbol is included, and the number of segmentation start characters is multiple, it will be judged whether the sum of the number of segmentation start characters of the single designated symbol and the corresponding segmentation continuation characters is the same as the sum of the number of other segmentation start characters and the corresponding segmentation continuation characters. If they are the same, the segmentation start character with the longest character length and the corresponding segmentation continuation character will be used for segmentation. In this way, segmentation can be achieved according to the content characteristics of the designated field of the letter of credit using the designated segmentation start character and the corresponding segmentation continuation character, thereby improving the accuracy of segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a method, device, medium, product and equipment for segmenting content in a letter of credit field. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims. No description herein is admitted to be prior art by inclusion in this section.

[0003] A letter of credit is a written document issued by a bank to an exporter (seller) at the request of an importer (buyer) to guarantee the payment of the goods. The intelligent review application of bank import and export letters of credit is the empowerment of traditional business by artificial intelligence technology. Through machine learning and artificial intelligence algorithms, business scenarios are abstractly modeled and algorithms are developed to complete the upgrade from manual review to machine review, so as to shorten the review cycle and save labor costs.

[0004] During the machine review process, in order to reduce the difficulty of understanding the letter of credit and improve the accuracy of letter of credit parsing, the columns containing multiple clauses can be segmented to split the multiple clauses into independent short clauses so that each short clause can be parsed separately.

[0005] The content of the letter of credit field can be segmented based on the rule of segmenting with sentence ending characters such as ". / ?" and line break "\n". However, the accuracy of segmenting the content of the letter of credit field based on this segmentation rule is low, and it is impossible to accurately split the multiple clauses included in the content of the complex terms field of the letter of credit into independent short clauses.

[0006] In addition, deep learning models can be used to convert the segmentation problem into a start symbol recognition problem based on sequence labeling. However, the length of the content of the complex terms of the letter of credit is often long, and existing networks such as convolutional neural networks (CNN) and bidirectional long short-term memory networks (BILSTM) often have poor effects on long-distance dependencies. In addition, deep learning models are also limited by the number of training samples and cannot guarantee the accuracy of segmenting the content of the letter of credit field. Summary of the invention

[0007] The embodiments of the present disclosure provide a method, device, medium, product and equipment for segmenting the content of a letter of credit field, which are used to solve the problem of low accuracy in segmenting the content of a letter of credit field.

[0008] In a first aspect, the present disclosure provides a method for segmenting content in a letter of credit field, the method comprising:

[0009] For the designated field of the letter of credit, determine whether the first line of the content includes at least one designated segment start symbol;

[0010] If it is determined that at least one designated segment start symbol is included, for each segment start symbol included at the start position of the first line of the content, starting from the second line of the content, determine the line where the segment continuation symbol corresponding to the segment start symbol appears at the start position, and determine the sum of the number of the segment start symbol and the corresponding segment continuation symbol;

[0011] If it is determined that the at least one designated segment start symbol includes a segment start symbol of a single designated symbol, and the number of the determined at least one designated segment start symbol is at least two, determining whether the sum of the number of the segment start symbol of the single designated symbol and the corresponding segment continuation symbol is the same as the sum of the number of other segment start symbols and the corresponding segment continuation symbols;

[0012] If they are the same, the segment start symbol with the longest character length among each segment start symbol included in the first line of the content is determined, and the designated field is segmented according to the row where the segment continuation symbol corresponding to the segment start symbol appears at the determined starting position.

[0013] Optionally, if it is determined that the sum of the number of segment start symbols and corresponding segment continuation symbols of the single designated symbol is different from the sum of the number of other segment start symbols and corresponding segment continuation symbols, then:

[0014] Determine the segment start symbol with the longest character length among each segment start symbol included in the starting position of the first line of the content, excluding the segment start symbol of the single designated symbol; segment the designated field according to the row of the segment continuation symbol corresponding to the segment start symbol appearing at the determined starting position.

[0015] Optionally, if it is determined that the at least one designated segment start character includes a segment start character of a single designated symbol, and the number of the at least one designated segment start character determined is one, then:

[0016] Determine whether the specified paragraph end symbol appears at the end position of the line where the segment continuation symbol corresponding to the segment start symbol appears at the starting position;

[0017] If it is determined that the designated paragraph end symbol appears, the designated field is segmented according to the row where the segment continuation symbol corresponding to the segment start symbol appears at the determined starting position.

[0018] Optionally, if it is determined that the at least one specified segment start character does not include a segment start character of a single specified symbol, then the segment start character with the longest character length among each segment start character included in the first line of the content is determined, and the specified field is segmented according to the row of the segment continuation character corresponding to the segment start character that appears at the determined starting position.

[0019] Optionally, if it is determined that the start position of the first line of the content does not include at least one specified segment start symbol, or if it is determined that the specified paragraph end symbol does not appear, then:

[0020] Determine the probability that a specified number of characters at the beginning of each line are the starting characters of a segment;

[0021] If the determined probability is greater than a predetermined threshold, the line is determined to be the first line of a segment.

[0022] Optionally, determining the probability that a specified number of characters at the beginning of a line are the starting characters of a segment includes:

[0023] Determine the probability that a line of characters whose starting position is not greater than the specified number is the starting character of a segment;

[0024] According to the determined probability that each character not greater than the specified number is the starting character of a segment, the probability that the specified number of characters at the starting position of the row is the starting character of a segment is determined.

[0025] Optionally, determining the probability that a line of characters whose starting position is not greater than the specified number is used as the starting character of a segment includes:

[0026] Based on the improved language model n-gram, determine the probability of the character at the beginning of each line being the starting character of a segment;

[0027] According to the determined probability, the probability of a row of characters whose starting positions are not greater than the specified number being the starting characters of a segment is determined.

[0028] In a second aspect, the present disclosure further provides a device for segmenting content in a letter of credit field, the device comprising:

[0029] A segment start character search module is used to determine whether the starting position of the first line of the content includes at least one designated segment start character for a designated field of the letter of credit;

[0030] A segment continuation symbol search module is used for, if it is determined that at least one specified segment start symbol is included, for each segment start symbol included at the start position of the first line of the content, starting from the second line of the content, determining the line where the segment continuation symbol corresponding to the segment start symbol appears at the start position, and determining the sum of the number of the segment start symbol and the corresponding segment continuation symbol;

[0031] A segmentation judgment module is used for determining whether the sum of the number of the segmentation start symbol and the corresponding segmentation continuation symbol of the single specified symbol is the same as the sum of the number of other segmentation start symbols and the corresponding segmentation continuation symbols if it is determined that the at least one specified segmentation start symbol includes a segmentation start symbol of a single specified symbol and the number of the at least one specified segmentation start symbol determined is at least two;

[0032] The segmentation module is used to determine the segmentation start symbol with the longest character length among each segmentation start symbol included in the first line of the content if they are the same, and segment the specified field according to the row of the segmentation continuation symbol corresponding to the segmentation start symbol appearing at the determined starting position.

[0033] Optionally, the segmentation module is also used to determine the segmentation start symbol with the longest character length among each segmentation start symbol included in the starting position of the first line of content, excluding the segmentation start symbol of the single specified symbol, if the sum of the segmentation start symbol and the corresponding segmentation continuation symbol of the single specified symbol is different from the sum of other segmentation start symbols and the corresponding segmentation continuation symbols; and segment the designated field according to the row in which the segment continuation symbol corresponding to the segmentation start symbol appears at the determined starting position.

[0034] Optionally, the segmentation module is also used to determine whether a specified paragraph end symbol appears at the end position of the row in which a segment continuation symbol corresponding to the segment start symbol appears at the starting position, if it is determined that the at least one specified segment start symbol includes a segment start symbol of a single specified symbol, and the number of the at least one specified segment start symbol determined is one; if it is determined that the specified paragraph end symbol appears, segment the specified field according to the row in which the segment continuation symbol corresponding to the segment start symbol appears at the determined starting position.

[0035] Optionally, the segmentation module is also used to determine the segmentation start character with the longest character length among each segmentation start character included in the first line of the content if it is determined that the at least one specified segmentation start character does not include a segmentation start character of a single specified symbol, and segment the specified field according to the row of the segmentation continuation character corresponding to the segmentation start character appearing at the determined starting position.

[0036] Optionally, the segmentation module is also used to: determine the probability of a specified number of characters at the starting position of each line being the starting characters of a segment if it is determined that the starting position of the first line of the content does not include at least one specified segmentation start symbol, or if it is determined that a specified paragraph end symbol does not appear; and if the determined probability is greater than a predetermined threshold, determine that the line is the first line of a segment.

[0037] Optionally, the segmentation module determines the probability of a specified number of characters at the starting position of a line as the starting characters of a segment, including: determining the probability of no more than the specified number of characters at the starting position of a line as the starting characters of a segment; and determining the probability of the specified number of characters at the starting position of the line as the starting characters of a segment based on the determined probability of each character no more than the specified number as the starting characters of a segment.

[0038] Optionally, the segmentation module determines the probability that a line with no more than the specified number of characters at the starting position is the starting character of a segment, including: based on the improved language model n-gram, determining the probability that the characters at the starting position of each line are the starting character of a segment; and determining the probability that a line with no more than the specified number of characters at the starting position is the starting character of a segment based on the determined probability.

[0039] In a third aspect, the present disclosure further provides a computer program product, wherein the computer program product comprises an executable program, and the executable program is executed by a processor to implement the method as described above.

[0040] In a fourth aspect, the present disclosure further provides a non-volatile computer storage medium, wherein the computer storage medium stores an executable program, and the executable program is executed by a processor to implement the method as described above.

[0041] In a fifth aspect, the present disclosure further provides a device for segmenting content in a letter of credit field, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;

[0042] The memory is used to store computer programs;

[0043] The processor is used to implement the above-mentioned method steps when executing the program stored in the memory.

[0044] According to the scheme provided by the embodiment of the present disclosure, it is possible to first determine whether the starting position of the first line of the content of the designated field includes the designated segment start character. If it is included, the corresponding segment continuation character will be further searched from the second line for each segment start character, and the sum of the number of segment start characters and segment continuation characters that appear will be determined. If it is determined that the segment start character includes a segment start character of a single designated symbol, and the number of segment start characters is multiple, it will be further determined whether the sum of the number of segment start characters of a single designated symbol and the corresponding segment continuation characters is the same as the sum of the number of other segment start characters and the corresponding segment continuation characters. If they are the same, the segment start character with the longest character length and the corresponding segment continuation character will be used for segmentation. In this way, the designated field content can be segmented according to the content characteristics of the designated field of the letter of credit using the designated segment start character and the corresponding segment continuation character, thereby improving the accuracy of segmentation.

[0045] Other features and advantages of the present disclosure will be described in the following description, and partly become apparent from the description, or be understood by practicing the present disclosure. The purpose and other advantages of the present disclosure can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0047] Figure 1 A flow chart of a method for segmenting the content of a letter of credit field provided in an embodiment of the present disclosure;

[0048] Figure 2 A schematic diagram showing the comparison of the f-value distribution of the positive and negative example corpora and the effect of the threshold δ on the false positive rate and the missed positive rate provided in the embodiments of the present disclosure;

[0049] Figure 3 A schematic diagram of the structure of a device for segmenting content in a letter of credit field provided in an embodiment of the present disclosure;

[0050] Figure 4 A schematic diagram of the structure of a device for segmenting content in a letter of credit field provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.

[0052] It should be noted that the "multiple or several" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0053] The terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in sequences other than those illustrated or described herein.

[0054] In addition, the terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.

[0055] The acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.

[0056] The applicant in this case found through research that due to the requirements of the issuance specifications of letters of credit on the number of characters to be displayed on a single line, line breaks are common in the column content, and since abbreviations are relatively common, there may also be periods and question marks at non-sentence ends. Directly using sentence end symbols to segment the column content will result in low segmentation accuracy, and it is impossible to accurately split the multiple clauses included in the complex terms column content of the letter of credit into independent short clauses.

[0057] The applicant in this case further discovered through data analysis that in the columns of complex terms in letters of credit, the starting position of each short clause generally contains a relatively clear and continuous identifier. If characters are classified into three categories: numbers, letters, and symbols (characters that are neither letters nor numbers), the continuous identifier corresponding to the starting position of each short clause in the letter of credit can be classified, but not limited to, as non-single designated symbols and single designated symbols. For the convenience of description and understanding, non-single designated symbols can be further divided into numbers or number symbol strings, letters or letter symbol strings. The identifier sequence formed by continuous identifiers can be, but is not limited to, as shown in Table 1 (it can be understood that the same single designated symbol is also continuous). For the columns of complex terms in letters of credit, the probability that the starting position of each short clause contains a relatively clear and continuous identifier is relatively high. It is based on the above data analysis research that the applicant in this case proposed that relevant rules can be set to segment the column content.

[0058] Table 1

[0059]

[0060] Furthermore, in view of the fact that the starting position of each short clause does not contain a relatively clear and continuous identifier, the present application proposes that segmentation can be achieved through a machine learning method based on the characteristics of the characters at the starting position of each short clause. Thus, segmentation can be achieved by using rules and machine learning together to solve the problem of incomplete coverage of segmentation using rules. Among them, segmentation can be achieved based on an improved n-gram model, achieving lightweight development, and the model is easy to understand and maintain.

[0061] Based on the above description, the present disclosure provides a method for segmenting the content of a letter of credit field. The steps of the method can be as follows: Figure 1 As shown, including:

[0062] Step 101: for a designated field of the letter of credit, determine whether the start position of the first line of the content includes at least one designated segment start character.

[0063] In this step, for a certain designated field in the letter of credit, based on the established identifier sequence library, it can be determined whether the starting position of the first line of the field content includes at least one designated segment start character. The starting position of a line can be understood as a range of characters starting from the first character and not more than a set number of characters backward.

[0064] The segment starter can be understood as the first identifier in the identifier sequence. The identifier sequence can be understood as being formed by the continuous identifiers included at the beginning of each short clause. The identifier sequence stored in the identifier sequence library can be understood as being obtained based on the analysis of a large number of letter of credit field contents.

[0065] For example, the segment start symbol can be understood as, but not limited to, "1", "1)", "1 / ", "01-", "+", "a.", "+1)" and "PART A" etc. as shown in Table 1.

[0066] The complex terms of the letter of credit may include, but are not limited to, the document requirement terms field 46A, the other requirement terms field 47A, and the guidance field 78 on related banks. Each of the above fields contains multiple requirement terms, and the expression and required content of each requirement term are independent of each other. Therefore, this embodiment may be understood, but is not limited to, as segmenting the content of any of the above fields into separate requirement terms. In other words, the designated field mentioned in this step may be understood, but is not limited to, as the document requirement terms field 46A, the other requirement terms field 47A, or the guidance field 78 on related banks.

[0067] Taking the first line of the designated column content as "+1ORIGINAL AND 3COPIES COMMERCIAL INVOICE..." as an example, based on the established identifier sequence library, the starting position of the first line of the column content can be determined, including three designated segment start characters, namely "+", "1" and "+1".

[0068] If it is determined that the starting position of the first line of the content includes at least one designated segment start character, step 102 may be continued to be executed; otherwise, step 108 may be skipped to be executed.

[0069] Step 102: for each segment start symbol included in the starting position of the first line of the content, starting from the second line of the content, determine the line where the segment continuation symbol corresponding to the segment start symbol appears at the starting position, and determine the sum of the number of the segment start symbol and the corresponding segment continuation symbol.

[0070] In this step, the corresponding segment continuation characters can be matched in a continuous state for each segment start character starting from the second line. If the corresponding segment continuation character is matched, the line number where the corresponding segment continuation character appears can be recorded, and the next corresponding segment continuation character can be matched in a continuous state for the next line. Otherwise, if the corresponding segment continuation character is not matched, the segment continuation character can be re-matched for the next line.

[0071] The segment continuation symbol can be understood as a non-first identifier in the identifier sequence. For example, when the segment start symbol is "1" as shown in Table 1, the corresponding segment continuation symbols are "2", "3", "4", etc. as shown in Table 1. For another example, when the segment start symbol is "a." as shown in Table 1, the corresponding segment continuation symbols are "b.", "c.", "d.", etc. as shown in Table 1.

[0072] Continuing with the example given in step 101, for "+", the corresponding segment continuation character "+" can be matched starting from the second line. If matched, the line number is recorded, and the next segment continuation character "+" is matched for the third line. If matched, the next segment continuation character "+" is matched for the fourth line, and so on, until each line in the field content is matched.

[0073] For "1", you can start matching the corresponding segment continuation character "2" from the second line. If not matched, continue matching the segment continuation character "2" on the third line. If matched, record the line number and continue matching the next segment continuation character "3" on the fourth line, ... until every line in the field content is matched.

[0074] For "+1", you can start matching the corresponding segment continuation character "+2" from the second line. If not matched, continue matching the segment continuation character "+2" on the third line. If matched, record the line number and continue matching the next segment continuation character "+3" on the fourth line, and so on, until every line in the field content is matched.

[0075] Assume that the column content includes six lines in total, and further assume that for "+", the recorded line number information is [('+',1),('+',2),('+',3),('+',4),('+',5),('+',6)], it can be understood that each line of the column starts with "+".

[0076] Assume that for "1", the row number information recorded is [('1',1),('2',3),('3',4),('4',6)], which means that the first row of the column starts with "1", the third row starts with "2", the fourth row starts with "3", and the sixth row starts with "4".

[0077] Assume that for "+1", the recorded row number information is [('+1',1),('+2',3),('+3',4),('+4',6)], which means that the first row of the column starts with "+1", the third row starts with "+2", the fourth row starts with "+3", and the sixth row starts with "+4".

[0078] Step 103: If it is determined that the at least one designated segment start character includes a segment start character of a single designated symbol, and the number of the at least one designated segment start character determined is at least two, then determine whether the sum of the segment start characters of the single designated symbol and the number of corresponding segment continuation characters is the same as the sum of other segment start characters and the number of corresponding segment continuation characters.

[0079] In this step, it can be further determined whether the segment start character includes a segment start character of a single designated symbol, and the number of types of the segment start characters can be determined.

[0080] If it is determined that the segment start code includes a segment start code of a single specified symbol, and the number of types of the segment start code is at least two, then it can be further determined whether the sum of the number of segment start codes of the single specified symbol and the corresponding segment continuation codes is the same as the sum of the number of other segment start codes and the corresponding segment continuation codes.

[0081] Continuing with the example given in step 102, if it is determined that the segment start character includes a segment start character "+" of a single specified symbol, and the number of types of segment start characters is three, it can be further determined whether the sum of the number of segment start characters "+" and the corresponding segment continuation characters "+" (6) is the same as the sum of the number of segment start characters "1" and the corresponding segment continuation characters "2", "3", and "4" (4), and the sum of the number of segment start characters "+1" and the corresponding segment continuation characters "+2", "+3", and "+4" (4).

[0082] If they are the same, then step 104 may be continued to be executed, and the present process ends after step 104 ; otherwise, step 105 may be jumped to be executed, and the present process ends after step 105 .

[0083] If it is determined that the segment start character includes a segment start character of a single designated symbol, and the segment start character has one type, the process may jump to step 106 .

[0084] If it is determined that the segment start character does not include a segment start character of a single designated symbol, step 104 may be continued to be executed, and the present process ends after step 104 .

[0085] Step 104: Determine the segment start symbol with the longest character length among each segment start symbol included in the first line of the content, and segment the designated field according to the row where the segment continuation symbol corresponding to the segment start symbol appears at the determined starting position.

[0086] That is, in this step, segmentation can be performed according to the identifier sequence with the longest character length to ensure the accuracy of the segmentation as much as possible.

[0087] If the sum of the number of segment start characters and the corresponding segment continuation characters of a single specified symbol is determined to be the same as the sum of the number of other segment start characters and the corresponding segment continuation characters, continuing with the example provided in step 103, assuming that the sum of the number of segment start characters "+" and the corresponding segment continuation characters "+" is also 4, then in this step, the segment start character "+1" with the longest character length among "+", "1" and "+1" can be determined, and the designated column is segmented according to the rows in which the segment continuation characters "+2", "+3" and "+4" appear at the determined starting positions, and rows 1 to 2 are divided into the first segment, row 3 is divided into the second segment, rows 4 to 5 are divided into the third segment, and row 6 is divided into the fourth segment. It can be understood that the segmentation starts from rows 1, 3, 4 and 6 respectively.

[0088] If the segment start character does not include a single designated symbol, continue with the example provided in step 103, assuming that the segment start character only includes "+1" and "1", then in this step, the segment start character "+1" with the longest character length between "1" and "+1" can be determined, and the designated column is segmented according to the rows where the segment continuation characters "+2", "+3", and "+4" appear in the determined starting position, and the 1st to 2nd rows are divided into the first segment, the 3rd row is divided into the second segment, the 4th to 5th rows are divided into the third segment, and the 6th row is divided into the fourth segment. It can be understood that the segmentation starts from the 1st, 3rd, 4th, and 6th rows respectively.

[0089] Step 105, determine the segmentation start symbol with the longest character length among each segmentation start symbol included in the starting position of the first line of the content, excluding the segmentation start symbol of the single designated symbol; and segment the designated field according to the row in which the segmentation continuation symbol corresponding to the segmentation start symbol appears at the determined starting position.

[0090] That is, in this step, the identifier sequence including a single designated character can be excluded, and the identifier sequence with the longest character length can be segmented to ensure the accuracy of the segmentation as much as possible.

[0091] Continuing with the example provided in step 103, in this step, the segment start character "1" that does not include "+" among "+", "1" and "+1" can be determined. At this time, "1" is the segment start character with the longest character length among the segment start characters that do not include the single specified symbol "+".

[0092] At this time, the specified field can be segmented according to the rows where the segment continuation characters "2", "3", and "4" appear in the determined starting positions, with rows 1 to 2 being divided into the first segment, row 3 being divided into the second segment, rows 4 to 5 being divided into the third segment, and row 6 being divided into the fourth segment.

[0093] Step 106: Determine whether the specified paragraph end symbol appears at the end position of the line where the segment continuation symbol corresponding to the segment start symbol appears at the start position.

[0094] In this step, the paragraph end mark can be further combined with the segment start mark and segment continuation mark of a single designated symbol to achieve segmentation, so as to avoid incorrect segmentation and further improve the accuracy of segmentation. The paragraph end mark can be, but is not limited to, a sentence end mark such as ".". The end position of a line can be understood as a range of characters not greater than a specified number from the last character forward.

[0095] Continuing with the example given in step 102, if it is determined that the segment start character only includes the segment start character of a single specified symbol "+", and for "+", the recorded line number information is [('+',1),('+',2),('+',3),('+',4),('+',5),('+',6)], and each line starts with "+", then in this step, the end position of each line can be further determined to see whether the specified paragraph end character "." appears.

[0096] If so, step 107 may be continued to be executed, and the process ends after step 107 is executed; otherwise, the process jumps to step 108.

[0097] Step 107: Segment the designated column according to the row where the segment continuation symbol corresponding to the segment start symbol appears at the determined starting position.

[0098] That is, in this step, segmentation can be performed based on the identifier sequence of a single specified character.

[0099] Still following the example given in step 106, in this step, the first line can be divided into sections, which can be understood as sections starting from the first, second, third, fourth, fifth, and sixth lines respectively.

[0100] It should be further pointed out that after segmentation according to the identifier sequence, in a possible implementation, the identifier sequence included in the content of the letter of credit field can also be removed, thereby removing the influence of the identifier sequence in the subsequent content parsing process and further improving the accuracy of the parsing.

[0101] Step 108: Determine the probability that a specified number of characters at the beginning of each line are the starting characters of a segment.

[0102] If it is determined in step 101 that the starting position of the first line of content does not include at least one specified segment start character, or if it is determined in step 106 that the specified paragraph end character does not appear, in this step, the probability of a specified number of characters at the starting position of each line being the starting characters of a segment can be determined by machine learning.

[0103] In a possible implementation, determining the probability that a specified number of characters at the start position of a line are the start characters of a segment may include:

[0104] Determine the probability that a line starting with no more than the specified number of characters is the starting character of a segment; and determine the probability that a specified number of characters at the starting position of the line is the starting character of a segment based on the determined probability that each character not more than the specified number is the starting character of a segment.

[0105] Further, in a possible implementation, determining the probability that a line of characters whose starting positions are not greater than the specified number are the starting characters of a segment may include:

[0106] Based on the improved language model n-gram, the probability of the character at the beginning of each line being the beginning character of a segment is determined; based on the determined probability, the probability of the characters at the beginning of a line not greater than the specified number being the beginning of a segment is determined.

[0107] For example, assuming that it is necessary to determine the probability of L (4) characters (the 4 characters are assumed to be denoted as w1, w2, w3, w4 in sequence) at the starting position of a line as the starting characters of a segment (assuming it is denoted as f(w1, w2, w3, w4)), then the probability of w1 being the starting character of a segment can be determined (which can be denoted as p(w1)), the probability of w1 and w2 being the starting characters of a segment can be denoted as p(w1, w2)), the probability of w1, w2, w3 being the starting characters of a segment can be denoted as p(w1, w2, w3)) and the probability of (w1, w2, w3, w4) being the starting characters of a segment can be determined (which can be denoted as p((w1, w2, w3, w4)).

[0108] And the probability f(w1, w2, w3, w4) that the four characters at the beginning of the line are the starting characters of a segment can be determined based on the determined p(w1), p(w1, w2), p(w1, w2, w3) and p((w1, w2, w3, w4). At this time, f(w1, w2, w3, w4) can be expressed as follows but is not limited to:

[0109] f(w1,w2,w3,w4)=α1p(w1)+α2p(w1,w2)+α3p(w1,w2,w3)+α4p(w1,w2,w3,w4)

[0110] Among them, α i is the harmonic coefficient, and

[0111] That is to say, in this embodiment, in order to avoid the problem of insufficient samples leading to the inability to accurately obtain the probability of a specified number of characters at the starting position of a line as the starting characters of a segment, not only the probability of a specified number of characters at the starting position of a line as the starting characters of a segment can be determined, but also the probability of a number of characters less than the specified number at the starting position of a line as the starting characters of a segment can be simultaneously determined, and these two methods are used together to determine the probability of a specified number of characters at the starting position of a line as the starting characters of a segment, so as to improve the accuracy of segmentation judgment as much as possible. At this time, f(w1, w2, w3, w4) can be understood as representing the joint probability of the first 1, ..., L characters appearing at the beginning of a segment. The more often the L characters appear at the beginning of a segment, the greater their output probability, otherwise the smaller their output probability.

[0112] The traditional n-gram model does not limit the occurrence position of characters. In this embodiment, the traditional n-gram model can be improved, and the probability of a line of characters whose starting position is not greater than a specified number as the starting character of a segment can be determined based on the improved n-gram model.

[0113] Assuming that n is 3 in the improved n-gram model, it can be understood as the current character w i Depends only on the first 2 characters w i-1 , w i-2 , and is not related to the previous characters.

[0114] Then, continuing with the previous example, at this time:

[0115] p(w1,w2,...,w k )=

[0116] p(w k |w k-1 , w k-2 )*p(w k-1 |w k-2 , w k-3 )*...*p(w3|w2,w1)*p(w2|w1)p(w1)

[0117] Among them, based on the law of large numbers, it can be estimated according to the frequency:

[0118]

[0119]

[0120]

[0121] Of course, in order to avoid the problem of out of vocab (OOV) in actual operation, the denominator Count k (wi , w i-2 )=0 build power Failure, the following back-off method can be used for estimation, at this time:

[0122]

[0123] It can be understood that p(w i ) represents the character w i The probability of appearing at the beginning of a segment, p(w i |w i-1 ) indicates that the previous character at the start position of a segment is w i-1 Under the condition that the current character is w i The probability, p(w i-1 , w i ) represents the character w i-1 , w i The joint probability of appearing in sequence at the starting position of a segment. The probability mentioned in this embodiment is consistent with the definition of probability in the traditional n-gram model, and only the constraint of the starting position of a segment is added in the calculation.

[0124] Among them, Count k (w1) indicates the expected frequency of character w1 in the first {k, k=1, ..., L} characters of each line. For example, if the character "COMMERCIAL" appears 120 times in the sentence start range defined by the first k characters, then Count k ("COMMERCIAL") = 120, N represents the total number of characters contained in the beginning of the corpus sentence (equal to the number of corpus lines multiplied by k). k (w i-1 , w i-2 ) represents the first {k, k=1, ..., l} characters of each line in the corpus. i-1 , w i-2 The frequency of two characters co-occurring. Count k (w i , w i-1 , w i-2 ) represents the first {k, k=1, ..., l} characters of each line in the corpus. i , w i-1 , w i-2 The frequency of three characters co-occurring. It should be noted that if the content of the letter of credit field is in English, then in a possible implementation, this step can be understood as, but not limited to, determining the probability of a specified number of words at the beginning of each line being the starting characters of a segment.

[0125] As can be seen from the above formula, the improved n-gram model needs to calculate the co-occurrence times of characters when k takes different values. In the traditional n-gram model, Count does not restrict the occurrence position of the word, and the co-occurrence times of characters only need to be calculated once.

[0126] Step 109: If the determined probability is greater than a predetermined threshold, determine that the line is the first line of a segment.

[0127] After determining the probability that a specified number of characters at the beginning of each line are the starting characters of a segment, compare it with a given threshold. If it is greater than the given threshold, the line can be considered to be the first line of a segment. Otherwise, the line can be considered to be a continuation of the previous segment and does not belong to a new segment.

[0128] For example, field 46A of a letter of credit may read:

[0129] :46A:DOCUMENTS REQUIRED

[0130] +SIGNED COMMERCIAL INVOICE IN 01 COPIES

[0131] +FULL SET OF CLEAN ON BOARD OCEAN BILLS OF LADING MAKE OUT TO

[0132] +THE ORDER OF KEB HANA BANK

[0133] +MARKED FREIGHT PREPAID AND NOTIFY APPLICANT

[0134] +PACKING LIST IN 01 COPIES

[0135] +CERTIFICATE OF ORIGIN IN 01 FOLD”

[0136] You can remove the "+" sign and add a new line header. <s>The following content is obtained:

[0137] <s>SIGNED COMMERCIAL INVOICE IN 01 COPIES

[0138] <s>FULL SET OF CLEAN ON BOARD OCEAN BILLS OF LADING MAKE OUT TO

[0139] <s>THE ORDER OF KEB HANA BANK

[0140] <s>MARKED FREIGHT PREPAID AND NOTIFY APPLICANT

[0141] <s>PACKING LIST IN 01 COPIES

[0142] <s>CERTIFICATE OF ORIGIN IN 01 FOLD

[0143] For the first row, we can calculate f(SIGNED, COMMERCIAL, INvOICE, IN)

[0144] =α1p(SIGNED)+α2p(SIGNED,COMMERICIAL)+…+α4p(SIGNED,COMMERCIAL,INVOICE,IN)

[0145] For the second row, calculate f(FULL, SET, OF, CLEAN)

[0146] =α1p(FULL)+α2p(FULL,sET)+…+α4p(FULL,SET,OF,CLEAN)

[0147] For the third row, calculate f(THE, OTHER, OF, KEB)

[0148] =α1p(THE)+α2p(THE, OTHER)+…+α4p(THE, OTHER, OF, KEB)

[0149] For the fourth row, calculate f(MARKED, FREIGHT, PREPAID, AND)

[0150] =α1p(MARKED)+α2p(MARKED,FREIGHT)+…+α4p(MARKED,FREIGHT,PREPAID,PREPAID)

[0151] For the fifth row, calculate f(PACKING, LIST, IN, 01)

[0152] =α1p(PACKING)+α2p(PACKING,LIST)+…+α4p(PACKING,LIST,IN,01)

[0153] For the sixth row, calculate f(CERTIFICATE,OF,ORIGIN,IN)

[0154] =a1p(CERTIFICATE)+α2p(CERTIFICATE,OF)+…+α2p(CERTIFICATE,OF,ORIGIN,IN)

[0155] The above six f probabilities are calculated respectively and compared with the given threshold δ. If f>δ, the line is the beginning of a paragraph; otherwise, if f≤δ, the line is considered to be a continuation of the previous line, not a new paragraph. Among them, a string of continuous characters consisting of at least one character of numbers or symbols can also be understood as a word for processing.

[0156] In general scenarios, the training set is obtained by manual annotation. In the process of machine learning for steps 108 and 109, considering the closed scenario of the complex clause field of the letter of credit, in most cases, the segments can be identified based on the rules according to the relatively clear and continuous identifiers at the beginning of each short clause. Therefore, in this embodiment, the training set corpus can be obtained based on the segments identified by the rules, so that the training set does not need to be manually annotated to obtain the training set.

[0157] For example, the contents of columns 46A, 47A, and 78 can be obtained from the letter of credit; segmentation can be performed according to steps 101 to 107; and positive and negative examples can be obtained according to the segmentation results. The positive example can be the first line of the segment after removing the identifier in the identifier sequence; and the negative example can be the non-first line of the segment after removing the identifier in the identifier sequence.

[0158] Continuing with the previous example, in the process of machine learning for steps 108 and 109, there are L+1 parameters, which are the harmonic coefficients α1, ..., α L and threshold δ. By giving parameters α1,...,α L The value of w1, w2, ..., w can be calculated L By comparing f with a given threshold δ, it can be determined whether the line is the first line of the segment.

[0159] Among them, we can give α based on artificial experience i = 1 / L , or, a certain growth sequence can be given so that the longer the character combination, the corresponding weight α i In addition, a set of {α i , i=1,...,l}, so that the distribution difference of the f-values ​​of the positive and negative example corpora is the largest. In addition, by comparing the distribution of the f-values ​​of the positive and negative example corpora, the format of δ can be selected to reconcile the missed detection rate and the false detection rate. The missed detection rate can be understood as the probability of the occurrence of positive samples in the samples classified as negative examples, and the false detection rate can be understood as the probability of the occurrence of negative samples in the samples classified as positive examples. The comparison of the f-value distribution of the positive and negative example corpora and the schematic diagram of the influence of the threshold δ on the false detection rate and the missed detection rate can be shown as follows Figure 2 As shown. Figure 2 It can be seen that the f-value distribution of positive and negative example corpus (in Figure 2 The intersection of the graph (denoted as positive example distribution and negative example distribution in the figure) and the graph formed by the horizontal axis is divided into two parts by the threshold δ (the vertical line parallel to the vertical axis). The area of ​​the left side of the separated graph represents the size of the missed detection rate, and the area of ​​the right side of the separated graph represents the size of the false detection rate.

[0160] It can be understood that the language of the letter of credit is a closed scene language, which has the characteristics of requiring elements and content in a limited set and relatively fixed sentence expressions. Therefore, the word expression of the paragraph head content is also relatively fixed, mostly expressing the requirements of document name, signature, number of copies, etc., but the expression methods under each requirement are different, and the arrangement is different. Other requirements such as display content and notification person are generally at a relatively late position in the sentence. Therefore, the present invention can adopt a joint solution based on rules and language models, and use rules to solve the scene with clear paragraph head symbols under a high probability to ensure its recognition accuracy. A model based on the language model to determine the probability of occurrence of the starting character of a paragraph is used to identify paragraphs with unclear starting characters, reducing the impact of error accumulation on the overall field resolution accuracy. Among them, for the paragraph head symbol with continuous features, rule matching can be used starting from the first line of the field. If there is a matching paragraph head symbol, the paragraph head symbol is identified based on the rule and segmentation is completed. If the known paragraph head symbol cannot be matched, the probability of occurrence of the paragraph starting character can be determined based on the language model, and the binary classification result of whether the first n characters of each line are the paragraph starting characters is output.

[0161] The present invention proposes a segmentation method combining rules and improved language models to solve the segmentation problem in clause parsing. It not only uses rules to ensure the parsing accuracy and interpretability of most clauses, but also makes model processing for the remaining small part of the difficult segmentation problems, reducing the transmission effect of segmentation errors on subsequent parsing, and the improved language model used has the advantages of strong interpretability, intuitive and easy to understand and maintain. It can be understood that the present invention has the characteristics of high accuracy, low resource consumption, strong interpretability, and intuitive and easy to understand.

[0162] In addition, the present invention applies the language model to the segmentation problem for the first time, and makes improvements and restrictions on the expression requirements of the language model, including requiring the position of characters to be restricted in the probability calculation, requiring only the co-occurrence of the first k characters to be considered, and limiting the co-occurrence probability of characters to the probability of the characters at the beginning of the paragraph. The model training corpus can be generated by rules, and the word co-occurrence probability in the calculation formula f can be directly calculated based on the massive corpus based on the n-gram language model and the law of large numbers, avoiding the resource requirements of manual annotation. The parameters α involved in the model i and δ, which have the characteristics of small number of parameters and easy training, and can be solved by simple optimization algorithms, such as grid search, etc., and by statistically analyzing the f distribution of positive and negative examples, the corresponding false positive rate and missed positive rate can be obtained. It should be pointed out that the model parameter training method mentioned in the present invention refers to a class of methods, including manual setting, grid search, ant colony algorithm, etc., which are all protected by this patent.

[0163] Corresponding to the provided method, the following device is further provided.

[0164] The present disclosure provides a device for segmenting the content of a letter of credit field. The structure of the device can be as follows: Figure 3 As shown, including:

[0165] The segment start character search module 11 is used to determine whether the starting position of the first line of the content includes at least one designated segment start character for the designated field of the letter of credit;

[0166] The segment continuation symbol search module 12 is used for, if it is determined that at least one specified segment start symbol is included, for each segment start symbol included at the start position of the first line of the content, starting from the second line of the content, determining the line where the segment continuation symbol corresponding to the segment start symbol appears at the start position, and determining the sum of the number of the segment start symbol and the corresponding segment continuation symbol;

[0167] The segmentation determination module 13 is configured to determine whether the sum of the number of the segmentation start symbol and the corresponding segmentation continuation symbol of the single specified symbol is the same as the sum of the number of other segmentation start symbols and the corresponding segmentation continuation symbols if it is determined that the at least one specified segmentation start symbol includes a segmentation start symbol of a single specified symbol, and the number of the at least one specified segmentation start symbol determined is at least two;

[0168] The segmentation module 14 is used to determine the segmentation start symbol with the longest character length among each segmentation start symbol included in the first line of the content if they are the same, and segment the designated field according to the row of the segmentation continuation symbol corresponding to the segmentation start symbol appearing at the determined starting position.

[0169] Optionally, the segmentation module 14 is also used to determine the segmentation start symbol with the longest character length among each segmentation start symbol included in the starting position of the first line of the content, excluding the segmentation start symbol of the single specified symbol, if the sum of the segmentation start symbol and the corresponding segmentation continuation symbol of the single specified symbol is different from the sum of other segmentation start symbols and the corresponding segmentation continuation symbols; and segment the designated field according to the row in which the segmentation continuation symbol corresponding to the segmentation start symbol appears at the determined starting position.

[0170] Optionally, the segmentation module 14 is also used to determine whether a specified paragraph end symbol appears at the end position of the row where a segment continuation symbol corresponding to the segment start symbol appears at the starting position, if it is determined that the at least one specified segment start symbol includes a segment start symbol of a single specified symbol, and the number of the at least one specified segment start symbol determined is one; if it is determined that the specified paragraph end symbol appears, the specified field is segmented according to the row where a segment continuation symbol corresponding to the segment start symbol appears at the determined starting position.

[0171] Optionally, if it is determined that the at least one specified segment start character does not include a segment start character of a single specified symbol, the segment start character with the longest character length among each segment start character included in the first line of the content is determined, and the specified field is segmented according to the row of the segment continuation character corresponding to the segment start character appearing at the determined starting position.

[0172] Optionally, the segmentation module 14 is also used to: determine the probability of a specified number of characters at the starting position of each line being the starting characters of a segment if it is determined that the starting position of the first line of the content does not include at least one specified segmentation start symbol, or if it is determined that a specified paragraph end symbol does not appear, then: determine the probability of a specified number of characters at the starting position of each line being the starting characters of a segment; if the determined probability is greater than a predetermined threshold, determine that the line is the first line of a segment.

[0173] Optionally, the segmentation module 14 determines the probability of a specified number of characters at the starting position of a line as the starting characters of a segment, including: determining the probability of no more than the specified number of characters at the starting position of a line as the starting characters of a segment; and determining the probability of the specified number of characters at the starting position of the line as the starting characters of a segment based on the determined probability of each character no more than the specified number as the starting characters of a segment.

[0174] Optionally, the segmentation module 14 determines the probability that a line whose starting position is not greater than the specified number of characters is the starting character of a segment, including: based on the improved language model n-gram, determining the probability that the characters at the starting position of each line are the starting character of a segment; and determining the probability that a line whose starting position is not greater than the specified number of characters is the starting character of a segment based on the determined probability.

[0175] The functions of each functional unit of each device provided in the above embodiments of the present disclosure can be implemented through the steps of the above corresponding methods. Therefore, the possible working processes and beneficial effects of each functional unit in each device provided in the embodiments of the present disclosure are not repeated here.

[0176] Based on the same inventive concept, the embodiments of the present disclosure provide the following devices and media.

[0177] The present disclosure provides a device for segmenting the content of a letter of credit field. The structure of the device can be as follows: Figure 4 As shown, it includes a processor 21, a communication interface 22, a memory 23 and a communication bus 24, wherein the processor 21, the communication interface 22, and the memory 23 communicate with each other through the communication bus 24;

[0178] The memory 23 is used to store computer programs;

[0179] The processor 21 is used to implement the steps described in the above method embodiment of the present disclosure when executing the program stored in the memory.

[0180] Optionally, the processor 21 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), may be one or more integrated circuits for controlling program execution, may be a hardware circuit developed using a field programmable gate array (FPGA), or may be a baseband processor.

[0181] Optionally, the processor 21 may include at least one processing core.

[0182] Optionally, the memory 23 may include a read-only memory (ROM), a random access memory (RAM) and a disk memory. The memory 23 is used to store data required by at least one processor 21 when running. The number of memories 23 may be one or more.

[0183] An embodiment of the present disclosure further provides a non-volatile computer storage medium, wherein the computer storage medium stores an executable program. When the executable program is executed by a processor, the method provided by the above method embodiment of the present disclosure is implemented.

[0184] In a possible implementation process, computer storage media may include: Universal Serial Bus Flash Drive (USB), mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other storage media that can store program code.

[0185] The embodiments of the present disclosure further provide a computer program product, which includes an executable program, and the executable program is executed by a processor to implement the method provided by the above method embodiment of the present disclosure.

[0186] In the embodiments of the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0187] The functional units in the embodiments of the present disclosure may be integrated into one processing unit, or the units may be independent physical modules.

[0188] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device, such as a personal computer, a server, or a network device, or a processor to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: a Universal Serial Bus Flash Drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk, and other media that can store program codes.

[0189] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0190] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, apparatus (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0191] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0192] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0193] Although the preferred embodiments of the present disclosure have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present disclosure.

[0194] Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is also intended to include these modifications and variations.< / s> < / s> < / s> < / s> < / s> < / s> < / s>

Claims

1. A method for segmenting the content of a letter of credit field, characterized in that: The method comprises: For the designated field of the letter of credit, determine whether the first line of the content includes at least one designated segment start symbol; If it is determined that at least one designated segment start symbol is included, for each segment start symbol included at the start position of the first line of the content, starting from the second line of the content, determine the line where the segment continuation symbol corresponding to the segment start symbol appears at the start position, and determine the sum of the number of the segment start symbol and the corresponding segment continuation symbol; If it is determined that the at least one designated segment start character includes a segment start character of a single designated symbol, and the number of the determined at least one designated segment start character is at least two, then determine whether the sum of the number of the segment start characters of the single designated symbol and the corresponding segment continuation characters is the same as the sum of the number of other segment start characters and the corresponding segment continuation characters; wherein the at least one designated segment start character includes a non-single designated symbol and a single designated symbol, wherein the non-single designated symbol includes a number or a number symbol string, a letter or a letter symbol string, and the single designated symbol is "+"; If they are the same, then determine the segment start symbol with the longest character length among each segment start symbol included in the first line of the content, and segment the designated field according to the row where the segment continuation symbol corresponding to the segment start symbol appears at the determined starting position; If it is determined that the starting position of the first line of the content does not include at least one specified segment start symbol, or if it is determined that the specified paragraph end symbol does not appear, then: based on the improved language model n-gram, determine the probability of the characters at the starting position of each line being the starting character of a segment; based on the determined probability, determine the probability that no more than the specified number of characters at the starting position of a line are the starting characters of a segment; based on the determined probability that each character no more than the specified number is the starting character of a segment, determine the probability that the specified number of characters at the starting position of the line are the starting characters of a segment; if the determined probability is greater than a predetermined threshold, determine that the line is the first line of a segment.

2. The method according to claim 1, characterized in that If it is determined that the sum of the number of segment start symbols and corresponding segment continuation symbols of the single designated symbol is different from the sum of the number of other segment start symbols and corresponding segment continuation symbols, then: Determine the segment start symbol with the longest character length among each segment start symbol included in the starting position of the first line of the content, excluding the segment start symbol of the single designated symbol; segment the designated field according to the row of the segment continuation symbol corresponding to the segment start symbol appearing at the determined starting position.

3. The method according to claim 1, characterized in that If it is determined that the at least one designated segment start character includes a segment start character of a single designated symbol, and the number of the at least one designated segment start character determined is one, then: Determine whether the specified paragraph end symbol appears at the end position of the line where the segment continuation symbol corresponding to the segment start symbol appears at the starting position; If it is determined that the designated paragraph end symbol appears, the designated field is segmented according to the row where the segment continuation symbol corresponding to the segment start symbol appears at the determined starting position.

4. The method according to claim 1, characterized in that If it is determined that the at least one designated segment start symbol does not include a segment start symbol of a single designated symbol, then the segment start symbol with the longest character length among each segment start symbol included in the first line of the content is determined, and the designated field is segmented according to the row of the segment continuation symbol corresponding to the segment start symbol appearing at the determined starting position.

5. A device for segmenting content of a letter of credit field, characterized in that: The device comprises: A segment start character search module is used to determine whether the starting position of the first line of the content includes at least one designated segment start character for a designated field of the letter of credit; A segment continuation symbol search module is used for, if it is determined that at least one specified segment start symbol is included, for each segment start symbol included at the start position of the first line of the content, starting from the second line of the content, determining the line where the segment continuation symbol corresponding to the segment start symbol appears at the start position, and determining the sum of the number of the segment start symbol and the corresponding segment continuation symbol; A segmentation judgment module, for determining whether the sum of the number of the segmentation start symbol and the corresponding segmentation continuation symbol of the single designated symbol is the same as the sum of the number of other segmentation start symbols and the corresponding segmentation continuation symbols if it is determined that the at least one designated segmentation start symbol includes a segmentation start symbol of a single designated symbol, and the number of the at least one designated segmentation start symbol determined is at least two; wherein the at least one designated segmentation start symbol includes a non-single designated symbol and a single designated symbol, wherein the non-single designated symbol includes a number or a number symbol string, a letter or a letter symbol string, and the single designated symbol is "+"; A segmentation module, for determining the segmentation start symbol with the longest character length among each segmentation start symbol included in the first line of the content, if they are the same, and segmenting the designated column according to the row of the segmentation continuation symbol corresponding to the segmentation start symbol appearing at the determined starting position; The segmentation module is also used for, if it is determined that the starting position of the first line of content does not include at least one specified segmentation start symbol, or if it is determined that the specified paragraph end symbol does not appear, then: based on the improved language model n-gram, determine the probability that the characters at the starting position of each line are the starting characters of a segment; based on the determined probability, determine the probability that no more than the specified number of characters at the starting position of a line are the starting characters of a segment; based on the determined probability that each character no more than the specified number is the starting character of a segment, determine the probability that the specified number of characters at the starting position of the line are the starting characters of a segment; if the determined probability is greater than a predetermined threshold, determine that the line is the first line of a segment.

6. The device according to claim 5, characterized in that The segmentation module is also used to determine the segmentation start symbol with the longest character length among each segmentation start symbol included in the starting position of the first line of the content, excluding the segmentation start symbol of the single specified symbol, if the sum of the number of segmentation start symbols and the corresponding segmentation continuation symbols of the single specified symbol is different from the sum of the number of other segmentation start symbols and the corresponding segmentation continuation symbols; and segment the specified field according to the row in which the segmentation continuation symbol corresponding to the segmentation start symbol appears at the determined starting position.

7. The device according to claim 5, characterized in that The segmentation module is also used to determine whether a specified paragraph end symbol appears at the end position of the row where a segment continuation symbol corresponding to the segment start symbol appears at the starting position, if it is determined that the at least one specified segment start symbol includes a segment start symbol of a single specified symbol, and the number of the at least one specified segment start symbol determined is one; if it is determined that the specified paragraph end symbol appears, the specified field is segmented according to the row where a segment continuation symbol corresponding to the segment start symbol appears at the determined starting position.

8. The device according to claim 5, characterized in that The segmentation module is also used to determine the segmentation start character with the longest character length among each segmentation start character included in the first line of the content if it is determined that the at least one specified segmentation start character does not include a segmentation start character of a single specified symbol, and segment the specified field according to the row of the segmentation continuation character corresponding to the segmentation start character appearing at the determined starting position.

9. A non-volatile computer storage medium, characterized in that: The computer storage medium stores an executable program, and the executable program is executed by a processor to implement any one of the methods of claims 1 to 4.

10. A computer program product, characterized in that The computer program product comprises an executable program, and the executable program is executed by a processor to implement the method according to any one of claims 1 to 4.

11. A device for segmenting content of a letter of credit field, characterized in that: The device comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is used to implement the method steps described in any one of claims 1 to 4 when executing the program stored in the memory.

Citation Information

Patent Citations

  • Method and system of segmentedly modifying validity periods of various phases of electronic certificates

    CN106164955A

  • Scanning character segmentation method and device, computer equipment and storage medium

    CN110135429A