Document typesetting method and system based on text matching
Through the official document layout method based on text matching, the keywords in the official document text are analyzed, the appropriate text template is selected, and the automatic matching and layout is carried out, the problem of inefficient official document layout in the existing technology is solved, and efficient and accurate official document processing is achieved.
Patent Information
- Application Number
- CN202510003680.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
There may be multiple formats and standards for official documents in the prior art. The existing typesetting tools lack support for the special needs of government official documents and can only rely on manual typesetting inefficiency, resulting in the inability to accurately match information, affecting work efficiency and accuracy.
The official document layout method based on text matching is adopted to analyze the samples to be processed, determine the text type of the sample to be processed based on the keywords in the parsed content, select the corresponding text template, and match the obtained text template with the parsed content to type, and generate the target official document text.
Through the automated analysis and matching process, text type identification and template matching of the to-process samples is quickly and accurately, reducing the need for manual intervention, greatly improving the efficiency of official document processing, and improving the accuracy of official document layout.
Smart Images

Figure CN119940334A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of text analysis and matching, and relates to a document typesetting method and system based on text matching. Background Art
[0002] With the continuous development of information technology, people will face a large number of official documents that need to be processed, and at this stage, the conversion and matching of official documents has increasingly relied on electronic and information means. By building an electronic archive management system, it is possible to quickly enter, query, modify and delete archival information. However, due to the wide range of sources of official document information and the possibility of multiple formats and standards, data inconsistencies are prone to occur during the conversion and matching process. Existing typesetting tools lack support for the special needs of government documents and can only rely on inefficient manual typesetting, which may lead to inaccurate matching of information, thus affecting work efficiency and accuracy. Summary of the invention
[0003] The purpose of the present invention is to solve the problem that official documents in the prior art may have multiple formats and standards, and the existing typesetting tools lack support for the special needs of government documents, and can only rely on manual typesetting, which is inefficient and leads to the inability to accurately match information. A document typesetting method and system based on text matching is provided.
[0004] In order to achieve the above object, the present invention adopts the following technical solutions:
[0005] The document typesetting method based on text matching includes:
[0006] Randomly extract an official document text from the text set to be processed as a sample to be processed;
[0007] Analyze the samples to be processed, and determine the text type of the samples to be processed based on the keywords in the analyzed content;
[0008] When the obtained text type corresponds to the text template in the database, the corresponding text template is called;
[0009] If the acquired text type corresponds to multiple text templates, all text templates are traversed to match the text until the most suitable text template corresponding to the text type is found;
[0010] The acquired text template is matched with the parsed content, and the matched content is typeset to generate the target official document text.
[0011] A further improvement of the present invention is:
[0012] Furthermore, the sample to be processed is parsed, and the text type of the sample to be processed is determined according to the keywords in the parsed content, specifically:
[0013] Identify the samples to be processed and obtain text information in the samples;
[0014] Based on the pre-defined dictionary, the character string in the text is matched with the dictionary summary words to complete the text segmentation;
[0015] Conduct contextual analysis on the segmented text to identify and remove typos in the text;
[0016] Determine whether the text contains the keyword set in the database. If so, determine the text type corresponding to the keyword.
[0017] Furthermore, the segmented text is subjected to context analysis to identify and remove typos in the text. Specifically, the context window is slid at a fixed step size to obtain a number of words in the window, each word in the window is converted into a corresponding word vector based on word embedding technology, and the similarity between the word vector of the current word and the word vector of each word in the correct vocabulary library is calculated based on cosine similarity. If the highest similarity is lower than a certain threshold, the current word is considered to be a suspected typo. According to the similarity calculation result, the word that best matches the suspected typo is selected, and the suspected typo is replaced with the best matching word.
[0018] Furthermore, it is determined whether the text contains the keywords set by the database. If so, the text type corresponding to the keywords is determined. Specifically, after removing typos and modal particles in the text, the acquired text is input into a pre-trained language model for text classification, and regular expressions and semantic parsing methods are used to extract fields. Then, the extracted specific fields are integrated to form a structured data output, and it is determined whether the output data corresponds to the keywords in the database. If so, the relevant text template is called.
[0019] Furthermore, if multiple keywords correspond to multiple text types, the position information and frequency of occurrence of different keywords in the text are counted to obtain the priority of each keyword, and the text type corresponding to the keyword with the highest priority is used as the target text type;
[0020] Count the location information and frequency of occurrence of different keywords in the text, and obtain the priority of each keyword, specifically:
[0021]
[0022] Among them, k is a keyword; t is a text sample; i is the position of the keyword in the text sample; PW(k,t,i) is the position weight of the keyword k at the i-th position in the text sample t; FW(k,t,i) is the frequency weight of the keyword k at the i-th occurrence in the text sample t.
[0023] Furthermore, all text templates are traversed to match the text, specifically: the fields extracted based on the regular expression are put into the text template, the similarity between the current template and the text to be matched is calculated based on the similarity algorithm, and text manuscripts are generated in the text template in turn according to the size of the similarity, so that people can judge whether there is the most suitable text template corresponding to the text type.
[0024] Furthermore, the acquired text template is matched with the parsed content, and the matched content is typeset to generate the target official document text. Specifically, the parsed content is placed in the text template, and the text character adjustment model and the secret level identification model corresponding to the text template are called from the database to adjust the size of the text characters and generate the secret level on the document respectively. The size of the text characters and the secret level correspond to the preset requirements of the text template.
[0025] Official document typesetting system based on text matching, including:
[0026] An extraction module, wherein the extraction module randomly extracts an official document text from the text set to be processed as a sample to be processed;
[0027] A parsing module, wherein the parsing module parses the sample to be processed and determines the text type of the sample to be processed according to keywords in the parsed content;
[0028] A calling module, wherein the calling module calls the corresponding text template based on the one-to-one correspondence between the obtained text type and the text template in the database;
[0029] A traversal module, if the acquired text type corresponds to multiple text templates, the traversal module traverses all the text templates to match the text until the most suitable text template corresponding to the text type is found;
[0030] A matching typesetting module matches the acquired text template with the parsed content, typesets the matched content, and generates a target official document text.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] The present invention parses the sample to be processed, determines the text type of the sample to be processed according to the keywords in the parsed content, selects the corresponding text template according to the text type, matches the obtained text template with the parsed content, and typesets the matched content to generate the target official document text. The present invention can quickly and accurately identify the text type and match the template of the sample to be processed through the automated parsing and matching process, reduces the need for manual intervention, and greatly improves the efficiency of official document processing. At the same time, the keyword recognition technology is used to determine the text type, and the matching is combined with a plurality of preset text templates, which can effectively avoid the text type error or template mismatch problem caused by human misjudgment, thereby improving the accuracy of official document typesetting. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0034] Figure 1 A schematic diagram of the flow of the document typesetting method based on text matching of the present invention;
[0035] Figure 2 It is a structural schematic diagram of the official document typesetting system based on text matching of the present invention. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0037] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0038] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0039] In the description of the embodiments of the present invention, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. indicate an orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed when in use, it is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0040] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", which does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0041] In the description of the embodiments of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0042] The present invention is further described in detail below in conjunction with the accompanying drawings:
[0043] See also Figure 1 The present invention discloses a document typesetting method based on text matching, comprising:
[0044] S101, randomly extracting an official document text from the text set to be processed as a sample to be processed;
[0045] S102, parsing the sample to be processed, and determining the text type of the sample to be processed according to the keywords in the parsed content;
[0046] Identify the samples to be processed and obtain text information in the samples;
[0047] Based on the pre-defined dictionary, the character string in the text is matched with the dictionary summary words to complete the text segmentation;
[0048] Conduct contextual analysis on the segmented text to identify and remove typos in the text;
[0049] Determine whether the text contains the keyword set in the database. If so, determine the text type corresponding to the keyword.
[0050] The method performs context analysis on the segmented text to identify and remove typos in the text, specifically by sliding the context window at a fixed step size to obtain a number of words in the window, converting each word in the window into a corresponding word vector based on word embedding technology, and calculating the similarity between the word vector of the current word and the word vector of each word in the correct vocabulary library based on cosine similarity. If the highest similarity is lower than a certain threshold, the current word is considered to be a suspected typo. According to the similarity calculation result, the word that best matches the suspected typo is selected, and the word suspected to be typo is replaced with the best matching word.
[0051] The method determines whether the keywords set in the database exist in the text. If so, the text type corresponding to the keywords is determined. Specifically, after removing typos and modal particles in the text, the obtained text is input into a pre-trained language model for text classification, and field extraction is performed by combining regular expressions and semantic parsing methods. Then, the extracted specific fields are integrated to form a structured data output, and it is determined whether the output data corresponds to the keywords in the database. If so, the relevant text template is called.
[0052] If there are multiple keywords corresponding to multiple text types, the position information and frequency of occurrence of different keywords in the text are counted to obtain the priority of each keyword, and the text type corresponding to the keyword with the highest priority is taken as the target text type;
[0053] Count the location information and frequency of occurrence of different keywords in the text, and obtain the priority of each keyword, specifically:
[0054]
[0055] Among them, k is a keyword; t is a text sample; i is the position of the keyword in the text sample; PW(k,t,i) is the position weight of the keyword k at the i-th position in the text sample t; FW(k,t,i) is the frequency weight of the keyword k at the i-th occurrence in the text sample t.
[0056] S103, when the obtained text type corresponds to a text template in the database, calling the corresponding text template;
[0057] S104, if the acquired text type corresponds to multiple text templates, traverse all the text templates to match the text until the most suitable text template corresponding to the text type is found;
[0058] The fields extracted based on the regular expression are put into the text template, and the similarity between the current template and the text to be matched is calculated based on the similarity algorithm. According to the size of the similarity, the text manuscripts are generated in the text template in turn, so that people can judge whether there is the most suitable text template corresponding to the text type. S105, the obtained text template is matched with the parsed content, and the matched content is typeset to generate the target official document text.
[0059] S105, matching the acquired text template with the parsed content, and typeset the matched content to generate a target official document text;
[0060] The parsed content is placed in a text template, and the text character adjustment model and the secret level identification model corresponding to the text template are called from the database to adjust the size of the text characters and generate the secret level on the document respectively. The size of the text characters and the secret level correspond to the preset requirements of the text model.
[0061] See also Figure 2 The present invention discloses a document typesetting system based on text matching, comprising:
[0062] An extraction module, wherein the extraction module randomly extracts an official document text from the text set to be processed as a sample to be processed;
[0063] A parsing module, wherein the parsing module parses the sample to be processed and determines the text type of the sample to be processed according to keywords in the parsed content;
[0064] A calling module, wherein the calling module calls the corresponding text template based on the one-to-one correspondence between the obtained text type and the text template in the database;
[0065] A traversal module, if the acquired text type corresponds to multiple text templates, the traversal module traverses all the text templates to match the text until the most suitable text template corresponding to the text type is found;
[0066] A matching typesetting module matches the acquired text template with the parsed content, typesets the matched content, and generates a target official document text.
[0067] Example:
[0068] Official documents include 15 forms, namely resolutions, decisions, orders, bulletins, announcements, circulars, opinions, notifications, circulars, reports, requests for instructions, replies, proposals, letters, and minutes. Each type of official document corresponds to a text template; the text template corresponds to different character sizes, line spacing, character fonts, etc., and contains different levels of confidentiality marks. All 15 types of official documents can be marked as top secret before being published;
[0069] The present invention discloses a document typesetting method based on text matching, which is specifically:
[0070] The obtained text samples are analyzed in context to remove typos in the text and identify keywords in the text, such as one or more of "announcement" and "resolution". If only one keyword is contained, the text type is considered to be "announcement" or "resolution"; if multiple keywords are contained, the position information and frequency of occurrence of different keywords in the text are counted to obtain the priority of each keyword, specifically:
[0071]
[0072] Among them, k is a keyword; t is a text sample; i is the position of the keyword in the text sample; PW(k,t,i) is the position weight of the keyword k at the i-th position in the text sample t; FW(k,t,i) is the frequency weight of the keyword k at the i-th occurrence in the text sample t.
[0073] The position information of the text includes the importance of keywords. The importance of keywords is as follows: the importance of keywords in the title is greater than that of keywords in the text; the importance of keywords at the beginning or end of a paragraph is greater than that of keywords in the middle; the importance of keywords in the introduction, conclusion and other parts has a higher priority.
[0074] When the text type corresponds to multiple text templates, it is necessary to traverse all text templates to match the text, put the fields extracted according to the regular expression into the text template, calculate the similarity between the current template and the text to be matched based on the similarity algorithm, and generate text manuscripts in the text templates according to the size of the similarity, so that people can judge whether there is the most suitable text template corresponding to the text type. The parsed content is placed in the text template. If the announcement template is matched with the announcement content, the text character adjustment model and the secret level identification model corresponding to the text template are called from the database to adjust the size of the announcement content and generate the secret level on the manuscript respectively. For example, the announcement is marked as a secret level. When the typesetting is completed, it is also necessary to detect whether the typesetting meets the national standards, prompt to revise the non-standard content, and finally generate a document that meets the requirements, support multiple formats for export (such as PDF, Word, etc.), and can be sent directly to the printing device or storage system.
[0075] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A document typesetting method based on text matching, characterized in that: include: Randomly extract an official document text from the text set to be processed as a sample to be processed; Analyze the samples to be processed, and determine the text type of the samples to be processed based on the keywords in the analyzed content; When the obtained text type corresponds to the text template in the database, the corresponding text template is called; If the acquired text type corresponds to multiple text templates, all text templates are traversed to match the text until the most suitable text template corresponding to the text type is found; The acquired text template is matched with the parsed content, and the matched content is typeset to generate the target official document text.
2. The document typesetting method based on text matching according to claim 1 is characterized in that: The sample to be processed is parsed, and the text type of the sample to be processed is determined according to the keywords in the parsed content, specifically: Identify the samples to be processed and obtain text information in the samples; Based on the pre-defined dictionary, the character string in the text is matched with the dictionary summary words to complete the text segmentation; Conduct contextual analysis on the segmented text to identify and remove typos in the text; Determine whether the text contains the keyword set in the database. If so, determine the text type corresponding to the keyword.
3. The document typesetting method based on text matching according to claim 2 is characterized in that: The method performs context analysis on the segmented text to identify and remove typos in the text, specifically by sliding the context window at a fixed step size to obtain a number of words in the window, converting each word in the window into a corresponding word vector based on word embedding technology, and calculating the similarity between the word vector of the current word and the word vector of each word in the correct vocabulary library based on cosine similarity. If the highest similarity is lower than a certain threshold, the current word is considered to be a suspected typo. According to the similarity calculation result, the word that best matches the suspected typo is selected, and the word suspected to be typo is replaced with the best matching word.
4. The document typesetting method based on text matching according to claim 3 is characterized in that: The method determines whether the keywords set in the database exist in the text. If so, the text type corresponding to the keywords is determined. Specifically, after removing typos and modal particles in the text, the obtained text is input into a pre-trained language model for text classification, and field extraction is performed by combining regular expressions and semantic parsing methods. Then, the extracted specific fields are integrated to form a structured data output, and it is determined whether the output data corresponds to the keywords in the database. If so, the relevant text template is called.
5. The document typesetting method based on text matching according to claim 4 is characterized in that: If there are multiple keywords corresponding to multiple text types, the position information and frequency of occurrence of different keywords in the text are counted to obtain the priority of each keyword, and the text type corresponding to the keyword with the highest priority is taken as the target text type; Count the location information and frequency of occurrence of different keywords in the text, and obtain the priority of each keyword, specifically: Among them, k is a keyword; t is a text sample; i is the position of the keyword in the text sample; PW(k,t,i) is the position weight of the keyword k at the i-th position in the text sample t; FW(k,t,i) is the frequency weight of the keyword k at the i-th occurrence in the text sample t.
6. The document typesetting method based on text matching according to claim 5 is characterized in that: The traversal of all text templates to match the text is specifically as follows: the fields extracted based on the regular expression are placed into the text template, the similarity between the current template and the text to be matched is calculated based on the similarity algorithm, and text manuscripts are generated in the text template in sequence according to the size of the similarity, so that people can judge whether there is the most suitable text template corresponding to the text type.
7. The document typesetting method based on text matching according to claim 6 is characterized in that: The obtained text template is matched with the parsed content, and the matched content is typeset to generate the target official document text, specifically: the parsed content is placed in the text template, and the text character adjustment model and the secret level identification model corresponding to the text template are called from the database to adjust the size of the text characters and generate the secret level on the document respectively, and the size of the text characters and the secret level correspond to the preset requirements of the text template.
8. The official document typesetting system based on text matching is characterized by: include: An extraction module, wherein the extraction module randomly extracts an official document text from the text set to be processed as a sample to be processed; A parsing module, wherein the parsing module parses the sample to be processed and determines the text type of the sample to be processed according to keywords in the parsed content; A calling module, wherein the calling module calls the corresponding text template based on the one-to-one correspondence between the obtained text type and the text template in the database; A traversal module, if the acquired text type corresponds to multiple text templates, the traversal module traverses all the text templates to match the text until the most suitable text template corresponding to the text type is found; A matching typesetting module matches the acquired text template with the parsed content, typesets the matched content, and generates a target official document text.