Information processing method, apparatus, device, and medium

By classifying and mapping Chinese and foreign patent documents of the same family according to similarity, constructing a parallel Chinese-English corpus, and performing data cleaning and standardization, the problem of scarce parallel patent text corpus resources was solved, and efficient and accurate cross-language patent retrieval support was achieved.

WO2025213738A1PCT designated stage Publication Date: 2025-10-16BEIJING AUGUST MELON TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/124152
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-08
Filing Date
2024-10-11
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

The lack of dedicated machine translation models for patent texts in existing technologies leads to a shortage of resources for building large-scale multilingual parallel corpora of patent texts, affecting the efficiency and accuracy of cross-language patent retrieval.

Method used

By acquiring Chinese and foreign patent documents of the same family, performing similarity classification and calculation, filtering out highly similar document pairs, using dictionary mapping technology to convert them into Chinese-English patent document pairs, and aligning Chinese and English parallel corpora, combined with batch parsing and normalization processing, finely aligned corpus pairs are formed.

Benefits of technology

A structured, multilingual parallel corpus of patent texts was generated, supporting cross-language information retrieval, improving retrieval efficiency and accuracy, and reducing operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024124152_16102025_PF_FP_ABST
    Figure CN2024124152_16102025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an information processing method, an apparatus, a device, and a medium. The method comprises: acquiring a set of first document pairs of Chinese-foreign patent family members; on the basis of a preset grading threshold, performing similarity calculation on document pairs in the set of first document pairs and performing similarity grading, and on the basis of a preset screening threshold, screening for a set of highly similar second document pairs from the set of first document pairs; converting non-Chinese-English patent document pairs in the set of second document pairs into corresponding Chinese-English patent document pairs on the basis of a dictionary mapping technology, then performing Chinese-English parallel corpus alignment processing on each document pair, and constructing memory database data on the basis of the aligned corpus pairs; performing data cleansing on the corpus pairs in the memory database by means of batch parsing; and performing normalization processing and screening on the cleansed corpus pairs to obtain finely aligned corpus pairs. The present application can achieve generation of a structured, reusable, large-scale multilingual parallel corpus of patent texts, supporting application scenarios such as cross-lingual information retrieval, and taking efficiency and accuracy into account.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, device, equipment and medium TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to an information processing method, device, equipment and medium. BACKGROUND

[0002] At present, parallel corpus has become an indispensable resource in the fields of machine translation, cross-language information retrieval, text classification, etc. With the continuous development of artificial intelligence technology, high-quality parallel corpus can improve the accuracy and fluency of machine translation, help build cross-language information retrieval systems, promote multilingual text classification, and promote the maturity and popularity of multilingual natural language processing technology, providing better support for global information exchange and cross-cultural exchange.

[0003] The research on cross-language information retrieval can be traced back to 1973, and has become a hot topic with the rise of the Internet. The cross-language information retrieval method based on machine translation, dictionary and corpus is the mainstream method at present. How to construct high-quality parallel corpus and train better machine learning model is a current hot issue.

[0004] The current effective method is still the cross-language information retrieval method based on machine translation, which is highly dependent on the accuracy of machine translation. Currently, there are also methods for cross-language information retrieval using knowledge graph, keyword mapping, concept mapping, etc., but the effect is not ideal. In different vertical fields, constructing large-scale corresponding parallel corpus to obtain more accurate machine translation results is an important means to solve specific application problems. At present, some researches introduce text matching algorithm into the corpus alignment task, trying to construct parallel corpus semi-automatically or automatically.

[0005] Advanced cross-language information retrieval technology can help various industries and systems improve work efficiency. For example, in the field of intellectual property, in the process of product infringement monitoring, through cross-language information retrieval technology, it can quickly identify patents with high similarity that do not pass the state, and through cross-language information retrieval in the patent database, it can quickly identify patents with similar characteristics, which can effectively reduce the scope of investigation and save resources.

[0006] However, there is still a lack of special machine translation model for patent text, and the resource of large-scale multilingual patent text parallel corpus for constructing related systems is scarce. Therefore, it is urgent to provide a method for efficiently constructing large-scale cross-language patent parallel corpus to support the application scenario of cross-language patent retrieval.

[0007] SUMMARY

[0008] To overcome the problems in the related art, the present disclosure provides an information processing method, device, equipment and medium to solve the technical problem of lack of special machine translation model for patent text in the related art.

[0009] One or more embodiments of the present specification provide an information processing method for multilingual patent text parallel corpus production, the method comprising the following steps:

[0010] Step S1, obtaining a first document pair set of Chinese and foreign same-family patents;

[0011] Step S2, calculating the similarity of the document pairs in the first document pair set according to a preset hierarchical threshold, and performing similarity classification, and screening out highly similar second document pairs in the first document pair set according to a preset screening threshold to form a second document pair set;

[0012] Step S3, converting non-Chinese-English patent document pairs in the second document pair set into corresponding Chinese-English patent document pairs based on dictionary mapping technology, and then performing Chinese-English parallel corpus alignment processing on each document pair in the second document pair set, and constructing a memory bank data based on the aligned corpus pairs; wherein the data in the memory bank contains the aligned sentences or paragraphs, and the corresponding relationship between the source Chinese text and the target English text;

[0013] Step S4, performing data cleaning on the corpus pairs in the memory bank through batch parsing;

[0014] Step S5, performing normalization processing and screening on the cleaned corpus pairs to obtain fine-aligned corpus pairs.

[0015] Further, the step S2 further comprises the following steps:

[0016] Step S21, performing similarity calculation iteration on the second document pair set through a reject method by using a preset screening threshold range and threshold step length, and determining the corresponding screening threshold at the lowest error rate on the ROC curve as the optimal screening threshold;

[0017] Step S22, screening the document pairs in the second document pair set by using the optimal screening threshold, thereby obtaining an optimized second document pair set.

[0018] Further, it further comprises the following steps:

[0019] determining whether the optimal screening threshold is greater than the preset screening threshold, if it is less than, screening the document pairs in the first document pair set that do not meet the preset screening threshold according to the optimal screening threshold, adding the screened document pairs to the second document pair set, and obtaining an optimized second document pair set.

[0020] Further, in the step S2, the similarity of the specification abstract content and the patent name of the document pair in the first document pair set is calculated according to a preset hierarchical threshold, and similarity grading is performed.

[0021] Further, after the non-Chinese-English patent document pair in the second document pair set is converted into a corresponding Chinese-English patent document pair based on the dictionary mapping technology, the method further comprises the steps of:

[0022] Based on the obtained optimal screening threshold, the Chinese-English corpus obtained by conversion is evaluated for similarity with the corresponding source Chinese text, and patent document pairs that do not meet the optimal screening threshold condition are screened out.

[0023] Further, the batch parsing method for data cleaning of the memory bank in the step S4 comprises the following steps:

[0024] Step S41, the first data cleaning of the corpus pair in the memory bank is performed by using the first regular expression, which is used to clean up and includes filtering illegal characters, including deleting Chinese and English square brackets and redundant spaces;

[0025] Step S42, the second data cleaning of the corpus pair in the memory bank is performed by using the second regular expression, which is used to clean up and includes patent specification expression symbols, including deleting Korean separation symbols, redundant semicolons and supplementing missing semicolons, etc., to obtain a rough alignment document pair.

[0026] Further, the normalization processing and screening in the step S5 comprise the following steps:

[0027] According to a preset character number threshold, the character number of each corpus pair in the memory bank is calculated line by line, and if the line character number exceeds the character number threshold, the line that does not meet the standard rule in the corpus pair is deleted or marked at the same time; wherein the character number threshold includes a preset minimum min value and a maximum max value.

[0028] One or more embodiments of the present specification provide an information processing apparatus, comprising:

[0029] A data acquisition module is configured to acquire a first document pair set of Chinese-foreign same-family patents.

[0030] A first screening module is configured to calculate the similarity of the document pairs in the first document pair set according to a preset hierarchical threshold and perform similarity grading, and screen out highly similar second document pairs in the first document pair set according to a preset screening threshold to form a second document pair set.

[0031] The first processing module is configured to convert non-Chinese-English patent document pairs in the second document pair set into corresponding Chinese-English patent document pairs based on a dictionary mapping technology, and perform Chinese-English parallel corpus alignment on each document pair in the second document pair set, and construct a memory database based on the aligned corpus pairs; wherein the data in the memory database includes aligned sentences or paragraphs, and the corresponding relationship between the source Chinese text and the target English text.

[0032] The cleaning module is configured to clean the corpus pairs in the memory database by batch parsing.

[0033] The second processing module is configured to perform normalization processing and screening on the cleaned corpus pairs to obtain fine-aligned corpus pairs.

[0034] One or more embodiments of the present specification provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the multilingual patent text parallel corpus production method according to any one of the above embodiments when executing the computer program.

[0035] One or more embodiments of the present specification provide a computer-readable storage medium, which stores a computer program, wherein the computer program is executable by a processor to implement the steps of the information processing method according to any one of the above embodiments.

[0036] The information processing method, device, equipment and medium provided by the present disclosure have the advantages that: for the obtained Chinese-foreign same-family patent document pairs, similarity grading and calculation are performed, a high threshold filtering method is combined to judge and screen high-similarity same-family patent document pairs, a non-Chinese-English patent document pair is mapped and converted into a Chinese-English patent document pair based on a language dictionary mapping technology, the obtained Chinese-English parallel corpus is aligned, data is cleaned by batch parsing, a coarse-aligned corpus pair is further processed based on a language dictionary inverse mapping technology, and finally, fine-aligned corpus pairs are formed through normalization processing and screening to be presented as results, so as to ensure the standardization and quality of data, quickly and accurately screen data meeting specific requirements, and provide a reliable data basis for applications; the structured and reusable large-scale multilingual patent text parallel corpus library generated by the embodiment method supports cross-language information retrieval and other application scenarios, takes into account efficiency and accuracy, and has low operation and maintenance costs. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to make one or more embodiments of the present specification or the prior art more clearly illustrate the technical solutions, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present specification, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0038] Figure 1 is a flowchart of an information processing method provided by one or more embodiments of the present specification.

[0039] Figure 2 is a block diagram of an information processing device provided by one or more embodiments of the present specification.

[0040] Figure 3 is a structural schematic diagram of a computer device provided by one or more embodiments of the present specification. DETAILED DESCRIPTION

[0041] In order to make one or more embodiments of the present specification or the prior art more clearly illustrate the technical solutions, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present specification, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0042] The present application will be described in detail below in conjunction with the specific embodiments and the drawings of the specification.

[0043] Method embodiment

[0044] According to the embodiment of the present application, an information processing method is provided. As shown in Figure 1, it is a flowchart of the information processing method provided by the present embodiment. According to the information processing method of the present embodiment, the information processing method comprises:

[0045] Step S1, obtaining a first document pair set of Chinese and foreign same family patents.

[0046] Step S2, calculating the similarity of the first document pair according to the preset hierarchical threshold and classifying the similarity, and screening out a second document pair set with high similarity in the first document pair according to the preset screening threshold.

[0047] Step S3, the non-Chinese-English patent document pairs in the second document pair set are converted into corresponding Chinese-English patent document pairs based on a dictionary mapping technology, and then the document pairs in the second document pair set are subjected to Chinese-English parallel corpus alignment processing, and a memory bank data is constructed based on the aligned corpus pairs. The data in the memory bank contains the aligned sentences or paragraphs, and the corresponding relationship between the source Chinese text and the target English text.

[0048] Step S4, data cleaning is performed on the corpus pairs in the memory bank through batch parsing.

[0049] Step S5, the cleaned corpus pairs are subjected to normalization processing and screening to obtain fine alignment corpus pairs.

[0050] The method provided in the embodiment is used to obtain Chinese-foreign same-family patent document pairs, and the obtained Chinese-foreign same-family patent document pairs are subjected to similarity grading and calculation, combined with a preset high threshold filtering method to judge and screen highly similar same-family patent document pairs. The non-Chinese-English patent document pairs are mapped and converted into Chinese-English patent document pairs based on a language dictionary mapping technology, and then the obtained Chinese-English parallel corpus is aligned. The data is cleaned through batch parsing, and the coarse alignment corpus pairs are further processed based on a language dictionary reverse mapping technology. Finally, fine alignment corpus pairs are formed through normalization processing and screening, and are presented as results, so as to ensure the normalization and quality of the data, quickly and accurately screen the data meeting specific requirements, and provide reliable data basis for application. The structured reusable large-scale multilingual patent text parallel corpus generated by the method supports cross-language information retrieval and other application scenarios, and takes into account efficiency and accuracy, and has low operation and maintenance cost.

[0051] In an embodiment, the similarity grading of the first document pairs according to the preset grading threshold in step S2 specifically includes the following grading settings.

[0052] For the Chinese-English same-family patent document pairs with large data volume, the similarity value of the patent document pairs can be determined by a similarity measurement method, such as a minimum edit distance method, a Euclidean distance method, etc. The patent document pairs are graded according to the set four levels, including four levels of “extremely similar”, “relatively similar”, “generally similar” and “not similar”. The specific grading standards are as follows:

[0053] The grading standard of “extremely similar” is that the similarity value is preset as 0.9-1, that is, a few individual subjects or part of nouns are replaced by synonymous words without affecting the semantics, and the overall semantics has no difference in understanding.

[0054] The grading standard of “relatively similar” is that the similarity value is preset as 0.75-0.89, that is, the key words are replaced by synonymous words, and part of the text is missing or added, so that the understanding of the semantics appears ambiguous.

[0055] The similarity value is preset as 0.5-0.74, i.e. a large amount of content is added or lost, but the detection part remains consistent, and the overall understanding and recognition of the text content are affected;

[0056] The similarity value is preset as 0.1-0.49, i.e. the keyword part is inconsistent, leading to different semantic understanding.

[0057] In this embodiment, the high-similarity document pairs are taken as the parallel corpus to be processed according to the above classification results, so that the value of the high-quality parallel corpus pairs can be fully utilized, and the cross-language information retrieval is more accurate and reliable. In this embodiment, the high-similarity second document pair set in the first document pair set that meets the screening threshold and the third document pair set that does not meet the screening threshold are screened out according to the preset screening threshold.

[0058] In this embodiment, in order to avoid the waste of corpus caused by the artificially set screening threshold or the poor reliability of the corpus caused by the unreasonable similarity of the corpus, the accuracy of the high-quality corpus pairs can be fully utilized in subsequent processing and analysis, and more valuable and reliable results can be obtained. In this embodiment, the screening threshold is determined by the following steps, which are as follows:

[0059] In step S21, the similarity of the second document pair set is calculated by the rejection method according to the preset screening threshold range and threshold step length. The screening threshold corresponding to the lowest error rate determined by the ROC curve is the best screening threshold.

[0060] In this embodiment, the best screening threshold is determined by the rejection method. The relationship among the preset screening threshold, the false rejection rate and the false acceptance rate is matched by calculating the false rejection rate and the false acceptance rate, and the ROC curve. When the ROC curve shows a lower error rate, the corresponding screening threshold is the most reasonable.

[0061] In step S22, the document pairs in the second document pair set are screened by the best screening threshold, so as to obtain the optimized second document pair set.

[0062] In the above step, the filtering performance under different thresholds is evaluated, including accuracy, recall rate and other indicators. By comparing the performance under different thresholds, the best screening threshold is selected and determined to ensure that the similarity of the screened corpus pairs reaches the optimal level. This will ensure that the value of these high-quality corpus pairs can be fully utilized in subsequent processing and analysis, and more accurate and reliable results can be obtained.

[0063] In this embodiment, there are still a certain proportion of high-similarity high-quality corpus pairs in the low-similarity same-family patent group obtained by preliminary screening through similarity classification and calculation. In order to fully utilize all data resources, the following steps are further included:

[0064] determining whether the optimal screening threshold is greater than the preset screening threshold, and if not, screening the third document pair set that does not satisfy the preset screening threshold according to the optimal screening threshold, and adding the screened document pair to the second document pair set optimized in step S22.

[0065] In this embodiment, since the Chinese-foreign same-family patent document pairs are obtained in step S1, the similarity of the specification, claims and abstract parts is relatively high, and mainly the patent name and the abstract part have some differences. Therefore, in order to save computer resources, the similarity can be graded according to the specification abstract content and the patent name in step S2.

[0066] In step S3 of this embodiment, the non-Chinese-English patent document pairs in the second document pair set are converted into corresponding Chinese-English patent document pairs based on the dictionary mapping technology, which specifically includes the following steps.

[0067] Step S31, screening the second document pair set to obtain non-Chinese-English patent document pairs;

[0068] Step S32, mapping the foreign language corpus in the document pair row by row and sentence by sentence based on the dictionary mapping technology to obtain the corresponding English corpus, recording the input and output in the mapping process, and saving them in the form of data dictionary, that is, the mapping relationship of the dictionary type.

[0069] In this embodiment, when screening the non-Chinese-English patent document pairs in the highly similar same-family patents, for example, Chinese-Japanese and Chinese-Korean parallel corpora, the foreign language corpus part in such non-Chinese-English patent document pairs is mapped row by row and sentence by sentence to the corresponding English corpus. In order to uniformly record the input and output in the mapping process, these records are saved in the form of data dictionary, which helps to accurately take the input and output of the mapping process as the input of the subsequent Chinese-English parallel corpus alignment process, and also helps to avoid the consistency problem that may occur in the translation process.

[0070] Preferably, step S32 of this embodiment further includes the following steps:

[0071] Step S33, based on the obtained optimal screening threshold, evaluating the similarity of the Chinese-English corpus obtained in step S32 and the corresponding source Chinese text, and screening out the patent document pairs that do not satisfy the optimal screening threshold condition.

[0072] In step S33, the similarity of the Chinese-English corpus and the corresponding source Chinese text can be evaluated by at least two similarity calculation methods in turn, for example, the first similarity evaluation uses the Levenshtein ratio calculation method, and on this basis, the Euclidean distance method or the Jaccard similarity calculation method is used to evaluate the similarity of the Chinese-English corpus and the corresponding source Chinese text to obtain more accurate evaluation results.

[0073] In order to ensure the alignment of the Chinese-English parallel corpus in step S3, the second document pair set file format is preprocessed in this embodiment, including converting the file format, removing special characters or marks, and the like.

[0074] In order to achieve the basic requirements of semi-automatic alignment, the file format is preprocessed to meet the basic requirements of semi-automatic alignment, so as to ensure that the source Chinese text and the target English text can be correctly read and processed by the tool subsequently, and the source Chinese text and the target English text are added to the tool as two kinds of input contents. When adding, it is necessary to ensure that the order of the source text and the target text is one-to-one corresponding, so it is important to standardize the naming in order to facilitate and accurately select the subsequent operation.

[0075] In this embodiment, the Chinese-English parallel corpus alignment processing of each document pair in the second document pair set is performed in step S3, and the memory bank data is constructed based on the aligned corpus pair. The specific implementation steps are as follows.

[0076] In step S301, the source Chinese text and the target English text are added to the ABBYY tool as two kinds of input contents. When adding, it is necessary to ensure that the order of the source text and the target text is one-to-one corresponding, so it is important to standardize the naming in order to facilitate and accurately select the subsequent operation;

[0077] After completing the two types of input, the corpus alignment is performed in a semi-automatic manner. Starting from the top, the source Chinese text and the target English text are aligned one by one, and the corresponding rows are aligned in order. The alignment process can be assisted by the functions and algorithms provided by the tool;

[0078] In step S302, the memory bank data is constructed based on the aligned corpus pair. The memory bank file contains the aligned sentences or paragraphs, and the corresponding relationship between the source Chinese text and the target English text, so as to support the subsequent translation and processing work, and improve the consistency and efficiency of translation.

[0079] In this embodiment, in order to facilitate the fine alignment operation of the corpus in the subsequent steps, the data cleaning batch parsing mode of the memory bank in step S4 includes the following steps:

[0080] In step S41, the corpus pair in the memory bank is cleaned for the first time by using the first regular expression, which is used to clean the illegal characters, including deleting Chinese and English square brackets, unnecessary spaces, and the like. Through appropriate regular expression patterns, these illegal characters can be matched and replaced to remove invalid contents in the data;

[0081] Step S42, the second regular expression is used to clean the corpus in the memory bank for the second time, which is used to clean the symbols of the patent specification, including deleting the Korean escape character, the redundant semicolon and supplementing the missing semicolon, etc., to obtain the coarse alignment document pair; this step also uses appropriate regular expressions to identify and process these specification symbols to ensure the accuracy and consistency of the data;

[0082] After cleaning, the cleaned output file is used as the optimized coarse alignment document pair. The processed and cleaned coarse alignment corpus data is convenient for subsequent further analysis and other processing tasks, and also ensures the specification and accuracy of the data.

[0083] In an embodiment, step S4 further includes further processing of the optimized coarse alignment document pair, specifically including:

[0084] Based on the language dictionary inverse mapping technology, the optimized coarse alignment document pair is processed for alignment to further improve and optimize the accuracy and quality of the alignment.

[0085] In this embodiment, step S5 needs to further standardize the corpus data obtained after the coarse alignment operation in step S4 to obtain fine corpus. The standardization processing and screening include the following steps:

[0086] According to the preset character number threshold, the character number of each corpus pair in the memory bank is calculated row by row, and if the row character number exceeds the character number threshold, the row that does not meet the standard rule in the corpus pair is deleted or marked; wherein the character number threshold includes a preset minimum min value and a maximum max value.

[0087] In a specific embodiment, the minimum min value and the maximum max value are set to 10 characters and 100 characters respectively, and the screening condition is to screen the original Chinese text of each corpus pair in the memory bank by the preset character number threshold. Since Chinese generally has fewer characters than corresponding English, only the original Chinese text needs to be screened in the screening process to realize the standardization processing of the document pair.

[0088] In addition, the standardization processing also includes deleting the entire row with an empty cell in the document pair. It is detected whether there is an empty cell or missing data in each row. For the entire row containing an empty cell, it can be deleted or marked to exclude these incomplete data rows.

[0089] In this embodiment, the data is checked and processed according to the set conditions and rules to ensure the specification and quality of the data, quickly and accurately screen the data meeting the specific requirements, and provide a reliable data basis for application.

[0090] Device embodiment

[0091] According to the embodiment of the present application, an information processing device is provided, as shown in FIG. 1, which is a block diagram of the information processing device provided in the embodiment, and the information processing device according to the embodiment of the present application comprises:

[0092] The data acquisition module 10 is configured to acquire a first document pair set of Chinese and foreign same-family patents.

[0093] The first screening module 20 is configured to perform similarity calculation on the first document pair according to a preset hierarchical threshold, and perform similarity grading, and screen out a second document pair set of highly similar document pairs in the first document pair according to a preset screening threshold.

[0094] The first processing module 30 is configured to convert non-Chinese-English patent document pairs in the second document pair set into corresponding Chinese-English patent document pairs based on a dictionary mapping technology, and then perform Chinese-English parallel corpus alignment processing on each document pair in the second document pair set, and construct a memory database based on the aligned corpus pairs; wherein the data in the memory database contains the aligned sentences or paragraphs, and the corresponding relationship between the source Chinese text and the target English text.

[0095] The cleaning module 40 is configured to clean the corpus pairs in the memory database by batch parsing.

[0096] The second processing module 50 is configured to perform normalization processing and screening on the cleaned corpus pairs to obtain fine-aligned corpus pairs.

[0097] The device provided in the embodiment filters and screens highly similar same-family patent document pairs through the first screening module 20 to perform similarity grading and calculation on the acquired Chinese and foreign same-family patent document pairs, and judges and screens the highly similar same-family patent document pairs by combining a preset high threshold filtering method, the first processing module 30 maps and converts non-Chinese-English patent document pairs into Chinese-English patent document pairs based on a language dictionary mapping technology, and then aligns the acquired Chinese-English parallel corpus, cleans the data by batch parsing of the cleaning module 40, further processes the coarse-aligned corpus pairs based on a language dictionary inverse mapping technology, and finally forms fine-aligned corpus pairs as a result by normalization processing and screening of the second processing module 50, so as to ensure the normalization and quality of the data, quickly and accurately screen out data meeting specific requirements, and provide a reliable data basis for application; the structured reusable large-scale multilingual patent text parallel corpus generated by the method of the embodiment supports cross-language information retrieval and other application scenarios, and takes into account efficiency and accuracy, with low operation and maintenance cost.

[0098] In the embodiment, the first screening module 20 screens out a second document pair set of highly similar document pairs in the first document pair meeting the screening threshold and a third document pair set not meeting the screening threshold according to a preset screening threshold, and the first screening module 20 comprises:

[0099] The screening threshold determination submodule 210 is used to perform similarity calculation on the second document pair set using a rejection method based on a preset screening threshold range and threshold step, and determine the screening threshold corresponding to the lowest error rate according to the ROC curve as the optimal screening threshold.

[0100] The first screening submodule 220 screens the second document pair set using an optimal screening threshold, thereby obtaining an optimized second document pair set.

[0101] In this preferred embodiment, the data acquisition module 10 acquires pairs of Chinese and foreign patent documents of the same family, resulting in relatively high similarity between the specifications, claims, and abstracts, primarily due to some differences between the patent titles and abstracts. Therefore, to conserve computer resources, the first screening module 20 can perform similarity grading based on the content of the specification abstract and the patent title.

[0102] In this embodiment, the first processing module 30 is configured to convert the non-Chinese-English patent document pairs in the second document pair set into corresponding Chinese-English patent document pairs based on the dictionary mapping technology, specifically including the following steps:

[0103] Step A31: Filter the second document pair set to obtain non-Chinese-English patent document pairs;

[0104] Step A32: Based on the dictionary mapping technology, the Chinese and foreign language materials are mapped line by line and sentence by sentence to obtain the corresponding English material, and the input and output of the mapping process are recorded and saved in the form of a data dictionary, that is, a dictionary-type mapping relationship.

[0105] In this embodiment, the first processing module 30 is configured to include a second screening submodule 310 for performing similarity evaluation on the Chinese and English corpora obtained in step A32 and the corresponding source Chinese texts based on the obtained optimal screening threshold, and screening out patent document pairs that do not meet the requirements.

[0106] In this embodiment, the first processing module 30 is configured to execute the following steps to implement Chinese-English parallel corpus alignment processing on each document pair in the second document pair set, and to construct memory database data based on the aligned corpus pairs, as follows.

[0107] Step A301: Add the source Chinese text and the target English text as two input contents to the ABBYY tool at the same time.

[0108] Step S302: construct memory database data based on the aligned corpus pairs. The memory database file contains the aligned sentences or paragraphs, and the correspondence between the source Chinese text and the target English text.

[0109] In this embodiment, the cleaning module 40 includes a first cleaning sub-module 410, a second cleaning sub-module 420, and an alignment sub-module 430, wherein,

[0110] The first cleaning sub-module 410 is configured to perform first data cleaning on the corpus pairs in the memory bank using a first regular expression, including filtering illegal characters, such as deleting Chinese and English brackets, extra spaces, etc. By using appropriate regular expression patterns, these illegal characters can be matched and replaced to remove invalid content in the data.

[0111] The second cleaning sub-module 420 is configured to perform second data cleaning on the corpus pairs in the memory bank using a second regular expression, including cleaning the patent specification expression symbols, such as deleting Korean escape symbols, extra semicolons, and supplementing missing semicolons, etc., to obtain a coarse alignment document pair. This step also uses appropriate regular expressions to identify and process these specification expression symbols to ensure the accuracy and consistency of the data.

[0112] The alignment sub-module 430 is configured to perform alignment processing on the optimized coarse alignment document pair based on a language dictionary inverse mapping technique to further improve and optimize the accuracy and quality of the alignment.

[0113] In this embodiment, the second processing module 50 needs to further standardize the corpus data obtained by the cleaning module 40 after the coarse alignment operation to obtain fine corpus. The standardization processing and screening include the following steps:

[0114] According to a preset character number threshold, the character number of each line of the original Chinese text or the target English text of each corpus pair in the memory bank is calculated, and if the line character number exceeds the character number threshold, the line that does not meet the standard rule in the corpus pair is deleted or marked at the same time. The character number threshold includes a preset minimum min value and a maximum max value.

[0115] In a specific embodiment, the minimum min value and the maximum max value are set to 10 characters and 100 characters respectively, and the screening condition is to screen the original Chinese text of each corpus pair in the memory bank by the preset character number threshold. Since Chinese generally has fewer characters than corresponding English, only the original Chinese text needs to be screened in the screening process to achieve the standardization processing of the document pair.

[0116] In addition, the standardization processing also includes deleting the entire row with an empty cell in the document pair. It is detected whether there is an empty cell or missing data in each line. For the entire line containing an empty cell, it can be deleted or marked to exclude these incomplete data lines.

[0117] The embodiment of the present application is a device corresponding to the above-mentioned method embodiment, and the specific operation of each module processing step can be understood with reference to the description of the method embodiment, which will not be repeated here.

[0118] As shown in FIG. 3, the present application also provides a computer readable storage medium, which stores a computer program, the computer program is executed by a processor to implement the information processing method in the above-mentioned embodiment, or the computer program is executed by a processor to implement the information processing method in the above-mentioned embodiment, and the computer program is executed by the processor to implement the following method steps:

[0119] Step S1, obtaining a first document pair set of Chinese and foreign same family patents.

[0120] Step S2, calculating the similarity of the first document pair according to the preset hierarchical threshold and performing similarity classification, and screening out a second document pair set of highly similar documents in the first document pair according to the preset screening threshold.

[0121] Step S3, converting the non-Chinese-English patent document pair in the second document pair set into a corresponding Chinese-English patent document pair based on a dictionary mapping technology, and then performing Chinese-English parallel corpus alignment processing on each document pair in the second document pair set, and constructing a memory database based on the aligned corpus pairs; wherein the data in the memory database contains the aligned sentences or paragraphs, and the corresponding relationship between the source Chinese text and the target English text.

[0122] Step S4, data cleaning of the corpus pairs in the memory database by batch parsing.

[0123] Step S5, normalizing and screening the cleaned corpus pairs to obtain fine-aligned corpus pairs.

[0124] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in each embodiment provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0125] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the device or system embodiment, since it is basically similar to the method embodiment, it is described more simply, and the relevant part can be referred to the part of the method embodiment. The above-described device and system embodiments are only illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0126] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can still be modified, or some or all of the technical features can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and the contents not described in detail in the specification of the present application are the known technology of those skilled in the art.

Claims

1. An information processing method, characterized in that The following steps are involved: Step S1, obtaining a first document pair set of Chinese and foreign patent families; Step S2: Calculating and grading similarities for the document pairs in the first document pair set according to a preset grading threshold, and screening out highly similar second document pairs in the first document pair set according to a preset screening threshold to form a second document pair set; Step S3: converting the non-Chinese-English patent document pairs in the second document pair set into corresponding Chinese-English patent document pairs based on a dictionary mapping technique, then performing Chinese-English parallel corpus alignment on each document pair in the second document pair set, and constructing memory database data based on the aligned corpus pairs; wherein the data in the memory database includes the aligned sentences or paragraphs, as well as the correspondence between the source Chinese text and the target English text; Step S4: performing data cleaning on the corpus pairs in the memory bank by batch parsing; and Step S5: normalize and filter the cleaned corpus pairs to obtain finely aligned corpus pairs.

2. The information processing method according to claim 1, wherein: The step S2 further comprises the following steps: Step S21: performing similarity calculation on the second document pair set using a rejection method within a preset screening threshold range and threshold step size, and determining, based on the ROC curve, that the screening threshold corresponding to the lowest error rate is the optimal screening threshold; and Step S22: Screen the document pairs in the second document pair set using the optimal screening threshold, thereby obtaining an optimized second document pair set.

3. The information processing method according to claim 2, wherein The following steps are also included: Determine whether the optimal screening threshold is greater than the preset screening threshold. If it is less than the optimal screening threshold, filter the document pairs in the first document pair set that do not meet the preset screening threshold according to the optimal screening threshold, add the filtered document pairs to the second document pair set, and obtain an optimized second document pair set.

4. The information processing method according to claim 1, wherein: In step S2, similarity calculation and similarity grading are performed on the description abstract contents and patent names of the document pairs in the first document pair set according to a preset grading threshold.

5. The information processing method according to claim 1, wherein: After converting the non-Chinese-English patent document pairs in the second document pair set into corresponding Chinese-English patent document pairs based on the dictionary mapping technology, the following steps are also included: Based on the obtained optimal screening threshold, the converted Chinese and English corpora are evaluated for similarity with the corresponding source Chinese texts, and patent document pairs that do not meet the optimal screening threshold conditions are screened out.

6. The information processing method according to claim 1, wherein: The batch parsing method for cleaning the memory bank data in step S4 includes the following steps: Step S41: using a first regular expression to perform a first data cleansing on the corpus in the memory, wherein the cleaning includes filtering out illegal characters, including deleting Chinese and English square brackets and redundant spaces; and Step S42: Use the second regular expression to perform a second data cleansing on the corpus pairs in the memory, including cleaning the patent standard expression symbols, including deleting Korean separators, redundant semicolons, and supplementing missing semicolons, etc., to obtain roughly aligned document pairs.

7. The information processing method according to claim 1, wherein: The normalization processing and screening in step S5 includes the following steps: According to the preset character count threshold, the character count of the original Chinese text or target English text of each corpus pair in the memory bank is calculated line by line. If the number of characters in a line exceeds the character count threshold, the lines in the corpus pair that do not meet the standard rules are deleted or marked at the same time; the character count threshold includes the preset minimum min value and maximum max value.

8. An information processing device, characterized in that include: A data acquisition module, used to acquire the first document pair set of Chinese and foreign patent families; A first screening module is configured to calculate and grade similarities for the document pairs in the first document pair set according to a preset grading threshold, and to screen out highly similar second document pairs in the first document pair set according to the preset screening threshold to form a second document pair set; The first processing module is used to map the non-Chinese-English patent document pairs in the second document pair set based on the dictionary The technology is converted into corresponding Chinese-English patent document pairs, and then each document pair in the second document pair set is aligned with the Chinese-English parallel corpus, and memory database data is constructed based on the aligned corpus pairs; wherein the data in the memory database includes the aligned sentences or paragraphs, as well as the correspondence between the source Chinese text and the target English text; A cleaning module, used to clean the data of the corpus pairs in the memory database through batch parsing; and The second processing module normalizes and filters the cleaned corpus pairs to obtain finely aligned corpus pairs.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the information processing method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the information processing method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Knowledge graph construction method and system based on multi-source heterogeneous data, and terminal

    CN113157930A

  • Full-automatic corpus alignment system and method

    CN114564970A

  • Automatic bilingual parallel corpus extraction method, electronic equipment and storage medium

    CN115937890A

  • Document similarity detection and classification system

    US20050060643A1