Assembly instruction quotation and body comparison method and system
Through the preprocessing of assembly instructions and systematic comparison methods, the problems of omissions in manual comparison of assembly instruction text consistency and limited tool comparison are solved, and automated and comprehensive assembly instruction comparison is achieved, which improves the accuracy and comprehensiveness of the comparison.
Patent Information
- Application Number
- CN202510941958.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-31
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-09
AI Technical Summary
In the existing technology, manual comparison of the consistency of aircraft assembly instructions is prone to omissions, and the use of tools for comparison is limited by file size and ambiguity, resulting in incomplete comparison differences.
By preprocessing the page attributes of the assembly instructions, a comparison object is generated, and character format separation, binarization, tilt correction and text feature extraction are performed. Combined with logical consistency check and formatting comparison, the system module is used for automatic comparison and difference annotation.
It realizes the automated and comprehensive comparison of assembly instructions, reduces the omission of differences caused by manual comparison negligence, and improves the comprehensiveness and accuracy of the comparison.
Smart Images

Figure CN120708232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and system for comparing an assembly instruction citation with a body text. Background Art
[0002] Aircraft assembly instructions are process documents used to define aircraft assembly and testing processes, implement quality control during the assembly process, and record production and inspection status. These documents are presented in PDF format. Due to the large variety and quantity of aircraft parts, completing an aircraft assembly requires a large number of assembly instructions. To optimize the compilation of assembly instructions and improve their practicality and operability, standardization of assembly instructions is particularly important.
[0003] PDF assembly instructions can be divided into three parts: the compilation instruction page, the process content page, and the supporting page. When compiling assembly instructions for aircraft production, it is necessary to ensure that the instruction page and the process content page, as well as the process content page and the supporting page, are consistent. The consistency between the two requires manual verification by the assembly instruction compiler. During manual comparison, the assembly instruction compiler reviews the PDF file page by page, carefully compares the citation with the main text, and records the differences. However, manual negligence often leads to text discrepancies. Using general text comparison tools such as DiffChecker and ComparePDF requires uploading two PDF files for comparison, and cannot achieve internal comparison within a single file. In addition, their use is limited by file size and file ambiguity. Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the present invention provides a method and system for comparing assembly instruction citations and texts, which solves the problems of omissions in manual document comparison and incomplete comparison differences caused by unclear documents.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method and system for comparing assembly instruction citations and texts, comprising:
[0006] According to the different attributes of the assembly instruction page, pre-processing preparation work is performed on the assembly instruction page, the process content page, and the supporting page of the assembly instruction before comparison;
[0007] Generate corresponding first comparison objects, second comparison objects, and third comparison objects for the assembly instruction page, the process content page, and the supporting page of the assembly instruction, respectively; and perform character format separation, binarization, image noise reduction, tilt correction, and text feature extraction on the first comparison objects, the second comparison objects, and the third comparison objects;
[0008] Compare the first comparison object with the second comparison object, and the second comparison object with the third comparison object respectively, use item-by-item comparison and cross-reference verification, and then perform logical consistency check, format comparison and annotation;
[0009] For the differences after comparison of the first, second and third comparison objects, the processed image is compared with the characters in the database again according to the different characteristics of the characters, and the differences in the comparison areas are marked uniformly by the depth of the same color.
[0010] Preferably, the pre-processing preparation work before comparing the compilation description page of the assembly instruction, the process content page of the assembly instruction, and the supporting page of the assembly instruction includes:
[0011] Scan the PDF file with a scanner to ensure it is legible, without any damage or encryption restrictions before comparison.
[0012] Before performing any comparison operation, back up the original PDF file by storing it in the file to prevent it from being tampered with by other programs during the data comparison process.
[0013] Preferably, according to the assembly instruction page attributes, the first comparison object and the second comparison object are generated into a first comparison combination, and the second comparison object and the third comparison object are generated into a second comparison combination, and the first comparison combination and the second comparison combination are respectively subjected to step decomposition, citation comparison and content verification. The steps of performing tilt correction on the first comparison object, the second comparison object and the third comparison object are: edge detection, Hough transform, straight line analysis, calculation of tilt angle and image rotation, and the tilt angle calculation formula is as follows:
[0014]
[0015] where m i is the slope of the line detected by Hough transform, and i is the index of the line.
[0016] Preferably, the step decomposes the document contents of the first comparison object and the third comparison object in the comparison combination of the assembly instruction into several specific comparison elements, the citation comparison is to search for another comparison object in the corresponding comparison combination for each comparison element, and the content verification comparison is to compare the descriptions of the two comparison objects in the comparison combination to check whether there are differences and omissions.
[0017] Preferably, the cross-reference verification includes:
[0018] Create an index and mark each comparison object in the assembly instruction to facilitate quick positioning;
[0019] Reverse search: according to the index mark of the comparison object, reverse search the corresponding content in the original assembly instruction compilation page, process content page, and supporting page files;
[0020] Compare the consistency and verify whether the contents in the assembly instruction preparation instructions page, process content page, and supporting page files are consistent with the description in the assembly instruction to ensure the accuracy of the reference.
[0021] Preferably, the logical consistency check includes:
[0022] Process analysis: analyze the assembly process in the assembly instructions to ensure that the logical relationship between each step is reasonable and coherent;
[0023] Exception investigation: For any abnormal situations that may occur in the process, check whether there are corresponding citations to explain them.
[0024] Preferably, the formatted comparison includes:
[0025] Standardized citations to ensure uniform and standardized citation formats in assembly instructions;
[0026] Version control, focusing on the version information of the technical documents referenced by the first, second, and third comparison objects;
[0027] Preferably, according to the assembly instruction page attributes, the recognized text is compared with its possible similar candidate characters, and logical words are found according to the recognized text before and after, and displayed in the sidebar. Corrections are then made and the recognized and corrected text structure is output to a file in txt, doc, or exl format.
[0028] A system for comparing assembly instruction citations and text includes a processor module, a network interface, a display screen, an input device, and a comparison block connected via a system bus. The processor module is configured to provide computing and control capabilities. The processor module is externally connected to a scanning component. The network interface, display screen, input device, and comparison module are all connected to the processor module. The comparison block includes:
[0029] A memory, the memory being plugged into an output end of the processor module;
[0030] An image-text conversion and recognition module, which is used for conversion between image documents and content recognition;
[0031] A comparison module, which is used for pairwise comparison of documents in the same group;
[0032] An analysis and verification module, which is used for content analysis and anomaly verification;
[0033] Output module, which is used to produce and output comparison results in multiple formats.
[0034] Preferably, the network interface is used to communicate with an external terminal via a network connection, the display screen is used to display data, and the input device is used to input data.
[0035] The present invention discloses a method and system for comparing assembly instruction citations and texts, which has the following beneficial effects:
[0036] By scanning and recognizing the PDF assembly instruction file image into an editable text format, and then dividing the text into the first comparison object, the second comparison object and the third comparison object, and comparing and marking the differences between the first comparison object and the second comparison object, and the second comparison object and the third comparison object respectively, the comparison differences are directly displayed in the comparison document. It is not limited by the file size and file ambiguity, and can correct the compared files, reducing the negligent difference omissions caused by manual document comparison, and improving the comprehensiveness of the difference comparison of the assembly instruction PDF file. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0038] Figure 1 This is a flow chart of a method for comparing assembly instruction citations and texts according to a first embodiment of the present invention;
[0039] Figure 2 This is a schematic diagram of the structure of a system for comparing assembly instruction citations and texts according to a first embodiment of the present invention;
[0040] Figure 3 This is a schematic diagram of the structure of the assembly instruction citation and text comparison system according to the second embodiment of the present invention;
[0041] In the figure: 1. Processor module; 2. Network interface; 3. Display; 4. Input device; 5. Comparison component; 51. Memory; 52. Image-text conversion and recognition module; 53. Comparison module; 54. Analysis and verification module; 55. Output module; 6. Scanning component. DETAILED DESCRIPTION
[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0043] The embodiment of the present application solves the problems of omissions in manual comparison of the consistency of the content of a single document and incomplete comparison differences due to file size and unclear file when uploading the comparison tool by providing a method and system for comparing the citation and text of the assembly instruction, thereby achieving the goal of scanning and identifying the PDF assembly instruction file image as an editable text format, and then dividing the text into a first comparison object, a second comparison object and a third comparison object, and comparing and annotating the differences between the first comparison object and the second comparison object, and between the second comparison object and the third comparison object respectively, directly displaying the comparison differences in the comparison document without being affected by file ambiguity, being able to correct the comparison document, reducing the phenomenon of omissions in manual document comparison due to negligence, and improving the comprehensiveness of the difference comparison of the assembly instruction PDF file.
[0044] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0045] The embodiment of the present invention discloses a method and system for comparing an assembly instruction citation with a main text.
[0046] Example 1
[0047] According to the attached Figure 1 -2, including:
[0048] According to the different attributes of the assembly instruction page, pre-processing preparation work is performed on the assembly instruction compilation instruction page, the assembly instruction process content page, and the assembly instruction supporting page before comparison. The process content page is used as the main text of the assembly instruction, and the compilation instruction page and the supporting page are used as the assembly instruction citation. The pre-processing preparation work before comparison of the assembly instruction compilation instruction page, process content page, and supporting page includes:
[0049] Scan the printed PDF file with a scanner and convert the paper document into editable electronic text, which is convenient for document storage, retrieval, comparison and sharing. Before comparison, ensure that the PDF file is clear and readable, without damage or encryption restrictions, to ensure the quality of the PDF file;
[0050] Before performing any comparison operation, back up the original PDF file by storing it in the file to prevent it from being tampered with by other programs during the data comparison process.
[0051] Generate the corresponding first comparison object, second comparison object, and third comparison object for the assembly instruction compilation description page, assembly instruction process content page, and assembly instruction supporting page respectively. Perform character format separation, binarization, image noise reduction, tilt correction, and text feature extraction on the first comparison object, second comparison object, and third comparison object to improve the accuracy of the image. The character format separation uses Python's split() function to support splitting strings by specifying delimiters. The basic usage of the split() method is as follows:
[0052] str.split(sep=None, maxsplit=-1)
[0053] str is the string to be split.
[0054] sep is an optional parameter that specifies the delimiter. If not specified or None, any whitespace character (such as space, newline \n, tab \t, etc.) will be used as the delimiter. If another string is specified as the delimiter, the method will split the original string at the occurrence of that string.
[0055] maxsplit is an optional parameter that specifies the maximum number of splits. If maxsplit is specified, the split operation will stop after reaching the maximum number of times, even if there are delimiters in the string. If not specified or specified as -1, the split operation will continue to the end, that is, split all possible substrings. For example, in Python
[0056] text="apple, banana, cherry"
[0057] result=text.split(",")
[0058] #result: ['apple', 'banana', 'cherry']
[0059] #Use space as separator (default)
[0060] text="apple banana cherry"
[0061] result = text.split()
[0062] #result: ['apple', 'banana', 'cherry']
[0063] #Specify the maximum number of splits
[0064] text="apple, banana, cherry, date"
[0065] result=text.split(",", maxsplit=1)
[0066] #result: ['apple', 'banana, cherry, date'];
[0067] Binarization is to set the grayscale value of the pixels on the first, second and third comparison objects to 0 or 255, thereby generating an image with only two colors: black and white. In other words, the entire image presents a clear visual effect of only black and white, which greatly simplifies the amount of information in the image while making it unaffected by the file size and improving the processing speed.
[0068] Image noise reduction and tilt correction can both improve the clarity and quality of documents for the first, second, and third comparison objects, while text feature extraction converts the text data of the first, second, and third comparison objects into structured data that computers can understand and process. By extracting key features from the text, scientific abstraction and mathematical modeling of the text content can be achieved, thereby supporting various text processing tasks such as text classification, clustering, and automatic summarization, improving the quality of content comparison between PDF assembly instruction citations and the main text, and reducing comparison errors. The steps for tilt correction of the first, second, and third comparison objects are as follows:
[0069] Edge detection,First, the edge detection algorithm of the Canny edge detector is used to identify the edges in the image, which usually correspond to the boundaries of text, lines or graphics;
[0070] Hough transform, which is applied to detect straight lines in an image. Hough transform is a feature extraction technique that converts straight lines in image space into points in parameter space, making it easier to detect straight lines.
[0071] Line analysis: In the results of the Hough transform, the detected lines are analyzed, especially those that may represent text lines. The slopes of these lines will be used to determine the tilt angle of the image;
[0072] Calculate the tilt angle. By calculating the average slope or main slope of the detected straight lines, the tilt angle of the image can be determined. This angle is measured in degrees and indicates how many degrees the image needs to be rotated to reach a horizontal or vertical state.
[0073] Image rotation,Finally, the image is rotated using the calculated tilt angle to,correct its orientation;
[0074] The formula for calculating the tilt angle is as follows:
[0075]
[0076] where m i is the slope of the line detected by Hough transform, i is the index of the line;
[0077] Generate a first comparison combination from the first comparison object and the second comparison object, and generate a second comparison combination from the second comparison object and the third comparison object. After item-by-item comparison and cross-reference verification, perform a logical consistency check, format comparison, and annotation.
[0078] The item-by-item comparison involves factor decomposition and content comparison of the first and second comparison combinations. Factor decomposition involves breaking down the contents of the first and third comparison objects in the comparison combination into several specific comparison elements. Content comparison involves searching for each comparison element in the corresponding comparison object in the other comparison combination to check for any differences or omissions.
[0079] Cross-reference verification includes: indexing, reverse lookup, and consistency comparison. In the assembly instructions, a tag is created for each comparison object to facilitate quick location. Based on the index tag of the comparison object, the corresponding content is reversely searched in the original assembly instruction compilation instructions page, process content page, and supporting page files to verify whether the content of the assembly instruction compilation instructions page, process content page, and supporting page files is consistent with the description in the assembly instructions, ensuring the accuracy of the reference.
[0080] Furthermore, logical consistency checks include process analysis and exception detection. This involves analyzing the assembly process in the assembly instructions to ensure the logical relationships between each step are reasonable and coherent. For any exceptions that may occur in the process, check whether there are corresponding citations to explain them. By using document analysis, the assembly instruction PDF text and citations are centrally compared to improve comparison efficiency and accuracy.
[0081] Furthermore, formatted comparison includes: standardized references and version control to ensure that the citation format in the assembly instructions is unified and standardized to facilitate comparison and review, pay attention to the version information of the technical documents referenced by the first comparison object and the second comparison object, and ensure that the versions used by the first comparison object and the second comparison object are consistent.
[0082] According to the assembly instruction page attributes, the differences between the first, second, and third comparison objects are compared. Based on the different characteristics of the characters, the processed image is compared with the characters in the database, and the differences in the comparison area are marked uniformly by the depth of the same color. The documents marked with differences after comparison are saved separately, for example:
[0083] When comparing the first comparison object and the second comparison object, and also comparing the technical document number, name, and version, if there is a slight difference in the technical document number between the first comparison object and the second comparison object, the difference will also be marked in red font color, and the technical document numbers on the first comparison object and the second comparison object will be marked in bright red and light red respectively, and the first comparison object and the second comparison object with the different markings will be saved separately;
[0084] By comparing the recognized text with its possible similar candidate words, find logical words based on the recognized text before and after, display them in the sidebar, and then make corrections. The recognized and corrected text structure is output to txt, doc, and exl format files for easy reference and sharing by operators, so as to enhance the accuracy of the comparison and unify the formatting of the revised documents under the unified format.
[0085] A system for comparing assembly instruction citations and body texts comprises a processor module 1, a network interface 2, a display screen 3, an input device 4, and a comparison module 5 connected via a system bus. The processor 1 is configured to provide computing and control capabilities. The processor module 1 is externally connected to a scanning component 6, which is a scanner or other scanning device. The network interface 2, the display screen 3, the input device 4, and the comparison module are all connected to the processor module 1. The network interface 2 is configured to communicate with an external terminal via a network connection. The display screen 3 is configured to display data. The display screen 3 can display the content of a PDF document. The input device 4 is configured to input data and includes a mouse and a keyboard.
[0086] Specifically, the comparison block 5 includes: a memory 51, an image-text conversion and recognition module 52, a comparison module 53, an analysis and verification module 54 and an output module 55. The memory 51 is plugged into the output end of the processor module 1. The memory 51 can store the original assembly instruction PDF document, and can also store the marked differences and corrected documents. The recognition module 52, the comparison module 53, the analysis and verification module 54 and the output module 55 are all installed on the processor module 1. The recognition module 52 is used for conversion between image documents and content recognition, which facilitates the recognition and extraction of text and digital information in the document and converts it into an editable and searchable electronic format. The comparison module 53 is used for pairwise comparison of documents in the same group, which facilitates the difference comparison and marking of the first comparison combination, the second comparison combination and the third comparison combination. The analysis and verification module 54 is used for analysis and anomaly verification of the comparison content, which is used to improve the accuracy of the comparison. The output module 55 is used for the production of comparison results and output in multiple formats, which facilitates the output of the comparison documents into files in txt, doc, and exl formats to enhance the accuracy and comprehensiveness of the comparison.
[0087] Example 2
[0088] According to the attached Figure 3 As shown, including:
[0089] Based on the different attributes of the assembly instruction pages, the pre-comparison pre-processing preparation work is performed on the assembly instruction compilation instructions page, the assembly instruction process content page, and the assembly instruction supporting page. The assembly instruction compilation instructions page, the assembly instruction process content page, and the assembly instruction supporting page are subjected to optical character recognition (OCR) to convert the image into text. The OCR text recognition adopts an end-to-end text recognition model based on deep learning, extracts image features through a convolutional neural network, and then performs sequence recognition through a recurrent neural network.
[0090] A BERT-based pre-trained language model is used to perform semantic analysis and key information extraction on the text to generate structured first, second, and third comparison objects. The ontology knowledge of the assembly field is introduced into the semantic analysis process to guide the extraction of key information.
[0091] A hybrid comparison algorithm based on rules and statistics is used to compare the first comparison object and the second comparison object, and the second comparison object and the third comparison object. The rule-based comparison uses a graph-based matching algorithm, and the statistics-based comparison uses a twin network structure.
[0092] Wherein the above hybrid alignment algorithm includes:
[0093] Rule-based pairwise comparison uses domain knowledge to set comparison rules, which are formalized through assembly process flow, assembly 3D model, etc.
[0094] Based on statistical comparison, the text semantic similarity is learned using a twin network structure. The twin network is trained by contrastive learning, using assembly instruction text pairs as training samples.
[0095] The DS evidence theory is used to fuse the rule-based and statistics-based comparison results, adaptively adjust the evidence weights, and generate the final comparison results, further limiting the details of the hybrid comparison algorithm.
[0096] Rule-based comparison uses domain knowledge to set comparison rules and guides the comparison process by formalizing assembly process flow, 3D models, etc.
[0097] The statistical comparison uses a twin network structure and is trained through contrastive learning to learn the semantic similarity of assembly instruction texts. When fusing the two comparison results, the DS evidence theory is used to adaptively adjust the evidence weights to improve the robustness of the comparison results. This hybrid comparison algorithm combines the advantages of rules and statistics, balancing accuracy and adaptability.
[0098] An adaptive color depth mapping mechanism is used to annotate the differences based on the type and degree of the differences. The adaptive mapping mechanism groups the differences using a clustering algorithm and dynamically assigns a color depth to each group.
[0099] The adaptive color depth mapping mechanism includes:
[0100] The K-means clustering algorithm was used to group the differences and calculate the cluster centers according to the type and degree of difference;
[0101] Dynamically assign a color tone to each cluster, and adjust the color depth according to the degree of difference under the color tone;
[0102] Supports user-defined parameters such as the number of clusters and color tones to improve the flexibility of difference annotation;
[0103] The implementation details of the adaptive color depth mapping mechanism have been further refined. First, a K-means clustering algorithm is used to group differences, and cluster centers are calculated based on the type and degree of difference. A color tone is then dynamically assigned to each cluster, and the color depth is adjusted based on the degree of difference. This mechanism also allows users to customize parameters such as the number of clusters and color tone, increasing the flexibility of difference annotation. Adaptive color mapping makes annotation more intuitive and easier to identify, enhancing the human-computer interaction experience.
[0104] The comparison results are organized into an assembly knowledge graph, allowing users to intuitively browse the assembly process and the causes of differences through the graph. The assembly knowledge graph uses an RDF-based quintuple representation and supports SPARQL queries and reasoning.
[0105] The steps of constructing the assembly knowledge graph include:
[0106] Define the assembly domain ontology, including core concepts and relationships such as parts, processes, procedures, and tooling;
[0107] Extract key information from assembly instructions as ontology instances, and build relationships between instances through semantic links;
[0108] Use ontology reasoning rules to discover implicit knowledge, such as assembly sequence relationships, difference propagation paths, etc.
[0109] A graph visualization algorithm based on force-directed layout is used to generate an interactive visualization interface for the assembly knowledge graph;
[0110] The detailed process of constructing the assembly knowledge graph is as follows: first, define the assembly domain ontology, including core concepts and relationships such as parts, processes, procedures, and tooling; then, extract the key information in the assembly instructions as ontology instances, and construct the relationships between instances through semantic links; then, use the ontology reasoning rules to discover implicit knowledge, such as assembly sequence relationships, difference propagation paths, etc.; finally, use a graph visualization algorithm based on force-directed layout to generate an interactive visualization interface for the assembly knowledge graph; the assembly knowledge graph organizes assembly knowledge in a structured and semantic manner, supports intelligent query and reasoning, and provides knowledge support for assembly decisions.
[0111] The above embodiment first adopts OCR text recognition technology based on deep learning, extracts image features through convolutional neural networks, and then performs sequence recognition through recurrent neural networks to convert the image of assembly instructions into text. Then, a pre-trained language model based on BERT is used, combined with assembly domain ontology knowledge, to perform semantic analysis and key information extraction on the text to generate structured comparison objects. Next, a hybrid comparison algorithm is adopted to combine rule-based graph matching and statistics-based twin networks to perform pairwise comparisons on the comparison objects. When annotating differences, an adaptive color depth mapping mechanism is adopted, and the differences are grouped through a clustering algorithm, and colors are dynamically assigned to intuitively highlight the differences. Finally, the comparison results are organized into an RDF-based assembly knowledge graph, which supports semantic query and reasoning and provides an interactive visualization interface. This method integrates technologies such as deep learning and knowledge graphs, and can intelligently and efficiently complete the comparison of assembly instructions, improving accuracy and interpretability.
[0112] A system for comparing assembly instruction citations and texts, comprising:
[0113] The OCR text recognition module uses an end-to-end text recognition model based on deep learning, and completes image-to-text conversion through the collaborative work of convolutional neural networks and recurrent neural networks;
[0114] The natural language processing module uses a BERT-based pre-trained language model and combines assembly domain ontology knowledge to achieve text semantic analysis and key information extraction;
[0115] The hybrid comparison module uses a graph-based matching algorithm and a twin network structure to perform rule-based comparison and statistical comparison respectively, and integrates the results through DS evidence theory;
[0116] The difference annotation module uses the K-means clustering algorithm to group differences and an adaptive color depth mapping mechanism to annotate differences, supporting user-defined clustering and color parameters;
[0117] The knowledge graph construction module uses RDF-based quintuple representation to construct an assembly knowledge graph, supports SPARQL queries and ontology reasoning, and provides an interactive graph visualization interface based on force-directed layout.
[0118] The system includes multiple functional modules, implementing functions such as OCR text recognition, natural language processing, hybrid comparison, difference annotation, and knowledge graph construction. The OCR text recognition module utilizes an end-to-end deep learning model, the natural language processing module employs a pre-trained language model and domain ontology, the hybrid comparison module integrates rule matching and twin network comparison, the difference annotation module employs a clustering algorithm and adaptive color mapping, and the knowledge graph construction module uses RDF quintuple representation and ontology reasoning. The system architecture utilizes microservices and containerized deployment, improving scalability and maintainability.
[0119] Furthermore, the system architecture adopts microservice architecture and containerized deployment, including:
[0120] Split functions such as OCR text recognition, natural language processing, mixed comparison, difference annotation, and knowledge graph construction into independent microservices, each of which is independently developed, tested, and deployed;
[0121] Use Docker container technology to encapsulate the microservice operating environment, and use the Kubernetes orchestration tool to achieve elastic scaling and load balancing of microservices;
[0122] Adopt an asynchronous communication mechanism based on message queues to achieve decoupling and collaboration between microservices;
[0123] The system's microservices architecture and containerized deployment are further explained. The system splits each functional module into independent microservices, each of which can be independently developed, tested, and deployed. Microservices are decoupled and collaborate with each other through asynchronous communication mechanisms based on message queues. The system uses Docker container technology to encapsulate the microservices runtime environment and utilizes the Kubernetes orchestration tool to achieve elastic scaling and load balancing for microservices. This microservices architecture and containerized deployment improve the system's flexibility, scalability, and fault tolerance, enabling it to efficiently meet the computational demands of assembly instruction comparison.
[0124] The system further includes:
[0125] The GPU acceleration module uses the CUDA parallel computing framework and the multi-core parallel processing capabilities of the GPU to accelerate computationally intensive tasks such as OCR text recognition, semantic analysis, and twin network training;
[0126] The distributed storage module uses the HDFS distributed file system and HBase distributed column storage to support high-concurrency reading and writing and real-time query of the assembly knowledge graph;
[0127] The human-computer interaction module uses Web front-end technology to provide a graphical assembly instruction comparison and knowledge graph display interface, supporting users to participate in the comparison process through interactive methods such as dragging and clicking.
[0128] Specifically, the GPU acceleration module uses the CUDA parallel computing framework, leveraging the GPU's multi-core parallel processing capabilities to accelerate compute-intensive tasks such as OCR text recognition, semantic analysis, and twin network training. The distributed storage module utilizes HDFS and HBase technologies to support high-concurrency reading and writing of assembly knowledge graphs and real-time queries. The human-computer interaction module utilizes web front-end technology to provide a graphical assembly instruction comparison and knowledge graph display interface, allowing users to participate in the comparison process through interactive operations. The introduction of these modules further enhances the system's performance, storage capacity, and user experience, making it a complete and practical assembly instruction comparison solution.
[0129] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for comparing assembly instruction citations and texts, characterized in that: include: According to the different attributes of the assembly instruction page, pre-processing preparation work is carried out on the assembly instruction compilation description page, process content page, and supporting page before comparison; Generate corresponding first, second, and third comparison objects for the assembly instruction compilation description page, the assembly instruction process content page, and the assembly instruction supporting page, respectively; and perform character format separation, binarization, image noise reduction, tilt correction, and text feature extraction on the first, second, and third comparison objects; Compare the first comparison object with the second comparison object, and the second comparison object with the third comparison object respectively, use item-by-item comparison and cross-reference verification, and then perform logical consistency check, format comparison and annotation; For the differences after comparison of the first, second and third comparison objects, the processed image is compared with the characters in the database again according to the different characteristics of the characters, and the differences in the comparison areas are marked uniformly by the depth of the same color.
2. A method for comparing assembly instruction citations and texts according to claim 1, characterized in that: The pre-processing preparation work before comparing the assembly instruction compilation instruction page, the assembly instruction process content page, and the assembly instruction supporting page includes: Scan the PDF file with a scanner to ensure it is legible, without any damage or encryption restrictions before comparison. Before performing any comparison operations, save the original PDF file as a backup. The method for comparing assembly instruction citations and text according to claim 2 is characterized in that, based on the assembly instruction page attributes, a first comparison combination is generated by combining the first comparison object and the second comparison object, and a second comparison combination is generated by combining the second comparison object and the third comparison object, and step decomposition, citation comparison, and content verification are performed on the first comparison combination and the second comparison combination, respectively. The steps of performing tilt correction on the first comparison object, the second comparison object, and the third comparison object include: edge detection, Hough transform, straight line analysis, calculation of tilt angle, and image rotation, and the tilt angle calculation formula is as follows: where m i is the slope of the line detected by Hough transform, and i is the index of the line.
3. A method for comparing assembly instruction citations and texts according to claim 3, characterized in that: The step decomposes the document contents of the first comparison object and the third comparison object in the comparison combination of the assembly instruction into several specific comparison elements. The citation comparison is to search for another comparison object in the corresponding comparison combination for each comparison element. The content verification is to compare the descriptions of the two comparison objects in the comparison combination to check whether there are differences and omissions.
4. A method for comparing assembly instruction citations and texts according to claim 3, characterized in that: The cross-reference verification includes: Create an index and mark each comparison object in the assembly instruction to facilitate quick positioning; Reverse search: according to the index mark of the comparison object, reverse search the corresponding content in the original assembly instruction compilation page, process content page, and supporting page files; Compare the consistency and verify whether the content descriptions of the assembly instruction compilation instructions page, the assembly instruction process content page, and the assembly instruction supporting page are consistent to ensure the accuracy of the reference.
5. The method for comparing assembly instruction citation and text according to claim 3, characterized in that: The logical consistency check includes: Process analysis: analyze the assembly process in the assembly instructions to ensure that the logical relationship between each step is reasonable and coherent; Exception investigation: For any abnormal situations that may occur in the process, check whether there are corresponding citations to explain them.
6. A method for comparing assembly instruction citations and texts according to claim 3, characterized in that: The formatting comparison includes: Standardized citations to ensure uniform and standardized citation formats in assembly instructions; Version control focuses on the version information of the technical documents referenced by the first comparison object, the second comparison object, and the third comparison object.
7. The method for comparing assembly instruction citation and text according to claim 3, characterized in that: According to the assembly instruction page attributes, the recognized text is compared with its possible similar candidate characters, and logical words are found according to the recognized text before and after, displayed in the sidebar, and then corrections are made, and the recognized and corrected text structure is output to a file in txt, doc, or exl format.
8. A system for comparing assembly instruction citations and texts, characterized in that: include: A processor module (1), a network interface (2), a display screen (3), an input device (4), and a comparison block (5) connected via a system bus, wherein the processor (1) is used to provide computing and control capabilities, the processor module (1) is externally connected to a scanning component (6), the network interface (2), the display screen (3), the input device (4), and the comparison module are all connected to the processor module (1), and the comparison block (5) includes: A memory (51), wherein the memory (51) is plugged into an output end of the processor module (1); An image-text conversion and recognition module (52), the recognition module (52) is used for conversion between image documents and content recognition; A comparison module (53), wherein the comparison module (53) is used for pairwise comparison of documents in the same group; An analysis and verification module (54), the analysis and verification module (54) is used for analysis and anomaly verification of the comparison content; The output module (55) is used for producing and outputting the comparison results in various formats.
9. The assembly instruction citation and text comparison system according to claim 9, characterized in that: The network interface (2) is used for communicating with an external terminal via a network connection, the display screen (3) is used for displaying data, and the input device (4) is used for inputting data.