Serialization processing and retrieval method for spaceflight control software table data

Through serialization processing and retrieval enhancement technology, the fusion representation problem of table data in the aerospace control software field is solved, and the efficiency and quality of software development are improved.

CN120372000APending Publication Date: 2025-07-25BEIJING INST OF CONTROL ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335194.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing field of embedded aerospace control software lacks effective tabular data fusion representation methods, making it difficult to construct and utilize a rich and diverse aerospace control software asset data, affecting the efficiency and quality of software development.

Method used

By serializing the table data, using the Large Language Model (LLM) to generate text summary and perform data cleaning and preprocessing, a data set in the field of aerospace control software is constructed, and combined with search enhancement technology, the search and utilization of table knowledge is realized.

Benefits of technology

It improves software developers' understanding and application of software assets in form forms, and improves the efficiency and quality of software development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372000A_ABST
    Figure CN120372000A_ABST
Patent Text Reader

Abstract

A serialization processing and retrieval method for spaceflight control software table data belongs to the technical field of aerospace, and comprises the following steps: reading spaceflight control software table data, and representing table elements with placeholders; according to the table data after data cleaning, constructing an LLM input, backfilling the LLM output to a corresponding placeholder, and storing the backfilled table data into a knowledge base; constructing a data set according to the preprocessed table data, and continuing pre-training the LLM on the basis of the data set; and generating a large model retrieval framework according to offline deployment retrieval enhancement of the knowledge base. According to the method, table data in software documents are extracted, large model reasoning is utilized to obtain fusion with original text data to train a field large model, and a retrieval enhancement technology is utilized to improve understanding and application of the large model to table knowledge; and the understanding of software developers on software asset contents borne in a table form in the field of spaceflight embedded control software is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a serialization processing and retrieval method for tabular data of aerospace control software, belonging to the field of aerospace technology. Background Art

[0002] Facing the increasingly complex business forms and continuously changing business requirements, aerospace embedded control software presents the characteristics of increasingly complex functions, growing scale, accelerating demand changes, shortening development cycles, and improving quality requirements, making it difficult to meet the software development needs of the rapid development of the national aerospace industry. In recent years, large language models (LLMs) such as GhatGPT have brought improvements in R & D efficiency for software developers. For example, Microsoft's Copilot has been used by one million users after one year of launch, and the productivity of each user has increased by 30%. In the long-term software development process in the field of aerospace control software, a large number of reusable software assets have been accumulated, including requirement documents, design reports, case libraries, and scientific research intelligence. These assets make it possible to build and apply large models in the field of aerospace embedded control software. The large model in the field of aerospace control software aims to utilize its powerful natural language processing capabilities to assist software engineers in processing and analyzing the natural language descriptions of aerospace tasks, helping to understand and clarify software requirements, thereby improving the efficiency and quality of software development. However, the existing methods for constructing knowledge models in fields such as finance, healthcare, and law rely on the integration of in-domain text data. Currently, there is a lack of a fusion representation method for the rich and diverse document data sources in the field of aerospace control software. How to effectively organize and utilize software asset data containing heterogeneous information such as tables, text, and pseudocode has become a difficult problem to overcome in building domain large models and realizing software asset reuse. Summary of the Invention

[0003] The technical problem solved by the present invention is: overcoming the deficiencies of the prior art, providing a serialization processing and retrieval method for tabular data of aerospace control software, extracting tabular data from software documents, using large model reasoning to fuse with the original text data for training the domain large model, and using retrieval enhancement technology to improve the understanding and application of tabular knowledge by the large model, and improving the understanding of software asset content carried in tabular form by software developers in the field of aerospace embedded control software.

[0004] The technical solution of the present invention is: a serialization processing and retrieval method for tabular data of aerospace control software, including:

[0005] Traverse all the collected aerospace control software tables, read the table data, and represent the table elements with placeholder symbols; clean the table data, construct the input of the pre-set large language model LLM for the purpose of outputting the text summary of the table data by the pre-set large language model LLM, fill the output of the pre-set large language model LLM back to the corresponding placeholder symbols, and store the table data after backfilling in the knowledge base;

[0006] Preprocess the table data, construct a dataset in the field of aerospace control software for training the pre-set large language model LLM based on the preprocessed table data, and continue to pre-train the pre-set large language model LLM on this dataset to construct a large model in the field of aerospace control software;

[0007] According to the knowledge base, offline deploy a retrieval-enhanced generation large model retrieval framework to realize the retrieval and utilization of the table knowledge in the knowledge base by the large model in the field of aerospace control software.

[0008] Further, the reading of the table data includes: using the Python-docx library to traverse all the collected software documents, and uniformly representing the extracted table elements with the placeholder symbol TABLE[N], where N represents the table number index.

[0009] Further, the data cleaning includes: those that meet the following rules simultaneously will be retained and stored in the table knowledge base, otherwise they will be deleted:

[0010] The table content is not empty;

[0011] The proportion of Chinese characters in the table is greater than or equal to the threshold.

[0012] Further, the input of the pre-set large language model LLM includes prompt words and table data.

[0013] Further, the preprocessing of the table data includes:

[0014] Based on regular expressions, denoise each data item in the aerospace control software field dataset, specifically including deleting HTML format pictures, table tags, reference symbols, book titles, and redundant empty numbers, to ensure that the text paragraphs after cleaning only contain natural language text;

[0015] Use the MinHash algorithm to calculate text similarity and remove text data with similarity exceeding the preset threshold.

[0016] Further, the text similarity calculation includes:

[0017] For text paragraph X1, K hash functions are specified, and the signature vector representation of the MinHash values under each hash function is calculated as signature(X1) = [h1(X1), h2(X1),..., h K (X1)]; The calculation method of h i (X1), i ∈ [1, K] is that for the characters in paragraph X1, they are mapped to a hash value of a fixed length through the hash function h i and then the smallest one is selected from the hash values as the MinHash value of this paragraph;

[0018] For the text similarity between paragraph X1 and paragraph X2, it is calculated using the Jaccord similarity, and the calculation formula is as follows:

[0019]

[0020] By pairwise comparing the Jaccord similarities of the paragraph texts, the text paragraphs with similarities greater than the preset threshold will be deleted.

[0021] Furthermore, the continued pre-training includes: designing a pre-training task of filling in words, randomly masking words using a masked language model, using an autoregressive model to predict the next word, and selecting ChatGLM3-6B as the base model, and using the continued pre-training method based on Low-Rank Adaptation (LoRA) to fine-tune the parameters of the linear layer of the model.

[0022] Furthermore, the retrieval-enhanced generation large model retrieval framework according to the knowledge base for offline deployment includes:

[0023] Vectorize the tabular data in the knowledge base. For text paragraph X1, use a pre-trained language model as the encoder to obtain the word vector representation of the text paragraph X1 = Encoder(X1); where the language model of the encoder selects bge-large-zh; the word vectors of the paragraph will be saved in the knowledge base together with the original text content;

[0024] For the question Q input by the user, use the same pre-trained language model as the encoder for question vectorization representation, Q = Encoder(Q); use the cosine similarity function to calculate the cosine similarities between the question vector and the text paragraph vectors one by one; select the three text paragraphs in the tabular knowledge base with the highest cosine similarity to the question Q as reference materials and add them to the input instruction of the large model together with the question, and return the final answer, so as to realize the retrieval and use of tabular knowledge.

[0025] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the serialization processing and retrieval method for table data of aerospace control software are implemented.

[0026] A serialization processing and retrieval device for table data of aerospace control software includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the serialization processing and retrieval method for table data of aerospace control software are implemented.

[0027] The advantages of the present invention compared with the prior art are as follows:

[0028] (1) Through the text summarization generation method based on a large model, the present invention serializes the extracted table data into text, which is convenient for fusion processing with the original text data.

[0029] (2) The present invention proposes a method for fusing and processing table and text data. By using data denoising based on regular expressions and data deduplication technology based on the MinHash algorithm, a unified knowledge representation and domain dataset covering multi-source heterogeneous data in the field of aerospace control software are established, and on this basis, continued pre-training is carried out to construct a large model in the field of aerospace control software.

[0030] (3) By using retrieval enhancement technology to assist the domain large model to complete the retrieval, analysis, and reasoning of table knowledge, the present invention improves the understanding of designers on the content and knowledge of software assets carried in the form of tables during the software development process. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0032] Figure 1 is the framework of the serialization processing and retrieval method for aerospace control software asset data of the present invention;

[0033] Figure 2 is the fusion and preprocessing process diagram of aerospace control software table and text data of the present invention;

[0034] Figure 3 is the process diagram of aerospace control software table knowledge retrieval and use of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] To better understand the above technical solution, the technical solution of the present invention will be described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.

[0036] The following further details the method for serializing and retrieving tabular data of aerospace control software provided by the embodiments of the present invention in conjunction with the accompanying drawings of the specification. As Figure 1 , the specific implementation manner may include:

[0037] Traverse all collected aerospace control software tables, read the table data, and represent the table elements with placeholder symbols; perform data cleaning on the table data. For the purpose of outputting the text of the table data summary by the preset large language model LLM, construct the input of the preset large language model LLM according to the table data after data cleaning, fill the output of the preset large language model LLM back to the corresponding placeholder, and store the table data after filling back into the knowledge base.

[0038] Preprocess the table data, construct a dataset in the field of aerospace control software for training the preset large language model LLM based on the preprocessed table data, and continue to pre-train the preset large language model LLM on the basis of this dataset to construct a large model in the field of aerospace control software.

[0039] Deploy a retrieval-enhanced generation large model retrieval framework offline according to the knowledge base to realize the retrieval and utilization of the table knowledge in the large model in the field of aerospace control software.

[0040] In the solution provided by the embodiments of the present invention, as Figure 1 shown, the steps of a method for serializing and retrieving tabular data of aerospace control software of the present invention are:

[0041] (1) Generate a summary of the table content based on the large language model to realize the serialization of the table data.

[0042] (2) Through data denoising and deduplication, fuse the table and text data to construct.

[0043] (3) Based on the retrieval enhancement technology, retrieve and use the table knowledge.

[0044] The implementation manner of step (1):

[0045] Use the Python-docx library to traverse all collected software documents, and uniformly represent the extracted table elements with the placeholder "TABLE[N]", where N represents the table number index.

[0046] The data storage structure of the table knowledge base is shown in the following table. Each row of the table will be read and its data content will be stored in JSON format.

[0047]

[0048] For all the extracted table data, data cleaning will be performed based on the rules formulated manually. Those that meet either of the following rules ① or ② will be retained and stored in the table knowledge base, otherwise they will be deleted:

[0049] ① The table content is not empty;

[0050] ② The proportion of Chinese characters in the table is greater than or equal to the threshold δ.

[0051] Here, the threshold δ is set at 40%.

[0052] For the cleaned table data, an instruction in the form of "prompt word + table data" will be constructed as the input to the ChatGLM3-6B large model.

[0053] Here, the content of the "prompt word" is as follows: "Describe the information expressed in the following table in a paragraph. Note that no extra content should be added randomly. Here is the table I provided:".

[0054] The inference result of the large model will be filled back to the corresponding placeholder in the original document data and saved in the table knowledge base, finally realizing the serialization and textification of the table data.

[0055] Implementation method of step (2):

[0056] As Figure 2 shown, through data denoising based on regular expressions and data deduplication based on the MinHash algorithm, the integration and preprocessing of the table and text data of the aerospace control software are realized. The serialized document data of the table is converted into a dataset suitable for large model training, and continued pre-training is carried out on the basis of this dataset to build a large model in the field of aerospace control software. The specific process is as follows:

[0057] Step one: Based on regular expressions, perform denoising processing on each data item in the dataset in the field of aerospace control software, specifically including deleting tags such as HTML format pictures, table tags, citation symbols, book titles, and redundant empty tags, to ensure that the cleaned text paragraph contains only natural language text as much as possible.

[0058] Step two: Use the MinHash algorithm to calculate the text similarity and remove the text data with high similarity.

[0059] For text paragraph X1, K hash functions are specified, and the signature vector representation of the MinHash values under each hash function is calculated as signature(X1) = [h1(X1), h2(X1),..., h K (X1)], where the calculation method of h i (X1), i ∈ [1, K] is that for the characters in paragraph X1, they are mapped to a hash value of a fixed length through the hash function h i , and then the smallest one is selected from the hash values as the MinHash value of this paragraph.

[0060] For the text similarity between paragraph X1 and paragraph X2, it is calculated using the Jaccord similarity, and the calculation formula is as follows:

[0061]

[0062] By comparing the Jaccord similarities of paragraph texts pairwise, the text paragraphs with a similarity greater than the threshold of 0.8 will be deleted.

[0063] Step 3: After the data preprocessing steps of denoising and duplicate removal, a dataset in the field of aerospace embedded control software is obtained. Design a pre-training task of filling in words, randomly mask words using a masked language model, use an autoregressive model to predict the next word, and select ChatGLM3-6B as the base model. Use the continued pre-training method based on Low-Rank Adaptation of LLMs (LoRA) to fine-tune the parameters of the linear layer of the model to build a large model in the field of aerospace embedded control software.

[0064] Implementation method of step (3):

[0065] As Figure 3 shown, based on the domain large model, use the LangChain-Chatchat framework to offline deploy a retrieval-enhanced generation large model knowledge base and retrieval framework to realize the retrieval and utilization of tabular knowledge. The specific process is as follows:

[0066] Step 1: Vectorize the text of the tabular knowledge base. For text paragraph X1, use a pre-trained language model as an encoder to obtain the word vector representation of the text paragraph X1 = Encoder(X1). The language model of the encoder selects bge-large-zh, which has the best support for Chinese. The word vectors of the paragraph will be saved in the tabular knowledge base together with the original text content.

[0067] Step 2: For the question Q input by the user, use the same pre-trained language model as the encoder to perform question vector representation, Q = Encoder(Q). Use the cosine similarity function to calculate the cosine similarity between the question vector and the text paragraph vectors one by one. Taking the text paragraph X1 of the question Q as an example, the cosine similarity calculation formula is as follows:

[0068]

[0069] Select the three text paragraphs with the highest cosine similarity to the question Q in the table knowledge base as "reference materials" and add them to the input instruction of the large model together with the question, and return to obtain the final answer, so as to realize the retrieval and use of table knowledge. The specific input instruction form is as follows: "Question Q + Please refer to the following content to answer this question + Reference materials".

[0070] The present invention provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions. When the computer instructions run on a computer, the computer is made to execute Figure 1 the method described above.

[0071] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program codes.

[0072] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the specified function in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0073] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the process in Figure 1one or more processes and / or blocks Figure 1 the functions specified in one or more blocks.

[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one Figure 1 one or more processes and / or blocks Figure 1 or more processes and / or blocks.

[0075] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

[0076] The content not described in detail in the specification of the present invention belongs to the well-known technology of those skilled in the art.

Claims

1. A serialization processing and retrieval method for tabular data of aerospace control software, characterized in that Including: Traverse all collected aerospace control software tables, read the table data, and represent the table elements with placeholders. Clean the table data. For the purpose of using a pre-set large language model (LLM) to output a text summary of the table data, construct the input of the pre-set LLM based on the cleaned table data, fill the output of the pre-set LLM back to the corresponding placeholder, and store the filled table data in the knowledge base. Preprocess the table data, construct a dataset for aerospace control software field training of the pre-set LLM based on the preprocessed table data, and continue to pre-train the pre-set LLM on this dataset to construct a large model for the aerospace control software field. According to the knowledge base, offline deploy a retrieval-enhanced generation large model retrieval framework to realize the retrieval and utilization of the table knowledge in the knowledge base by the large model in the aerospace control software field.

2. The serialization processing and retrieval method for tabular data of aerospace control software according to claim 1, wherein, The reading of the table data includes: using the Python-docx library to traverse all collected software documents, and uniformly represent the extracted table elements with the placeholder TABLE[N], where N represents the table number index.

3. A serialization processing and retrieval method for tabular data of aerospace control software according to claim 1, characterized in that, The data cleaning includes: those that meet the following rules simultaneously will be retained and stored in the table knowledge base, otherwise they will be deleted: The table content is not empty. The proportion of Chinese characters in the table is greater than or equal to the threshold.

4. A serialization processing and retrieval method for tabular data of aerospace control software according to claim 1, characterized in that, The input of the pre-set LLM includes prompt words and table data.

5. A serialization processing and retrieval method for tabular data of aerospace control software according to claim 1, characterized in that, The preprocessing of the table data includes: Based on regular expressions, denoise each data item in the aerospace control software field dataset, specifically including deleting HTML format pictures, table tags, reference symbols, book titles, and redundant empty numbers to ensure that the cleaned text paragraph only contains natural language text. Use the MinHash algorithm to calculate text similarity and remove text data with similarity exceeding the pre-set threshold.

6. A serialization processing and retrieval method for tabular data of aerospace control software according to claim 5, characterized in that The calculation of the text similarity includes: For text paragraph X1, K hash functions are specified, and the signature vector representation of the MinHash value under each hash function is calculated as signature(X1) = [h1(X1), h2(X1),..., h K (X1)]; The calculation method of h i (X1), i ∈ [1, K] is that for the characters in paragraph X1, through the hash function h i they are mapped to a hash value of a fixed length, and then the smallest one is selected from the hash values as the MinHash value of this paragraph; For the text similarity between paragraph X1 and paragraph X2, it is calculated using the Jaccord similarity, and the calculation formula is as follows: By comparing the Jaccord similarities of paragraph texts pairwise, text paragraphs with similarities greater than the pre-set threshold will be deleted.

7. A serialization processing and retrieval method for tabular data of aerospace control software according to claim 1, characterized in that, The continued pre-training includes: designing a pre-training task of filling in the blanks, randomly masking words using a masked language model, using an autoregressive model to predict the next word, selecting ChatGLM3-6B as the base model, and using the continued pre-training method based on Low-Rank Adaptation (LoRA) to fine-tune the parameters of the linear layer of the model.

8. A serialization processing and retrieval method for tabular data of aerospace control software according to claim 1, characterized in that The offline deployment of the retrieval-enhanced generation large model retrieval framework according to the knowledge base includes: Vectorize the text of the table data in the knowledge base. For text paragraph X1, use a pre-trained language model as the encoder to obtain the word vector representation of the text paragraph X1 = Encoder(X1); where the language model of the encoder selects bge-large-zh; the word vector of the paragraph will be saved in the knowledge base together with the original text content. For the question Q input by the user, use the same pre-trained language model as the encoder to represent the question in vector form, Q = Encoder(Q); use the cosine similarity function to calculate the cosine similarity between the question vector and the text paragraph vectors one by one; select the three text paragraphs with the highest cosine similarity to the question Q in the table knowledge base as reference materials and add them to the input instruction of the large model together with the question, and return to obtain the final answer, so as to realize the retrieval and use of table knowledge.

9. A computer-readable storage medium storing a computer program, characterized in that, When the described computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

10. A serialization processing and retrieval device for tabular data of aerospace control software, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the described processor executes the described computer program, it implements the steps of the method according to any one of claims 1 to 8.