Satellite document table data retrieval method and device, equipment and medium
By formatting the table information of satellite documents into JSON structures and semantic compression, building node clusters, and using large language models for vectorization, the problems of missing and unclear semantics in the existing technology are solved, and more efficient tabular data retrieval and semantic understanding are achieved.
Patent Information
- Application Number
- CN202510149630.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art causes missing retrieval and unclear semantic problems when processing satellite document table information, mainly because after converting two-dimensional structured data into one-dimensional data, the context topological structured relationship of the table is lost, and it is difficult for large language models to understand specially defined characters.
By formatting the table information in the satellite document into JSON structure tabular data, and constructing a node cluster with table content as the core through semantic compression, using a large language model to vectorize the JSON structure tabular data and context compressed information to obtain an embedded vector representing the table information.
It effectively solves the problems of missing searches and unclear semantics, and improves the search accuracy and semantic understanding of tabular data.
Smart Images

Figure CN120216548A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a method, device, equipment and medium for retrieving satellite document table data. Background Art
[0002] In the satellite field, the retrieval function of table content not only requires high accuracy, but also needs to cope with flexible and changeable questioning methods. The formatting method of table content also has certain differences from traditional document information, with characteristics such as a standardized data information representation form and a lack of context association descriptions of element domain information. Some existing technologies directly process table information as document information, that is, extract the data information in the table as natural language characters, convert the two-dimensional structured data in the table into one-dimensional character data, and then provide it to a large language question-and-answer model such as GPT as prompt information and combine it with the question asked, and finally obtain the required answer. Similarly, there are also those that use the table data format as part of the data source in Retrieval-Augmented Generation (RAG). The document is segmented by a fixed length, and in the processing of table data, the context topological structure of the table format is retained, and the division relationship of the element domain is provided with characters such as "|", and it is expected to utilize the context semantic understanding ability of the large model to let the large model understand the expression form of the table constructed under "|".
[0003] When converting the two-dimensional structured data of the table into one-dimensional, the context topological structured relationship provided by the table is lost. In addition, the table data is constructed in a horizontal and vertical manner, and the vertical column data should have similar semantic meanings. However, after directly converting it into one-dimensional data, the previous standardized topological data relationship structure is broken, making it more difficult for the large model to understand the data.
[0004] In addition, when reconstructing the structure of the table information again, filling the table content in characters such as "|", although humans can understand the original data structure relationship. However, because in the early training of the pre-trained language model, the special-defined "|" may have different semantic meanings, which will cause the pre-trained large language model to not be able to well understand the meaning of this kind of character. Although the structured data format of the original table is maintained in form, the large model cannot understand the meaning of the structured data semantically, which will cause problems such as retrieval missing and semantic ambiguity. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, device, equipment and medium for retrieving satellite document table data. By formatting the table information in the satellite document into JSON structured table data and constructing a node cluster with the table content as the core through semantic compression, the technical problems of direct processing of table information as document information in the prior art, resulting in retrieval omission and unclear semantics, are solved.
[0006] To solve the above technical problems, the present invention is realized through the following technical solutions:
[0007] The present invention provides a method for retrieving satellite document table data, which includes:
[0008] Obtain the table information in the satellite document, where the table information includes header information, table body information, and table context document material information;
[0009] Perform vectorization processing on the table information to obtain an embedding vector representing the table information;
[0010] Compare and analyze the user's question with the embedding vector of the table information to screen out the table data that matches the user's question;
[0011] Feed the screened table data back to the user as the retrieval result.
[0012] In an embodiment of the present invention, the obtaining of the table information in the satellite document includes:
[0013] Locate the table according to the satellite document to determine the position of the table in the satellite document;
[0014] Extract the header information, table body information, and table context document material information of the table according to the position of the table in the satellite document.
[0015] In an embodiment of the present invention, the performing of vectorization processing on the table information to obtain an embedding vector representing the table information includes:
[0016] Format the header information and the table body information to construct nested JSON structured table data;
[0017] Perform offline compression processing on the header information and the table context document material information to obtain the context compression information of the table;
[0018] Use a large language model to perform vectorization processing on the JSON structured table data and the context compression information to obtain an embedding vector representing the table information.
[0019] In one embodiment of the present invention, the process of using a large language model to perform vectorization processing on the JSON-structured table data and the context compression information to obtain an embedding vector representing the table information includes:
[0020] Using a large language model to perform semantic information compression on the JSON-structured table data and the context compression information to extract the central condensed expression of the table information;
[0021] According to the central condensed expression of the table information and in combination with the original table information, perform mapping processing in a high-dimensional space to obtain an embedding vector representing the table information.
[0022] In one embodiment of the present invention, the process of comparing and analyzing the user's question with the embedding vector of the table information to filter out the table data that matches the user's question includes:
[0023] Calculate the similarity between the user's question and the embedding vector of the table information to obtain a similarity comparison result;
[0024] According to the similarity comparison result, use the prompts technique to search to determine the table data that best matches the user's question.
[0025] In one embodiment of the present invention, the process of calculating the similarity between the user's question and the embedding vector of the table information to obtain a similarity comparison result includes:
[0026] Perform vectorization processing on the user's question to obtain an embedding vector corresponding to the user's question;
[0027] Calculate the similarity between the embedding vector corresponding to the user's question and the embedding vector of the table information to obtain a similarity comparison result.
[0028] In one embodiment of the present invention, the process of using the prompts technique to search according to the similarity comparison result to determine the table data that best matches the user's question includes:
[0029] If the similarity comparison result is similar, calculate the token length of the natural language description of the table;
[0030] When the token length exceeds the maximum limit that the large language model can handle, split the table information into multiple parts to ensure that the token length of each part does not exceed the maximum acceptable value of the large language model;
[0031] Use a large language model to perform information integration processing on each part of the segmented table information, and splice the processed table information of each part again to restore the integrity of the table information;
[0032] Use the prompts technique to search for and filter out the table data that best matches the user's question.
[0033] Based on the same inventive concept, another embodiment of the present invention further provides a satellite document table data retrieval device, which includes:
[0034] An information acquisition module for acquiring table information in a satellite document, where the table information includes header information, table body information, and table context document material information;
[0035] An information processing module for performing vectorization processing on the table information to obtain an embedding vector representing the table information;
[0036] An information matching module for comparing and analyzing the user's question with the embedding vector of the table information to filter out the table data that matches the user's question;
[0037] An information output module for feeding back the filtered table data as a retrieval result to the user.
[0038] Based on the same inventive concept, another embodiment of the present invention further provides an electronic device, which includes:
[0039] One or more processors;
[0040] A storage device for storing one or more programs, which when executed by the one or more processors, causes the electronic device to implement the satellite document table data retrieval method as described in any of the above embodiments.
[0041] Based on the same inventive concept, another embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, which when executed by a processor of a computer, causes the computer to execute the satellite document table data retrieval method as described in any of the above embodiments.
[0042] As described above, a method for retrieving satellite document table data provided by the present invention obtains table information in a satellite document, where the table information includes header information, table body information, and table context document data information, vectorizes the table information to obtain an embedding vector representing the table information, compares and analyzes a user's question with the embedding vector of the table information to screen out table data matching the user's question, and feeds back the screened table data as a retrieval result to the user. The method decodes the table content according to the document XML format and recursively constructs a JSON formatted structure for complex nested tables to accurately describe the information content in the table. Subsequently, using the header information, table body information, and table context document data information, combined with the Prompts technology of a large language model, core information condensation and high-dimensional vector numerical compression are performed on the foregoing content. Aiming at the problem that the number of tokens after table JSON formatting may exceed the processing limit of the large language model, an iterative information integration method is designed to gradually condense and summarize the table information, thereby ensuring the accuracy of the retrieved content. Of course, any product implementing the present invention does not necessarily need to achieve all the above-mentioned advantages at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0044] Figure 1 It is a flowchart of a method for retrieving satellite document table data provided by an exemplary embodiment of the present application.
[0045] Figure 2 It is a schematic diagram of JSON structured table data generated for table nesting situations provided by an exemplary embodiment of the present application.
[0046] Figure 3 It is a flowchart of vectorizing table information provided by an exemplary embodiment of the present application.
[0047] Figure 4 It is a flowchart of iterative information integration retrieval provided by an exemplary embodiment of the present application.
[0048] Figure 5 It is a schematic diagram of the structure of a device for retrieving satellite document table data provided by another exemplary embodiment of the present application.
[0049] Figure 6A schematic structural diagram of an electronic device provided by another exemplary embodiment of the present application. Detailed implementation manners
[0050] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0051] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0052] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0053] To solve the problems of retrieval loss and semantic ambiguity caused by directly treating tabular information as document information in the prior art, the present invention proposes a method for retrieving tabular data in satellite documents. This method recursively generates a structured format based on the table content, using the table header and element field data content, thereby constructing a JSON structure suitable for tables in complex scenarios. At the same time, in order to achieve effective retrieval, all the content information in the table is integrated, and a node cluster with the table content as the core is constructed by using the method of semantic vector compression. In addition, for the ultra-long tabular information after structuring, according to the number of tokens, some content is compressed for key information using a large language model, and the effective information is gradually integrated.
[0054] It should be noted that the method of directly converting tabular data into JSON structure data aims to integrate tabular information in a unified data format, store the element fields in the table in the form of a dictionary, and then use the converted data resource as a part of the document in the Retrieval-Augmented Generation (RAG) operation to calculate the similarity with the content related to the question and perform the retrieval task.
[0055] Please refer to Figure 1 As shown, the satellite document table data retrieval method includes the following steps:
[0056] S100: Obtain the table information in the satellite document, where the table information includes header information, table body information, and table context document information;
[0057] S200: Perform vectorization processing on the table information to obtain an embedding vector representing the table information;
[0058] S300: Compare and analyze the user's question with the embedding vector of the table information to filter out the table data that matches the user's question;
[0059] S400: Feed back the filtered table data as the retrieval result to the user.
[0060] First, execute step S100, that is, obtain the table information in the satellite document, where the table information includes header information, table body information, and table context document information.
[0061] In an exemplary embodiment of the present application, in step S100, obtaining the table information in the satellite document further includes:
[0062] S110: Locate the table according to the satellite document to determine the position of the table in the satellite document;
[0063] S120: Extract the header information, table body information, and table context document information of the table according to the position of the table in the satellite document.
[0064] Specifically, for the document information with tabular data, the first thing to do is to locate the table. According to the WordprocessingML standard, the tables, natural language characters, etc. in the satellite document are extracted, and their corresponding position relationships in the satellite document are determined. It should be noted that the WordprocessingML standard is part of the Microsoft Office Open XML standard, which defines the storage format of Word documents. After determining the position of the table in the satellite document, the header information, body information, and table context document information of the table are extracted. In this embodiment, as shown in Table 1, the header information is the theme of the table. For example, Table 1. Energy Monitoring Telemetry Content. The body information is the actual data part in the table, including the row and column data in the table. The table context document information is the additional background information provided in the document paragraphs before or after the table to provide information about the table content or usage. When the format of the table is parsed, the python-docx library function is used to further split and parse the table content, where <w:tbl>Represents a table, <w:tr>Indicates one of the lines, <w:tc>Represents a cell.
[0065] Table 1. Energy Monitoring Telemetry Content
[0066]
[0067]
[0068] Next, perform step S200, that is, vectorize the table information to obtain an embedding vector representing the table information.
[0069] In an exemplary embodiment of the present application, step S200, vectorizing the table information to obtain an embedding vector representing the table information, further includes:
[0070] S210: Format the header information and the body information to construct nested JSON-structured table data;
[0071] S220: Perform offline compression processing on the header information and the table context document information to obtain the context compression information of the table;
[0072] S230: Use a large language model to vectorize the JSON-structured table data and the context compression information to obtain an embedding vector representing the table information.
[0073] In an exemplary embodiment of the present application, in step S230, using a large language model to vectorize the JSON-structured table data and the context compression information to obtain an embedding vector representing the table information, further includes:
[0074] S231: Use a large language model to compress the semantic information of the JSON-structured table data and the context compression information to extract the central condensed expression of the table information;
[0075] S232: According to the central condensed expression of the table information and in combination with the original table information, perform mapping processing in a high-dimensional space to obtain an embedding vector representing the table information.
[0076] Specifically, please refer to Figure 2 as shown Figure 2 Shows the JSON structure table data generated for the table nesting situation. Search the table content in a left-to-right and top-to-bottom manner, and construct a dictionary structure with text, cell, and children as key values. Among them, the value associated with text is the character in the element field, the value associated with cell is a list containing all the element fields in a row, and children is a special structure. Since nested forms may appear in the actual table, in order to express this nested data structure, the children field is needed to recursively search for text and cell in the fields containing children.
[0077] Please refer to Figure 3 As shown, after successfully constructing the nested JSON structure, the table data is stored in a structured form. However, for the retrieval task, since the table data may contain overly long characters and sparse information expressions, it cannot be directly retrieved efficiently. In order to condense the core content of the table, offline information compression processing is performed on the table using the header information and the table context document information. Next, with the help of the information organization ability of the large language model, further compress and refine the key information of the table. Finally, using the method of natural language vectorization, map the central condensed expression of the table given by the large language model to a high-dimensional space in combination with the original table information, so as to obtain the embedding vector representing the table information.
[0078] Then, perform step S300, that is, compare and analyze the user's question with the embedding vector of the table information to filter out the table data that matches the user's question.
[0079] In an exemplary embodiment of the present application, in step S300, the comparing and analyzing the user's question with the embedding vector of the table information to filter out the table data that matches the user's question further includes:
[0080] S310: Calculate the similarity between the user's question and the embedding vector of the table information to obtain a similarity comparison result;
[0081] S320: According to the similarity comparison result, use the prompts technology to search to determine the table data that best matches the user's question.
[0082] In an exemplary embodiment of the present application, in step S310, the calculating the similarity between the user's question and the embedding vector of the table information to obtain a similarity comparison result further includes:
[0083] S311: Vectorize the user's question to obtain the embedding vector corresponding to the user's question;
[0084] S312: Calculate the similarity between the embedding vector corresponding to the user's question and the embedding vector of the table information to obtain a similarity comparison result.
[0085] In an exemplary embodiment of the present application, step S320, according to the similarity comparison result, use the prompts technology to search to determine the table data that best matches the user's question, further includes:
[0086] S321: If the similarity comparison result is similar, calculate the token length of the natural language description of the table;
[0087] S322: When the token length exceeds the maximum limit that the large language model can handle, split the table information into multiple parts to ensure that the token length of each part does not exceed the maximum acceptable value of the large language model;
[0088] S323: Use the large language model to perform information integration processing on each part of the split table information, and splice the processed table information of each part again to restore the integrity of the table information;
[0089] S324: Use the prompts technology to search to filter out the table data that best matches the user's question.
[0090] Specifically, please refer to Figure 4 As shown, the vectorized numerical values of the table information constructed offline not only include the global gist information of the table, but also achieve information compression of the table content through natural language vectorized embedding. However, due to the flexibility of the questions asked and the possible information loss after information compression, simply comparing the similarity between the questions asked and the vectorized table information often makes it difficult to accurately locate the specific results. In addition, considering the particularity of aerospace table information, after most tables are JSON-formatted, the token values will exceed the maximum value M that the large model can accept. Therefore, an iterative information integration method is adopted to perform progressive key information compression and effective information integration on the formatted table data. When the token length exceeds the maximum limit that the large language model can handle, the table information is split into multiple parts to ensure that the token length of each part does not exceed the maximum acceptable value M of the large language model. The large language model is used to perform information integration processing on each part of the split table information, and the processed table information of each part is re-stitched to restore the integrity of the table information. Subsequently, the prompts technology is used for searching, and the table data that best matches the user's question is screened out from the processed table information.
[0091] Finally, step S400 is executed to feedback the screened table data as the retrieval result to the user.
[0092] It should be noted that if the similarity result calculated in step S310 is dissimilar, an empty retrieval result will be feedback to the user.
[0093] In summary, the present invention provides a method for retrieving satellite document table data. By obtaining the table information in the satellite document, the table information includes header information, table body information, and table context document information, the table information is vectorized to obtain an embedding vector representing the table information, and the user's question is compared and analyzed with the embedding vector of the table information to screen out the table data that matches the user's question, and the screened table data is feedback to the user as the retrieval result. The method decodes the table content according to the document XML format and recursively constructs a JSON formatting structure for complex nested tables to accurately describe the information content in the table. Subsequently, using the header information, table body information, and table context document information, combined with the Prompts technology of the large language model, the foregoing content is condensed into core information and compressed into high-dimensional vectorized numerical values. Aiming at the problem that the number of tokens after table JSON formatting may exceed the processing upper limit of the large language model, an iterative information integration method is designed to ensure the accuracy of the retrieved content by progressively condensing and summarizing the table information.
[0094] Based on the same inventive concept, please refer to Figure 5 As shown, another embodiment of the present invention further provides a satellite document table data retrieval device 11, and the device includes:
[0095] An information acquisition module 111, configured to acquire table information in a satellite document, where the table information includes header information, table body information, and table context document material information;
[0096] An information processing module 112, configured to perform vectorization processing on the table information to obtain an embedding vector representing the table information;
[0097] An information matching module 113, configured to perform comparative analysis on a user's question and the embedding vector of the table information to screen out table data matching the user's question;
[0098] An information output module 114, configured to feedback the screened table data to the user as a retrieval result.
[0099] Based on the same inventive concept, please refer to Figure 6 As shown, another embodiment of the present invention further provides an electronic device 1. The electronic device 1 may include a memory 12, a processor 13, and a bus, and may further include a computer program stored in the memory 12 and executable on the processor 13, such as a satellite document table data retrieval program.
[0100] Among them, the memory 12 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. The memory 12 may be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 12 may also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 12 may also include both an internal storage unit and an external storage device of the electronic device 1. The memory 12 can be used not only to store application software installed on the electronic device 1 and various types of data, such as the code for satellite document table data retrieval, etc., but also to temporarily store data that has been output or will be output.
[0101] In some embodiments, the processor 13 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control core of the electronic device 1. It uses various interfaces and circuits to connect all components of the entire electronic device 1. By running or executing programs or modules (such as satellite document table data retrieval programs, etc.) stored in the memory 12, and by calling the data stored in the memory 12, it executes various functions of the electronic device 1 and processes data.
[0102] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above satellite document table data retrieval method.
[0103] Exemplarily, the computer program may be divided into one or more modules. The one or more modules are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into an information acquisition module 111, an information processing module 112, an information matching module 113, and an information output module 114.
[0104] The above integrated units implemented in the form of software function modules can be stored in a computer-readable storage medium. The computer-readable storage medium may be non-volatile or volatile. The above software function modules are stored in a storage medium and include several instructions to enable a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute some functions of the satellite document table data retrieval method described in various embodiments of this application.
[0105] The above embodiments only illustrate the principles and effects of the present invention by way of example, and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.< / w:tc> < / w:tr> < / w:tbl>
Claims
1. A satellite document table data retrieval method, characterized in that: include: Acquire table information in a satellite document, wherein the table information includes table header information, table body information, and table context document material information; Performing vectorization processing on the table information to obtain an embedding vector representing the table information; Comparing and analyzing the user's question with the embedding vector of the table information to filter out the table data that matches the user's question; The screened table data is fed back to the user as a search result.
2. The satellite document table data retrieval method according to claim 1, characterized in that: The step of obtaining the table information in the satellite document includes: Positioning the table according to the satellite document to determine the position of the table in the satellite document; According to the position of the table in the satellite document, the table header information, table body information and table context document data information of the table are extracted.
3. The satellite document table data retrieval method according to claim 1, characterized in that: The vectorizing the table information to obtain an embedding vector representing the table information includes: Formatting the table header information and the table body information to construct nested JSON structure table data; Offline compression processing is performed on the table header information and the table context document information to obtain context compression information of the table; The JSON structured table data and the contextual compression information are vectorized using a large language model to obtain an embedding vector representing the table information.
4. The satellite document table data retrieval method according to claim 3, characterized in that: The using of a large language model to vectorize the JSON structure table data and the context compression information to obtain an embedding vector representing the table information includes: Using a large language model to perform semantic information compression on the JSON structure table data and the context compression information, so as to extract a central condensed expression of the table information; According to the central condensed expression of the table information and in combination with the original table information, a mapping process of a high-dimensional space is performed to obtain an embedding vector representing the table information.
5. The satellite document table data retrieval method according to claim 1, characterized in that: The comparing and analyzing the embedded vectors of the user's question and the table information to screen out the table data matching the user's question includes: Calculate the similarity between the question asked by the user and the embedding vector of the table information to obtain a similarity comparison result; Based on the similarity comparison results, prompts technology is used to search to determine the table data that best matches the question asked by the user.
6. The satellite document table data retrieval method according to claim 5, characterized in that: The calculating the similarity between the user's question and the embedding vector of the table information to obtain a similarity comparison result includes: Vectorize the user's question to obtain an embedding vector corresponding to the user's question; The similarity between the embedding vector corresponding to the question asked by the user and the embedding vector of the table information is calculated to obtain a similarity comparison result.
7. The satellite document table data retrieval method according to claim 5, characterized in that: The method of searching based on the similarity comparison result using prompts technology to determine the table data that best matches the question asked by the user includes: If the similarity comparison result is similar, then calculate the token length of the natural language description of the table; When the token length exceeds the maximum limit that the large language model can handle, the table information is divided into multiple parts to ensure that the token length of each part does not exceed the maximum acceptable value of the large language model; Use the large language model to integrate the information of each part of the table after segmentation, and reassemble each part of the table information after processing to restore the integrity of the table information; Prompts technology is used to search and filter out the table data that best matches the question asked by the user.
8. A satellite document table data retrieval device, characterized in that: The device comprises: An information acquisition module is used to acquire table information in a satellite document, wherein the table information includes table header information, table body information and table context document material information; An information processing module, used for performing vectorization processing on the table information to obtain an embedded vector representing the table information; An information matching module, used for comparing and analyzing the questions asked by the user with the embedding vectors of the table information, so as to screen out the table data matching the questions asked by the user; The information output module is used to feed back the screened table data as search results to the user.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the satellite document table data retrieval method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the satellite document table data retrieval method according to any one of claims 1 to 7.