Retrieval method and device and electronic equipment
By using the matching mechanism between search vectors and actual slice vectors in the search model, combined with text completion technology, the problem of low accuracy of search files caused by users entering different search information is solved, and higher search accuracy is achieved.
Patent Information
- Application Number
- CN202411945508.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing search model results in a low accuracy rate of search files when users enter different search information.
By obtaining the search vector corresponding to the search information, filtering and determining the actual slice vector matching the search vector in the pre-configured vector library, then querying the actual slice text in the database, and performing a complete operation on it to generate the complete text, and finally determining the target text based on the search information and the complete text.
The accuracy of the search file search of the search model is improved, so that the target text in the displayed search results is more in line with the needs of users.
Smart Images

Figure CN120067053A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to a retrieval method, apparatus, and electronic device. Background Art
[0002] Currently, with the development of enterprises, the amount of data stored in the enterprise database is also increasing. To facilitate the management of data in the database, historical retrieval information and the retrieval files corresponding to the retrieval information are collected, and a retrieval model can be established based on the collected historical retrieval information and the retrieval files corresponding to the retrieval information. Subsequently, a user can input corresponding retrieval information into the retrieval model to obtain the corresponding retrieval files.
[0003] However, when using the above retrieval model for data retrieval, since the retrieval information input by different users for retrieving the same retrieval file is different, the accuracy of the retrieval files actually retrieved by the retrieval model is relatively low.
[0004] Therefore, how to improve the accuracy of the retrieval files retrieved by the retrieval model has become an urgent problem to be solved. Summary of the Invention
[0005] To solve the above technical problems, the present disclosure provides a retrieval method, apparatus, and electronic device.
[0006] In a first aspect, the present disclosure provides a retrieval method, including: in response to a retrieval operation, obtaining retrieval information; based on the retrieval vector corresponding to the retrieval information, screening in a pre-configured vector library to determine at least one actual slice vector that matches the retrieval vector; based on the actual slice vector, querying in a pre-configured database to determine the actual slice text corresponding to the actual slice vector; performing a complement operation on the actual slice text to obtain a complement text; based on the retrieval information and the complement text, determining a target text; where the target text includes one or more of the complement text; and displaying a retrieval result including the target text and the retrieval information.
[0007] Second aspect, the present disclosure provides a retrieval device, including: a processing unit, configured to control an acquisition unit to acquire retrieval information in response to a retrieval operation; the processing unit is further configured to screen in a pre-configured vector library based on a retrieval vector corresponding to the retrieval information acquired by the acquisition unit, and determine at least one actual slice vector that matches the retrieval vector; the processing unit is further configured to query in a pre-configured database based on the actual slice vector, and determine an actual slice text corresponding to the actual slice vector; the processing unit is further configured to perform a completion operation on the actual slice text to obtain a completed text; the processing unit is further configured to determine a target text based on the retrieval information and the completed text; wherein, the target text includes one or more of the completed texts; the processing unit is further configured to control a display unit to display a retrieval result including the target text and the retrieval information.
[0008] Third aspect, the present invention provides an electronic device, characterized by including: a memory and a processor, the memory is configured to store a computer program; the processor is configured to cause the electronic device to implement the retrieval method according to any one of the first aspect when executing the computer program.
[0009] Fourth aspect, the present invention provides a computer-readable storage medium, including: a computer program stored on the computer-readable storage medium, and the computer program is executed by a controller to implement the retrieval method according to any one of the first aspect.
[0010] Fifth aspect, the present invention provides a computer program product, when the computer program product runs on a computer, it causes the computer to execute the retrieval method according to any one of the first aspect.
[0011] These aspects or other aspects of the present disclosure will be more clearly understood in the following description.
[0012] The technical solution provided by the present disclosure has the following advantages compared with the prior art:
[0013] As can be seen from the above, for the retrieval method provided by the present disclosure, when receiving a retrieval operation input by a user, by obtaining the retrieval information of the retrieval operation, and then based on the retrieval vector corresponding to the retrieval information, screening in a pre-configured vector library to determine at least one actual slice vector that matches the retrieval vector. Then, based on the actual slice vector, querying in a pre-configured database to determine the actual slice text corresponding to the actual slice vector; performing a completion operation on the actual slice text to obtain a completed text; then, based on the retrieval information and the completed text, determining a target text; for example, by calculating the similarity between the retrieval information and the completed text, and then using the completed text with a similarity greater than a similarity threshold as the target text, so that the target text displayed in the retrieval result better meets the user's needs. Then, displaying a retrieval result including the target text and the retrieval information. Since the target text in the displayed retrieval result better meets the user's needs, the accuracy of the retrieval documents retrieved by the retrieval model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure.
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 FIG. 1 schematically shows one of the flowcharts of a retrieval method provided in Embodiment 1 of the present disclosure;
[0017] Figure 2 FIG. 2 schematically shows another flowchart of a retrieval method provided in Embodiment 1 of the present disclosure;
[0018] Figure 3 FIG. 3 schematically shows yet another flowchart of a retrieval method provided in Embodiment 1 of the present disclosure;
[0019] Figure 4 FIG. 4 schematically shows still another flowchart of a retrieval method provided in Embodiment 1 of the present disclosure;
[0020] Figure 5 FIG. 5 schematically shows yet another flowchart of a retrieval method provided in Embodiment 1 of the present disclosure;
[0021] Figure 6 FIG. 6 schematically shows still another flowchart of a retrieval method provided in Embodiment 1 of the present disclosure;
[0022] Figure 7 The seventh schematic flow chart of a retrieval method provided in the first embodiment is exemplarily shown;
[0023] Figure 8 The eighth schematic flow chart of a retrieval method provided in the first embodiment is exemplarily shown;
[0024] Figure 9 The ninth schematic flow chart of a retrieval method provided in the first embodiment is exemplarily shown;
[0025] Figure 10 The tenth schematic flow chart of a retrieval method provided in the first embodiment is exemplarily shown;
[0026] Figure 11 The eleventh schematic flow chart of a retrieval method provided in the first embodiment is exemplarily shown;
[0027] Figure 12 The twelfth schematic flow chart of a retrieval method provided in the first embodiment is exemplarily shown;
[0028] Figure 13 The thirteenth schematic flow chart of a retrieval method provided in the first embodiment is exemplarily shown;
[0029] Figure 14 The fourteenth schematic flow chart of a retrieval method provided in the first embodiment is exemplarily shown;
[0030] Figure 15 The schematic structural diagram of a retrieval device provided in the second embodiment is exemplarily shown;
[0031] Figure 16 The schematic structural diagram of an electronic device provided in the second embodiment is exemplarily shown. Detailed implementation manners
[0032] In order to more clearly understand the above objects, features, and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0033] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.
[0034] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0035] Embodiment 1
[0036] Figure 1 The flowchart of the retrieval method is exemplarily shown. The execution subject of this example can be a server, such as a television set, as Figure 1 shown, the method includes:
[0037] S11. In response to the retrieval operation, obtain the retrieval information.
[0038] In some examples, after the client establishes a communication connection with the server, the user can input the retrieval information to be retrieved on the client. At this time, the client responds to the retrieval operation and sends the retrieval information to the server. Then, the server can generate a corresponding retrieval result based on the retrieval information and send the retrieval result to the client for display. In this way, the user can learn the information to be retrieved based on the retrieval result.
[0039] In some examples, the user can input the retrieval information to be retrieved on the server. At this time, the server responds to the retrieval information in the retrieval operation, generates a corresponding retrieval result, and displays the retrieval result. In this way, the user can learn the information to be retrieved based on the retrieval result.
[0040] S12. Based on the retrieval vector corresponding to the retrieval information, filter in the pre-configured vector library to determine at least one actual slice vector that matches the retrieval vector.
[0041] In some examples, multiple target files are pre-stored in the memory of the server; among them, the target files include one or more of files in PDF text format, files in word text format, files in txt text format, files in doc text format, and files in docx text format. After that, by performing structured parsing on the target files, at least one theoretical slice text is obtained; the theoretical slice text is vector-converted to obtain a theoretical slice vector; where one theoretical slice vector corresponds to one vector identifier; based on the hierarchical relationship of the theoretical slice text, a node tree diagram is constructed; where one theoretical slice text corresponds to one node position, and the node position belongs to the nodes in the node tree diagram; the theoretical slice vectors are stored in a vector library, and the theoretical slice text and the corresponding node tree diagram of the theoretical slice text are stored in a database.
[0042] In some examples, the node tree diagram can be generated based on reference relationships, paragraph relationships, etc. in the target file.
[0043] Exemplarily, the content included in the target file is as follows:
[0044] 1.AA
[0045] 1.1bbb
[0046] Cccccccc.
[0047] 1.2dd
[0048] Xxxxx.
[0049] 1.2.1xxx
[0050] Qqqq.
[0051] It can be seen that by performing structured parsing on the target file, at least one theoretical slice text is obtained, such as: theoretical slice text (1, AA), theoretical slice text (1.1bbb), theoretical slice text (Cccccccc.), theoretical slice text (1.2dd), theoretical slice text (Xxxxx.), theoretical slice text (1.2.1xxx), theoretical slice text (Qqqq.). Since there are reference relationships between the various theoretical slice texts in the target file, the corresponding node tree diagram can be generated based on the reference relationships between the theoretical slice texts.
[0052] Exemplarily, the content included in the target file is as follows:
[0053] Bbbb, cccc. Aaa, dddd.
[0054] It can be seen that by performing structured parsing on the target file, at least one theoretical slice text is obtained, such as: theoretical slice text (Bbbb, cccc.), theoretical slice text (Aaa, dddd.). Since there is a paragraph relationship (the sequential relationship of sentences in the same paragraph) among the various theoretical slice texts in the target file, a corresponding node tree diagram can be generated based on the paragraph relationship among the theoretical slice texts.
[0055] Exemplarily, the content included in the target file is as follows:
[0056] AAA BBB ccc ddd
[0057] It can be seen that by performing structured parsing on the target file, at least one theoretical slice text is obtained, such as: theoretical slice text (AAA), theoretical slice text (BBB), theoretical slice text (ccc), theoretical slice text (ddd). Since there is a corresponding positional relationship among the various theoretical slice texts in the target file in the table, a corresponding node tree diagram can be generated based on the positional relationship among the theoretical slice texts.
[0058] In some examples, after the server obtains the retrieval information, it can perform vector conversion on the retrieval information to obtain a retrieval vector, so that at least one actual slice vector matching the retrieval vector can be determined by screening in a pre-configured vector library based on the retrieval vector corresponding to the retrieval information.
[0059] In some examples, when determining at least one actual slice vector matching the retrieval vector, the first similarity (such as: cosine similarity, Euclidean distance, etc.) between the retrieval vector and the theoretical slice vectors stored in the vector library can be calculated. Then, the theoretical slice vectors with the first similarity greater than the first similarity threshold are used as at least one actual slice vector matching the retrieval vector.
[0060] S13. Query in a pre-configured database based on the actual slice vector to determine the actual slice text corresponding to the actual slice vector.
[0061] In some examples, the theoretical slice text stored in the database corresponds to a vector identifier. The theoretical slice vectors stored in the vector library correspond to a vector identifier. Then, since an actual slice vector corresponds to a theoretical slice vector and a theoretical slice vector corresponds to a vector identifier. In this way, the theoretical slice text with the same vector identifier can be determined by querying in a pre-configured database based on the vector identifier of the theoretical slice vector corresponding to the actual slice vector. Then, the theoretical slice text is used as the actual slice text corresponding to the actual slice vector.
[0062] S14. Perform a complement operation on the actual slice text to obtain a complemented text.
[0063] In some examples, since the actual sliced text is scattered data, in order to make the obtained retrieval results more complete, a complement operation is performed on the actual sliced text, such as supplementing data based on the context information in the target file where the actual sliced text is located, so as to obtain a complemented text. Since the obtained complemented text is supplemented with data based on the context information in the target file where the actual sliced text is located, the obtained complemented text is more complete, thereby improving the readability of the retrieval results.
[0064] S15. Determine the target text based on the retrieval information and the complemented text. Among them, the target text includes one or more items in the complemented text.
[0065] In some examples, when determining the target text, the second similarity between the retrieval information and the complemented text can be calculated; for example: perform vector conversion on the complemented text to obtain a complemented vector. Then, calculate the similarity between the complemented vector and the retrieval vector corresponding to the retrieval information (such as: cosine similarity, Euclidean distance, etc.). Then, use the similarity between the complemented vector and the retrieval vector corresponding to the retrieval information as the second similarity. Then, use the complemented text with the second similarity greater than the second similarity threshold as the target text.
[0066] In some examples, the retrieval information and the complemented text can be input into a similarity model for similarity calculation to obtain the second similarity between the retrieval information and the complemented text. Among them, the training process of the similarity model includes:
[0067] Obtain training sample data and the labeled results of the training sample data; among them, the training sample data includes historical retrieval information and historical supplementary text, and the labeled results include the corresponding similarity between the historical retrieval information and the historical supplementary text.
[0068] Input the training sample data into a neural network model for learning to obtain the prediction results of the neural network model for the training sample data.
[0069] Based on the prediction results and the labeled results, adjust the network parameters of the neural network model until the neural network model converges to obtain a similarity model.
[0070] S16. Display the retrieval results including the target text and the retrieval information.
[0071] As can be seen from the above, for the retrieval method provided by the embodiments of the present disclosure, when receiving a retrieval operation input by a user, by obtaining the retrieval information of the retrieval operation, and then screening in a pre-configured vector library based on the retrieval vector corresponding to the retrieval information, at least one actual slice vector matching the retrieval vector is determined. Thus, when different users retrieve the same retrieval file, although the retrieval information input by the users may be different, by converting the retrieval information into a retrieval vector and screening in the pre-configured vector library based on the retrieval vector, at least one actual slice vector matching the retrieval vector can be screened. Then, based on the actual slice vector, a query is performed in the pre-configured database to determine the actual slice text corresponding to the actual slice vector; a completion operation is performed on the actual slice text to obtain a completed text; then, based on the retrieval information and the completed text, a target text is determined; for example, by calculating the similarity between the retrieval information and the completed text, and then using the completed text with a similarity greater than the similarity threshold as the target text, so that the target text displayed in the retrieval result better meets the needs of the user. Then, the retrieval result including the target text and the retrieval information is displayed. Since the target text in the displayed retrieval result better meets the needs of the user, the accuracy of the retrieval files retrieved by the retrieval model can be improved.
[0072] In some feasible examples, in combination with Figure 1 , as Figure 2 shown, before executing S11, the retrieval method provided by the embodiments of the present disclosure further includes S17 - S21.
[0073] S17. In response to a file upload operation, obtain the uploaded target file.
[0074] S18. Perform a structured parsing on the target file to obtain at least one theoretical slice text.
[0075] S19. Perform a vector conversion on the theoretical slice text to obtain a theoretical slice vector. Among them, one theoretical slice vector corresponds to one vector identifier.
[0076] S20. Based on the hierarchical relationship of the theoretical slice text, construct a node tree diagram. Among them, one theoretical slice text corresponds to one node position, and the node position belongs to the nodes in the node tree diagram.
[0077] S21. Store the theoretical slice vector in the vector library, and store the theoretical slice text and the node tree diagram corresponding to the theoretical slice text in the database.
[0078] In some examples, by performing a structured parsing on the target file, the theoretical slice vector and the hierarchical information of the theoretical slice vector can be obtained. A node tree diagram is constructed through the hierarchical information, which is mainly divided into text nodes and table nodes, and the relationships between the nodes are recorded, including before, after, parent node, child node, left, and right.
[0079] In some examples, the user can upload the required target file to the server. After that, after receiving the uploaded target file, the server performs a structured parsing on the target file to obtain at least one theoretical slice text. Then, a vector transformation is performed on the theoretical slice text to obtain a theoretical slice vector. A node tree diagram is constructed based on the hierarchical relationship of the theoretical slice text. And the theoretical slice vector is stored in the vector library, and the theoretical slice text and the corresponding node tree diagram of the theoretical slice text are stored in the database.
[0080] As can be seen from the above, the retrieval method provided by the embodiments of the present disclosure pre-stores the theoretical slice text, the theoretical slice vector, and the node tree diagram of at least one uploaded target file in the memory of the server. Thus, when the server performs a retrieval, it can find the target text that matches the retrieval information based on the theoretical slice text, the theoretical slice vector, and the node tree diagram stored in the memory, ensuring the retrieval accuracy of the retrieval.
[0081] In some feasible examples, in combination with Figure 2 , as Figure 3 shown, the above S12 can be specifically implemented by the following S120 and S121.
[0082] S120. Calculate the first similarity between the retrieval vector corresponding to the retrieval information and each theoretical slice vector in the pre-configured vector library.
[0083] S121. Use the theoretical slice vectors with the first similarity greater than the first similarity threshold as at least one actual slice vector that matches the retrieval vector.
[0084] As can be seen from the above, the retrieval method provided by the embodiments of the present disclosure calculates the first similarity between the retrieval vector corresponding to the retrieval information and each theoretical slice vector in the pre-configured vector library. Furthermore, based on the magnitude relationship between the first similarity and the first similarity threshold, the theoretical slice vectors are screened, so that the unnecessary theoretical slice vectors can be eliminated, reducing the occupation of computing resources and improving the user experience.
[0085] In some feasible examples, in combination with Figure 2 , as Figure 4 shown, the above S13 can be specifically implemented by the following S130 and S131.
[0086] S130. Query in a pre-configured database based on the vector identifier of the theoretical slice vector corresponding to the actual slice vector to determine the theoretical slice text identical to the vector identifier.
[0087] S131. Use the theoretical slice text as the actual slice text corresponding to the actual slice vector.
[0088] As can be seen from the above, in the retrieval method provided by the embodiments of the present disclosure, when obtaining the theoretical slice vector corresponding to the actual slice vector, since the corresponding content cannot be displayed by the theoretical slice vector, it is necessary to query in a pre-configured database based on the vector identifier of the theoretical slice vector to determine the theoretical slice text identical to the vector identifier. In this way, the theoretical slice text can be used as the actual slice text corresponding to the actual slice vector for display, thus ensuring the accuracy of the display.
[0089] In some feasible examples, in combination with Figure 1 , as Figure 5 shown, the above S14 can be specifically implemented through the following S140 - S142.
[0090] S140. Obtain the node position of the actual slice text and the text type of the actual slice text; wherein, the node position includes any one of the parent node and the child node, and the text type includes any one of character text and table text.
[0091] In some examples, the node position of the actual slice text is equal to the node position of the corresponding theoretical slice text in the corresponding node tree diagram of the actual slice position.
[0092] In some examples, when the character corresponding to the actual slice file is not a character in the table, it is determined that the text type of the node position is character text, and when the character corresponding to the actual slice file is a character in the table, it is determined that the text type of the node position is table text.
[0093] S141. Determine the character completion method based on the node position and the text type.
[0094] In some examples, the corresponding relationship between the node position and the text type and the character completion method is pre-stored in the memory of the server. After that, when the server determines the character completion method, it can read the corresponding relationship from the memory and query the corresponding relationship based on the node position and the text type, so as to determine the character completion method corresponding to the node position and the text type.
[0095] In some examples, the server determines whether there is a parent node and a child node at the node position based on the node position, and determines the character completion method based on whether there is a parent node and a child node at the node position and the text type.
[0096] S142. Character-complement the actual sliced text according to the character-completion method to obtain the complemented text.
[0097] As can be seen from the above, since the actual sliced text only records part of the content in the target file, when the user reads the actual sliced text, there will be a problem that the content recorded in the actual sliced text cannot be understood. Therefore, by character-complementing the actual sliced file according to the character-completion method, the complemented text is obtained. Since more content is recorded in the complemented text, the user can read better when reading the complemented text, improving the user experience.
[0098] In some implementable examples, the text type includes character text; combined Figure 5 , as Figure 6 shown, the above S141 can be specifically implemented through the following S1410 - S1413.
[0099] S1410. Based on the node position, determine the node tree diagram where the node position is located.
[0100] S1411. When the text type is character text and the node position has a parent node in the node tree diagram, obtain the node link corresponding to the node position.
[0101] S1412. Starting from the node position, in the order from bottom to top, obtain the first number of characters corresponding to the nth parent node and the second total number of characters corresponding to the (n - 1)th parent node.
[0102] In some examples, assume that the node link contains 3 parent nodes, namely parent node 1, parent node 2, and parent node 3. Parent node 1 is the parent node of the node position, parent node 2 is the parent node of parent node 1, and parent node 3 is the parent node of parent node 2. Then when calculating in the order from bottom to top starting from the node position, the 1st parent node is parent node 1, the 2nd parent node is parent node 2, and the 3rd parent node is parent node 3.
[0103] In some examples, the total number of characters corresponding to a parent node is equal to the sum of the total number of characters contained in the theoretical sliced text corresponding to the parent node and the total number of characters contained in all child nodes under the parent node. For example: There are 3 child nodes under parent node 3, namely child node 1, child node 2, and child node 3. Then the total number of characters corresponding to parent node 3 is equal to the sum of the total number of characters contained in the theoretical sliced text corresponding to parent node 3, the total number of characters contained in the theoretical sliced text corresponding to child node 1, the total number of characters contained in the theoretical sliced text corresponding to child node 2, and the total number of characters contained in the theoretical sliced text corresponding to child node 3.
[0104] In some examples, the total number of characters corresponding to a child node is equal to the total number of characters contained in the theoretical slice text corresponding to that child node.
[0105] S1413. When the first number of characters is greater than the maximum total number of characters and the second total number of characters is less than or equal to the maximum total number of characters, determining the character completion method includes: generating a completion text based on the theoretical slice text corresponding to the (n - 1)-th parent node and the theoretical slice texts corresponding to all child nodes under the (n - 1)-th parent node; wherein, the second total number of characters corresponding to the completion text is less than or equal to the maximum total number of characters.
[0106] In some examples, the process of obtaining the maximum total number of characters includes: obtaining the eighth total number of characters contained in the retrieval information. Based on the first maximum total number of characters contained in the pre-configured retrieval information, the ninth total number of characters contained in the pre-configured retrieval result that includes the target text, the second maximum total number of characters contained in the pre-configured target text, and the eighth total number, determining the maximum total number of characters.
[0107] In some examples, combining with the example given in S1412 above, the first number of characters of the 1st parent node is equal to the sum of the total number of characters contained in the theoretical slice text corresponding to parent node 1 and the total number of characters contained in all child nodes under parent node 1. The first number of characters of the 2nd parent node is equal to the sum of the total number of characters contained in the theoretical slice text corresponding to parent node 2, the first total number of characters, and the total number of characters contained in the theoretical slice texts corresponding to other child nodes under parent node 2 except parent node 1. The first number of characters of the 3rd parent node is equal to the sum of the total number of characters contained in the theoretical slice text corresponding to parent node 3, the second total number of characters, and the total number of characters contained in the theoretical slice texts corresponding to other child nodes under parent node 3 except parent node 2.
[0108] If the first number of characters of the 1st parent node is less than or equal to the maximum total number of characters and the first number of characters of the 2nd parent node is greater than the maximum total number of characters, determining the character completion method is to use the theoretical slice text corresponding to parent node 1 and the theoretical slice texts corresponding to all child nodes under parent node 1 as the completion text.
[0109] If the first number of characters of the 2nd parent node is less than or equal to the maximum total number of characters and the first number of characters of the 3rd parent node is greater than the maximum total number of characters, determining the character completion method is to use the theoretical slice text corresponding to parent node 2 and the theoretical slice texts corresponding to all child nodes under parent node 2 as the completion text.
[0110] In some examples, if the number of the first characters of the nth parent node is less than or equal to the maximum total number of characters, and the number of the first characters of the (n + 1)th parent node is greater than the maximum total number of characters, and the difference between the number of the first characters and the maximum total number of characters is greater than the word number threshold, obtain the total number of characters included in the theoretical slice text corresponding to each child node in the parent node n + 1. Then, in the order of the child nodes in the parent node n + 1, accumulate the total number of characters included in the theoretical slice text corresponding to the a-th child node to the number of the first characters in sequence to obtain the target total number. When the sum of the total number of characters included in the theoretical slice text corresponding to the a-th child node accumulated to the number of the first characters is less than or equal to the maximum total number of characters, and the sum of the total number of characters included in the theoretical slice text corresponding to the (a + 1)-th child node accumulated to the number of the first characters is greater than the maximum total number of characters, determine the character completion method as using the theoretical slice text corresponding to the first a child nodes in the parent node n + 1, the theoretical slice text corresponding to the parent node n, and the theoretical slice text corresponding to all child nodes under the parent node n as the completion text.
[0111] As can be seen from the above, since the actual slice text only records part of the content in the target file, when the user reads the actual slice text, there will be a problem that the content recorded in the actual slice text cannot be understood. For this reason, character completion is performed on the actual slice file according to the character completion method to obtain the completion text. Since more content is recorded in the completion text, when the user reads the completion text, better reading can be carried out, improving the user experience.
[0112] In some feasible examples, the text type includes character text; combined with Figure 5 , as Figure 7 shown, the above S141 can be specifically implemented by the following S1414.
[0113] S1414. When the text type is character text, and there are no child nodes and parent nodes at the node position in the node tree diagram, and the first total number of characters included in the actual slice text is less than or equal to the maximum total number of characters, determine that the character completion method includes: using the actual slice text as the completion text.
[0114] In some examples, since there are no child nodes and parent nodes at the node position in the node tree diagram, and the first total number of characters included in the actual slice text is less than or equal to the maximum total number of characters, this indicates that there are no characters that can be used for character completion at this time. Therefore, the actual slice text can be used as the completion text.
[0115] As described above, since the actual sliced text only records part of the content in the target file, when the user reads the actual sliced text, there will be a problem that the content recorded in the actual sliced text cannot be understood. For this reason, the actual sliced file is character-complemented according to the character complement method to obtain the complemented text. Since more content is recorded in the complemented text, the user can read better when reading the complemented text, improving the user experience.
[0116] In some implementable examples, the text type includes character text; combined Figure 7 , such as Figure 8 shown, the retrieval method provided by the embodiments of the present disclosure further includes S22 and S23.
[0117] S22. When the first total number is greater than the maximum character total number, the actual sliced text is sliced according to the order of characters in the actual sliced text to obtain a first text and a second text; wherein, the second total number of characters in the first text is equal to the maximum character total number, and the sum of the second total number and the third total number of characters in the second text is equal to the first total number.
[0118] S23. Determining the character complement method includes: using the first text as the complemented text.
[0119] In some examples, since the first total number is greater than the maximum character total number, it means that the content contained in the actual sliced text is too much. Therefore, it is necessary to slice the actual sliced text. When slicing, in order to ensure the continuity of reading, it is necessary to slice according to the order of characters in the actual sliced text, so as to obtain a first text with the total number of characters equal to the maximum character total number, and then use the first text as the complemented text.
[0120] As described above, since the actual sliced text only records part of the content in the target file, when the user reads the actual sliced text, there will be a problem that the content recorded in the actual sliced text cannot be understood. For this reason, the actual sliced file is character-complemented according to the character complement method to obtain the complemented text. Since more content is recorded in the complemented text, the user can read better when reading the complemented text, improving the user experience.
[0121] In some implementable examples, the text type includes character text; combined Figure 5 , such as Figure 9 shown, the above S141 can be specifically implemented by the following S1415 and S1416.
[0122] S1415. When the text type is character text, and the node position has no parent node in the node tree diagram and has child nodes in the node tree diagram, obtain the total number of second characters corresponding to the node position; wherein, the total number of second characters is equal to the sum of the fourth total number of characters included in the actual sliced text and the total number of characters included in the theoretical sliced texts corresponding to all child nodes under the node position.
[0123] S1416. When the total number of second characters is less than or equal to the maximum total number of characters, determine that the character completion method includes: generating a completed text based on the actual sliced text and the theoretical sliced texts corresponding to all child nodes under the node position.
[0124] As can be seen from the above, since the actual sliced text only records part of the content in the target file, when the user reads the actual sliced text, there will be a problem that the content recorded in the actual sliced text cannot be understood. For this reason, character completion is performed on the actual sliced file according to the character completion method to obtain a completed text. Since more content is recorded in the completed text, when the user reads the completed text, better reading can be carried out, improving the user experience.
[0125] In some feasible examples, the text type includes character text; combined Figure 9 , as Figure 10 shown, the retrieval method provided by the embodiments of the present disclosure further includes S24 - S26.
[0126] S24. When the total number of second characters is greater than the maximum total number of characters, obtain the node link corresponding to the node position.
[0127] S25. Starting from the node position, in sequence, obtain the total number of third characters corresponding to the (m - 1)th child node and the total number of fourth characters corresponding to the mth child node from the node link.
[0128] S26. When the difference between the fourth total number and the total number of fourth characters is greater than the maximum total number of characters, and the difference between the fourth total number and the total number of third characters is less than or equal to the maximum total number of characters, determine that the character completion method includes: generating a completed text based on all child nodes between the mth child node and the node position, and the actual sliced text.
[0129] In some examples, assume that there are 3 child nodes under the node position in the node link, namely child node 1, child node 2, and child node 3. Then, when calculating in sequence starting from the node position, the 1st child node is child node 1, the 2nd child node is child node 2, and the 3rd child node is child node 3.
[0130] In some examples, the first number of characters of the first child node is equal to the actual total number of characters contained in the theoretical sliced text corresponding to child node 1, the first number of characters of the second child node is equal to the total number of characters contained in the theoretical sliced text corresponding to child node 2, and the first number of characters of the third child node is equal to the total number of characters contained in the theoretical sliced text corresponding to child node 3.
[0131] When the sum of the fourth total and the actual total corresponding to child node 1 is less than or equal to the maximum total number of characters, and the sum of the fourth total and the actual total corresponding to child node 2 is greater than the maximum total number of characters, determine the character completion method as using the theoretical sliced text and the actual sliced text corresponding to child node 1 as the completion text.
[0132] When the sum of the fourth total, the actual total corresponding to child node 1, and the actual total corresponding to child node 2 is less than or equal to the maximum total number of characters, and the sum of the fourth total, the actual total corresponding to child node 1, the actual total corresponding to child node 2, and the actual total corresponding to child node 3 is greater than the maximum total number of characters, determine the character completion method as using the theoretical sliced text corresponding to child node 1, the theoretical sliced text corresponding to child node 2, and the actual sliced text as the completion text.
[0133] As can be seen from the above, since the actual sliced text only records part of the content in the target file, when the user reads the actual sliced text, there will be a problem of not being able to understand the content recorded in the actual sliced text. For this reason, character completion is performed on the actual sliced file according to the character completion method to obtain the completion text. Since more content is recorded in the completion text, the user can read better when reading the completion text, improving the user experience.
[0134] In some feasible examples, the text type includes table text; combined with Figure 5 , as Figure 11 shown, the above S141 can be specifically implemented through the following S1417 - S1419 - 1.
[0135] S1417. When the text type is table text and there is a parent node at the node position, obtain the node link corresponding to the node position;
[0136] S1418. Based on the node link, determine the fifth total number of characters contained in the table corresponding to the actual sliced text;
[0137] S1419. When the fifth total number of characters is less than the maximum total number of characters, starting from the node position, obtain the fifth number of characters corresponding to the xth parent node in the order from bottom to top; where the fifth number of characters is equal to the sum of the total number of characters contained in the theoretical sliced text corresponding to the xth parent node and the total number of characters contained in the theoretical sliced texts corresponding to all child nodes under the xth parent node, and x is an integer greater than or equal to 1.
[0138] S1419-1. When the number of the fifth characters is greater than the total number of the maximum characters, the character completion method is determined as follows: generating a completion text based on the theoretical slice text corresponding to the (x - 1)-th parent node and the theoretical slice texts corresponding to all child nodes under the (x - 1)-th parent node; wherein, the total number of the sixth characters corresponding to the completion text is less than or equal to the total number of the maximum characters.
[0139] As can be seen from the above, since the actual slice text only records part of the content in the target file, when the user reads the actual slice text, there will be a problem that the content recorded in the actual slice text cannot be understood. Therefore, the actual slice file is character-completed according to the character completion method to obtain a completion text. Since more content is recorded in the completion text, when the user reads the completion text, the user can read better and improve the user experience.
[0140] In some feasible examples, the text type includes character text; combined with Figure 11 , such as Figure 12 shown, the retrieval method provided by the embodiments of the present disclosure further includes S27 - S29.
[0141] S27. When the fifth total number is greater than the total number of the maximum characters, determining the table header corresponding to the identification node based on the table node corresponding to the node position;
[0142] S28. If the total number of the sixth characters contained in the table row corresponding to the table header is less than the total number of the maximum characters, centering on the table row corresponding to the table header, sequentially accumulating each corresponding theoretical slice text in each row of the table, and calculating the sum of the total number of the sixth characters and the total number of characters contained in each accumulated theoretical slice text to obtain a seventh total number;
[0143] Exemplarily, the table is shown in Table 1.
[0144] Table 1
[0145] 111 222 333 444 555 666
[0146] It can be seen that after the structured parsing of Table 1, 6 theoretical slice vectors will be obtained, namely theoretical slice vector (111), theoretical slice vector (222), theoretical slice vector (333), theoretical slice vector (444), and theoretical slice vector (555). When the actual slice vector is the theoretical slice vector (444), it can be determined that the header corresponding to the actual slice vector is the cell where the theoretical slice vector (333) is located. Then, since the sixth total number of characters contained in the table row where the theoretical slice vector (333) is located is less than the maximum total number of characters, taking the table row where the theoretical slice vector (333) is located as the center, the corresponding theoretical slice texts in each row of the table are accumulated in turn. For example, in the up and down order, that is, first accumulate each theoretical slice text in the row above the table row where the theoretical slice vector (333) is located. After accumulating all the theoretical slice texts in the row above the table row where the theoretical slice vector (333) is located, if the total number of accumulated characters is less than or equal to the maximum total number of characters, then continue to accumulate each theoretical slice text in the row below the table row where the theoretical slice vector (333) is located; after accumulating all the theoretical slice texts in the row below the table row where the theoretical slice vector (333) is located, if the total number of accumulated characters is less than or equal to the maximum total number of characters, then continue to accumulate all the theoretical slice texts in the row above the row above the table row where the theoretical slice vector (333) is located, and so on in a loop until the total number of accumulated characters is greater than the maximum total number of characters, then stop accumulating.
[0147] Exemplarily, taking the 1st theoretical slice text as the theoretical slice text corresponding to the theoretical slice vector (111), the 2nd theoretical slice text as the theoretical slice text corresponding to the theoretical slice vector (222), the 3rd theoretical slice text as the theoretical slice text corresponding to the theoretical slice vector (555), and the 4th theoretical slice text as the theoretical slice text corresponding to the theoretical slice vector (666) as examples, the specific implementation process is described as follows:
[0148] Calculate the sum of the sixth total number and the total number of characters contained in the theoretical slice text corresponding to the theoretical slice vector (111) to obtain the first seventh total number; and calculate the sum of the first seventh total number and the total number of characters contained in the theoretical slice text corresponding to the theoretical slice vector (222) to obtain the second seventh total number. If the first seventh total number is less than or equal to the maximum total number of characters, and the second seventh total number is greater than the maximum total number of characters, then take the theoretical slice text corresponding to the theoretical slice vector (111), the theoretical slice text corresponding to the theoretical slice vector (333), and the theoretical slice text corresponding to the theoretical slice vector (444) as the target text. Exemplarily, the target text is shown in Table 2.
[0149] Table 2
[0150] 111 333 444
[0151] If the second seventh total is less than or equal to the maximum character total, calculate the sum of the second seventh total and the total number of characters contained in the theoretical slice text corresponding to the theoretical slice vector (555) to obtain three seventh totals. If the second seventh total is less than or equal to the maximum character total and the third seventh total is greater than the maximum character total, then use the theoretical slice texts corresponding to the theoretical slice vectors (111), (333), and (444) as the target text. Exemplarily, the target text is shown in Table 3.
[0152] Table 3
[0153] 111 222 333 444
[0154] If the third seventh total is less than or equal to the maximum character total, calculate the sum of the third seventh total and the total number of characters contained in the theoretical slice text corresponding to the theoretical slice vector (666) to obtain four seventh totals. If the third seventh total is less than or equal to the maximum character total and the fourth seventh total is greater than the maximum character total, then use the theoretical slice texts corresponding to the theoretical slice vectors (111), (333), (444), and (555) as the target text. Exemplarily, the target text is shown in Table 4.
[0155] Table 4
[0156] 111 222 333 444 555
[0157] In some examples, when completing the table headers and rows, it is necessary to format the rows and columns in the completed table according to markdown to ensure the correctness of the output table.
[0158] S29. When the seventh total corresponding to accumulating the y-th theoretical slice text is less than or equal to the maximum character total, and the seventh total corresponding to accumulating the (y + 1)-th theoretical slice text is greater than or equal to the maximum character total, the character completion method is determined to include: generating a completion text based on the accumulated y theoretical slice texts and the actual slice text, where y is an integer greater than or equal to 1.
[0159] As described above, since the actual sliced text only records part of the content in the target file, when the user reads the actual sliced text, there will be a problem that the content recorded in the actual sliced text cannot be understood. Therefore, by performing character completion on the actual sliced file in a character completion manner, a completed text is obtained. Since the completed text records more content, when the user reads the completed text, they can read better and improve the user experience.
[0160] In some implementable examples, in combination with Figure 5 , such as Figure 13 shown, the retrieval method provided by the embodiments of the present disclosure further includes: S30 and S31.
[0161] S30. Obtain the eighth total number of characters included in the retrieval information.
[0162] S31. Based on the first maximum total number of characters included in the pre-configured retrieval information, the ninth total number of target texts included in the pre-configured retrieval result, the second maximum total number of characters included in the pre-configured target text, and the eighth total number, determine the maximum total number of characters.
[0163] In some examples, since the retrieval information is too long, it will cause the server to be unable to perform retrieval quickly. Therefore, it is necessary to limit the eighth total number of characters included in the input retrieval information, such as the eighth total number being less than or equal to the first maximum total number.
[0164] In some examples, in order to facilitate the user to consult, it is necessary to limit the first quantity of the target text displayed in the retrieval result, such as the first quantity being less than or equal to the ninth total number.
[0165] In some examples, when the number of characters included in the target text is too large, it will cause the user to not easily understand quickly when reading. Therefore, it is necessary to limit the second quantity of characters included in the target text, such as the second quantity being less than or equal to the second maximum total number. In some examples, the maximum total number of characters is equal to where q represents the first maximum total number, w represents the eighth total number, e represents the second maximum total number, and j represents the ninth total number.
[0166] As described above, the retrieval method provided by the embodiments of the present disclosure obtains the eighth total number of characters included in the retrieval information. Based on the first maximum total number of characters included in the pre-configured retrieval information, the ninth total number of target texts included in the pre-configured retrieval result, the second maximum total number of characters included in the pre-configured target text, and the eighth total number, determine the maximum total number of characters, so that the maximum total number of characters can be calculated in real time based on the total number of characters included in the retrieval information input by the user each time, thereby ensuring the accuracy of the obtained target text.
[0167] In some implementable examples, in combination with Figure 1 , such as Figure 14 shown, the above S15 can be specifically implemented by the following S150 and S151.
[0168] S150. Calculate the second similarity between the retrieved information and the supplemented text;
[0169] S151. Use the supplemented text with the second similarity greater than the second similarity threshold as the target text.
[0170] As can be seen from the above, since there may be differences between the supplemented text and the actual retrieved information, in order to avoid the situation where the relevance between the supplemented text retrieved by the user and the retrieved information is not high, the retrieval method provided by the embodiments of the present disclosure calculates the second similarity between the retrieved information and the supplemented text; and filters the supplemented text based on the magnitude relationship between the second similarity and the second similarity threshold, thereby ensuring the relevance between the supplemented text and the retrieved information and improving the user experience.
[0171] Embodiment 2
[0172] The structural schematic diagram of the retrieval device provided by the second embodiment of the present application is shown in Figure 15 shown. The retrieval device includes: a processing unit 201, an acquisition unit 202, and a display unit 203.
[0173] The processing unit 201 is configured to, in response to a retrieval operation, control the acquisition unit to acquire retrieval information; the processing unit 201 is further configured to screen in a pre-configured vector library based on the retrieval vector corresponding to the retrieval information acquired by the acquisition unit to determine at least one actual slice vector that matches the retrieval vector; the processing unit 201 is further configured to query in a pre-configured database based on the actual slice vector to determine the actual slice text corresponding to the actual slice vector; the processing unit 201 is further configured to perform a supplementation operation on the actual slice text to obtain a supplemented text; the processing unit 201 is further configured to determine a target text based on the retrieval information and the supplemented text; wherein the target text includes one or more of the supplemented texts; the processing unit 201 is further configured to control the display unit 203 to display a retrieval result including the target text and the retrieval information.
[0174] In some implementable examples, the processing unit 201 is further configured to control the acquisition unit to acquire the target file uploaded in response to a file upload operation; the processing unit 201 is further configured to perform a structured parsing on the target file acquired by the acquisition unit to obtain at least one theoretical slice text; the processing unit 201 is further configured to perform a vector conversion on the theoretical slice text to obtain a theoretical slice vector; wherein, one theoretical slice vector corresponds to one vector identifier; the processing unit 201 is further configured to construct a node tree diagram based on the hierarchical relationship of the theoretical slice texts; wherein, one theoretical slice text corresponds to one node position, and the node position belongs to a node in the node tree diagram; the processing unit 201 is further configured to store the theoretical slice vector in a vector library, and store the theoretical slice text and the node tree diagram corresponding to the theoretical slice text in a database.
[0175] In some implementable examples, the processing unit 201 is specifically configured to calculate a first similarity between the retrieval vector corresponding to the retrieval information acquired by the acquisition unit and each theoretical slice vector in a pre-configured vector library; the processing unit 201 is specifically configured to use the theoretical slice vectors with the first similarity greater than a first similarity threshold as at least one actual slice vector that matches the retrieval vector.
[0176] In some implementable examples, the processing unit 201 is specifically configured to query in a pre-configured database based on the vector identifier of the theoretical slice vector corresponding to the actual slice vector to determine the theoretical slice text with the same vector identifier; the processing unit 201 is specifically configured to use the theoretical slice text as the actual slice text corresponding to the actual slice vector.
[0177] In some implementable examples, the acquisition unit is specifically configured to acquire the node position of the actual slice text and the text type of the actual slice text; wherein, the node position includes any one of a parent node and a child node, and the text type includes any one of a character text and a table text; the processing unit 201 is specifically configured to determine a character completion method based on the node position and the text type acquired by the acquisition unit; the processing unit 201 is specifically configured to perform character completion on the actual slice text according to the character completion method to obtain a completed text.
[0178] In some implementable examples, the text type includes character text; the processing unit 201 is specifically configured to determine the node tree diagram where the node position is located based on the node position obtained by the obtaining unit; the processing unit 201 is specifically configured to, when the text type is character text and the node position has a parent node in the node tree diagram, obtain the node link corresponding to the node position; the processing unit 201 is specifically configured to, starting from the node position obtained by the obtaining unit and in the order from bottom to top, obtain the first number of characters corresponding to the nth parent node and the total number of the second characters corresponding to the (n - 1)th parent node; the processing unit 201 is specifically configured to, when the first number of characters is greater than the maximum total number of characters and the total number of the second characters is less than or equal to the maximum total number of characters, determine that the character completion method includes: generating a completion text based on the theoretical sliced text corresponding to the (n - 1)th parent node and the theoretical sliced texts corresponding to all child nodes under the (n - 1)th parent node; wherein, the total number of the second characters corresponding to the completion text is less than or equal to the maximum total number of characters.
[0179] In some implementable examples, the text type includes character text; the processing unit 201 is specifically configured to, when the text type is character text, the node position has no child nodes and parent nodes in the node tree diagram, and the first total number of characters included in the actual sliced text is less than or equal to the maximum total number of characters, determine that the character completion method includes: using the actual sliced text as the completion text.
[0180] In some implementable examples, the processing unit 201 is further configured to, when the first total number of characters is greater than the maximum total number of characters, split the actual sliced text in the order of the characters in the actual sliced text to obtain a first text and a second text; wherein, the second total number of characters included in the first text is equal to the maximum total number of characters, and the sum of the second total number and the third total number of characters included in the second text is equal to the first total number; the processing unit 201 is further configured to determine that the character completion method includes: using the first text as the completion text.
[0181] In some implementable examples, the text type includes character text; the processing unit 201 is specifically configured to, when the text type is character text, the node position has no parent node in the node tree diagram, and the node position has child nodes in the node tree diagram, obtain the total number of the second characters corresponding to the node position; wherein, the total number of the second characters is equal to the sum of the fourth total number of characters included in the actual sliced text and the total number of characters included in the theoretical sliced texts corresponding to all child nodes under the node position; the processing unit 201 is specifically configured to, when the total number of the second characters is less than or equal to the maximum total number of characters, determine that the character completion method includes: generating a completion text based on the actual sliced text and the theoretical sliced texts corresponding to all child nodes under the node position.
[0182] In some implementable examples, the processing unit 201 is further configured to, when the total number of the second characters is greater than the maximum total number of characters, obtain the node link corresponding to the node position; the processing unit 201 is further configured to, starting from the node position, sequentially obtain the total number of the third characters corresponding to the (m - 1)-th child node and the total number of the fourth characters corresponding to the m-th child node from the node link; the processing unit 201 is further configured to, when the difference between the fourth total number and the total number of the fourth characters is greater than the maximum total number of characters, and the difference between the fourth total number and the total number of the third characters is less than or equal to the maximum total number of characters, determine that the character completion method includes: generating a completion text based on all child nodes between the m-th child node and the node position, and the actual sliced text.
[0183] In some implementable examples, the text type includes tabular text; specifically, the processing unit 201 is configured to, when the text type is tabular text and there is a parent node at the node position, obtain the node link corresponding to the node position; the processing unit 201 is specifically configured to determine the total number of the fifth characters included in the table corresponding to the actual sliced text based on the node link; the processing unit 201 is specifically configured to, when the total number of the fifth characters is less than the maximum total number of characters, starting from the node position, obtain the number of the fifth characters corresponding to the x-th parent node in the order from bottom to top; wherein, the number of the fifth characters is equal to the sum of the total number of characters included in the theoretical sliced text corresponding to the x-th parent node and the total number of characters included in the theoretical sliced texts corresponding to all child nodes under the x-th parent node, and x is an integer greater than or equal to 1; the processing unit 201 is specifically configured to, when the number of the fifth characters is greater than the maximum total number of characters, determine that the character completion method includes: generating a completion text based on the theoretical sliced text corresponding to the (x - 1)-th parent node and the theoretical sliced texts corresponding to all child nodes under the (x - 1)-th parent node; wherein, the total number of the sixth characters corresponding to the completion text is less than or equal to the maximum total number of characters.
[0184] In some implementable examples, the processing unit 201 is further configured to, when the total number of the fifth characters is greater than the maximum total number of characters, determine the table header corresponding to the identification node based on the table node corresponding to the node position; the processing unit 201 is further configured to, if the total number of the sixth characters included in the table row corresponding to the table header is less than the maximum total number of characters, centered on the table row corresponding to the table header, sequentially accumulate each corresponding theoretical sliced text in the table, and calculate the sum of the total number of the sixth characters and the total number of characters included in each accumulated theoretical sliced text to obtain the total number of the seventh characters; the processing unit 201 is further configured to, when the total number of the seventh characters corresponding to the accumulation of the y-th theoretical sliced text is less than or equal to the maximum total number of characters, and the total number of the seventh characters corresponding to the accumulation of the (y + 1)-th theoretical sliced text is greater than or equal to the maximum total number of characters, determine that the character completion method includes: generating a completion text based on the accumulated y theoretical sliced texts and the actual sliced text, where y is an integer greater than or equal to 1.
[0185] In some implementable examples, the obtaining unit is further configured to obtain a fifth total number of characters included in the retrieval information; the processing unit 201 is further configured to determine the maximum total number of characters based on a first maximum total number of characters included in the pre-configured retrieval information, a sixth total number of target texts included in the pre-configured retrieval result, a second maximum total number of characters included in the pre-configured target text, and the fifth total number obtained by the obtaining unit.
[0186] In some implementable examples, the processing unit 201 is specifically configured to calculate a second similarity between the retrieval information and the supplemented text; the processing unit 201 is specifically configured to use the supplemented text with the second similarity greater than the second similarity threshold as the target text.
[0187] Among them, all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and its function will not be elaborated here.
[0188] Of course, the retrieval device provided by the embodiment of the present invention includes but is not limited to the above modules. For example, the retrieval device may further include a storage unit 204. The storage unit 204 can be used to store the program code of the retrieval device, and can also be used to store the data generated during the operation of the retrieval device, such as diagnostic data, etc.
[0189] Embodiment III
[0190] A schematic structural diagram of an electronic device provided by an embodiment of the present invention is as Figure 16 shown. The electronic device may include: at least one processor 51, a memory 52, a communication interface 53, a communication bus 54, and a display 55.
[0191] The following is a specific introduction to each component of the electronic device:
[0192] Among them, the processor 51 is the control center of the electronic device, which can be a single processor or a collective term for multiple processing elements. For example, the processor 51 is a central processing unit (CPU), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more DSPs, or one or more field programmable gate arrays (FPGAs).
[0193] In a specific implementation, as an example, the processor 51 may include one or more CPUs. For example, the CPU includes CPU0 and CPU1. Also, as an example, the electronic device may include multiple processors. For example, the CPU includes processor 51 and processor 55. Each of these processors may be a single-core processor (Single-CPU) or a multi-core processor (Multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0194] The memory 52 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or it may be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory 52 may exist independently and be connected to the processor 51 through the communication bus 54. The memory 52 may also be integrated with the processor 51.
[0195] In a specific implementation, the memory 52 is used to store the data in the present invention and execute the software program of the present invention. The processor 51 can execute various functions of the air conditioner by running or executing the software program stored in the memory 52 and calling the data stored in the memory 52.
[0196] The communication interface 53 uses any device such as a transceiver for communicating with other devices or communication networks, such as a radio access network (RAN), a wireless local area network (WLAN), a terminal, the cloud, etc. The communication interface 53 may include an acquisition unit to implement the acquisition function.
[0197] The communication bus 54 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used to represent it, but it does not mean that there is only one bus or one type of bus.
[0198] The display screen 55 is used to display images, videos, etc. The display screen 55 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc.
[0199] As an example, in combination with Figure 15 , the function implemented by the acquisition unit 202 of the retrieval device is the same as the function of the communication interface 53, the function implemented by the processing unit 201 in the retrieval device is the same as the function of the processor 51, the function implemented by the processing unit 203 in the retrieval device is the same as the function of the display 55, and the function implemented by the storage unit 204 in the retrieval device is the same as the function of the memory 54.
[0200] Embodiment 4
[0201] The embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method in any one of the embodiments.
[0202] Embodiment 5
[0203] In a fifth aspect, the present invention provides a computer program product, when the computer program product runs on a computer, it enables the computer to execute the method as described in any one of the embodiments.
[0204] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A search method, characterized in that: include: In response to the search operation, obtaining search information; Based on the search vector corresponding to the search information, screening is performed in a pre-configured vector library to determine at least one actual slice vector that matches the search vector; Based on the actual slice vector, query in a pre-configured database to determine the actual slice text corresponding to the actual slice vector; Performing a completion operation on the actual slice text to obtain a completed text; Determine a target text based on the search information and the completion text; wherein the target text includes one or more items of the completion text; The search result including the target text and the search information is displayed.
2. The search method according to claim 1, characterized in that: Before obtaining the search information in response to the search operation, the method further includes: In response to the file upload operation, obtaining the uploaded target file; Performing structural analysis on the target file to obtain at least one theoretical slice text; Performing vector conversion on the theoretical slice text to obtain a theoretical slice vector; wherein one theoretical slice vector corresponds to one vector identifier; Based on the hierarchical relationship of the theoretical slice text, a node tree diagram is constructed; wherein one theoretical slice text corresponds to one node position, and the node position belongs to a node in the node tree diagram; The theoretical slicing vector is stored in the vector library, and the theoretical slicing text and the node tree diagram corresponding to the theoretical slicing text are stored in the database.
3. The search method according to claim 2, characterized in that: The step of screening a pre-configured vector library based on a search vector corresponding to the search information to determine at least one actual slice vector matching the search vector includes: Calculating a first similarity between a search vector corresponding to the search information and each theoretical slice vector in a pre-configured vector library; The theoretical slice vector whose first similarity is greater than a first similarity threshold is used as at least one actual slice vector that matches the search vector.
4. The search method according to claim 2, characterized in that: The querying in a pre-configured database based on the actual slice vector to determine the actual slice text corresponding to the actual slice vector includes: Based on the vector identifier of the theoretical slice vector corresponding to the actual slice vector, a query is performed in a pre-configured database to determine the theoretical slice text identical to the vector identifier; The theoretical slice text is used as the actual slice text corresponding to the actual slice vector.
5. The search method according to claim 1, characterized in that: The performing a completion operation on the actual slice text to obtain a completed text includes: Obtaining a node position of the actual slice text and a text type of the actual slice text; wherein the node position includes any one of a parent node and a child node, and the text type includes any one of a character text and a table text; Determining a character completion mode based on the node position and the text type; The actual slice text is completed with characters according to the character completion method to obtain a completed text.
6. The search method according to claim 5, characterized in that: The text type includes character text; The determining of the character completion mode based on the node position and the text type includes: Based on the node position, determine a node tree diagram where the node position is located; When the text type is character text and the node position has a parent node in the node tree diagram, obtaining a node link corresponding to the node position; Taking the node position as the starting point, in order from bottom to top, obtain the number of first characters corresponding to the nth parent node and the total number of second characters corresponding to the n-1th parent node; When the first number of characters is greater than the maximum total number of characters, and the second total number of characters is less than or equal to the maximum total number of characters, determining the character completion method includes: generating a completion text based on the theoretical slice text corresponding to the n-1th parent node and the theoretical slice text corresponding to all child nodes under the n-1th parent node; wherein the second total number of characters corresponding to the completion text is less than or equal to the maximum total number of characters.
7. The search method according to claim 5, characterized in that: The text type includes character text; The determining of the character completion mode based on the node position and the text type includes: When the text type is character text, and the node position has no child nodes and parent nodes in the node tree diagram, and the first total number of characters contained in the actual slice text is less than or equal to the maximum total number of characters, determining the character completion method includes: using the actual slice text as the completion text.
8. The search method according to claim 7, characterized in that: The method further comprises: When the first total number is greater than the maximum total number of characters, segmenting the actual slice text according to the order of characters in the actual slice text to obtain a first text and a second text; wherein the second total number of characters in the first text is equal to the maximum total number of characters, and the sum of the second total number and the third total number of characters in the second text is equal to the first total number; Determining the character completion method includes: using the first text as the completion text.
9. The search method according to claim 5, characterized in that: The text type includes character text; The determining of the character completion mode based on the node position and the text type includes: When the text type is character text, and the node position does not have a parent node in the node tree diagram, and the node position has a child node in the node tree diagram, obtain a second total number of characters corresponding to the node position; wherein the second total number of characters is equal to the sum of a fourth total number of characters contained in the actual slice text and a total number of characters contained in theoretical slice texts corresponding to all child nodes under the node position; When the second total number of characters is less than or equal to the maximum total number of characters, determining the character completion method includes: generating a completion text based on the actual slice text and the theoretical slice text corresponding to all child nodes under the node position.
10. The search method according to claim 9, characterized in that: The method further comprises: When the second total number of characters is greater than the maximum total number of characters, obtaining a node link corresponding to the node position; Taking the node position as the starting point, in order, obtain the total number of third characters corresponding to the m-1th child node and the total number of fourth characters corresponding to the mth child node from the node link; When the difference between the fourth total and the fourth total number of characters is greater than the maximum total number of characters, and the difference between the fourth total number and the third total number of characters is less than or equal to the maximum total number of characters, determining the character completion method includes: generating a completion text based on all child nodes between the mth child node to the node position, and the actual slice text.