Information processing method and apparatus, device and medium
By establishing a terminology database and information database, the problem of low accuracy in machine translation of terminology in specific professional fields has been solved, resulting in more efficient information processing.
Patent Information
- Application Number
- PCT/CN2025/082442
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-17
- Filing Date
- 2025-03-13
- Publication Date
- 2025-12-26
AI Technical Summary
Existing machine translation technology has low accuracy when dealing with terminology and special vocabulary in specific professional fields, requiring manual proofreading, which increases time and cost.
Establish terminology and information databases, and improve the accuracy of information processing by matching keywords and providing reference information.
This improves the accuracy and effectiveness of information processing results and reduces the need for manual proofreading.
Smart Images

Figure CN2025082442_26122025_PF_FP_ABST
Abstract
Description
Methods, apparatus, equipment and media for information processing
[0001] This application claims priority to Chinese Patent Application No. 202410781053.5, filed on June 17, 2024, entitled "Method, Apparatus, Device and Medium for Information Processing", the entire contents of which are incorporated herein by reference. Technical Field
[0002] The exemplary embodiments disclosed herein relate generally to computer technology, and in particular to methods, apparatus, devices and computer-readable storage media for information processing. Background Technology
[0003] Today, with economic globalization and the development of the internet, people need to process various kinds of information from different languages in all aspects of their lives. Information between different languages can usually be translated manually or by machine. Human translation relies on professional translators to complete the task, which is costly in terms of both time and money, and is also difficult to handle real-time translation tasks. Machine translation results are usually not directly usable and still require human editing and proofreading, making it difficult to obtain accurate and efficient translations. Summary of the Invention
[0004] In a first aspect of this disclosure, a method for information processing is provided. In this method, in response to receiving user input, matching results are determined from a terminology database associated with the input content for at least one keyword in the content, the terminology database including multiple pairs of terms, each pair including a first term in a first language and a second term in a second language; information related to the content is determined from an information database associated with the content, the information database including at least one reference information item and its feature representation; and a processing result for the input content is provided based on the input content, the matching results, and the information.
[0005] In a second aspect of this disclosure, an apparatus for information processing is provided. The apparatus includes: a matching result determination module configured to, in response to receiving user input, determine matching results for at least one keyword in the content from a terminology library associated with the input content, the terminology library including multiple pairs of terms, each pair including a first term in a first language and a second term in a second language; an information determination module configured to determine information related to the content from an information library associated with the content, the information library including at least one reference information item and its feature representation; and a processing result generation module configured to provide a processing result for the input content based on the input content, the matching results, and the information. The apparatus also includes other modules configured to implement other steps in the above method.
[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processor.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, which, when executed by a processor, cause the processor to implement the method according to a first aspect of this disclosure.
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent in the following detailed description, taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0011] Figure 2 illustrates a schematic diagram of an information processing procedure according to some embodiments of the present disclosure;
[0012] Figure 3 shows a schematic diagram of an example structure of user input according to some embodiments of the present disclosure;
[0013] Figure 4 illustrates an exemplary flowchart of a process for constructing a terminology database according to some embodiments of the present disclosure;
[0014] Figure 5 shows an exemplary flowchart of the process for determining matching results according to some embodiments of the present disclosure;
[0015] Figure 6 illustrates a schematic diagram of determining content-related information according to some embodiments of this disclosure;
[0016] Figure 7 shows a schematic diagram of an example structure of the processing result according to some embodiments of the present disclosure;
[0017] Figure 8 shows a flowchart of a method for information processing according to some embodiments of the present disclosure;
[0018] Figure 9 shows a block diagram of an apparatus for information processing according to some embodiments of the present disclosure; and
[0019] Figure 10 shows a block diagram of an apparatus capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0020] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0021] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationship between various data. For example, the above-mentioned relationship can be obtained based on various technical solutions that are currently known and / or will be developed in the future.
[0022] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0023] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0024] For example, upon receiving user input, a prompt message is sent to the user to explicitly inform them that the operation performed in response to their input will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution.
[0025] As an optional but non-limiting embodiment, the method of sending a prompt message to the user in response to receiving user input can be, for example, a pop-up window, in which the prompt message can be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0026] It is understood that the above notification and user authorization acquisition process is merely illustrative and does not constitute a limitation on the embodiments of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the embodiments of this disclosure.
[0027] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
[0028] Traditionally, machine translation methods include rule-based translation systems (RBMT), statistical machine translation (SMT), and neural machine translation (NMT), among others.
[0029] Rule-based machine translation (RBC) refers to machine translation techniques that utilize linguistic rules and dictionaries to perform translations. Typical RBC methods include transfer-based, interlingual, and dictionary-based machine translation. While RBC offers advantages such as interpretability and lower complexity, its high cost stems from the need to pre-set numerous linguistic rules. Furthermore, because RBC relies on dictionary information, its accuracy in translating the same word across different contexts is relatively low.
[0030] Statistical machine translation (SMT) primarily involves analyzing large amounts of bilingual text to obtain translation rules and patterns, and then translating the source text into the target language based on these rules and patterns. The effectiveness of SMT heavily relies on the construction of these translation rules and patterns. However, SMT translation rules often struggle to utilize contextual information, making reordering difficult during the translation process. Furthermore, because the translation rules and patterns are built upon general translation texts from common domains, their accuracy in translating domain-specific terminology and vocabulary is relatively low.
[0031] Neural machine translation generates translation results based on trained neural network models. The accuracy of neural machine translation highly depends on the amount of training samples. For languages with less available training data, neural machine translation usually performs poorly. This is because neural network models generally require a large amount of data to learn effective features and patterns. For example, compared with the large amount of training data in the general domain, the training data in specific professional fields such as biology, medicine, and industry are less, which makes the translation results provided by the neural network of neural machine translation less accurate.
[0032] Specifically, in fields such as law, medicine, engineering, or finance, the translation results corresponding to certain terms or technical vocabulary are different from those in the general domain. The labeled data received by machine learning models during training usually comes from the general domain. When the input includes terms or technical vocabulary in these fields, the accuracy of machine learning models in processing information from these fields is relatively low. For example, in the field of energy engineering, there is a term "white coal" which refers to water resources used for hydroelectric power generation. In a conventional machine learning model, due to the presence of "white" and "coal" in the term, it may translate this term as white coal, deviating from the original meaning in the text and resulting in a low translation accuracy. In the business field, "tender" refers to "bid" in commercial activities in some cases; in the financial field, "tender" refers to "offer to purchase (such as stocks, etc.)" in some cases; while in the general domain, "tender" means "gentle". For such polysemous words, machine learning models may not be able to provide accurate translation results and usually require more information to determine their meanings in the text.
[0033] When the input contains terms in specific fields, rare special vocabulary, or rapidly changing external information, machine learning models lack an effective term translation mechanism and may result in low-quality translation results. As a result, it may be necessary to further use human translation or machine translation for proofreading and modification, increasing the time and cost of information processing, which are all undesirable.
[0034] To at least partially solve the above problems and other potential problems, the present disclosure proposes a solution for information processing. The solution of the present disclosure improves the accuracy and effectiveness of information processing results (such as translation results) by establishing a term library and an information library to perform term matching on the input and providing reference information for machine learning models.
[0035] For ease of explanation, the environment of Figure 1 will be used as an example for discussion below. Specifically, Figure 1 shows a schematic diagram of an example environment 100 in which the embodiments of this disclosure can be implemented. In this example environment 100, user 110 can provide input 120 to information processing device 101. After receiving the input 120 from the user, information processing device 101 provides it to machine learning model 130 and provides the processing result of machine learning model 130, that is, output 140.
[0036] In some exemplary embodiments, the domain of the content of input 120 can encompass many fields such as law, medicine, engineering, or finance. The form of the content of input 120 can include text, images, audio, video, etc.
[0037] In some exemplary embodiments, the machine learning model 130 may be trained using samples performing different tasks and may perform a variety of tasks, such as translation, semantic analysis, text processing, etc. Furthermore, the machine learning model 130 may be configured to implement any or all of the techniques disclosed herein, and the embodiments of this disclosure do not impose any limitations on this.
[0038] For example, input 120 may include a translation task and a corresponding text to be translated. Information processing device 101 may use machine learning model 130 to translate the text to be translated in input 120 and provide a corresponding translation result, i.e., output 140.
[0039] Specifically, in response to receiving user input 120, machine learning model 130 determines matching results for at least one keyword in the content from a terminology database associated with the content of input 120. Furthermore, machine learning model 130 determines information related to the content from an information database associated with it, thereby obtaining domain-specific information related to the content, such as contextual information, descriptive information, etc. Then, based on the content of input 140, the aforementioned matching results, and the determined information, machine learning model 130 provides a processing result for the content of input 140 as output 150.
[0040] The technical solution of this disclosure is further described below with reference to Figures 2 to 9.
[0041] Figure 2 illustrates a schematic diagram of a process 200 for information processing according to some exemplary embodiments of the present disclosure. Process 200 may be performed, for example, by the information processing device 101 in Figure 1 or other suitable device. Hereinafter, process 200 will be described with reference to Figure 1 for illustrative purposes.
[0042] As shown in Figure 2, in response to receiving user input 210, information processing device 101 determines a matching result 222 for at least one keyword in the content from a terminology database 220 associated with the content of input 210. The terminology database 220 includes multiple pairs of terms, each pair including a first term in a first language and a second term in a second language. Furthermore, information processing device 101 also determines content-related information 232 from an information database 230 associated with the content, the information database 230 including at least one reference information item and its feature representation. Then, based on the content of input 210, the matching result 222, and the information 232, information processing device 101 provides a processing result 240 for the content of input 210.
[0043] By utilizing exemplary embodiments of this disclosure, a pre-built terminology database and information database can be used to generate processing results by determining matching results and information associated with the input content, thereby improving the effectiveness of terms and reference information in the information processing results and thus enhancing the accuracy of information processing.
[0044] The technical solutions according to the embodiments of this disclosure can be applied to information processing scenarios involving professional technical fields. For example, when the input content involves a specific field, such as customer service records in e-commerce, the solution according to the embodiments of this disclosure can determine the field involved in the input content based on specific instructions in the input. Furthermore, the solution according to the embodiments of this disclosure can also automatically determine the field involved in the input content based on the input content, and then perform information processing using a pre-set terminology database and information database related to that field.
[0045] Referring again to Figure 2, in one example, information processing device 101 searches terminology database 220 based on the content of input 210 to obtain matching results for keywords in the content. As mentioned above, terminology database 220 may include multiple pairs of terms, each pair including a first term in a first language and a second term in a second language. It should be understood that the first language in this document may be the language corresponding to the content of input 120 (hereinafter referred to as the "input language"), and the second language may be the language corresponding to the content of output 140 (hereinafter referred to as the "output language"). Alternatively, the first language in this document may be the output language, and the second language may be the input language.
[0046] In some embodiments, during the process of determining the matching result, the input content can be segmented to obtain at least one keyword, which corresponds to a first language. Then, if a first term matching at least one keyword is found in the terminology database, a second term corresponding to the first term can be included in the matching result. In this case, the matching result implicitly indicates that the keyword can be found in the terminology database; that is, the matching result is an implicit hit indication. Alternatively, if no first term matching at least one keyword is found in the terminology database, a miss indication can be included in the matching result. The miss indication indicates that the keyword in the input content cannot be found in the terminology database.
[0047] In the keyword matching process described above, the word segmentation of the content of input 210 includes dividing the text into multiple words. This can be achieved, for example, through word segmentation methods based on string matching, understanding-based word segmentation, and statistical word segmentation. It should be understood that other word segmentation methods are also available in other embodiments, and this disclosure does not impose any limitations on them. After word segmentation, one or more keywords are extracted from the content of input 120. These keywords may include keywords in the first language or keywords in the second language.
[0048] The word segmentation process will be described in further detail with reference to the embodiments shown in Figure 3 below. Figure 3 shows a schematic diagram of an example input structure 300 according to some embodiments of the present disclosure. As shown in Figure 3, the example structure 300 includes the content of input 210. The content of input 210 includes content related to e-commerce customer service chat records. During the word segmentation process, for the content of input 210, each sentence in the text is divided into multiple keywords. For example, "My left earbud is not working" can be segmented into the following keywords after word segmentation: "my", "earphone", "left", "earbud", "not", "working", "is not working".
[0049] It should be understood that, although the above example of Chinese as a natural language was used to describe the specific task of text classification, alternatively and / or additionally, text can be written in other languages such as English, French, Japanese, etc.
[0050] Then, a search can be performed on the terminology database 220, which includes searching for each keyword in the segmented content within the terminology database 220. If a first term matching the keyword of the segmented content is found in the terminology database 220, the second term corresponding to the first term is included in the matching result.
[0051] Additionally, the terminology database 220 can also store other terminology information, such as the terminology's definition, context, usage status, usage instructions, grammatical information, synonyms, abbreviations, etc.
[0052] The terminology database 220 can be constructed in various ways. In some embodiments, the terminology database 220 can be constructed, for example, using a trie structure. In this tree structure, each node can represent a lexical unit, and the path from the root node to any given node can represent a complete term. Specifically, a first reference term can be determined first by segmenting the reference text in the first language, the first reference term comprising at least one terminology unit. Next, each terminology unit of the first reference term can be stored in a node of the trie structure. Then, a second reference term corresponding to the first reference term can be associatedly stored in the terminology database.
[0053] Table 1 shows an example of a prefix tree structure.
[0054] Table 1
[0055] In the data structure shown in Table 1, the prefix tree is instantiated from a GlossaryTrie, which contains a root node and a property that records the length of the tree. Each node, a glossaryTrieNode, stores its set of child nodes (next) and a list of terms (data) associated with those nodes.
[0056] Each term in the input 120 is inserted into the prefix tree using its word segmentation result as the key. If the current node does not have a corresponding child node, a new node is created and added to the tree, until the entire lexical unit of the term is completely added to the tree. In this way, the terminology base 220 (also known as the "dictionary") can be dynamically constructed.
[0057] Therefore, the embodiments of this disclosure utilize prefix tree technology to efficiently retrieve and match terms in a specific domain, specifically involving the construction of a terminology database, dynamic updates, and a call strategy in conjunction with a translation model, thereby significantly improving the retrieval efficiency of the terminology database.
[0058] It should be understood that although the terminology library 220 discussed above has a prefix tree structure, this is merely exemplary and not restrictive, and embodiments of this disclosure may also use terminology libraries with other structures.
[0059] The following is a schematic diagram of a process 400 for constructing a terminology database according to some embodiments of the present disclosure, with reference to FIG4. Specifically, the embodiment of FIG4 constructs a terminology database with a prefix tree structure. It should be understood that this is merely exemplary and is not intended to limit the embodiments of the present disclosure in any way. Process 400 may be performed, for example, by the information processing device 101 in FIG1 or other suitable device. Hereinafter, for illustrative purposes, process 400 will be described with reference to the embodiments of FIG1 and FIG2.
[0060] At 402, the root node of the prefix tree is initialized. Then, at 404, it is checked whether there is reference text input. In one example, the reference text can be terminology information text from various fields, such as a glossary. If no reference text input is detected at 404, the current process ends at 418. If reference text input is detected at 404, the reference text is segmented at 406 to determine the first reference term. In one example, segmenting the reference text can yield multiple first reference terms. When the segmented reference text includes multiple first reference terms, the following process is performed on each first reference term. At 408, among these reference terms, it is checked whether there are any first reference terms not stored in the prefix tree, that is, whether there is any newly added terminology information in the reference text.
[0061] In one example, the terminology database can also be updated based on reference text. At step 408, if a first reference term already stored in the prefix tree is detected, it continues to check whether the second reference term corresponding to the first reference term given in the reference text matches the second reference term stored in the terminology database. If they do not match, the second reference term stored in the terminology database is updated using the second reference term given in the reference text. In one example, the update time and historical versions can also be stored in the terminology database.
[0062] If a first reference term not stored in the prefix tree is detected at 408, then at 410 it is determined whether a term unit of the first reference term was not retrieved in the prefix tree and whether there are term units that have not been saved. At 410, term "retrieval" means searching in the child nodes of the current node in the prefix tree, for example, it means detecting in the child nodes of the root node.
[0063] Furthermore, the determination process of 410 is performed on a per-terminal basis for each term unit of the first reference term. For example, this could be performed on the current term unit of the current first reference term. For instance, if the first reference term is in English, then its term units are English letters. For example, if the first reference term is "model", its term units are "m", "o", "d", "e", and "l".
[0064] If at 410 it is determined that the current term unit of the first reference term is not retrieved in the child nodes of the current node in the prefix tree and that there are unsaved term units of the first reference term, then process 400 proceeds to box 412 via the "Yes" branch of 410, and creates a new child node at 412. Creating a new child node involves including the current term unit not detected in 410 in the newly created child node. For example, if the first reference term is "model", and at 410, the term unit "m" is not detected, then a new child node is created at 412 and the term unit "m" is included in that child node. Then, at 414, the process moves to the newly created child node. Then, it returns to 410, continues to retrieve the next term unit of the first reference term in the prefix tree, and continues the above process until the term unit of the first reference term is retrieved at 410, or all of its term units are saved.
[0065] If it is determined at 410 that a term unit of the first reference term is retrieved in the prefix tree, or that all of its term units are saved, then process 400 proceeds to 416 via the "No" branch of 410. At 416, the second reference term corresponding to the first reference term is stored in the terminal node to associate the second reference term corresponding to the first reference term with the terminology database.
[0066] In some embodiments, to save storage space or for other purposes, only term tags associated with the second reference term may be stored. These term tags have a mapping relationship with the second reference term, and the corresponding second reference term can be determined through the term tags. It should be understood that, for structural simplicity, the second reference term corresponding to the first reference term is only stored in the end node, while in other embodiments of this disclosure, the second reference term may be stored in other locations in the terminology library. The embodiments of this disclosure do not impose any limitations here.
[0067] In one example, at step 416, the length of the tree associated with the first reference term can also be stored in the end node to facilitate the retrieval of the first reference term during term matching. In another example, the first reference term and its corresponding second reference term can also be stored together in the end node. Furthermore, at step 416, other relevant information about the first reference term can be stored in the end node as supplementary information.
[0068] Then, process 400 repeats from 416 to 408. If no first reference term not stored in the prefix tree is detected at 408, that is, all first reference terms in the reference text are stored in the terminology database, then process 400 ends. In one example, the reference text is segmented at 406 to obtain multiple first reference terms, and the above process continues to be performed on the next first reference term until all first reference terms have been processed.
[0069] Optionally, in some embodiments, after the prefix tree structure is constructed, the length of the tree can also be stored in the terminology library 220.
[0070] Additionally, in some embodiments, one or more terminology databases may be constructed. For example, multiple terminology databases may be constructed for multiple fields respectively. This facilitates the management and use of the corresponding terminology databases. It should be understood that this is merely exemplary, and in other embodiments of this disclosure, a single terminology database may be constructed for multiple fields.
[0071] An exemplary flow of a process 500 for determining a matching result according to some embodiments of the present disclosure will now be described with reference to FIG. 5. FIG. 5 illustrates, for example, a process 500 for searching a terminology database constructed according to an embodiment of FIG. 4. Process 500 may be performed, for example, by the information processing device 101 of FIG. 1 or other suitable device. Hereinafter, for illustrative purposes, process 500 will be described with reference to embodiments of FIG. 1 and FIG. 2.
[0072] At step 502, the user input "120" is retrieved. Then, at step 504, the input is segmented to obtain keywords. In one example, segmenting the input yields multiple keywords. When the segmented input includes multiple keywords, the following process is performed on each keyword individually. At step 506, the search begins from the root node of the prefix tree. Next, at step 508, a search is performed at the root node to check for any unsearched keywords, thus avoiding duplicate keyword searches and improving processing efficiency. If no unsearched keywords are found at step 508, the current process ends at step 522.
[0073] If there are unsearched keywords at step 508, then retrieve those unsearched keywords at step 510. In one example, multiple keywords can be retrieved at step 510. When multiple keywords are retrieved, each keyword is processed as follows: At step 512, determine whether the keyword unit of the current keyword is a child node of the current node. The determination process at step 512 is performed on each term unit of the keyword, specifically checking the current keyword unit of the current keyword. When the keyword is in English, its keyword units are English letters. For example, when the keyword is "model", its keyword units are "m", "o", "d", "e", and "l".
[0074] If, at step 512, the current keyword unit is detected to be not in the child nodes of the current node (i.e., the current keyword is not included in the terminology database), then at step 514, a miss indication is included in the matching results. The current process then ends at step 522.
[0075] If a keyword unit of the keyword is detected as a child node of the current node at step 512, then at step 516, the process moves to that child node. For example, when the keyword is "model", if the first keyword unit "m" of the keyword is detected as a child node of the root node, then the process moves to that child node. Then at step 518, it checks whether the node contains a second reference term corresponding to the keyword. If the current node contains a second reference term corresponding to the keyword (i.e., the terminology database stores a term associated with the keyword), and the current node is the last keyword unit of the keyword, then at step 520, the second reference term is included in the matching result. In one example, other terminology information related to the second reference term can also be included in the matching result, such as the term's definition, context, usage status, usage instructions, grammatical information, synonyms, abbreviations, etc.
[0076] If no second reference term corresponding to the keyword is stored in node 518, then return to 512 and check if the keyword unit is a child node of the current node. Here, the keyword unit is the next untested keyword unit. For example, when the keyword is "model", if the first keyword unit "m" is detected as a child node of the root node, then move to that child node and check if no second reference term corresponding to "model" is stored there. Then return to 512 and check if the keyword unit "o" is a child node of the current node.
[0077] The above process continues until a second reference term corresponding to the keyword is detected in node 518, or the keyword unit for the keyword is detected in node 512 and is not in the child node of the current node. If the keyword unit for the keyword is detected in node 512 and is not in the child node of the current node, then a miss indication is included in the matching result in node 514, and the current process ends in node 522. If a second reference term corresponding to the keyword is detected in node 518, then the second reference term is included in the matching result in node 520, and the current process ends in node 522. In one example, the input content is segmented at node 504 to obtain multiple keywords, and the above process is continued for the next keyword until all keywords have been processed.
[0078] Referring again to Figure 2, as described above, the information base 230 includes at least one reference information item and its feature representation. In one example, a reference information item is, for example, an information entry, including text or information obtained from a specific data source. Such a reference information item can be constructed as a feature representation, such as a vector. In this way, the information base can be constructed as a domain-specific knowledge base, such as cross-border product translation, comment translation, or real-time conversation translation. In some embodiments, one or more information bases can be constructed, each associated with a specific domain. For example, multiple information bases can be constructed for multiple domains separately, or a single information base can be constructed for multiple domains.
[0079] In some embodiments, determining content-related information from the content-associated information database 230 can be done in various ways. For example, it can be based on matching the feature information of the content with the feature information of reference information items in the information database.
[0080] Specifically, first, a first feature representation of the content can be determined. Then, a second feature representation can be determined from the information database based on the first feature representation, the second feature representation corresponding to a reference information item in the information database. Finally, the reference information item corresponding to the second feature representation is identified as information related to the content.
[0081] In some embodiments, the information base 230 may be pre-built. In some embodiments, the information base 230 may be built based on text processing and feature extraction of pre-set reference information about one or more fields. For example, reference information about one or more fields may first be added to the information base, then the reference information may be processed to obtain reference information items, and then feature extraction may be performed on these reference information items to obtain feature representations of the reference information items.
[0082] Reference information items can be obtained by preprocessing the reference information. In some implementations, the text content of the reference information can be converted to a consistent format, such as converting PDF or HTML files into plain text. Syntax correction and synonym replacement can also be performed to ensure data quality.
[0083] Additionally, in some embodiments, the reference information may be segmented into paragraphs, and sentence boundary detection (SBD) may be applied to identify natural boundaries of the reference information and segment the text at appropriate locations to accommodate the input length limitations of the machine learning model.
[0084] An example process 600 for determining content-related information based on a database 230 according to some embodiments of the present disclosure will now be described with reference to FIG6.
[0085] First, a first feature representation 610 of the content of the user's input 210 can be determined based on the user's input 210. For example, multiple features representing the content can be extracted from the content, and a first vector including these multiple features is determined as the first feature representation 610. In one example, multiple features representing the content can be extracted from the content through text feature extraction, such as the bag-of-words model, TF-IDF (term frequency-inverse document frequency), or word embeddings. Alternatively, machine learning models such as BERT, RoBERTa, and Sentence-BERT can be used to convert the content into a high-dimensional vector representation, and these high-dimensional vectors can be used as the first feature representation 610.
[0086] In one example, before determining the first feature representation 610 of the content of input 210, the content of input 210 can be preprocessed. The text content of input 210 can be converted to a consistent format, such as converting a PDF or HTML file to plain text. Syntax correction, synonym replacement, and other methods can also be performed to ensure data quality.
[0087] In one example, the content of input 210 can also be segmented into paragraphs. Sentence Boundary Detection (SBD) is applied to identify natural boundaries in the content and to segment the text at appropriate locations to accommodate the input length limitations of the machine learning model.
[0088] Next, a second feature representation 620 is determined from the information database 230 based on the first feature representation 610. The second feature representation 620 corresponds to a reference information item 622 in the information database 230. Exemplarily, the similarity between the feature representations of the first feature representation 610 and at least one reference information item 622 in the information database 230 can be determined; and the second feature representation 620 is determined based on the similarity. In one example, a similarity threshold can be set, and a second feature representation 620 with a similarity greater than the similarity threshold to the first feature representation 610 can be searched in the information database 230. In another example, the similarity between the first feature representation 610 and the feature representations of all reference information items 622 in the information database 230 can be obtained and sorted, and the feature representation with the highest similarity can be determined as the second feature representation 620.
[0089] Then, the reference information item corresponding to the second feature representation is determined as content-related information. Since the reference information items and their feature representations are already stored in the information database, the corresponding reference information item can be found based on the second feature representation.
[0090] As shown in Figure 6, reference information item 622 can include reference information about one or more fields to provide background information in the relevant fields, thereby improving the accuracy of information processing results. For example, in an e-commerce customer service scenario, reference information item 622 as shown in Figure 6 includes the following: "The Lenovo Thinkplus LP40 headphones include touch control functionality and need to be fully charged before first use to ensure optimal performance. If the battery is completely depleted, the charging time is approximately 2 hours. It is recommended to check the charging indicator light; a steady light indicates a full charge. Lenovo Thinkplus LP40 specifications: The LP40 model is designed with an IPX5 waterproof rating and is equipped with noise cancellation, making it suitable for wearing during exercise." This reference information item 622 can be used to provide background information and can also be included in the output of information processing.
[0091] Based at least on the terminology database 220 and information database 230 discussed above, the information processing device 101 can output corresponding processing results (i.e., output 140). These processing results may include a variety of content, such as translation results for the content, prompts for providing to a machine learning model, and so on. It should be understood that this is merely exemplary and is not intended to limit the embodiments of this disclosure in any way.
[0092] Therefore, the embodiments of this disclosure combine Retrieval-Augmented Generation (RAG) technology to retrieve information from internal and / or external knowledge bases, effectively improving the accuracy and timeliness of translation results.
[0093] The following is a schematic diagram of an example structure of a processing result 700 according to some embodiments of the present disclosure, with reference to FIG7. As shown in FIG7, the processing result 700 may include a matching result 222, information 232, input 210, and processing result 240. In some embodiments, the processing result 700 may include a translation result, prompt information, etc.
[0094] Where the processing result includes prompts, the processing result 700 may also include a background description 710. The background description 710 refers to the overall context of the current task (e.g., a translation task), such as whether the task is translating a cross-border product title or an instant messaging conversation. The background description 710 may be part of the user's input 210 or may be generated by a machine learning model based on the content of input 210. For example, a translation result can be generated by inputting prompts into an information processing model.
[0095] In some embodiments, a background description 710 related to the content of input 120 can be obtained, and then a processing result 700 can be obtained based on the input content, the matching result, the relevant information determined from the information database 230, and the background description 710, such as generating a prompt message.
[0096] The processing result 700 can have various structures or forms. Table 2 shows an exemplary structure of a prompt message according to an embodiment of this disclosure. This prompt message can, for example, be a prompt provided to a machine learning model.
[0097] Table 2
[0098] Table 3 shows one form of the processing result 700 of Figure 7 that is specifically implemented based on the exemplary structure shown in Table 2.
[0099] Table 3
[0100] In one example, inputting the above prompt information into the information processing model can yield a processing result similar to Figure 7, 700.
[0101] Therefore, the embodiments of this disclosure provide a scheme for automatically constructing a structured Prompt involving context, terminology, and knowledge information. This scheme improves the output performance of the language model by adjusting and optimizing aspects such as terminology management, knowledge retrieval, and parameter configuration involved in the translation process, thereby automatically adjusting the Prompt structure according to different scenarios.
[0102] As discussed above, the embodiments of this disclosure provide a scenario-based dynamic selection system for knowledge bases and terminology databases. The solution of this disclosure can automatically select the most suitable knowledge base according to user needs or specific scenarios, thereby enhancing the relevance of translated content.
[0103] Furthermore, in some additional embodiments, the processing results can also be provided to the user through a pre-configured interface. For example, the processing results can be displayed to the user through, for instance, a human-computer interaction interface or any other means that can implement the embodiments of this disclosure. In this way, the user can conveniently and quickly understand the processing result 700, significantly improving the user experience.
[0104] Figure 8 shows a flowchart of a method 800 for information processing according to some embodiments of the present disclosure. Method 800 may be performed, for example, by the information processing device 101 in Figure 1 or other suitable device.
[0105] At 810, in response to receiving user input, information processing device 101 determines a matching result for at least one keyword in the content from a terminology database associated with the input content. The terminology database includes multiple pairs of terms, each pair including a first term in a first language and a second term in a second language. At 820, information processing device 101 determines information related to the content from an information database associated with the content. The information database includes at least one reference information item and its feature representation. At 830, information processing device 101 provides a processing result for the input content based on the input content, the matching result, and the information.
[0106] In some exemplary embodiments, determining a matching result includes: performing word segmentation on the input content to obtain at least one keyword, the at least one keyword corresponding to a first language; in response to finding a first term matching the at least one keyword in a terminology database, determining that the matching result includes a second term corresponding to the first term; and in response to not finding a first term matching the at least one keyword in the terminology database, determining that the matching result includes a miss indication.
[0107] In some exemplary embodiments, the terminology library has a prefix tree structure and is constructed as follows: by segmenting the reference text in the first language, a first reference term is determined, the first reference term including at least one term unit; each term unit of the first reference term is stored in a node of the prefix tree structure; and a second reference term corresponding to the first reference term is stored in the terminology library in association.
[0108] In some exemplary embodiments, determining content-related information from a content-associated information database includes: determining a first feature representation of the content; determining a second feature representation from the information database based on the first feature representation, the second feature representation corresponding to a reference information item in the information database; and determining the reference information item corresponding to the second feature representation as content-related information.
[0109] In some exemplary embodiments, determining the first feature representation includes: extracting multiple features representing the content from the content; and determining a first vector including the multiple features as the first feature representation.
[0110] In some exemplary embodiments, determining the second feature representation includes: determining the similarity between the first feature representation and the feature representation of at least one reference information item in the information database; and determining the second feature representation based on the similarity.
[0111] In some exemplary embodiments, the processing result includes at least one of the following: a translation result of the content, and prompting information for providing to the information processing model.
[0112] In some exemplary embodiments, providing the processing result includes generating a translation result by inputting prompt information into an information processing model.
[0113] In some exemplary embodiments, providing the processing result includes: obtaining a background description related to the input content; and generating a prompt message based on the input content, the matching result, the information, and the background description.
[0114] Figure 9 shows a block diagram of an apparatus 900 for information processing according to some embodiments of the present disclosure. The apparatus 900 may be implemented as or included in the information processing device 101 of Figure 1. The various modules / components in the apparatus 900 may be implemented by hardware, software, firmware, or any combination thereof.
[0115] As shown in the figure, the device 900 includes a matching result determination module 910, configured to determine, in response to receiving user input, a matching result for at least one keyword in the content from a terminology database associated with the input content, the terminology database including multiple pairs of terms, each pair including a first term in a first language and a second term in a second language; an information determination module 920, configured to determine information related to the content from an information database associated with the content, the information database including at least one reference information item and its feature representation; and a processing result generation module 930, configured to provide a processing result for the input content based on the input content, the matching result, and the information.
[0116] In some exemplary embodiments, the matching result determination module 910 is further configured to perform word segmentation on the input content to obtain at least one keyword, the at least one keyword corresponding to a first language; in response to finding a first term matching the at least one keyword in the terminology database, determine that the matching result includes a second term corresponding to the first term; and in response to not finding a first term matching the at least one keyword in the terminology database, determine that the matching result includes a miss indication.
[0117] In some exemplary embodiments, the terminology library has a prefix tree structure, and the apparatus 900 further includes a terminology library construction module configured to determine a first reference term by performing word segmentation processing on the reference text of the first language, the first reference term including at least one term unit; store each term unit of the first reference term in a node of the prefix tree structure; and store a second reference term corresponding to the first reference term in the terminology library in association.
[0118] In some exemplary embodiments, the information determination module 920 is further configured to determine a first feature representation of the content; determine a second feature representation from an information database based on the first feature representation, the second feature representation corresponding to a reference information item in the information database; and determine the reference information item corresponding to the second feature representation as information related to the content.
[0119] In some exemplary embodiments, the feature determination module is further configured to extract multiple features representing the content from the content; and to determine a first vector including the multiple features as a first feature representation.
[0120] In some exemplary embodiments, the feature determination module is further configured to determine the similarity between a first feature representation and a feature representation of at least one reference information item in the information database; and to determine a second feature representation based on the similarity.
[0121] In some exemplary embodiments, the processing result generation module further includes a translation and information prompting module, configured to provide translation results for the content, as well as prompting information for providing to the information processing model.
[0122] In some exemplary embodiments, the information prompting module is further configured to generate translation results by inputting prompting information into the information processing model.
[0123] In some exemplary embodiments, the processing result generation module is further configured to obtain a background description related to the input content; and to generate prompt information based on the input content, matching results, information, and background description.
[0124] In some exemplary embodiments, the processing result generation module is further configured to provide the processing result to the user through a pre-configured interface.
[0125] Figure 10 shows a block diagram of an electronic device 1000 capable of implementing embodiments of the present disclosure. It should be understood that the electronic device 1000 shown in Figure 10 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 1000 shown in Figure 10 can be used to implement the methods described above. The electronic device 1000 shown in Figure 10 can be used to implement the information processing device 101 of Figure 1 or the information processing apparatus 900 of Figure 9.
[0126] As shown in Figure 10, the electronic device 1000 is in the form of a general-purpose computing device. Components of the electronic device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. The processor 1010 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1020. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 1000.
[0127] Electronic device 1000 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1030 can be a removable or non-removable medium and may include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within electronic device 1000.
[0128] Electronic device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 1020 may include computer program product 1025 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0129] The communication unit 1040 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the electronic device 1000 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0130] Input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1060 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 1000 can also communicate with one or more external devices (not shown) via communication unit 1040 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 1000, or with any device that enables electronic device 1000 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0131] According to exemplary embodiments of the present disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary embodiments of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary embodiments of the present disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
[0132] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0133] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0134] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0136] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various embodiments disclosed herein.
Claims
1. A method for information processing, comprising: In response to receiving user input, a matching result for at least one keyword in the content is determined from a terminology database associated with the input content, the terminology database including multiple pairs of terms, each pair including a first term in a first language and a second term in a second language; Information related to the content is determined from an information database associated with the content, the information database including at least one reference information item and its feature representation; as well as Based on the input content, the matching result, and the information, a processing result for the input content is provided.
2. The method according to claim 1, wherein determining the matching result comprises: The input content is segmented to obtain at least one keyword, and the at least one keyword corresponds to a first language; In response to finding a first term matching the at least one keyword from the terminology database, it is determined that the matching result includes a second term corresponding to the first term; as well as In response to the first term not being found to match the at least one keyword from the terminology database, the matching result is determined to include a miss indication.
3. The method of claim 2, wherein the terminology database has a prefix tree structure and is constructed as follows: By segmenting the reference text in the first language, a first reference term is determined, which includes at least one term unit. Each term unit of the first reference term is stored in a node of the prefix tree structure; and A second reference term corresponding to the first reference term is stored in the terminology library.
4. The method of claim 1, wherein determining information related to the content from an information database associated with the content comprises: Determine the first feature representation of the content; A second feature representation is determined from the information database based on the first feature representation, the second feature representation corresponding to a reference information item in the information database; and The reference information item corresponding to the second feature representation is determined as information related to the content.
5. The method of claim 4, wherein determining the first feature representation comprises: Extract multiple features representing the content from the content; as well as The first vector comprising the plurality of features is determined as the first feature representation.
6. The method of claim 4, wherein determining the second feature representation comprises: Determine the similarity between the first feature representation and the feature representation of at least one reference information item in the information database; as well as The second feature representation is determined based on the similarity.
7. The method according to claim 1, wherein the processing result includes at least one of the following: The translation results for the aforementioned content, and Prompt information provided to the information processing model.
8. The method of claim 7, wherein providing the processing result comprises: The translation result is generated by inputting the prompt information into the information processing model.
9. The method of claim 7, wherein providing the processing result comprises: Obtain a background description related to the input content; as well as The prompt message is generated based on the input content, the matching result, the information, and the background description.
10. The method according to any one of claims 7-9, wherein providing the processing result comprises: The processing results are provided to the user through a pre-configured interface.
11. An information processing apparatus, comprising: The matching result determination module is configured to, in response to receiving user input, determine matching results for at least one keyword in the content from a terminology library associated with the content of the input, the terminology library including multiple pairs of terms, each pair including a first term in a first language and a second term in a second language; An information determination module is configured to determine information related to the content from an information database associated with the content, the information database including at least one reference information item and its feature representation; as well as The processing result generation module is configured to provide processing results for the input content based on the input content, the matching result, and the information.
12. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the device to perform the method according to any one of claims 1 to 10 when executed by the at least one processor.
13. A computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a processor, implement the method according to any one of claims 1 to 10.
14. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1-10.
Citation Information
Patent Citations
Translation method and device, electronic equipment and storage medium
CN113935339A
Term recognition method for multi-language translation
CN116822517A
Method and device for determining output data based on universal language model
CN117688166A
Information processing method and device, equipment and medium
CN118627521A
Translation using related term pairs
US20170300475A1