Language processing method, electronic equipment and computer readable storage medium
Through multilingual representation alignment and target language model analysis, the problems of language scope limitations and analysis scope limitations in large-scale language models in multilingual environments are solved, and more efficient and accurate multilingual problem handling is achieved, and the accuracy of cross-language understanding and small language problems are enhanced.
Patent Information
- Application Number
- CN202510330903.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
Existing large-scale language models (LLMs) have language scope limitations and analysis scope limitations when dealing with multilingual environments, especially in low-resource language environments, resulting in the generation of incorrect or unsafe answers.
The first language problem is transformed into a second language spatial representation through multilingual representation alignment, and the target language model is used for analysis to generate target reply, which enhances the cross-language understanding ability of large language models, especially the accuracy and responsiveness when dealing with small language problems.
It improves the knowledge boundary distinction ability of large language models in multilingual environments, reduces hallucination generation, enhances cross-language understanding and communication, expands the scope of analysis, and improves the accuracy and security of handling problems in small languages.
Smart Images

Figure CN120256571A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of large model technology and language processing technology. Specifically, it relates to a language processing method, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) as the core tools in the field of natural language processing have deeply influenced many applications such as information retrieval, intelligent dialogue, and text generation. However, when dealing with problems outside the knowledge boundary, LLMs often exhibit the "hallucination" phenomenon, that is, generating inaccurate or unsafe answers based on false premises.
[0003] Currently, for the analysis of the internal representations of LLMs, it has been proposed to analyze the internal representations of "correct and incorrect factual statements" in LLMs through probing techniques. Probing techniques can reveal the structure of the internal knowledge representation of LLMs, thereby evaluating their perception ability of factuality. However, the above analysis mainly focuses on the English language, which limits the application and evaluation of LLMs in multilingual environments. Moreover, the above analysis mainly focuses on the distinction between correct and incorrect factual statements by LLMs, and the analysis scope is relatively limited, lacking multi-domain generalization.
[0004] To address the above problems, no effective solutions have been proposed yet. Summary of the Invention
[0005] Embodiments of this application provide a language processing method, an electronic device, and a computer-readable storage medium to at least solve the technical problems of language scope limitations and limited analysis scope in the related art when analyzing the internal representations of large models through probing techniques.
[0006] According to one aspect of the embodiments of this application, a language processing method is provided, including: obtaining a language processing task, where the language processing task includes: a first language question, and the first language question uses a first language to describe the question to be answered; based on a multilingual representation alignment method, converting the first language question into a second language spatial representation, where the multilingual representation alignment method is used to obtain the ability to distinguish the knowledge boundary of the second language corresponding to the second language spatial representation; using a target language model to analyze the second language spatial representation and generate a target reply.
[0007] According to one aspect of the embodiments of the present application, a language processing method is provided, including: obtaining a cross-lingual text classification task, where the cross-lingual text classification task includes: a first language question, and the first language question uses a minority language to describe the question to be answered; based on a multi-lingual representation alignment method, converting the first language question into a second language spatial representation, where the multi-lingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the international common language corresponding to the second language spatial representation; using a target language model to analyze the second language spatial representation to generate a cross-lingual text classification result.
[0008] According to one aspect of the embodiments of the present application, a language processing method is provided, including: obtaining a language processing request through a first application programming interface, where the request data carried in the language processing request includes: a first language question, and the first language question uses a first language to describe the question to be answered; returning a language processing response through a second application programming interface, where the response data carried in the language processing response includes: a target reply, and the target reply is generated by using a target language model to analyze a second language spatial representation, and the second language spatial representation is obtained by converting the first language question based on a multi-lingual representation alignment method, and the multi-lingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language spatial representation.
[0009] According to one aspect of the embodiments of the present application, a language processing method is provided, including: obtaining a current input language processing dialogue request, where the request data carried in the language processing dialogue request includes: a first language question, and the first language question uses a first language to describe the question to be answered; in response to the language processing dialogue request, returning a language processing dialogue reply, where the information carried in the language processing dialogue reply includes: a target reply, and the target reply is generated by using a target language model to analyze a second language spatial representation, and the second language spatial representation is obtained by converting the first language question based on a multi-lingual representation alignment method, and the multi-lingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language spatial representation; displaying the target reply in a graphical user interface.
[0010] According to one aspect of the embodiments of the present application, a language processing method is provided, including: in response to an input instruction acting on an operation interface, displaying a first language question on the operation interface; in response to a processing instruction acting on the operation interface, displaying a target reply on the operation interface; where the first language question uses a first language to describe the question to be answered, and the target reply is generated by using a target language model to analyze a second language spatial representation, and the second language spatial representation is obtained by converting the first language question based on a multi-lingual representation alignment method, and the multi-lingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language spatial representation.
[0011] According to one aspect of the embodiments of the present application, a language processing system is provided, including: a client for sending a first-language question, where the first-language question describes the question to be answered in a first language; a server connected to the client, for converting the first-language question into a second-language space representation based on a multilingual representation alignment method, and analyzing the second-language space representation using a target language model to generate a target reply, where the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second-language space representation; the client is further used to output the target reply.
[0012] According to another aspect of the embodiments of the present application, an electronic device is further provided, including: a memory storing an executable program; a processor connected to the memory through a bus for running the program, where when the program runs, it executes any one of the above language processing methods.
[0013] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored executable program, where when the executable program runs, it controls the device where the computer-readable storage medium is located to execute any one of the above language processing methods.
[0014] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program that implements any one of the above language processing methods when executed by a processor.
[0015] In the embodiments of the present application, by obtaining a language processing task, the language processing task includes: a first-language question described in a first language, then converting the first-language question into a second-language space representation based on a multilingual representation alignment method, the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second-language space representation, and finally analyzing the second-language space representation using a target language model to generate a target reply, thereby achieving the purpose of efficiently and accurately generating answers to multilingual questions, thus enhancing the cross-language understanding ability of the large language model, improving the accuracy and response ability of the large model when processing questions in minority languages, and expanding the analysis scope of the large model, and further solving the technical problems of language range limitations and analysis scope limitations in the related art when analyzing the internal representation of the large model through probe technology.
[0016] It is easy to notice that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0018] Figure 1 is a schematic diagram of an application scenario of a language processing method according to an embodiment of the present application;
[0019] Figure 2 is a flowchart of a language processing method according to an embodiment of the present application;
[0020] Figure 3 is a schematic flowchart of the method effect according to an embodiment of the present application;
[0021] Figure 4 is a flowchart of a language processing method according to an embodiment of the present application;
[0022] Figure 5 is a flowchart of a language processing method according to an embodiment of the present application;
[0023] Figure 6 is a flowchart of a language processing method according to an embodiment of the present application;
[0024] Figure 7 is a flowchart of a language processing method according to an embodiment of the present application;
[0025] Figure 8 is a schematic diagram of the structure of a language processing system according to an embodiment of the present application;
[0026] Figure 9 is a schematic diagram of the structure of a language processing device according to an embodiment of the present application;
[0027] Figure 10 is a schematic diagram of the structure of another language processing device according to an embodiment of the present application;
[0028] Figure 11 is a schematic diagram of the structure of another language processing device according to an embodiment of the present application;
[0029] Figure 12 is a schematic diagram of the structure of another language processing device according to an embodiment of the present application;
[0030] Figure 13 is a schematic diagram of the structure of another language processing device according to an embodiment of the present application;
[0031] Figure 14 is a block diagram of the structure of a computing device according to an embodiment of the present application;
[0032] Figure 15It is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0033] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0034] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] The technical solutions provided by the present application are mainly implemented using large model technologies. Here, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. A large model can also be called a foundation model. Through large-scale pre-training of the large model with unlabeled corpora, a pre-trained model with more than hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability, such as large language models (LLMs), multi-modal pre-training models, etc.
[0036] It should be noted that in actual applications, the large model can be fine-tuned with a small number of samples for the pre-trained model, enabling the large model to be applied to different tasks. For example, the large model can be widely applied in fields such as Natural Language Processing (NLP), computer vision, and speech processing. Specifically, it can be applied to tasks in the field of computer vision such as Visual Question Answering (VQA), Image Caption (IC), and image generation. It can also be widely applied to tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. In the embodiments of this application, the language processing by the target language model proposed in this application in the language processing scenario is taken as an example for explanation.
[0037] First, some nouns or terms that appear during the description of the embodiments of this application are applicable to the following explanations:
[0038] Knowledge Boundary: The model's perception of whether knowledge exceeds the scope of answers. Common types include: 1) unanswerable questions; 2) questions beyond the model's knowledge; 3) questions with false premises.
[0039] Representation: The internal representation of text in the model, often an n-dimensional vector, also known as Embedding.
[0040] Probing: A methodology for analyzing the internal representation of the model. Common specific methods include training a linear classifier with the model representation and judging the richness and accuracy of the information contained in the model's internal representation based on the results.
[0041] The large model often produces hallucinations, generates incorrect and unsafe answers when facing questions beyond the knowledge boundary. This problem has received extensive attention and research in the English context, revealing the limitations of LLM when facing unanswerable questions, questions beyond the knowledge scope, and questions based on false premises. However, for the multilingual processing ability of LLM, especially the knowledge boundary perception in low-resource language environments, current research is relatively scarce.
[0042] Currently, in the field of research on this problem, for the analysis of the internal representation of LLM, the internal discrimination ability of LLM for correct and incorrect factual statements has been analyzed through probing techniques. Probing techniques can reveal the structure of the internal knowledge representation of LLM, thereby evaluating its perception ability of factual authenticity.
[0043] Although the above research provides some valuable insights in the English context, there are still the following deficiencies.
[0044] Deficiency 1: Limitation of language scope. Currently, the analysis mainly focuses on English, ignoring the ability of LLM to perceive the knowledge boundaries when dealing with other languages, especially low-resource languages. Therefore, it limits the application and evaluation of LLM in multilingual environments, especially in languages with scarce resources.
[0045] Deficiency 2: Limitation of analysis scope. Currently, the work mainly focuses on the discrimination of correct and incorrect factual statements by LLM, but fails to cover the broader concept of "knowledge boundaries", including but not limited to the identification of false premises and the unanswerability to non-existent entities. This limitation of the analysis scope makes the understanding of the true capabilities of LLM when dealing with various problems one-sided.
[0046] Deficiency 3: Lack of improvement strategies. Although probing techniques can reveal the deficiencies within LLM, the current research has not proposed effective methods to enhance the perception and processing ability of LLM for multilingual knowledge boundaries. This means that even if problems are discovered, there is a lack of systematic solutions to improve the performance of LLM.
[0047] In response to the above deficiencies, no effective solution has been proposed prior to this application.
[0048] According to the embodiments of this application, a language processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0049] Considering that the number of model parameters of large models is huge and the computing resources of mobile terminals are limited, the above method provided by the embodiments of this application can be applied to Figure 1 the application scenarios shown, but not limited thereto. In Figure 1In the application scenario shown, the large model is deployed in server 10, and server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Here, the client devices 20 can include, but are not limited to: smartphones, tablets, laptops, palmtop computers, personal computers, smart home devices, in-vehicle devices, etc. The client device 20 can interact with the user through a graphical user interface to implement the invocation of the large model, thereby implementing the method provided in the embodiments of the present application.
[0050] In the embodiments of the present application, the system composed of the client device and the server can perform the following steps: The client device performs obtaining a language processing task, where the language processing task includes steps such as describing a first language question of the question to be answered in a first language and sending the language processing task to the server. The server performs converting the first language question into a second language space representation based on a multilingual representation alignment method, where the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation, and then analyzes the second language space representation using a target language model to generate a target reply and return the target reply to the client device. It should be noted that in the case where the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiments of the present application can be performed in the client device.
[0051] It should be noted that with the rapid development of high-performance computing units, in other application scenarios, the above method provided in the embodiments of the present application can also be applied to a model all-in-one machine. In an optional embodiment, multiple models are built into the model all-in-one machine, and the user can select and adjust one model according to needs to obtain the user's own model. Thus, the high-performance computing unit built into the model all-in-one machine can directly call the adjusted model to execute the above method provided in the embodiments of the present application. In another optional embodiment, a trained model is built into the large model all-in-one machine. Thus, the high-performance computing unit built into the model all-in-one machine can directly call the model to execute the above method provided in the embodiments of the present application.
[0052] Furthermore, when the user needs to train their own model, they can also upload their own dataset through the client. The dataset is sent from the client to the server, enabling the server to adjust the pre-trained model with the dataset to obtain the user's own model, which is then deployed to the production environment. To facilitate the user's model adjustment requirements, the server can provide complete adjustment tools, development frameworks, and processes, and can support multiple adjustment strategies, enabling the adjusted model to better adapt to different field applications and achieve high customization.
[0053] In the above operating environment, the present application provides asFigure 2 The language processing method shown. Figure 2 is a flow chart of a language processing method according to an embodiment of the present application. Figure 2 As shown, the method may include the following steps:
[0054] Step S21, obtaining a language processing task, wherein the language processing task includes: a first language question, the first language question uses a first language to describe a question to be answered;
[0055] Step S22, based on the multilingual representation alignment method, converting the first language question into a second language space representation, wherein the multilingual representation alignment method is used to obtain the knowledge boundary distinguishing ability of the second language corresponding to the second language space representation;
[0056] Step S23: Analyze the second language spatial representation using the target language model to generate a target response.
[0057] In the embodiment of the present application, the language processing task can be understood as a task involving natural language understanding and generation, such as cross-language text classification, machine translation, question answering, language retrieval, text generation, etc., which are not limited here. In the embodiment of the present application, the language processing task can be a specific question raised by the user that needs to be understood and answered by the target language model.
[0058] The first language is the language used in the first language question. The first language may be, for example, Thai, Vietnamese or other minority languages, which is not limited here. That is, it can be understood that the first language is a low-resource language.
[0059] The first language problem is a problem to be solved described in the first language, such as a cross-language text classification problem described in Thai, or a machine translation problem described in Vietnamese, which is not restricted here.
[0060] Acquiring a language processing task can be understood as acquiring the unanswered questions described by the user in the first language (ie, a minority language) as a basis for subsequent processing and analysis.
[0061] The second language is another language different from the first language. The second language may be, for example, an internationally used language such as English or Chinese, which is not limited here. That is, it can be understood that the second language is a high-resource language.
[0062] It should be noted that the target language model mentioned later in the embodiments of the present application can effectively handle problems described in a second language, that is, the second language can be understood as a language that can be effectively processed by the target language model.
[0063] The multilingual representation alignment method can be understood as a method of transforming the representations of different languages into a shared or aligned representation space. The second-language space representation can be understood as the result obtained by transforming the representation of the first-language problem through the multilingual representation alignment method. The representation of the first-language problem is transformed from the subspace of the first language to the subspace of the second language through the multilingual representation alignment method, thereby obtaining the second-language space representation.
[0064] The multilingual representation alignment method is used to obtain the ability to distinguish the knowledge boundaries of the second language corresponding to the second-language space representation. Through the multilingual representation alignment method, low-resource languages can obtain an internal knowledge boundary discrimination ability equivalent to that of high-resource languages. This ensures that even in different languages, the internal processing mechanisms of the model can remain consistent or similar, thereby improving the cross-lingual performance and understanding ability of the model.
[0065] Exemplarily, taking the first language as Thai and the second language as English as an example, in the conversion from Thai to English, the representation of the Thai problem (i.e., the first-language problem) is transformed into a representation in the English context to better utilize the advantages of the English model.
[0066] Based on the multilingual representation alignment method, transforming the first-language problem into the second-language space representation can be understood as transforming the representation of the first-language problem into the representation space of another language (i.e., the second language) through the multilingual representation alignment method. Exemplarily, the multilingual representation alignment method can include Mean Shifting and Linear Projection, and can also include Multilingual Self-Alignment, Learning Shared Subspaces, etc., which are not limited here. This ensures that even if the problem is initially posed in a language different from the main training language of the target language model, the target language model can effectively understand and process this problem.
[0067] The target language model can be understood as a language model used to process and answer questions in the second-language space representation, that is, the target language model can be understood as a model optimized specifically for the second language, having a powerful processing ability in the second language (i.e., an internationally common language), but may perform poorly in the first language (i.e., a minority language).
[0068] The target reply can be understood as the answer or reply generated by the target language model after analyzing the question in the second-language space representation. Exemplarily, taking the first language as Thai and the second language as English as an example, the target language model is the English model, and the target reply can be understood as the answer to the Thai question generated based on the English model's understanding of the Thai question, presented in the expression manner that the English model is good at.
[0069] The target language model is used to analyze the second - language spatial representation to generate a target response. It can be understood that the target language model analyzes the second - language spatial representation, and the target language model will process the problem of the second - language spatial representation and generate a response based on its understanding of the knowledge boundary of the second language. By aligning the representation of the first - language question to a language that the target language model is better at processing, the target language model can show better performance and accuracy when generating answers, especially when dealing with questions about knowledge boundaries, and can effectively reduce the generation of hallucinations.
[0070] Exemplarily, taking the first language as Thai, the second language as English, and the target language model as an English model as an example, the problem of the English spatial representation is input into the English model for analysis. The English model will process this problem based on its understanding of the knowledge boundary of the English representation and generate a target response. Thus, even if the original question (i.e., the first - language question) is posed in Thai, the English model can understand and answer this Thai question in a way similar to how it processes English questions, thereby improving the accuracy and quality of the answer.
[0071] In the embodiments of the present application, by obtaining a language - processing task, the language - processing task includes: a first - language question described in the first language, and then based on a multi - language representation alignment method, the first - language question is transformed into a second - language spatial representation. The multi - language representation alignment method is used to obtain the ability to distinguish the knowledge boundary of the second language corresponding to the second - language spatial representation. Finally, the target language model is used to analyze the second - language spatial representation to generate a target response.
[0072] It can be seen that the present application adopts a multi - language representation alignment method. By aligning the internal representation of the small - language question (i.e., the first - language question) to the representation space of the target language model (i.e., the second - language spatial representation), the accuracy and response ability of the target language model in dealing with small - language questions are significantly improved, the performance of the model in dealing with small - language questions is enhanced, and it helps the target language model to better identify the knowledge boundary when dealing with questions, thereby reducing the generation of responses based on false premises, that is, reducing the generation of hallucinations and improving the safety and reliability of the generated answers. In addition, the present application uses the multi - language representation alignment method to enhance the cross - language understanding ability of the large - language model, especially the process of dealing with small - language questions. That is, the solution of the present application promotes the knowledge unity between different languages, enables the target language model to more consistently understand and process questions from different languages, and enhances cross - language understanding and communication.
[0073] The above language processing method provided by the embodiments of the present application can be, but is not limited to, applied to application scenarios involving language processing including minority languages in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example: language processing scenarios in e-commerce services, language processing scenarios in education services, language processing scenarios in legal services, language processing scenarios in medical services, etc., which are not limited here.
[0074] By adopting the embodiments of the present application, by obtaining a language processing task, the language processing task includes: a first language problem described in a first language, and then based on a multi-language representation alignment method, the first language problem is transformed into a second language space representation. The multi-language representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation. Finally, a target language model is used to analyze the second language space representation to generate a target response, thereby achieving the purpose of efficiently and accurately generating answers to multi-language problems, thus enhancing the cross-language understanding ability of the large language model, improving the accuracy and response ability of the large model when processing minority language problems, and expanding the analysis scope of the large model, and further solving the technical problems of language range limitation and analysis scope limitation in the related art by using the probe technology to perform internal representation analysis on the large model.
[0075] In an optional embodiment, the language processing method further includes the following method steps:
[0076] Step S24, using a probe verification method to select a second language from multiple candidate languages.
[0077] In the embodiments of the present application, a probe verification method can be used to select a second language from multiple candidate languages. Among them, the probe verification method can be understood as a method for evaluating the internal representation of a model. By training a simple classifier (i.e., a probe model) to determine whether the internal representation of the model for a specific concept or attribute is rich and accurate enough. Exemplarily, the probe model can be a linear classifier, which is used to evaluate the distinguishability and information content of different language representations, which is not limited here.
[0078] Multiple candidate languages can be understood as various languages that can be considered for alignment with the first language during the multi-language representation alignment process. The multiple candidate languages may include internationally common languages such as English, Chinese, Japanese, etc., and may also include other minority languages different from the first language, which is not limited here.
[0079] Using the probe verification method, a second language is selected from multiple candidate languages. It can be understood as using the probe verification method to evaluate the alignment effect between the internal representations of multiple candidate languages and the first language problem, so as to determine which language's representation has the best alignment effect with the first language problem, and then determine the second language.
[0080] Thus, through the probe verification method, a second language with the best alignment effect with the first language problem representation can be selected, ensuring that the subsequent representation conversion process can maximize the model's understanding and processing ability for the first language problem.
[0081] In an optional embodiment, in step S24, using the probe verification method to select a second language from multiple candidate languages includes the following method steps:
[0082] Step S241, using the target language model to perform feature encoding on multiple candidate languages to obtain knowledge boundary representations corresponding to the multiple candidate languages;
[0083] Step S242, using the target probe model to perform probe verification on the knowledge boundary representations to determine the second language.
[0084] In the embodiment of the present application, when using the probe verification method to select a second language from multiple candidate languages, the target language model can be used to perform feature encoding on multiple candidate languages to obtain knowledge boundary representations corresponding to the multiple candidate languages. Among them, feature encoding can be understood as the process of the model converting the input text into its internal knowledge representation. Exemplarily, feature encoding usually appears as a high-dimensional vector or matrix, and this representation contains the semantic information of the text and potential knowledge boundary features, which is not limited here.
[0085] Knowledge Boundary Representation can be understood as the internal representation generated by the target language model for the text. The knowledge boundary representation may include information about whether the text touches the model's knowledge boundary, that is, whether the text is answerable, based on a false premise, or contains non-existent entities, etc., which is not limited here.
[0086] Using the target language model to perform feature encoding on multiple candidate languages to obtain knowledge boundary representations corresponding to the multiple candidate languages can be understood as using the target language model to perform feature encoding on the text data of multiple candidate languages, so as to convert the text data of each candidate language into the internal representation form of the target language model. This internal representation form contains information related to the knowledge boundary, that is, to obtain the knowledge boundary representations corresponding to each candidate language.
[0087] Thus, the knowledge boundary representations corresponding to each candidate language can be obtained, and at the same time, it is also possible to check the representation of the knowledge boundary in the target language model in multiple languages, so as to analyze the target language model's understanding of the knowledge boundaries of different languages, providing a basis for subsequent analysis and verification.
[0088] After obtaining the knowledge boundary representations corresponding to multiple candidate languages, use the target probe model to perform probe verification on the knowledge boundary representations to determine the second language. Among them, the target probe model can be understood as a model specifically designed for analyzing and verifying the internal representations of a model. Exemplarily, the target probe model can be a simple linear classifier, which is used to determine whether the knowledge boundary representation can accurately reflect the knowledge boundary state of the text.
[0089] Probe verification can be understood as using the target probe model to evaluate the effect of the knowledge boundary representation in distinguishing problem types such as answerable, false premise, or non-existent entity, so as to judge the quality and alignment effect of the knowledge boundary representation.
[0090] Using the target probe model to perform probe verification on the knowledge boundary representations to determine the second language can be understood as using the target probe model to try to predict whether the text touches the knowledge boundary based on the representation. Thus, based on the performance of the target probe model, it is possible to evaluate which language's representation is most effective in distinguishing the knowledge boundary, and further determine which candidate language is used as the second language for subsequent representation alignment processes.
[0091] In an alternative embodiment, the language processing method further includes the following method steps:
[0092] Step S243, training the initial probe model with a multilingual knowledge boundary dataset to generate the target probe model, where the multilingual knowledge boundary dataset includes: correct and incorrect fact classification data, question answerability classification data, and false premise classification data.
[0093] In the embodiment of the present application, when training the target probe model, the initial probe model can be trained with a multilingual knowledge boundary dataset to generate the target probe model. Among them, the multilingual knowledge boundary dataset can be understood as a set of text data containing multiple languages, where each text is marked as whether it touches the knowledge boundary of the model, including correct and incorrect fact classification data, question answerability classification data, and false premise classification data. The multilingual knowledge boundary dataset provides necessary information for training the probe model, enabling it to distinguish different types of texts.
[0094] The correct and incorrect fact classification data can be understood as containing fact statements that are correct or incorrect for the model, and are used to train the model to distinguish between true and untrue information.
[0095] The question answerability classification data can be understood as questions being marked as answerable or unanswerable, which is used to train the model to identify questions within its knowledge scope and questions beyond the existing knowledge.
[0096] The false premise classification data can be understood as containing statements or questions based on false premises, which is used to train the model to identify and avoid false assumptions during reasoning.
[0097] Use the multilingual knowledge boundary dataset to train the initial probe model. Through training, the initial probe model can learn how to judge whether the text belongs to the known or unknown domain based on internal representations, and whether it is based on false premises or contains unanswerable questions. After training, the obtained target probe model has the cross - language knowledge boundary analysis ability, so as to accurately classify and identify texts in different languages.
[0098] Thus, by using the multilingual knowledge boundary dataset to train the initial probe model, it is ensured that the obtained target probe model has broad generalization ability, can adapt to texts in different languages, and can also maintain good performance on languages that have not been directly trained. In addition, the obtained target probe model can quickly judge which texts touch the knowledge boundary when processing multiple languages, thus avoiding hallucinations when dealing with unanswerable or false - premise - based questions.
[0099] In an alternative embodiment, in step S243, the initial probe model is trained with the multilingual knowledge boundary dataset to generate the target probe model, including the following method steps:
[0100] Step S2431, determine multiple in - domain linear classifiers included in the initial probe model based on the number of model layers of the target language model and the number of languages of multiple candidate language species;
[0101] Step S2432, determine the target representation of the question to be trained according to the multilingual knowledge boundary dataset;
[0102] Step S2433, use the target representation of the question to be trained to train the multiple in - domain linear classifiers to generate the target probe model.
[0103] In the embodiments of the present application, when training an initial probe model with a multilingual knowledge boundary dataset to generate a target probe model, multiple in-domain linear classifiers included in the initial probe model can be determined based on the number of model layers of the target language model and the number of languages of multiple candidate languages. Herein, the number of model layers of the target language model can be understood as the number of levels (Layers) in the internal architecture of the target language model. The target language model usually consists of multiple layers, and each layer is responsible for processing specific aspects of the text, such as grammar, semantics, and knowledge boundaries, etc., which are not limited herein.
[0104] Multiple in-domain linear classifiers can be understood as classifiers for analyzing the internal representations of the target language model. An in-domain linear classifier can be understood as a linear classifier for analyzing a specific language. Exemplarily, k×m in-domain linear classifiers can be trained, where k is the number of model layers of the target language model and m is the number of languages of multiple candidate languages.
[0105] Determining multiple in-domain linear classifiers included in the initial probe model based on the number of model layers of the target language model and the number of languages of multiple candidate languages can be understood as determining the layout and number of linear classifiers in the initial probe model according to the number of layers of the target language model and the number of multiple candidate languages. Thus, each linear classifier will focus on analyzing the internal representation of the text of a specific language in the target language model to detect whether the knowledge boundary is touched. In this way, the differences in the internal representations of texts in different languages in the target language model can be distinguished, laying a foundation for subsequent training and analysis.
[0106] Meanwhile, the target representation of the problem to be trained can be determined based on the multilingual knowledge boundary dataset. Herein, the problem to be trained can be understood as a text problem selected from the multilingual knowledge boundary dataset for training the probe model. The target representation of the problem to be trained can be understood as the representation form of the problem to be trained in the target language model, which contains potential information about whether the text touches the knowledge boundary.
[0107] Determining the target representation of the problem to be trained based on the multilingual knowledge boundary dataset can be understood as selecting a series of problems, i.e., the problems to be trained, from the multilingual knowledge boundary dataset as the input for training the probe model. Each problem is clearly marked with the status of its knowledge boundary, such as whether it is based on known or unknown facts, whether it is answerable or based on a false premise, etc. By determining the target representation of the problem to be trained, specific guidance is provided for training the target probe model, enabling it to learn how to identify and distinguish different types of texts.
[0108] Finally, the target representations of the problems to be trained are used to train multiple in-domain linear classifiers to generate a target probe model. It can be understood that the target representations of the problems to be trained are used to train multiple in-domain linear classifiers in the initial probe model. The multiple in-domain linear classifiers will learn how to determine whether a text touches the knowledge boundary based on the internal representation of the text. After training, the target probe model will be obtained. The target probe model can accurately analyze the knowledge boundary state of texts in multiple languages. Thus, it is ensured that the target probe model can not only handle languages with rich resources such as English and Chinese, but also effectively handle languages with fewer resources such as Vietnamese, Thai, Khmer, Indonesian, Malay, and Lao, and it is also ensured that the target probe model can maintain a high accuracy when dealing with low-resource languages.
[0109] In an optional embodiment, in step S22, based on the multi-language representation alignment method, converting the first-language problem into a second-language space representation includes the following method steps:
[0110] Step S221, determining the first-language subspace corresponding to the first-language problem and the second-language subspace corresponding to the second-language space representation;
[0111] Step S222, based on the multi-language representation alignment method, performing multi-language subspace linear alignment on the first-language subspace and the second-language subspace to obtain the second-language space representation.
[0112] In the embodiment of the present application, when converting the first-language problem into a second-language space representation based on the multi-language representation alignment method, the first-language subspace corresponding to the first-language problem and the second-language subspace corresponding to the second-language space representation can be determined. Among them, the subspace can be understood as a part of the model embedding space, and each dimension or vector in the subspace represents a specific attribute or information of the text.
[0113] The first-language subspace can be understood as the feature space occupied by the first-language problem in the internal representation of the target language model. The second-language subspace corresponds to the first-language subspace and can be understood as the distribution area of the second-language space representation in the model embedding space.
[0114] Determining the first-language subspace corresponding to the first-language problem and the second-language subspace corresponding to the second-language space representation can be understood as determining the subspace occupied by the internal representation (i.e., the embedding vector) corresponding to the first-language problem after it passes through the target language model, and the subspace of the second-language problem to be aligned. Since the two subspaces may show different distributions in the embedding space of the model, it is necessary to clarify the positions and distributions of these subspaces for subsequent multi-language representation alignment.
[0115] Then, based on the multi - language representation alignment method, perform a linear alignment of the first - language subspace and the second - language subspace in the multi - language subspace to obtain the second - language space representation. It can be understood that the representation subspaces of different languages have a linear structure, so they can be aligned by linear methods. The linear alignment of the multi - language subspace usually involves mathematical linear transformations (such as matrix multiplication, mean alignment, linear projection, etc., which are not limited here) to ensure that the distributions of the first - language subspace and the second - language subspace in the embedding space are similar or aligned.
[0116] Based on the multi - language representation alignment method, perform a linear alignment of the first - language subspace and the second - language subspace in the multi - language subspace to obtain the second - language space representation. It can be understood that by mathematical means, the two subspaces are adjusted so that they have similar distributions in the embedding space of the model to ensure that the second - language space representation can accurately reflect the knowledge boundary state of the first - language problem. The second - language space representation obtained after alignment will be closer to the internal representation of the first - language problem, thus ensuring that the target language model can maintain an accurate perception of the knowledge boundary when processing different languages, and improving the security and reliability of the model in a multi - language environment.
[0117] In an optional embodiment, the multi - language representation alignment method includes: the multi - language subspace mean - shifting alignment method. In step S222, based on the multi - language representation alignment method, perform a linear alignment of the first - language subspace and the second - language subspace in the multi - language subspace to obtain the second - language space representation, including the following method steps:
[0118] Step S2221, based on the multi - language subspace mean - shifting alignment method, calculate the first representation vector mean of the first - language subspace and the second representation vector mean of the second - language subspace respectively, and calculate the mean difference between the first representation vector mean and the second representation vector mean;
[0119] Step S2222, align the vector dimensions of the first - language subspace and the second - language subspace according to the mean difference to obtain the second - language space representation.
[0120] In the embodiment of the present application, the multi - language representation alignment method includes: the multi - language subspace mean - shifting alignment method (mean - shifting). The multi - language subspace mean - shifting alignment method is used to adjust the internal representation of the model to ensure that the feature distributions of different - language texts are similar, and it is achieved by calculating the means of different - language subspaces and adjusting the differences between the means.
[0121] When performing multi - language subspace linear alignment on the first - language subspace and the second - language subspace based on the multi - language representation alignment method to obtain the second - language space representation, the mean value of the first representation vectors of the first - language subspace and the mean value of the second representation vectors of the second - language subspace can be calculated respectively based on the multi - language subspace mean alignment method, and the mean difference between the mean value of the first representation vectors and the mean value of the second representation vectors can be calculated.
[0122] Among them, the mean value of the first representation vectors can be understood as the average value of all problem representation vectors in the first - language subspace, which is used to represent the central tendency of the first - language problems in the internal representation of the model.
[0123] The mean value of the second representation vectors corresponds to the mean value of the first representation vectors and can be understood as the average value of all problem representation vectors in the second - language subspace, which is used to represent the general characteristics of the second - language problems in the internal representation of the model.
[0124] The mean difference can be understood as the difference between the mean value of the first representation vectors and the mean value of the second representation vectors, which is used to measure the deviation of different language subspaces in the internal representation of the model.
[0125] Based on the multi - language subspace mean alignment method, the mean value of the first representation vectors of the first - language subspace and the mean value of the second representation vectors of the second - language subspace are calculated respectively, and the mean difference between the mean value of the first representation vectors and the mean value of the second representation vectors is calculated. Thus, the mean difference is used to assist in determining how the first - language subspace needs to be adjusted to be closer to the second - language subspace.
[0126] After that, vector - dimension alignment is performed on the first - language subspace and the second - language subspace according to the mean difference to obtain the second - language space representation. Among them, vector - dimension alignment can be understood as aligning different subspaces by adjusting the values of the vectors to ensure that the first - language subspace and the second - language subspace have similar positions and distributions in the embedding space of the model.
[0127] Performing vector - dimension alignment on the first - language subspace and the second - language subspace according to the mean difference to obtain the second - language space representation can be understood as adjusting the representation vectors of all problems in the first - language subspace according to the calculated mean difference so that they are aligned with the representation vectors of the second - language subspace in the embedding space of the model. Exemplarily, translation can be performed in the vector space to eliminate or reduce the mean difference, so as to ensure that the text representations of the two languages have similar distributions in the embedding space, which is not limited here. The adjusted representation vectors of the first - language problems will be transformed into the second - language space to obtain the second - language space representation.
[0128] By mean alignment, the consistency of texts in different languages in the internal representation of the model is ensured, the accuracy of cross-lingual knowledge boundary analysis is improved, and the knowledge boundary recognition error caused by language differences is reduced. At the same time, the mean alignment method also enables the model to better process low-resource languages. Even without a large amount of training data in these languages, the performance of knowledge boundary analysis can be improved by aligning with the subspace of resource-rich languages, enhancing the multi-lingual processing ability of the model.
[0129] In an optional embodiment, the multi-lingual representation alignment method includes: the multi-lingual subspace linear projection method. In step S222, based on the multi-lingual representation alignment method, the first language subspace and the second language subspace are linearly aligned in the multi-lingual subspace to obtain the second language space representation, including the following method steps:
[0130] Step S2223, based on the multi-lingual subspace linear projection method, perform singular value decomposition on the representation vectors of the first language subspace to obtain a pseudo matrix, and perform matrix multiplication on the pseudo matrix and the representation vectors of the second language subspace to obtain a linear projection matrix;
[0131] Step S2224, according to the linear projection matrix, perform vector representation conversion on the first language subspace and the second language subspace to obtain the second language space representation.
[0132] In the embodiment of the present application, the multi-lingual representation alignment method includes: the multi-lingual subspace linear projection method (linearprojection). The multi-lingual subspace linear projection method is used to adjust the internal representations (or called embedding vectors) generated by the model when processing different languages to ensure the consistency and comparability of these representations in the mathematical space, especially when it comes to knowledge boundary analysis and recognition.
[0133] When, based on the multi-lingual representation alignment method, the first language subspace and the second language subspace are linearly aligned in the multi-lingual subspace to obtain the second language space representation, the representation vectors of the first language subspace can be subjected to singular value decomposition based on the multi-lingual subspace linear projection method to obtain a pseudo matrix, and matrix multiplication can be performed on the pseudo matrix and the representation vectors of the second language subspace to obtain a linear projection matrix.
[0134] Among them, singular value decomposition is a decomposition method in linear algebra that can decompose a matrix into the product of three matrices, used to reveal the internal structure and properties of the matrix, and is commonly used in scenarios such as data dimensionality reduction, feature extraction, and model optimization.
[0135] The pseudo matrix can be understood as the intermediate matrix generated after singular value decomposition, containing the key feature information of the representation vectors of the first language subspace, and is used for subsequent linear projection operations.
[0136] A linear projection matrix can be understood as a matrix used to transform a vector space. Through matrix multiplication, the representation vectors in the first language subspace can be projected onto the second language subspace to achieve the alignment and transformation of representations.
[0137] Based on the linear projection method of the multilingual subspace, perform singular value decomposition on the representation vectors in the first language subspace to obtain a pseudo-matrix, and perform matrix multiplication on the pseudo-matrix and the representation vectors in the second language subspace to obtain a linear projection matrix. Thus, the obtained linear projection matrix will be used to transform the vectors in the first language subspace, aiming to make the representation of the first language problem better match the distribution of the second language subspace mathematically and achieve the alignment of the two language representations.
[0138] Then, perform vector representation transformation on the first language subspace and the second language subspace according to the linear projection matrix to obtain the second language space representation. It can be understood that the linear projection matrix is used to transform the representation vectors in the first language subspace so that they are projected onto the coordinate system of the second language subspace mathematically. Through vector representation transformation, the problem representation in the first language obtains a new representation in the second language space, and this new representation is closer to the distribution of the second language subspace, which helps to accurately identify and analyze the knowledge boundaries between different languages.
[0139] It can be seen that through singular value decomposition and the generation of the linear projection matrix, the structure of the representation vectors is optimized, enabling key feature information to be retained and aligned. Moreover, through linear projection, the problem representation in the first language subspace is remapped onto the coordinate system of the second language subspace, enhancing the model's ability to identify the knowledge boundaries of different languages, which thus helps the model better process texts in low-resource languages. Because even in the case of limited resources, by aligning with the representations of high-resource languages, the performance of knowledge boundary analysis for low-resource languages can be improved.
[0140] In an optional embodiment, the language processing method further includes the following method steps:
[0141] Step S25, obtain sample problem translation pairs;
[0142] Step S26, fine-tune the initial language model based on the sample problem translation pairs to obtain the target language model.
[0143] In the embodiments of the present application, when training a target language model, sample question-translation pairs can be obtained. Herein, a sample question-translation pair can be understood as a data pair consisting of a question text in a first language and its accurate translation in a second language. Such sample question-translation pairs are the basis for fine-tuning a multilingual model and are used to guide the model in learning and adjusting how to correctly process texts in different languages. Exemplarily, the sample question-translation pairs can cover various types of questions, including questions based on correct and incorrect assumptions, questions about known and unknown entities, and questions that can and cannot be answered, etc., without limitation here. In addition, the sample question-translation pairs can be obtained through manual translation, professional translation services, or by using existing translation datasets, without limitation here.
[0144] Then, based on the sample question-translation pairs, the initial language model is fine-tuned to obtain the target language model. Herein, fine-tuning can be understood as, based on a pre-trained model, using a more specific and relevant dataset for additional training to optimize the model's performance on a specific task or domain.
[0145] By obtaining sample question-translation pairs consisting of questions described in a first language and their corresponding translations in a second language, the sample question-translation pairs are used for subsequent model fine-tuning to guide the model on how to maintain the ability to perceive knowledge boundaries when processing texts in different languages, enabling it to learn how to accurately identify and distinguish different states of knowledge boundaries when dealing with questions in different languages. It can be understood that during the fine-tuning process, the model updates its weights and parameters to better adapt to the requirements of multilingual knowledge boundary recognition. The finally obtained target language model will have improved multilingual capabilities, especially in terms of knowledge boundary analysis and avoiding hallucination generation.
[0146] Thus, using sample question-translation pairs for fine-tuning can prompt the model to more accurately identify knowledge boundaries when processing different languages, enhancing the model's multilingual generalization ability. Especially for low-resource languages, the model's performance will be significantly improved. At the same time, the fine-tuning process ensures the consistency and accuracy of the model's knowledge boundary analysis in different languages, reducing the knowledge boundary recognition errors caused by language differences.
[0147] In an alternative embodiment, in step S25, obtaining the sample question-translation pairs includes the following method steps:
[0148] Step S251, obtaining a third-language data subset from a multilingual knowledge boundary dataset;
[0149] Step S252, performing machine translation on the third-language data subset to obtain a fourth-language translation result, where the scale of the training data of the language used for the third-language data subset is higher than the scale of the training data of the language used for the fourth-language translation result;
[0150] Step S253: Determine sample problem translation pairs based on the third - language data subset and the fourth - language translation result.
[0151] In the embodiments of the present application, when obtaining sample problem translation pairs, a third - language data subset can be obtained from a multilingual knowledge boundary dataset. The third - language data subset can be understood as a dataset of a specific language selected from the above - mentioned multilingual knowledge boundary dataset for machine translation. Exemplarily, the third - language data subset can be a dataset composed of languages with relatively large training data scales, such as English, Chinese, etc., and there is no limitation here.
[0152] Select a data subset of a specific language, that is, the third - language data subset, from a multilingual knowledge boundary dataset containing problems in multiple languages. The third - language data subset contains a large number of problem instances, covering various knowledge boundary situations, including but not limited to correct and incorrect factual statements, correct and incorrect premises, and existing and non - existing entities, etc., and there is no limitation here. By selecting languages with relatively rich resources to form the third - language data subset, it can ensure high - quality translation and provide more sufficient training data in the subsequent fine - tuning process.
[0153] After that, perform machine translation on the third - language data subset to obtain a fourth - language translation result, where the training data scale of the language used in the third - language data subset is higher than that of the language used in the fourth - language translation result. It can be understood as using a machine translation tool to translate the third - language data subset into the language used in the fourth - language translation result to obtain the fourth - language translation result, so as to generate training data for the language used in the fourth - language translation result and supplement the problem of insufficient data volume of the language used in the fourth - language translation result.
[0154] It can be understood that since the language used in the third - language data subset usually has more training data, the accuracy and reliability of the translation result can be guaranteed. By performing machine translation, additional training resources can be created for the language used in the fourth - language translation result with less resources, which will play a key role in the subsequent model fine - tuning step.
[0155] Finally, determine sample problem translation pairs based on the third - language data subset and the fourth - language translation result. It can be understood as pairing the problems in the third - language data subset with the translation results in the language used in the fourth - language translation result to form sample problem translation pairs for subsequent fine - tuning training to help the model learn how to accurately identify and analyze knowledge boundaries when processing texts in the language used in the fourth - language translation result.
[0156] As can be seen from the above, the embodiments of this application have for the first time extended the analysis of knowledge boundaries to multiple languages, and for the first time broadened the types of knowledge boundary analysis beyond "correct and incorrect factual statements" to include "correct and incorrect premises" and the issue of "unanswerability of non-existent entities".
[0157] This application proposes a multilingual representation alignment method that does not require training, including multilingual subspace mean alignment and linear projection. Through the above two methods, low-resource languages can obtain knowledge boundary discrimination capabilities comparable to those of English and Chinese. At the same time, this application also proposes a fine-tuning training method based on bilingual translation, and through experiments, it is proved that only training on bilingual question translation can bring an improvement in multilingual knowledge boundary perception ability.
[0158] Figure 3 It is a schematic flowchart of the method effect according to the embodiments of this application. As Figure 3 shown, this application has for the first time analyzed multiple languages in terms of the model hallucination problem and analyzed types of knowledge boundaries that have not been analyzed before. Through analysis, it is found that the representation subspaces of different languages have a linear structure and can be aligned by linear methods. Therefore, this application proposes multilingual subspace mean alignment (mean-shifting) and linear projection (linear projection). As Figure 3 shown by the model experiment, the method of this application can well align multilingual subspaces. After linear projection, the probe model (line A) trained in one language can basically achieve a comparable level (line B) in other languages, bringing a very large improvement compared to the baseline (line C) without any processing. And the mean shift (line D) greatly enhances the transfer ability across levels compared to the original representation.
[0159] In addition, this application also verifies the practical significance of probe analysis. On the data subset of "incorrect premises", this application compares the performance of directly asking the model questions (Baseline) and hinting the model in advance that there may be incorrect premises (False premise hinted, FP-hinted). Experiments prove that this application has achieved an improvement of 10-20 points in multiple languages, and in Vietnamese, the improvement of the model has reached more than 24 points, greatly reducing the hallucination of the model for questions about incorrect premises.
[0160] It can be seen that this application has systematically analyzed the multilingual knowledge boundaries of large language models through the internal representations of the models. This application enables low-resource languages to obtain internal knowledge boundary discrimination capabilities comparable to those of high-resource languages through the representation alignment method, and simultaneously improves multilingual capabilities through additional training.
[0161] It is easy to understand that the beneficial effects of the language processing method provided by this application include the following points.
[0162] Beneficial effect (1): For the first time, this application applies the knowledge boundary analysis method to multi-language scenarios, not limited to English. This enables the model to more accurately identify problems beyond its knowledge scope and problems based on false premises when processing different languages, thereby avoiding generating inaccurate or unsafe answers.
[0163] Beneficial effect (2): Through the alignment method of the internal representation of the model, low-resource languages (i.e., languages with less training data) can obtain knowledge boundary recognition capabilities equivalent to those of high-resource languages (such as English and Chinese). Thus, even in a language environment with scarce data, the model can effectively identify and process knowledge boundary problems.
[0164] Beneficial effect (3): This application proposes a method for fine-tuning training through bilingual question translations. Experiments have proven that this method can significantly improve the model's knowledge boundary perception ability in different languages. This method is not only applicable to specific models but also provides a new idea for the training of multi-language models.
[0165] Beneficial effect (4): The method of this application enables the model to detect in real time whether a question may exceed the knowledge boundary or be based on a false premise, greatly reducing the hallucination phenomenon of the model.
[0166] Beneficial effect (5): Through the proposed multi-language representation space alignment and bilingual fine-tuning training, the performance of the model in multiple languages has been generally significantly improved, including but not limited to Vietnamese, Thai, Khmer, Indonesian, Malay, etc., especially when dealing with problems in languages with less resources.
[0167] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0168] In addition, it should also be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0169] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0170] According to an embodiment of the present application, there is also provided a Figure 4 language processing method as shown in Figure 4 is a flowchart of a language processing method according to an embodiment of the present application. As shown in Figure 4 it, the method includes:
[0171] Step S41, obtain a cross-language text classification task, where the cross-language text classification task includes: a first language question, and the first language question uses a minority language to describe the question to be answered;
[0172] Step S42, based on a multi-language representation alignment method, convert the first language question into a second language space representation, where the multi-language representation alignment method is used to obtain the knowledge boundary discrimination ability of the international common language corresponding to the second language space representation;
[0173] Step S43, analyze the second language space representation by using a target language model to generate a cross-language text classification result.
[0174] The language processing method proposed in the embodiment of the present application can be applied to a cross-language text classification scenario. The cross-language text classification task can be understood as a classification task involving texts in different languages, requiring the model to classify the texts into different categories. The possible types of questions may include correct or incorrect factual statements, questions based on correct or incorrect premises, and unanswerable questions involving non-existent entities, which are not limited here.
[0175] Specifically, by obtaining a cross-language text classification task, where the cross-language text classification task includes: a first language question, and the first language question uses a minority language to describe the question to be answered. Then, based on a multi-language representation alignment method, convert the first language question into a second language space representation, where the multi-language representation alignment method is used to obtain the knowledge boundary discrimination ability of the international common language corresponding to the second language space representation. Finally, analyze the second language space representation by using a target language model to generate a cross-language text classification result.
[0176] For specific descriptions, reference may be made to the descriptions of the foregoing embodiments, which will not be limited herein.
[0177] The above language processing method provided by the embodiments of the present application can be but is not limited to being applied to application scenarios involving language processing including minority languages in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example: language processing scenarios in e-commerce services, language processing scenarios in education services, language processing scenarios in legal services, language processing scenarios in medical services, etc., which will not be limited herein.
[0178] By adopting the embodiments of the present application, a cross-lingual text classification task is obtained, where the cross-lingual text classification task includes: a first language question, and the first language question uses a minority language to describe the question to be answered. Then, based on the multi-lingual representation alignment method, the first language question is transformed into a second language space representation, where the multi-lingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the international common language corresponding to the second language space representation. Finally, a target language model is used to analyze the second language space representation to generate a cross-lingual text classification result, thereby achieving the purpose of efficiently and accurately generating answers to multi-lingual questions, thus enhancing the cross-lingual understanding ability of the large language model, improving the accuracy and response ability of the large model when dealing with minority language questions, and expanding the analysis scope of the large model, and further solving the technical problems of language scope limitation and analysis scope limitation in the related art by using probe technology to perform internal representation analysis on the large model.
[0179] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiments, which will not be elaborated herein.
[0180] According to the embodiments of the present application, there is also provided a Figure 5 language processing method as shown in Figure 5 is a flowchart of a language processing method according to the embodiments of the present application, as shown in Figure 5 shown, the method includes:
[0181] Step S51, obtaining a language processing request through a first application programming interface, where the request data carried in the language processing request includes: a first language question, and the first language question uses a first language to describe the question to be answered;
[0182] Step S52, return a language processing response through a second application programming interface. The response data carried in the language processing response includes: a target reply, which is generated by analyzing a second-language spatial representation using a target language model. The second-language spatial representation is obtained by transforming a first-language question based on a multilingual representation alignment method, and the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second-language spatial representation.
[0183] The above first application programming interface and second application programming interface can be either the same application programming interface or different application programming interfaces. In an alternative embodiment, the interface parameters in the above first application programming interface and second application programming interface may include, but are not limited to: interface global identifier, interface signature key, interface timestamp, interface request identifier, system call credential identifier, etc. The above first application programming interface can use GET or POST as the interface request method to obtain a file processing request. The above second application programming interface can use the JSON format to feedback a file processing response.
[0184] In the embodiments of the present application, a language processing request can be understood as a request received from a user through a first application programming interface, which contains request data. The language processing request is used to enable the system to understand and process a first-language question and give an appropriate response based on its knowledge boundary.
[0185] A language processing response can be understood as a response to a language processing request generated by the system after receiving a first-language question.
[0186] For specific descriptions, reference can be made to the descriptions of the foregoing embodiments, which are not limited herein.
[0187] The above language processing method provided by the embodiments of the present application can be, but is not limited to, applied to application scenarios involving language processing including minority languages in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example: language processing scenarios in e-commerce services, language processing scenarios in education services, language processing scenarios in legal services, language processing scenarios in medical services, etc., which are not limited herein.
[0188] By adopting the embodiment of the present application, a language processing request is obtained through a first application programming interface. Among them, the request data carried in the language processing request includes: a first language question, and the first language question uses a first language to describe the question to be answered. Then, a language processing response is returned through a second application programming interface. Among them, the response data carried in the language processing response includes: a target reply, and the target reply is generated after analyzing a second language space representation by a target language model. The second language space representation is obtained after transforming the first language question based on a multilingual representation alignment method. The multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation. Thus, the purpose of efficiently and accurately generating answers to multilingual questions is achieved, thereby enhancing the cross-language understanding ability of the large language model, improving the accuracy and response ability of the large model when dealing with questions in minority languages, and expanding the technical effect of the analysis scope of the large model. Furthermore, the technical problem that there are language range limitations and analysis scope limitations in the related art when analyzing the internal representation of the large model through probe technology is solved.
[0189] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, and will not be elaborated here.
[0190] According to the embodiment of the present application, there is also provided a Figure 6 language processing method as shown. Figure 6 is a flowchart of a language processing method according to the embodiment of the present application, as Figure 6 shown, and the method includes:
[0191] Step S61, obtaining a currently input language processing dialogue request. Among them, the request data carried in the language processing dialogue request includes: a first language question, and the first language question uses a first language to describe the question to be answered;
[0192] Step S62, in response to the language processing dialogue request, returning a language processing dialogue reply. Among them, the information carried in the language processing dialogue reply includes: a target reply, and the target reply is generated after analyzing a second language space representation by a target language model. The second language space representation is obtained after transforming the first language question based on a multilingual representation alignment method. The multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation;
[0193] Step S63, displaying the target reply in the graphical user interface.
[0194] In the embodiment of the present application, the language processing dialogue request can be understood as a request sent by the user to the system through the graphical user interface or a certain interface. The data carried in the request mainly includes the "first language question".
[0195] A language processing dialogue response can be understood as the system's response to a language processing dialogue request, which contains the "target response" generated by the system based on its analysis and knowledge.
[0196] For specific descriptions, reference can be made to the descriptions of the foregoing embodiments, which are not limited herein.
[0197] The above language processing method provided by the embodiments of the present application can be, but is not limited to, applied to application scenarios involving language processing including minority languages in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example: language processing scenarios in e-commerce services, language processing scenarios in education services, language processing scenarios in legal services, language processing scenarios in medical services, etc., which are not limited herein.
[0198] By adopting the embodiments of the present application, a language processing dialogue request currently input is obtained. Among them, the request data carried in the language processing dialogue request includes: a first language question, and the first language question uses a first language to describe the question to be answered. Then, in response to the language processing dialogue request, a language processing dialogue response is returned. Among them, the information carried in the language processing dialogue response includes: a target response, and the target response is generated after analyzing a second language space representation by using a target language model. The second language space representation is obtained after transforming the first language question based on a multi-language representation alignment method, and the multi-language representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation. Finally, the target response is displayed in the graphical user interface, thereby achieving the purpose of efficiently and accurately generating answers to multi-language questions, thus realizing the technical effects of enhancing the cross-language understanding ability of the large language model, improving the accuracy and response ability of the large model when processing minority language questions, and expanding the analysis scope of the large model, and further solving the technical problems of limited language scope and limited analysis scope in the related art by using probe technology to perform internal representation analysis on the large model.
[0199] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, which will not be elaborated herein.
[0200] According to the embodiments of the present application, there is also provided a language processing method as Figure 7 shown. Figure 7 is a flowchart of a language processing method according to the embodiments of the present application, as Figure 7 shown, and the method includes:
[0201] Step S71, in response to an input instruction acting on the operation interface, display a first language question on the operation interface;
[0202] Step S72: In response to a processing instruction on the operation interface, display a target reply on the operation interface; wherein, the first-language question describes the question to be answered in the first language, the target reply is generated by analyzing the second-language spatial representation using a target language model, the second-language spatial representation is obtained by transforming the first-language question based on a multilingual representation alignment method, and the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second-language spatial representation.
[0203] In the embodiments of the present application, the input instruction can be understood as an instruction issued by the user on the operation interface to submit a question that needs to be answered.
[0204] The processing instruction can be understood as an instruction requesting the system to process the first-language question.
[0205] For specific descriptions, reference can be made to the descriptions of the foregoing embodiments, which are not limited herein.
[0206] The above language processing method provided by the embodiments of the present application can be but is not limited to being applied to application scenarios involving language processing including minority languages in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example: language processing scenarios in e-commerce services, language processing scenarios in education services, language processing scenarios in legal services, language processing scenarios in medical services, etc., which are not limited herein.
[0207] By adopting the embodiments of the present application, by responding to an input instruction on the operation interface, display the first-language question on the operation interface, and then respond to a processing instruction on the operation interface to display the target reply on the operation interface; wherein, the first-language question describes the question to be answered in the first language, the target reply is generated by analyzing the second-language spatial representation using a target language model, the second-language spatial representation is obtained by transforming the first-language question based on a multilingual representation alignment method, and the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second-language spatial representation, thereby achieving the purpose of efficiently and accurately generating answers to multilingual questions, thus enhancing the cross-language understanding ability of the large language model, improving the accuracy and response ability of the large model in dealing with minority language questions, and expanding the analysis scope of the large model, and further solving the technical problems of limited language scope and limited analysis scope in the related art when analyzing the internal representation of the large model through probe technology.
[0208] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, which will not be elaborated herein.
[0209] According to the embodiments of the present application, there is also provided as Figure 8A language processing system as shown. Figure 8 It is a schematic structural diagram of a language processing system according to an embodiment of the present application, as Figure 8 shown. The system includes:
[0210] A client for sending a first language question, where the first language question describes the question to be answered in a first language;
[0211] A server connected to the client, for converting the first language question into a second language space representation based on a multi-language representation alignment method, and analyzing the second language space representation using a target language model to generate a target reply, where the multi-language representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation;
[0212] The client is also used to output the target reply.
[0213] In an embodiment of the present application, the language processing system may include a client and a server connected to the client.
[0214] For specific descriptions, reference may be made to the descriptions of the foregoing embodiments, which are not limited herein.
[0215] The above language processing method provided by the embodiments of the present application may be, but is not limited to, applied to application scenarios involving language processing including minority languages in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example: language processing scenarios in e-commerce services, language processing scenarios in education services, language processing scenarios in legal services, language processing scenarios in medical services, etc., which are not limited herein.
[0216] By adopting the embodiments of the present application, through the above language processing system, the purpose of efficiently and accurately generating answers to multi-language questions is achieved, thereby enhancing the cross-language understanding ability of the large language model, improving the accuracy and response ability of the large model in dealing with minority language questions, and expanding the analysis scope of the large model. Furthermore, the technical problems of language range limitation and analysis scope limitation existing in the related art when analyzing the internal representation of the large model through probe technology are solved.
[0217] It should be noted that the preferred implementation manners of this embodiment may be referred to the relevant descriptions in the embodiment, which will not be elaborated herein.
[0218] According to an embodiment of the present application, an apparatus embodiment for implementing the above language processing method is also provided. Figure 9 It is a schematic structural diagram of a language processing apparatus according to an embodiment of the present application, as Figure 9 shown. The apparatus includes:
[0219] The first acquisition module 901 is configured to acquire a language processing task, where the language processing task includes: a first language question, and the first language question uses a first language to describe a question to be answered;
[0220] The first conversion module 902 is configured to convert the first language question into a second language space representation based on a multi-language representation alignment method, where the multi-language representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation;
[0221] The first analysis module 903 is configured to analyze the second language space representation by using a target language model to generate a target reply.
[0222] Optionally, the apparatus further includes: a selection module, configured to select a second language from multiple candidate languages by using a probe verification method.
[0223] Optionally, the above selection module is further configured to: perform feature encoding on multiple candidate languages by using a target language model to obtain knowledge boundary representations corresponding to the multiple candidate languages; perform probe verification on the knowledge boundary representations by using a target probe model to determine the second language.
[0224] Optionally, the apparatus further includes: a training module, configured to train an initial probe model by using a multi-language knowledge boundary data set to generate a target probe model, where the multi-language knowledge boundary data set includes: correct and wrong fact classification data, question answerability classification data, and wrong premise classification data.
[0225] Optionally, the above training module is further configured to: determine multiple in-domain linear classifiers included in the initial probe model based on the number of model layers of the target language model and the number of languages of multiple candidate languages; determine a target representation of a question to be trained according to the multi-language knowledge boundary data set; train the multiple in-domain linear classifiers by using the target representation of the question to be trained to generate a target probe model.
[0226] Optionally, the above first conversion module 902 is further configured to: determine a first language subspace corresponding to the first language question and a second language subspace corresponding to the second language space representation; perform multi-language subspace linear alignment on the first language subspace and the second language subspace based on the multi-language representation alignment method to obtain the second language space representation.
[0227] Optionally, the multi - language representation alignment method includes: the multi - language subspace mean alignment method. The above - mentioned first conversion module 902 is further configured to: based on the multi - language subspace mean alignment method, calculate the mean of the first representation vectors of the first - language subspace and the mean of the second representation vectors of the second - language subspace respectively, and calculate the mean difference between the mean of the first representation vectors and the mean of the second representation vectors; align the vector dimensions of the first - language subspace and the second - language subspace according to the mean difference to obtain the second - language space representation.
[0228] Optionally, the multi - language representation alignment method includes: the multi - language subspace linear projection method. The above - mentioned first conversion module 902 is further configured to: based on the multi - language subspace linear projection method, perform singular value decomposition on the representation vectors of the first - language subspace to obtain a pseudo - matrix, and perform matrix multiplication on the pseudo - matrix and the representation vectors of the second - language subspace to obtain a linear projection matrix; perform vector representation conversion on the first - language subspace and the second - language subspace according to the linear projection matrix to obtain the second - language space representation.
[0229] Optionally, the apparatus further includes: a fine - tuning module, configured to obtain sample question - translation pairs; fine - tune the initial language model based on the sample question - translation pairs to obtain the target language model.
[0230] Optionally, the above - mentioned fine - tuning module is further configured to: obtain a third - language data subset from the multi - language knowledge boundary dataset; perform machine translation on the third - language data subset to obtain a fourth - language translation result, where the training data scale of the language used in the third - language data subset is higher than the training data scale of the language used in the fourth - language translation result; determine the sample question - translation pairs based on the third - language data subset and the fourth - language translation result.
[0231] By adopting the embodiments of the present application, by obtaining a language processing task, which includes: a first - language question described in a first language, and then based on the multi - language representation alignment method, converting the first - language question into a second - language space representation, where the multi - language representation alignment method is used to obtain the knowledge - boundary discrimination ability of the second language corresponding to the second - language space representation, and finally analyzing the second - language space representation using the target language model to generate a target response, the purpose of efficiently and accurately generating answers to multi - language questions is achieved, thereby enhancing the cross - language understanding ability of the large - language model, improving the accuracy and response ability of the large model in dealing with questions in minority languages, and expanding the analysis scope of the large model. Furthermore, the technical problem in the related art that there are limitations in the language scope and analysis scope in the internal representation analysis of the large model through the probe technology is solved.
[0232] It should be noted here that the above first acquisition module 901, first conversion module 902, and first analysis module 903 correspond to steps S21 to S23 in the embodiment. The instances and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above modules may also run in the server 10 provided in the embodiment.
[0233] According to an embodiment of the present application, there is also provided another device embodiment for implementing the above language processing method. Figure 10 It is a schematic structural diagram of another language processing device according to an embodiment of the present application, as Figure 10 shown. The device includes:
[0234] A second acquisition module 1001, configured to acquire a cross-language text classification task, where the cross-language text classification task includes: a first language question, and the first language question uses a minority language to describe the question to be answered;
[0235] A second conversion module 1002, configured to convert the first language question into a second language space representation based on a multi-language representation alignment method, where the multi-language representation alignment method is used to obtain the knowledge boundary discrimination ability of the international common language corresponding to the second language space representation;
[0236] A second analysis module 1003, configured to analyze the second language space representation using a target language model to generate a cross-language text classification result.
[0237] By using the embodiment of the present application, a cross-language text classification task is acquired, where the cross-language text classification task includes: a first language question, and the first language question uses a minority language to describe the question to be answered. Then, based on the multi-language representation alignment method, the first language question is converted into a second language space representation, where the multi-language representation alignment method is used to obtain the knowledge boundary discrimination ability of the international common language corresponding to the second language space representation. Finally, the second language space representation is analyzed using a target language model to generate a cross-language text classification result, thereby achieving the purpose of efficiently and accurately generating answers to multi-language questions, thus enhancing the cross-language understanding ability of the large language model, improving the accuracy and response ability of the large model when processing minority language questions, and expanding the analysis scope of the large model, and further solving the technical problems of language range limitations and analysis scope limitations in the related art by using probe technology to perform internal representation analysis on the large model.
[0238] It should be noted here that the above-mentioned second acquisition module 1001, second conversion module 1002, and second analysis module 1003 correspond to steps S41 to S43 in the embodiment. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above-mentioned module or unit can be a hardware component or a software component stored in the memory and processed by one or more processors, and the above-mentioned module can also run in the server 10 provided in the embodiment.
[0239] According to an embodiment of the present application, there is also provided another device embodiment for implementing the above language processing method. Figure 11 It is a schematic structural diagram of another language processing device according to an embodiment of the present application, as Figure 11 shown, the device includes:
[0240] A third acquisition module 1101, configured to acquire a language processing request through a first application programming interface, where the request data carried in the language processing request includes: a first language question, and the first language question uses a first language to describe the question to be answered;
[0241] A first return module 1102, configured to return a language processing response through a second application programming interface, where the response data carried in the language processing response includes: a target reply, and the target reply is generated by analyzing a second language space representation using a target language model, and the second language space representation is obtained by transforming the first language question based on a multi-language representation alignment method, and the multi-language representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation.
[0242] Adopting the embodiment of the present application, a language processing request is acquired through a first application programming interface, where the request data carried in the language processing request includes: a first language question, and the first language question uses a first language to describe the question to be answered, and then a language processing response is returned through a second application programming interface, where the response data carried in the language processing response includes: a target reply, and the target reply is generated by analyzing a second language space representation using a target language model, and the second language space representation is obtained by transforming the first language question based on a multi-language representation alignment method, and the multi-language representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation, thereby achieving the purpose of efficiently and accurately generating answers to multi-language questions, thus realizing the enhancement of the cross-language understanding ability of the large language model, improving the accuracy and response ability of the large model in processing small language questions, and expanding the analysis range of the large model, and further solving the technical problems of language range limitation and analysis range limitation in the related art by using probe technology to perform internal representation analysis on the large model.
[0243] It should be noted here that the above-mentioned third acquisition module 1101 and the first return module 1102 correspond to steps S51 and S52 in the embodiment. The instances and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above-mentioned module or unit can be a hardware component or a software component stored in the memory and processed by one or more processors, and the above-mentioned module can also run in the server 10 provided in the embodiment.
[0244] According to an embodiment of the present application, there is also provided another device embodiment for implementing the above-mentioned language processing method. Figure 12 It is a schematic structural diagram of another language processing device according to an embodiment of the present application, as Figure 12 shown. The device includes:
[0245] A fourth acquisition module 1201, configured to acquire a currently input language processing dialogue request, where the request data carried in the language processing dialogue request includes: a first language question, and the first language question uses a first language to describe a question to be answered;
[0246] A second return module 1202, configured to respond to the language processing dialogue request and return a language processing dialogue reply, where the information carried in the language processing dialogue reply includes: a target reply, and the target reply is generated after analyzing a second language space representation by using a target language model, and the second language space representation is obtained by transforming the first language question based on a multilingual representation alignment method, and the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation;
[0247] A display module 1203, configured to display the target reply in a graphical user interface.
[0248] By adopting the embodiments of the present application, a language processing dialogue request is obtained for the current input. Among them, the request data carried in the language processing dialogue request includes: a first language question, and the first language question uses a first language to describe the question to be answered. Then, in response to the language processing dialogue request, a language processing dialogue reply is returned. Among them, the information carried in the language processing dialogue reply includes: a target reply, and the target reply is generated after analyzing a second language space representation using a target language model. The second language space representation is obtained by transforming the first language question based on a multilingual representation alignment method, and the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation. Finally, the target reply is displayed in the graphical user interface, thereby achieving the purpose of efficiently and accurately generating answers to multilingual questions, thus enhancing the cross-language understanding ability of the large language model, improving the accuracy and response ability of the large model when processing minority language questions, and expanding the technical effect of the analysis scope of the large model. Furthermore, the technical problem of limited language scope and limited analysis scope in the related art when analyzing the internal representation of the large model through probe technology is solved.
[0249] It should be noted here that the above-mentioned fourth acquisition module 1201, second return module 1202, and display module 1203 correspond to steps S61 to S63 in the embodiment. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in the memory and processed by one or more processors, and the above-mentioned modules can also run in the server 10 provided in the embodiment.
[0250] According to the embodiments of the present application, another device embodiment for implementing the above language processing method is also provided. Figure 13 It is a schematic structural diagram of another language processing device according to the embodiments of the present application, as Figure 13 shown. The device includes:
[0251] A first display module 1301, configured to display a first language question on the operation interface in response to an input instruction acting on the operation interface;
[0252] A second display module 1302, configured to display a target reply on the operation interface in response to a processing instruction acting on the operation interface; wherein, the first language question uses a first language to describe the question to be answered, and the target reply is generated after analyzing a second language space representation using a target language model. The second language space representation is obtained by transforming the first language question based on a multilingual representation alignment method, and the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation.
[0253] By adopting the embodiment of the present application, by responding to an input instruction acting on an operation interface, a first-language question is displayed on the operation interface, and then by responding to a processing instruction acting on the operation interface, a target reply is displayed on the operation interface; wherein, the first-language question describes the question to be answered in a first language, the target reply is generated after analyzing a second-language spatial representation by a target language model, the second-language spatial representation is obtained after transforming the first-language question based on a multilingual representation alignment method, and the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second-language spatial representation, thereby achieving the purpose of efficiently and accurately generating answers to multilingual questions, thus enhancing the cross-language understanding ability of the large language model, improving the accuracy and response ability of the large model when dealing with questions in minority languages, and expanding the analysis scope of the large model, and further solving the technical problems of language range limitation and analysis scope limitation in the related art when analyzing the internal representation of the large model through probe technology.
[0254] It should be noted here that the above first display module 1301 and second display module 1302 correspond to steps S71 and S72 in the embodiment. The functions of the two modules are the same as those of the corresponding steps in terms of the implemented examples and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the above modules can also run in the server 10 provided in the embodiment.
[0255] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in the embodiments, but are not limited to the schemes provided in the embodiments.
[0256] An embodiment of the present application can provide a computing device. Figure 14 It is a structural block diagram of a computing device according to an embodiment of the present application. As Figure 14 shown, the computing device A may include: one or more ( Figure 14 only one is shown in the figure) processors 1402, a memory 1404, a storage controller, and a peripheral interface. The peripheral interface can be connected to a radio frequency module, an audio module, a display screen, etc., which are not limited here.
[0257] The above computing device A can be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a PC (Personal Computer), a model all-in-one machine, etc. And the model described in the above embodiments of the present application can be preset in the computing device.
[0258] Specifically, the computing device A can pre-install various types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multi-modal task processing, etc., so as to provide diverse model selections. In different product forms, the computing device A can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference and application, etc. In some product forms, the computing device A also supports model management, including but not limited to multi-type model management (supporting the management of various types of models such as discriminative and generative models), model version control (supporting the control of different model versions), model evaluation (evaluating the performance and effect of the model based on model evaluation tools), etc. In other product forms, the computing device A can also create applications based on the model, provide API invocation capabilities, and can call the model into the created application through the API interface. At the same time, an application management tool is provided to realize the management and monitoring of the application.
[0259] Furthermore, the computing device A can also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master AI technologies), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive and integrated AI development, training, deployment, and application device is provided.
[0260] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, to implement the methods in the above embodiments. The memory can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory can further include a memory remotely set relative to the processor, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.
[0261] The processor can call the executable program stored in the memory through the transmission device to execute the method described in any one of the above embodiments.
[0262] Those of ordinary skill in the art can understand that the structure shown Figure 14 is only schematic, and the computing device A can also be a terminal device such as a smart phone, a tablet computer, a handheld computer, and a Mobile Internet Device (MID), a PAD, etc. The Figure 14It does not limit the structure of the above computing device. For example, computing device A may also include more or fewer components (such as network interfaces, display devices, etc.) than those shown in the Figure 14 and may have a configuration different from that shown in the Figure 14 .
[0263] Embodiments of the present application may provide an electronic device. Figure 15 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 15 shown, the electronic device may include: an input / output device 152; a memory 154, and a processor 156, where the processor 156 is connected to the input / output device 152 and the memory 154 through a bus 158.
[0264] Among them, the memory may be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0265] The processor may call the executable program stored in the memory through a transmission device to execute the method described in any one of the above embodiments.
[0266] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0267] Embodiments of the present application also provide a computer-readable storage medium. Optionally, in this embodiment, the above computer-readable storage medium may be used to store the program code executed by the language processing method provided in the above embodiment.
[0268] Optionally, in this embodiment, the above computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.
[0269] Embodiments of the present application also provide a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements any of the above-described language processing methods.
[0270] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0271] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the units or modules can be in an electrical or other form.
[0272] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0273] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0274] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0275] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A language processing method, characterized in that, Including: Obtain a language processing task, where the language processing task includes: a first language question, and the first language question uses a first language to describe a question to be answered; Based on a multilingual representation alignment method, convert the first language question into a second language space representation, where the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation; Analyze the second language space representation using a target language model to generate a target response.
2. The language processing method according to claim 1, wherein The language processing method further includes: Use the target language model to perform feature encoding on multiple candidate languages to obtain knowledge boundary representations corresponding to the multiple candidate languages; Use a target probe model to perform probe verification on the knowledge boundary representations to determine the second language.
3. The language processing method according to claim 2, wherein The language processing method further includes: Based on the number of model layers of the target language model and the number of languages of the multiple candidate languages, determine multiple in-domain linear classifiers included in the initial probe model; Determine the target representation of the question to be trained according to a multilingual knowledge boundary data set, where the multilingual knowledge boundary data set includes: correct and wrong fact classification data, question answerability classification data, and wrong premise classification data; Use the target representation of the question to be trained to train the multiple in-domain linear classifiers to generate the target probe model.
4. The language processing method according to claim 1, characterized in that Based on the multilingual representation alignment method, converting the first language question into the second language space representation includes: Determine a first language subspace corresponding to the first language question and a second language subspace corresponding to the second language space representation; Based on the multilingual representation alignment method, perform multilingual subspace linear alignment on the first language subspace and the second language subspace to obtain the second language space representation.
5. The language processing method according to claim 4, characterized in that The multilingual representation alignment method includes: a multilingual subspace mean alignment method. Based on the multilingual representation alignment method, performing multilingual subspace linear alignment on the first language subspace and the second language subspace to obtain the second language space representation includes: Based on the multilingual subspace mean alignment method, calculate the first representation vector mean of the first language subspace and the second representation vector mean of the second language subspace respectively, and calculate the mean difference between the first representation vector mean and the second representation vector mean; Align the vector dimensions of the first language subspace and the second language subspace according to the mean difference to obtain the second language space representation.
6. The language processing method according to claim 4, wherein The multilingual representation alignment method includes: a multilingual subspace linear projection method. Based on the multilingual representation alignment method, performing multilingual subspace linear alignment on the first language subspace and the second language subspace to obtain the second language space representation includes: Based on the multilingual subspace linear projection method, perform singular value decomposition on the representation vectors of the first language subspace to obtain a pseudo matrix, and perform matrix multiplication on the pseudo matrix and the representation vectors of the second language subspace to obtain a linear projection matrix; Perform vector representation conversion on the first language subspace and the second language subspace according to the linear projection matrix to obtain the second language space representation.
7. The language processing method according to claim 1, characterized in that The language processing method further includes: Obtain a third language data subset from a multilingual knowledge boundary dataset; Perform machine translation on the third language data subset to obtain a fourth language translation result, where the training data scale of the language used in the third language data subset is higher than the training data scale of the language used in the fourth language translation result; Determine a sample question translation pair based on the third language data subset and the fourth language translation result; Fine-tune an initial language model based on the sample question translation pair to obtain the target language model.
8. A language processing method, characterized in that, Include: Obtain a cross-lingual text classification task, where the cross-lingual text classification task includes: a first language question, and the first language question uses a minority language to describe the question to be answered; Convert the first language question into a second language space representation based on a multilingual representation alignment method, where the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the international common language corresponding to the second language space representation; Analyze the second language space representation using the target language model to generate a cross-lingual text classification result.
9. A language processing method, characterized in that, Include: Obtain a language processing request through a first application programming interface, where the request data carried in the language processing request includes: a first language question, and the first language question uses a first language to describe the question to be answered; Return a language processing response through a second application programming interface, where the response data carried in the language processing response includes: a target reply, and the target reply is generated by analyzing the second language space representation using the target language model, and the second language space representation is obtained by converting the first language question based on a multilingual representation alignment method, and the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation.
10. A language processing method, characterized in that, Include: Obtain a current input language processing dialogue request, where the request data carried in the language processing dialogue request includes: a first language question, and the first language question uses a first language to describe the question to be answered; In response to the language processing dialogue request, return a language processing dialogue reply, where the information carried in the language processing dialogue reply includes: a target reply, and the target reply is generated by analyzing the second language space representation using the target language model, and the second language space representation is obtained by converting the first language question based on a multilingual representation alignment method, and the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language space representation; Display the target reply in a graphical user interface.
11. A language processing method, characterized in that, Include: In response to an input instruction acting on an operation interface, display a first language question on the operation interface; In response to a processing instruction acting on the operation interface, display a target reply on the operation interface; Among them, the first language problem describes the problem to be answered in the first language, the target reply is generated after the target language model analyzes the second language spatial representation, the second language spatial representation is obtained after transforming the first language problem based on the multilingual representation alignment method, and the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language spatial representation.
12. A language processing system, characterized in that, It includes: A client for sending a first language problem, where the first language problem describes the problem to be answered in the first language; A server connected to the client, configured to transform the first language problem into a second language spatial representation based on the multilingual representation alignment method, and analyze the second language spatial representation using a target language model to generate a target reply, where the multilingual representation alignment method is used to obtain the knowledge boundary discrimination ability of the second language corresponding to the second language spatial representation; The client is further configured to output the target reply.
13. An electronic device, characterized in that, It includes: A memory storing an executable program; A processor for running the program, where when the program runs, it executes the language processing method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, where when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the language processing method according to any one of claims 1 to 11.
15. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements the language processing method according to any one of claims 1 to 11.