Multi-language task execution method and device, equipment and medium
By constructing a cross-lingual word vector semantic space and cultural knowledge graph, the problem of lack of semantic understanding and cultural background in multilingual task processing is solved, and the deep fusion of multimodal information and accurate task execution are achieved.
Patent Information
- Application Number
- CN202511055840.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
Existing models suffer from inaccurate semantic understanding and action mapping, lack of cultural background knowledge integration, and fragmentation of multimodal information in multilingual and cultural contexts, leading to task execution bias.
By collecting text data and visual action data in multiple languages, a cross-lingual word vector semantic space is constructed after preprocessing. Multilingual semantic alignment is then performed, and a cultural knowledge graph is constructed for encoding. Finally, a gating fusion mechanism is used to achieve the fusion of multilingual and cultural knowledge.
It achieves deep integration of multimodal information in a multilingual environment, improves cultural awareness, and ensures the accuracy and applicability of task execution.
Smart Images

Figure CN120951244A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and medium for performing multilingual tasks. Background Technology
[0002] Currently, in VLA (Vision-Language-Action) and other task scenarios, the model has the following obvious shortcomings in handling tasks with multilingual and cultural backgrounds:
[0003] (1) Most models are trained based on single-language data. When processing multilingual instructions or descriptions, due to the large differences in semantic expression and grammatical structure between different languages, it is difficult to achieve accurate semantic understanding and action mapping.
[0004] For example, in a financial services hall, when facing foreign visitors, the robot may make mistakes in its actions due to the inability to align the semantics of the two languages.
[0005] (2) The model lacks consideration of cultural background knowledge. In different cultures, the same action may have different meanings, and the same semantic expression may correspond to different action norms. For example, in social scenarios, the ways of expressing "welcome" are completely different between Eastern and Western cultures. Because the model does not incorporate cultural knowledge, it cannot accurately execute actions that conform to local cultural habits in multicultural interaction scenarios, resulting in poor interaction effects.
[0006] For example, in the medical field, when it is necessary to hire foreign experts for consultation, the operating room's assistive robots may misunderstand the actions of foreign experts and give incorrect responses due to a lack of cultural knowledge integration.
[0007] (3) Existing multilingual processing technologies mostly adopt a simple machine translation and then processing mode, which severs the collaborative relationship between language and visual and action modalities, and cannot achieve deep integration of multimodal information in a multilingual environment. Summary of the Invention
[0008] In view of the above, it is necessary to provide a multilingual task execution method, apparatus, device and medium, which aims to solve the problem of task execution deviation caused by the failure to consider language alignment and cultural background.
[0009] A multilingual task execution method, the multilingual task execution method comprising:
[0010] In response to the execution instructions for the target multilingual task, text data and visual motion data in multiple languages are collected according to the execution instructions;
[0011] The text data and the visual motion data are preprocessed to obtain multilingual features;
[0012] Construct a cross-lingual word vector semantic space based on adversarial training according to the multilingual features;
[0013] Based on the cross-linguistic word vector semantic space, the multilingual features are semantically aligned to obtain multilingual aligned features.
[0014] Construct cultural knowledge graphs for the multiple languages and encode the cultural knowledge graphs to obtain graph embedding features;
[0015] The multilingual features, multilingual alignment features, and atlas embedding features are fused using a gating fusion mechanism to obtain fused features;
[0016] Retrieve the target model corresponding to the target multilingual task, and input the fusion features into the target model to obtain the target execution data of the target multilingual task;
[0017] The target multilingual task is executed based on the target execution data.
[0018] A multilingual task execution device, the multilingual task execution device comprising:
[0019] The acquisition unit is used to acquire text data and visual motion data in multiple languages in response to the execution instructions for the target multilingual task.
[0020] The preprocessing unit is used to preprocess the text data and the visual motion data to obtain multilingual features;
[0021] The construction unit is used to construct a cross-lingual word vector semantic space based on adversarial training according to the multilingual features;
[0022] An alignment unit is used to perform multilingual semantic alignment on the multilingual features according to the cross-lingual word vector semantic space to obtain multilingual aligned features.
[0023] The encoding unit is used to construct a cultural knowledge graph of the multiple languages and encode the cultural knowledge graph to obtain graph embedding features;
[0024] The fusion unit is used to fuse the multilingual features, the multilingual alignment features, and the graph embedding features through a gated fusion mechanism to obtain fused features;
[0025] An input unit is used to retrieve the target model corresponding to the target multilingual task and input the fused features into the target model to obtain the target execution data of the target multilingual task.
[0026] An execution unit is used to execute the target multilingual task based on the target execution data.
[0027] A computer device, the computer device comprising:
[0028] Memory, storing at least one instruction; and
[0029] The processor executes instructions stored in the memory to implement the multilingual task execution method.
[0030] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the multilingual task execution method.
[0031] As can be seen from the above technical solutions, this invention can collect text data and visual-action data in multiple languages according to execution instructions, and preprocess the text data and visual-action data to achieve standardized processing of multimodal data; construct a cross-language word vector semantic space based on adversarial training according to multilingual features to achieve preliminary word vector alignment; perform multilingual semantic alignment on multilingual features according to the cross-language word vector semantic space to obtain multilingual alignment features, further breaking down language barriers and achieving deep mapping and alignment between semantics of different languages; construct a multilingual cultural knowledge graph and encode the cultural knowledge graph to obtain graph embedding features, thereby introducing cultural knowledge and improving cultural perception ability; fuse multilingual features, multilingual alignment features and graph embedding features through a gating fusion mechanism to obtain fused features that have both linguistic and cultural attributes; retrieve the target model corresponding to the target multilingual task and input the fused features into the target model to obtain the target execution data of the target multilingual task; execute the target multilingual task according to the target execution data, thereby comprehensively integrating multilingual information and cultural knowledge to more accurately execute the multilingual task. Attached Figure Description
[0032] Figure 1 This is a flowchart of a preferred embodiment of the multilingual task execution method of the present invention.
[0033] Figure 2 This is a functional block diagram of a preferred embodiment of the multilingual task execution device of the present invention.
[0034] Figure 3 This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the multilingual task execution method of the present invention. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the multilingual task execution method of the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.
[0037] The multilingual task execution method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0038] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0039] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0040] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0041] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0042] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0043] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0044] S10, in response to the execution instruction for the target multilingual task, collect text data and visual motion data in multiple languages according to the execution instruction.
[0045] In this embodiment, the target multilingual task may include tasks such as VLA (Vision-Language-Action).
[0046] For example, the target multilingual task can be a robot guidance task in a financial scenario, or a voice response task for a smart diagnosis and treatment robot in a medical scenario.
[0047] In this embodiment, the execution instruction can be triggered by relevant personnel or automatically when data is detected to be uploaded to a designated platform or interface.
[0048] In this embodiment, the text data may include various text data in each language, such as instruction text, description text, etc.
[0049] In this embodiment, the visual motion data can be video, images, sensor data, etc., of the robot's motion execution process.
[0050] S11, preprocess the text data and the visual motion data to obtain multilingual features.
[0051] In this embodiment, the preprocessing of the text data and the visual motion data to obtain multilingual features includes:
[0052] We acquire natural language processing tools corresponding to each language, and use these tools to perform word segmentation, part-of-speech tagging, and morphological analysis on the text data of each language to obtain the analysis and processing results of the text data of each language. We then use pre-trained language models to encode the analysis and processing results of the text data of each language to obtain the language feature vectors of the text data of each language.
[0053] The visual feature vectors in the visual motion data are extracted using the Swin Transformer model, and the motion feature vectors in the visual motion data are extracted using a Temporal Convolutional Network (TCN).
[0054] The multilingual features are obtained by concatenating the language feature vectors of the text data for each language with the corresponding visual feature vectors and action feature vectors.
[0055] The natural language processing tools may include, but are not limited to, tools such as NLTK (Natural Language Toolkit) and spaCy (Special Language Processing Tool).
[0056] The language model may include, but is not limited to: BERT (Bidirectional Encoder Representations from Transformers) model for English, ERNIE (Enhanced Representation through kNowledge IntEgration) model for Chinese, and Japanese-BERT model for Japanese.
[0057] For example, in cross-border financial video conferences, conference speech texts in different languages (such as English, Chinese, and Japanese) are collected and processed using the tools and models mentioned above. At the same time, visual data such as facial expressions and gestures of the participants are collected through cameras, and motion data of the participants (such as signing actions) are collected through sensors. After preprocessing, this data is prepared for subsequent semantic understanding and motion analysis.
[0058] For example, in international telemedicine consultations, text data such as medical record descriptions and diagnostic suggestions in different languages are collected and processed to obtain language feature vectors. At the same time, patient image data (visual data) and medical staff operation data (such as surgical gestures) are collected and preprocessed for subsequent multimodal information fusion analysis.
[0059] Through the above embodiments, the standardized processing of multilingual text, visual, and action data has been achieved, transforming different types of data into a unified feature vector form. This lays the data foundation for subsequent multilingual semantic alignment and multimodal fusion, ensuring data consistency and processability.
[0060] S12, Construct a cross-lingual word vector semantic space based on adversarial training according to the multilingual features.
[0061] In this embodiment, the construction of a cross - language word vector semantic space based on adversarial training according to the multi - language features includes:
[0062] Construct a cross - language word vector mapping model based on adversarial training; wherein, the cross - language word vector mapping model includes an encoder and a decoder;
[0063] Map each language feature in the multi - language features to a shared semantic space through the encoder, and use the decoder to restore each language feature mapped to the shared semantic space to the word vector representation of the original language, thereby obtaining the cross - language word vector semantic space.
[0064] Specifically, an adversarial training mechanism can be introduced to train a discriminator to distinguish the source of word vectors, prompting the encoder to learn language - independent semantic representations, and optimizing the model parameters by minimizing the reconstruction error and the adversarial loss.
[0065] For example: through word vector representation alignment, "Apple" can be aligned with "苹果".
[0066] Through the above - mentioned embodiment, preliminary word vector representation alignment can be achieved.
[0067] S13. Perform multi - language semantic alignment on the multi - language features according to the cross - language word vector semantic space to obtain multi - language alignment features.
[0068] In this embodiment, the performing multi - language semantic alignment on the multi - language features according to the cross - language word vector semantic space to obtain multi - language alignment features includes:
[0069] Based on the word vector representations in the cross - language word vector semantic space, use the multi - head attention mechanism to calculate the attention coefficients between every two language features in the cross - language word vector semantic space;
[0070] Normalize the attention coefficients between every two language features to obtain the alignment weights between every two language features;
[0071] Perform multi - language semantic alignment on the multi - language features according to the alignment weights between every two language features to obtain the multi - language alignment features.
[0072] For example: through multi - language semantic alignment, the semantic feature vectors of the entire sentence "Je veux une pomme" in French and "我想要一个苹果" in Chinese can be aligned to understand the correspondence of the overall semantics of "want an apple" in the two language texts.
[0073] Through the above embodiments, based on the alignment of word vector representations, it is possible to further realize the deep mapping and alignment of semantics between different languages, break down language barriers, enable the model to more accurately understand the same semantics expressed in different languages, and provide a data foundation for multimodal information fusion in a multilingual environment.
[0074] S14, construct a cultural knowledge graph of the multiple languages, and encode the cultural knowledge graph to obtain graph embedding features.
[0075] In this embodiment, constructing the cultural knowledge graph of the multiple languages includes:
[0076] Collect cultural knowledge related to the context of each language;
[0077] Cultural concepts, cultural meanings of actions, and cultural meanings of language are extracted from the cultural knowledge as nodes, and the relationships between the nodes are extracted from the cultural knowledge as edges;
[0078] The cultural knowledge graph is obtained by constructing a graph based on the extracted nodes and edges.
[0079] The cultural knowledge mentioned may include, but is not limited to, knowledge of behavioral norms, semantic connotations, and social etiquette under different cultural backgrounds.
[0080] For example, the cultural concept may include "Eastern etiquette culture" and "Western social culture"; the cultural meaning of the action may include "nodding to indicate approval in the East"; the cultural meaning of the language may include the differences in the reference of specific words in different cultures; the relationship between nodes may include "inclusion" and "correspondence".
[0081] Through the above embodiments, cultural knowledge can be incorporated, and cultural perception capabilities can be acquired. When processing multimodal information, by integrating the influence of different cultural backgrounds, the applicability and accuracy in cross-cultural scenarios can be improved.
[0082] In this embodiment, the cultural knowledge graph can be encoded using a graph attention network, and the present invention does not limit the encoding method.
[0083] S15, the multilingual features, the multilingual alignment features, and the graph embedding features are fused through a gating fusion mechanism to obtain fused features.
[0084] In this embodiment, feature fusion can be used to embed multilingual and cultural knowledge, thereby improving the ability to perceive different languages and cultural backgrounds during task execution and improving the accuracy of task execution.
[0085] S16, retrieve the target model corresponding to the target multilingual task, and input the fusion features into the target model to obtain the target execution data of the target multilingual task.
[0086] In this embodiment, a model corresponding to each multilingual task can be pre-trained and stored locally for easy and quick retrieval.
[0087] In this embodiment, before inputting the fused features into the target model, the method further includes:
[0088] Training samples are constructed according to the data dimensions of the multilingual features;
[0089] Construct a multi-task loss function; wherein, the multi-task loss function includes multilingual semantic alignment loss, action generation error loss, and cultural knowledge matching loss;
[0090] Obtain the initial model corresponding to the target model;
[0091] Based on the multi-task loss function, the initial model is trained using the training samples to obtain the target model;
[0092] During training, the Adam optimizer with weight decay is used to update the model parameters through the backpropagation algorithm.
[0093] The training samples may include cross-cultural social scene videos, multilingual task instruction execution cases, etc., and are labeled with information such as language semantics, action tags, and cultural attributes.
[0094] During training, the weights of different loss terms are dynamically adjusted. For example, in scenarios with significant cultural differences, the weight of cultural knowledge matching loss can be increased.
[0095] Through the above embodiments, the model can learn the distribution patterns of multimodal data under different languages and cultures, which greatly improves its generalization ability in diverse scenarios and expands the application boundaries of the model in global multilingual and multicultural environments.
[0096] S17, Execute the target multilingual task according to the target execution data.
[0097] In this embodiment, performing the target multilingual task based on the target execution data includes:
[0098] When the target execution data is an action execution sequence corresponding to a language instruction execution task, control the target object corresponding to the execution instruction to execute the action execution sequence; or
[0099] When the target execution data is the voice response content corresponding to a multilingual dialogue task, response data in the language corresponding to the multilingual dialogue task is generated based on the voice response content.
[0100] For example, in an intelligent financial customer service system, when a customer inquires about financial products in different languages (e.g., asking about the returns of wealth management products in English and asking about the loan process in Chinese), the model uses multimodal reasoning to combine the customer's voice, facial expressions (visual data), and cultural knowledge in the financial field (e.g., different regions' ways of expressing risk) to generate accurate language responses that are in line with the customer's cultural background, or to generate corresponding operation instructions (e.g., actions to display product details).
[0101] For example, in remote surgical guidance, experts issue surgical instructions in different languages. The model combines images (visual data) from the surgical site, instrument motion data, and cultural norms in the medical field (such as differences in standard surgical procedures in different regions) to generate a detailed sequence of surgical action instructions, ensuring that the surgical assistant executes the operation accurately.
[0102] Through the above embodiments, reasoning can be performed by integrating multilingual information and cultural knowledge to generate execution data such as action instructions or language responses that meet task requirements and cultural habits. This achieves deep integration and application of multilingual and multimodal information, and improves the model's interactive effect and task execution capability in real-world scenarios.
[0103] As can be seen from the above technical solutions, this invention can collect text data and visual-action data in multiple languages according to execution instructions, and preprocess the text data and visual-action data to achieve standardized processing of multimodal data; construct a cross-language word vector semantic space based on adversarial training according to multilingual features to achieve preliminary word vector alignment; perform multilingual semantic alignment on multilingual features according to the cross-language word vector semantic space to obtain multilingual alignment features, further breaking down language barriers and achieving deep mapping and alignment between semantics of different languages; construct a multilingual cultural knowledge graph and encode the cultural knowledge graph to obtain graph embedding features, thereby introducing cultural knowledge and improving cultural perception ability; fuse multilingual features, multilingual alignment features and graph embedding features through a gating fusion mechanism to obtain fused features that have both linguistic and cultural attributes; retrieve the target model corresponding to the target multilingual task and input the fused features into the target model to obtain the target execution data of the target multilingual task; execute the target multilingual task according to the target execution data, thereby comprehensively integrating multilingual information and cultural knowledge to more accurately execute the multilingual task.
[0104] like Figure 2The diagram shown is a functional block diagram of a preferred embodiment of the multilingual task execution device of the present invention. The multilingual task execution device 11 includes a data acquisition unit 110, a preprocessing unit 111, a construction unit 112, an alignment unit 113, an encoding unit 114, a fusion unit 115, an input unit 116, and an execution unit 117. The module / unit referred to in this invention is a series of computer program segments that can be executed by a processor and perform a fixed function, and which are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0105] The acquisition unit 110 is used to acquire text data and visual motion data in multiple languages in response to the execution instruction of the target multilingual task.
[0106] In this embodiment, the target multilingual task may include tasks such as VLA (Vision-Language-Action).
[0107] For example, the target multilingual task can be a robot guidance task in a financial scenario, or a voice response task for a smart diagnosis and treatment robot in a medical scenario.
[0108] In this embodiment, the execution instruction can be triggered by relevant personnel or automatically when data is detected to be uploaded to a designated platform or interface.
[0109] In this embodiment, the text data may include various text data in each language, such as instruction text, description text, etc.
[0110] In this embodiment, the visual motion data can be video, images, sensor data, etc., of the robot's motion execution process.
[0111] The preprocessing unit 111 is used to preprocess the text data and the visual motion data to obtain multilingual features.
[0112] In this embodiment, the preprocessing unit 111 preprocesses the text data and the visual motion data to obtain multilingual features, including:
[0113] We acquire natural language processing tools corresponding to each language, and use these tools to perform word segmentation, part-of-speech tagging, and morphological analysis on the text data of each language to obtain the analysis and processing results of the text data of each language. We then use pre-trained language models to encode the analysis and processing results of the text data of each language to obtain the language feature vectors of the text data of each language.
[0114] The visual feature vectors in the visual motion data are extracted using the Swin Transformer model, and the motion feature vectors in the visual motion data are extracted using a Temporal Convolutional Network (TCN).
[0115] The multilingual features are obtained by concatenating the language feature vectors of the text data for each language with the corresponding visual feature vectors and action feature vectors.
[0116] The natural language processing tools may include, but are not limited to, tools such as NLTK (Natural Language Toolkit) and spaCy (Special Language Processing Tool).
[0117] The language model may include, but is not limited to: BERT (Bidirectional Encoder Representations from Transformers) model for English, ERNIE (Enhanced Representation through kNowledge IntEgration) model for Chinese, and Japanese-BERT model for Japanese.
[0118] For example, in cross-border financial video conferences, conference speech texts in different languages (such as English, Chinese, and Japanese) are collected and processed using the tools and models mentioned above. At the same time, visual data such as facial expressions and gestures of the participants are collected through cameras, and motion data of the participants (such as signing actions) are collected through sensors. After preprocessing, this data is prepared for subsequent semantic understanding and motion analysis.
[0119] For example, in international telemedicine consultations, text data such as medical record descriptions and diagnostic suggestions in different languages are collected and processed to obtain language feature vectors. At the same time, patient image data (visual data) and medical staff operation data (such as surgical gestures) are collected and preprocessed for subsequent multimodal information fusion analysis.
[0120] Through the above embodiments, the standardized processing of multilingual text, visual, and action data has been achieved, transforming different types of data into a unified feature vector form. This lays the data foundation for subsequent multilingual semantic alignment and multimodal fusion, ensuring data consistency and processability.
[0121] The construction unit 112 is used to construct a cross-language word vector semantic space based on adversarial training according to the multilingual features.
[0122] In this embodiment, the construction unit 112 constructs a cross - language word - vector semantic space based on adversarial training according to the multilingual features, including:
[0123] Construct a cross - language word - vector mapping model based on adversarial training; wherein, the cross - language word - vector mapping model includes an encoder and a decoder;
[0124] Map each language feature in the multilingual features to a shared semantic space through the encoder, and use the decoder to restore each language feature mapped to the shared semantic space to the word - vector representation of the original language, so as to obtain the cross - language word - vector semantic space.
[0125] Specifically, an adversarial training mechanism can be introduced to train a discriminator to distinguish the source of word vectors, prompting the encoder to learn language - independent semantic representations, and optimizing the model parameters by minimizing the reconstruction error and the adversarial loss.
[0126] For example: through word - vector representation alignment, "Apple" can be aligned with "苹果".
[0127] Through the above - mentioned embodiment, preliminary word - vector representation alignment can be achieved.
[0128] The alignment unit 113 is used to perform multilingual semantic alignment on the multilingual features according to the cross - language word - vector semantic space, so as to obtain multilingual alignment features.
[0129] In this embodiment, the alignment unit 113 performs multilingual semantic alignment on the multilingual features according to the cross - language word - vector semantic space, so as to obtain multilingual alignment features, including:
[0130] Based on the word - vector representations in the cross - language word - vector semantic space, use the multi - head attention mechanism to calculate the attention coefficients between every two language features in the cross - language word - vector semantic space;
[0131] Normalize the attention coefficients between every two language features to obtain the alignment weights between every two language features;
[0132] Perform multilingual semantic alignment on the multilingual features according to the alignment weights between every two language features to obtain the multilingual alignment features.
[0133] For example: through multilingual semantic alignment, the semantic feature vectors of the entire sentence "Je veux une pomme" in French and "我想要一个苹果" in Chinese can be aligned, and the corresponding relationship of the overall semantics of "want an apple" in the two language texts can be understood.
[0134] Through the above embodiments, based on the alignment of word vector representations, it is possible to further realize the deep mapping and alignment of semantics between different languages, break down language barriers, enable the model to more accurately understand the same semantics expressed in different languages, and provide a data foundation for multimodal information fusion in a multilingual environment.
[0135] The encoding unit 114 is used to construct a cultural knowledge graph of the multiple languages and to encode the cultural knowledge graph to obtain graph embedding features.
[0136] In this embodiment, the encoding unit 114 constructs the cultural knowledge graph of the multiple languages by:
[0137] Collect cultural knowledge related to the context of each language;
[0138] Cultural concepts, cultural meanings of actions, and cultural meanings of language are extracted from the cultural knowledge as nodes, and the relationships between the nodes are extracted from the cultural knowledge as edges;
[0139] The cultural knowledge graph is obtained by constructing a graph based on the extracted nodes and edges.
[0140] The cultural knowledge mentioned may include, but is not limited to, knowledge of behavioral norms, semantic connotations, and social etiquette under different cultural backgrounds.
[0141] For example, the cultural concept may include "Eastern etiquette culture" and "Western social culture"; the cultural meaning of the action may include "nodding to indicate approval in the East"; the cultural meaning of the language may include the differences in the reference of specific words in different cultures; the relationship between nodes may include "inclusion" and "correspondence".
[0142] Through the above embodiments, cultural knowledge can be incorporated, and cultural perception capabilities can be acquired. When processing multimodal information, by integrating the influence of different cultural backgrounds, the applicability and accuracy in cross-cultural scenarios can be improved.
[0143] In this embodiment, the cultural knowledge graph can be encoded using a graph attention network, and the present invention does not limit the encoding method.
[0144] The fusion unit 115 is used to fuse the multilingual features, the multilingual alignment features and the graph embedding features through a gating fusion mechanism to obtain fused features.
[0145] In this embodiment, feature fusion can be used to embed multilingual and cultural knowledge, thereby improving the ability to perceive different languages and cultural backgrounds during task execution and improving the accuracy of task execution.
[0146] The input unit 116 is used to retrieve the target model corresponding to the target multilingual task and input the fusion features into the target model to obtain the target execution data of the target multilingual task.
[0147] In this embodiment, a model corresponding to each multilingual task can be pre-trained and stored locally for easy and quick retrieval.
[0148] In this embodiment, before inputting the fused features into the target model, training samples are constructed according to the data dimensions of the multilingual features;
[0149] Construct a multi-task loss function; wherein, the multi-task loss function includes multilingual semantic alignment loss, action generation error loss, and cultural knowledge matching loss;
[0150] Obtain the initial model corresponding to the target model;
[0151] Based on the multi-task loss function, the initial model is trained using the training samples to obtain the target model;
[0152] During training, the Adam optimizer with weight decay is used to update the model parameters through the backpropagation algorithm.
[0153] The training samples may include cross-cultural social scene videos, multilingual task instruction execution cases, etc., and are labeled with information such as language semantics, action tags, and cultural attributes.
[0154] During training, the weights of different loss terms are dynamically adjusted. For example, in scenarios with significant cultural differences, the weight of cultural knowledge matching loss can be increased.
[0155] Through the above embodiments, the model can learn the distribution patterns of multimodal data under different languages and cultures, which greatly improves its generalization ability in diverse scenarios and expands the application boundaries of the model in global multilingual and multicultural environments.
[0156] The execution unit 117 is used to execute the target multilingual task according to the target execution data.
[0157] In this embodiment, the execution unit 117 performs the target multilingual task according to the target execution data, including:
[0158] When the target execution data is an action execution sequence corresponding to a language instruction execution task, control the target object corresponding to the execution instruction to execute the action execution sequence; or
[0159] When the target execution data is the voice response content corresponding to a multilingual dialogue task, response data in the language corresponding to the multilingual dialogue task is generated based on the voice response content.
[0160] For example, in an intelligent financial customer service system, when a customer inquires about financial products in different languages (e.g., asking about the returns of wealth management products in English and asking about the loan process in Chinese), the model uses multimodal reasoning to combine the customer's voice, facial expressions (visual data), and cultural knowledge in the financial field (e.g., different regions' ways of expressing risk) to generate accurate language responses that are in line with the customer's cultural background, or to generate corresponding operation instructions (e.g., actions to display product details).
[0161] For example, in remote surgical guidance, experts issue surgical instructions in different languages. The model combines images (visual data) from the surgical site, instrument motion data, and cultural norms in the medical field (such as differences in standard surgical procedures in different regions) to generate a detailed sequence of surgical action instructions, ensuring that the surgical assistant executes the operation accurately.
[0162] Through the above embodiments, reasoning can be performed by integrating multilingual information and cultural knowledge to generate execution data such as action instructions or language responses that meet task requirements and cultural habits. This achieves deep integration and application of multilingual and multimodal information, and improves the model's interactive effect and task execution capability in real-world scenarios.
[0163] As can be seen from the above technical solutions, this invention can collect text data and visual-action data in multiple languages according to execution instructions, and preprocess the text data and visual-action data to achieve standardized processing of multimodal data; construct a cross-language word vector semantic space based on adversarial training according to multilingual features to achieve preliminary word vector alignment; perform multilingual semantic alignment on multilingual features according to the cross-language word vector semantic space to obtain multilingual alignment features, further breaking down language barriers and achieving deep mapping and alignment between semantics of different languages; construct a multilingual cultural knowledge graph and encode the cultural knowledge graph to obtain graph embedding features, thereby introducing cultural knowledge and improving cultural perception ability; fuse multilingual features, multilingual alignment features and graph embedding features through a gating fusion mechanism to obtain fused features that have both linguistic and cultural attributes; retrieve the target model corresponding to the target multilingual task and input the fused features into the target model to obtain the target execution data of the target multilingual task; execute the target multilingual task according to the target execution data, thereby comprehensively integrating multilingual information and cultural knowledge to more accurately execute the multilingual task.
[0164] like Figure 3 The diagram shown is a schematic representation of the structure of a computer device that implements the multilingual task execution method of the present invention.
[0165] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a multilingual task execution program.
[0166] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.
[0167] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.
[0168] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as the code of multilingual task execution programs, but also to temporarily store data that has been output or will be output.
[0169] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., executing multilingual task execution programs) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.
[0170] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the various multilingual task execution method embodiments described above, for example... Figure 1 The steps are shown.
[0171] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into an acquisition unit 110, a preprocessing unit 111, a construction unit 112, an alignment unit 113, an encoding unit 114, a fusion unit 115, an input unit 116, and an execution unit 117.
[0172] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the multilingual task execution methods described in the various embodiments of this invention.
[0173] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.
[0174] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.
[0175] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.
[0176] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0177] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.
[0178] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0179] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish a communication connection between the computer device 1 and other computer devices.
[0180] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.
[0181] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0182] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0183] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement a multilingual task execution method, and the processor 13 can execute the multiple instructions to achieve:
[0184] In response to the execution instructions for the target multilingual task, text data and visual motion data in multiple languages are collected according to the execution instructions;
[0185] The text data and the visual motion data are preprocessed to obtain multilingual features;
[0186] Construct a cross-lingual word vector semantic space based on adversarial training according to the multilingual features;
[0187] Based on the cross-linguistic word vector semantic space, the multilingual features are semantically aligned to obtain multilingual aligned features.
[0188] Construct cultural knowledge graphs for the multiple languages and encode the cultural knowledge graphs to obtain graph embedding features;
[0189] The multilingual features, multilingual alignment features, and atlas embedding features are fused using a gating fusion mechanism to obtain fused features;
[0190] Retrieve the target model corresponding to the target multilingual task, and input the fusion features into the target model to obtain the target execution data of the target multilingual task;
[0191] The target multilingual task is executed based on the target execution data.
[0192] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0193] It should be noted that all data involved in this case was legally obtained. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0194] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0195] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0196] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0197] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0198] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0199] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0200] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for performing multilingual tasks, characterized in that, The multilingual task execution method includes: In response to the execution instructions for the target multilingual task, text data and visual motion data in multiple languages are collected according to the execution instructions; The text data and the visual motion data are preprocessed to obtain multilingual features; Construct a cross-lingual word vector semantic space based on adversarial training according to the multilingual features; Based on the cross-linguistic word vector semantic space, the multilingual features are semantically aligned to obtain multilingual aligned features. Construct cultural knowledge graphs for the multiple languages and encode the cultural knowledge graphs to obtain graph embedding features; The multilingual features, multilingual alignment features, and atlas embedding features are fused using a gating fusion mechanism to obtain fused features; Retrieve the target model corresponding to the target multilingual task, and input the fusion features into the target model to obtain the target execution data of the target multilingual task; The target multilingual task is executed based on the target execution data.
2. The multilingual task execution method as described in claim 1, characterized in that, The preprocessing of the text data and the visual motion data to obtain multilingual features includes: We acquire natural language processing tools corresponding to each language, and use these tools to perform word segmentation, part-of-speech tagging, and morphological analysis on the text data of each language to obtain the analysis and processing results of the text data of each language. We then use pre-trained language models to encode the analysis and processing results of the text data of each language to obtain the language feature vectors of the text data of each language. The visual feature vectors in the visual motion data are extracted using the Swing Transformer model, and the motion feature vectors in the visual motion data are extracted using a temporal convolutional network. The multilingual features are obtained by concatenating the language feature vectors of the text data for each language with the corresponding visual feature vectors and action feature vectors.
3. The multilingual task execution method as described in claim 1, characterized in that, The construction of a cross-lingual word vector semantic space based on adversarial training according to the multilingual features includes: A cross-language word vector mapping model is constructed based on adversarial training; wherein, the cross-language word vector mapping model includes an encoder and a decoder; The encoder maps each language feature in the multilingual features to a shared semantic space, and the decoder restores each language feature mapped to the shared semantic space to the word vector representation of the original language, thus obtaining the cross-lingual word vector semantic space.
4. The multilingual task execution method as described in claim 1, characterized in that, The step of performing multilingual semantic alignment on the multilingual features based on the cross-lingual word vector semantic space to obtain multilingual aligned features includes: Based on the word vector representation in the cross-linguistic word vector semantic space, the attention coefficient between each two language features in the cross-linguistic word vector semantic space is calculated using a multi-head attention mechanism; The attention coefficients between each pair of language features are normalized to obtain the alignment weights between each pair of language features. The multilingual features are semantically aligned based on the alignment weights between each pair of language features to obtain the multilingual aligned features.
5. The multilingual task execution method as described in claim 1, characterized in that, The construction of the cultural knowledge graph of the multiple languages includes: Collect cultural knowledge related to the context of each language; Cultural concepts, cultural meanings of actions, and cultural meanings of language are extracted from the cultural knowledge as nodes, and the relationships between the nodes are extracted from the cultural knowledge as edges; The cultural knowledge graph is obtained by constructing a graph based on the extracted nodes and edges.
6. The multilingual task execution method as described in claim 1, characterized in that, Before inputting the fused features into the target model, the method further includes: Training samples are constructed according to the data dimensions of the multilingual features; Construct a multi-task loss function; wherein, the multi-task loss function includes multilingual semantic alignment loss, action generation error loss, and cultural knowledge matching loss; Obtain the initial model corresponding to the target model; Based on the multi-task loss function, the initial model is trained using the training samples to obtain the target model; During training, the Adam optimizer with weight decay is used to update the model parameters through the backpropagation algorithm.
7. The multilingual task execution method as described in claim 1, characterized in that, The step of performing the target multilingual task based on the target execution data includes: When the target execution data is an action execution sequence corresponding to a language instruction execution task, control the target object corresponding to the execution instruction to execute the action execution sequence; or When the target execution data is the voice response content corresponding to a multilingual dialogue task, response data in the language corresponding to the multilingual dialogue task is generated based on the voice response content.
8. A multilingual task execution device, characterized in that, The multilingual task execution device includes: The acquisition unit is used to acquire text data and visual motion data in multiple languages in response to the execution instructions for the target multilingual task. The preprocessing unit is used to preprocess the text data and the visual motion data to obtain multilingual features; The construction unit is used to construct a cross-lingual word vector semantic space based on adversarial training according to the multilingual features; An alignment unit is used to perform multilingual semantic alignment on the multilingual features according to the cross-lingual word vector semantic space to obtain multilingual aligned features. The encoding unit is used to construct a cultural knowledge graph of the multiple languages and encode the cultural knowledge graph to obtain graph embedding features; The fusion unit is used to fuse the multilingual features, the multilingual alignment features, and the graph embedding features through a gated fusion mechanism to obtain fused features; An input unit is used to retrieve the target model corresponding to the target multilingual task and input the fused features into the target model to obtain the target execution data of the target multilingual task. An execution unit is used to execute the target multilingual task based on the target execution data.
9. A computer device, characterized in that, The computer device includes: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to implement the multilingual task execution method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the multilingual task execution method as described in any one of claims 1 to 7.
Citation Information
Cited By
Cross-language content generation method and device based on semantic enhanced knowledge graph
CN121146099A
Cross-language text classification and processing method and system based on deep transfer learning
CN121478978A