Hidden space communication method and model training method, device, storage medium and program product

CN121658930BActive Publication Date: 2026-09-25TAOBAO CHINA SOFTWARE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511649686.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-09-25
Estimated Expiration
2045-11-11

AI Technical Summary

Technical Problem

[0003]其中,将自然语言作为通信媒介,便于智能体理解与执行,但是,自然语言的信息承载能力有限,可能会造成信息损失,难易实现智能体之间意图的精准对齐,影响任务执行效果

Benefits of technology

[0011]在本申请实施例中,第一AI模型基于待处理任务的任务描述信息生成表征执行计划信息的隐空间消息,并将表征执行计划信息的隐空间消息直接传输给第二AI模型,第二AI模型基于表征执行计划信息的隐空间消息直接进行任务处理。由此,AI模型之间可以在隐空间直接进行通信,有效避免了将向量空间中的隐状态压缩为离散、自然语言时不可避免的信息损失,实现AI模型之间意图的精准对齐,提升了任务执行效果,大幅降低通信延迟与计算开销。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658930B_ABST
    Figure CN121658930B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a latent space communication method and a model training method, equipment, a storage medium and a program product. For the latent space communication method, a first AI model generates a latent space message representing execution plan information based on task description information of a to-be-processed task, and directly transmits the latent space message representing the execution plan information to a second AI model, and the second AI model directly processes the task based on the latent space message representing the execution plan information. Thus, the AI models can directly communicate in the latent space, effectively avoiding the inevitable information loss when compressing the latent state in the vector space into discrete natural language, achieving precise alignment of intentions between AI models, improving task execution effect, and reducing communication delay and computing overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a hidden space communication method and model training method, device, storage medium and program product. Background Technology

[0002] Intelligent agents are AI applications that use artificial intelligence (AI) models (such as large language models) as their core technology. They are capable of perceiving the environment, making decisions, and executing actions to complete specific tasks. Multi-agent collaboration, as a cutting-edge direction in the current AI field, aims to enable multiple intelligent agents to interact and coordinate through natural language to address the challenges of complex tasks.

[0003] Using natural language as a communication medium facilitates understanding and execution by intelligent agents. However, natural language has limited information carrying capacity, which may lead to information loss and make it difficult to achieve accurate alignment of intentions between intelligent agents, thus affecting the performance of tasks. Summary of the Invention

[0004] This application provides a latent space communication method, model training method, device, storage medium, and program product to achieve precise alignment of intentions between intelligent agents and improve task execution performance.

[0005] This application embodiment also provides a latent space communication method, including: obtaining task description information of a target task; invoking a first artificial intelligence (AI) model to generate a target latent space message in the latent space based on the task description information of the target task, wherein the target latent space message represents the execution plan information of the target task; and sending the target latent space message to a second AI model so that the second AI model can execute the target task based on the execution plan information represented by the target latent space message.

[0006] This application also provides a model training method, including: obtaining a first training sample, the first training sample including: a first latent space message and context information matching a first sample task; based on the first latent space message and context information, calling a second AI model to perform task processing on the first sample task to obtain at least one task result corresponding to the first training sample; updating the model parameters of the second AI model according to at least one task result corresponding to the first training sample, until the model training termination condition is met.

[0007] This application embodiment also provides a model training method, including: obtaining a second training sample based on a third latent space message, a fourth latent space message, and sample task description information matching a second sample task, wherein the length of the third latent space message is less than the length of the fourth latent space message, and the second training sample includes: first input information at least related to the third latent space message, second input information at least related to the fourth latent space message, and third input information only related to the sample task description information; calling a second AI model to perform task processing on the second sample task according to the first input information, the second input information, and the second input information respectively, to obtain at least one task result corresponding to the second training sample; updating the model parameters of the first AI model according to the at least one task result corresponding to the second training sample, until the model training termination condition is met.

[0008] This application also provides an electronic device, including: a memory and a processor; the memory for storing a computer program; the processor coupled to the memory for executing the computer program to perform steps in the latent space communication method and the model training method.

[0009] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the latent space communication method and the model training method.

[0010] This application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, enable the processor to implement the steps in the latent space communication method and the model training method.

[0011] In this embodiment, the first AI model generates a latent space message representing the execution plan information based on the task description information of the task to be processed, and directly transmits the latent space message representing the execution plan information to the second AI model. The second AI model then directly processes the task based on the latent space message representing the execution plan information. Thus, AI models can communicate directly in the latent space, effectively avoiding the unavoidable information loss when compressing the latent states in the vector space into discrete, natural language. This achieves precise alignment of intents between AI models, improves task execution performance, and significantly reduces communication latency and computational overhead. Attached Figure Description

[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a hidden space communication method provided in an embodiment of this application; Figure 2 This is an example of a communication principle diagram; Figure 3 A flowchart illustrating a model training method provided in this application embodiment; Figure 4 A flowchart illustrating another model training method provided in this application embodiment; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.

[0015] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0016] Figure 1 A flowchart illustrating a hidden space communication method provided in an embodiment of this application. See also... Figure 1 The method may include the following steps: 101. Obtain the task description information of the target task.

[0017] 102. Call the first artificial intelligence (AI) model to generate a target latent space message in the latent space based on the task description information of the target task. The target latent space message represents the execution plan information of the target task.

[0018] 103. Send the target latent space message to the second AI model so that the second AI model can execute the target task based on the execution plan information represented by the target latent space message.

[0019] The implicit space communication method provided in this application can be executed by a task processing system. This system can handle various tasks. For example, in a customer service scenario, it can understand user inquiries, make autonomous decisions, invoke service tools, and complete customer service tasks, achieving efficient 24 / 7 response. Similarly, in a code development scenario, it can perform tasks such as code generation, code completion, code testing, code release, or code repair, or simultaneously involve integrated tasks such as code development, testing, repair, and release. Furthermore, in the e-commerce field, it can perform order processing tasks, inventory query tasks, and promotion and marketing tasks.

[0020] See Figure 2 The task processing system can provide a first communication intelligent module and a second communication intelligent module that communicate with each other. The first communication intelligent module refers to an AI application that relies on a first AI model. This AI application needs to communicate with other AI applications, hence the name first communication intelligent module. Optionally, the AI ​​application based on the first AI model can also be called a first intelligent agent. Similarly, the second communication intelligent module refers to an AI application that relies on a second AI model. This AI application also needs to communicate with other AI applications, hence the name second communication intelligent module. Optionally, the AI ​​application based on the second AI model can also be called a second intelligent agent. The terms "first" and "second" are merely distinguishing names and do not constitute a limitation on order or quantity. In this application embodiment, the implementation form of the first or second AI model is not limited; it can be various neural network models based on deep learning. Optionally, the first or second AI model can be a deep learning model with a relatively small number of model parameters or a deep learning model with a relatively large number of model parameters. The large model is merely an example; this application embodiment does not limit the number of model parameters supported by the deep learning model used, aiming to meet actual needs. For example, the first AI model or the second AI model can be a language model (LM) or a multimodal model (MM) based on artificial intelligence, and there are no restrictions on this.

[0021] In this embodiment, the first AI model is primarily responsible for intent understanding and task planning; the second AI model is primarily responsible for task processing based on the execution plan information obtained from the task planning. For example, the first AI model (which can be considered a reasoning model) can perform intent understanding based on the task description information to be processed, and perform task planning based on the intent understanding result to obtain execution plan information. This execution plan information is used to describe the reasoning process required to execute the task to be processed, including but not limited to: the reasoning steps required to execute the task to be processed, the logic between the reasoning steps, and the decision conditions. In an optional embodiment, the execution plan information can be implemented as chain-of-thought (CoT) information required to execute the task to be processed, but is not limited to this. The chain-of-thought is a cue word technology, the purpose of which is to guide the AI ​​model to simulate the human thinking process to solve complex problems by generating a series of intermediate reasoning steps. In this embodiment, the first AI model provides the execution plan information (e.g., chain-of-thought information) to the second AI model (which can be considered a task execution model), and the second AI model simulates the human thinking process to process the task according to the execution plan information (e.g., chain-of-thought information) to obtain the task result.

[0022] In traditional approaches, the first and second AI models can communicate with each other in a language space. In this language space, the first and second AI models communicate using a Natural Language Communication (NLP) paradigm. For example, the first AI model generates and outputs a structured or unstructured natural language text (such as a question, instruction, reasoning steps, or conclusion). This structured or unstructured natural language text is then input into the second AI model, which performs tasks based on the input natural language text. See also... Figure 2 The task description information of the task to be processed is input into the first AI model. The first AI model outputs the execution plan information of the language space. The execution plan information of the language space is input into the second AI model. The second AI model performs task processing based on the execution plan information of the language space and outputs the task result of the task to be processed. However, using natural language as a communication medium, although convenient for human understanding and supervision, brings several limitations to the intelligent agent or communication intelligent module: (1) The AI ​​model needs to compress its rich hidden state into discrete, one-dimensional natural language text. This process inevitably loses a lot of information, which hinders the accurate alignment of intentions between intelligent agents or communication intelligent modules; (2) Part of the natural language text generated by the AI ​​model is for maintaining the coherence and readability of the language, rather than conveying the core information for solving the task. This leads to redundant computational overhead and communication bandwidth occupation.

[0023] In this embodiment, the first AI model and the second AI model can communicate in the latent space, a process referred to as latent space communication. The latent space is a continuous vector space within the AI ​​model used to represent and process information. In this embodiment, the dimension of this vector space is not limited and can be determined according to the application scenario. In some optional embodiments, the latent space can be implemented as a high-dimensional continuous vector space, while in other optional embodiments, it can be implemented as a low-dimensional continuous vector space. In this embodiment, the AI ​​model's thinking process mainly takes place in the latent space; correspondingly, latent space communication is a new communication paradigm proposed in this embodiment for communication between intelligent modules. In this communication paradigm, the intelligent modules directly transmit vector information (such as hidden states) from their latent spaces, rather than first converting the hidden states into natural language text before transmission. Taking the communication process between the first and second AI models in the latent space as an example, the first and second AI models directly exchange their internal hidden states (such as hidden vectors of a certain layer), instead of converting these hidden states into natural language text and communicating through natural language text, thereby achieving a transformation from the communication paradigm of the language space to the communication paradigm of the latent space. See also Figure 2 After the task description information of the task to be processed is input into the first AI model, the first AI model processes the task description information in the latent space and outputs the execution plan information of the latent space. The execution plan information of the latent space is the high-dimensional or low-dimensional vector information of the first AI model in the latent space, such as the hidden state sequence output by the last hidden layer of the first AI model. The execution plan information of the latent space is input into the second AI model, and the second AI model performs task processing based on the execution plan information of the latent space and outputs the task result of the task to be processed.

[0024] In this embodiment, compared to the execution plan information in the language space, the communication intelligent modules directly transmit the execution plan information in the latent space, instead of transmitting the discrete, one-dimensional natural language text formed by compressing the execution plan information in the latent space. This preserves a richer set of latent states in the vector space, eliminates information loss, and offers the advantage of high information fidelity, facilitating precise alignment of intentions between the communication intelligent modules. Furthermore, by directly transmitting the execution plan information in the latent space, the continuity and readability of the natural language do not need to be maintained. Therefore, the information required to maintain the continuity and readability of the natural language is not transmitted, reducing redundant information transmitted, saving computational overhead and communication bandwidth, and improving communication efficiency.

[0025] In this embodiment, for ease of understanding, the task to be processed is referred to as the target task. When the target task needs to be processed, firstly, the task description information of the target task is obtained, for example, "Why is there no liquid water on Mars?". Next, the task description information of the target task is input into the first communication intelligent module. The first communication intelligent module calls the first AI model to perform intent recognition and task planning, obtaining the target latent space message generated by the first AI model in the latent space. The target latent space message belongs to the latent space message category. For ease of partitioning and description, the latent space message generated for the target task is called the target latent space message. This target latent space message represents the execution plan information of the target task. Then, the first communication intelligent module sends the target latent space message to the second communication intelligent module. The second communication intelligent module calls the second AI model to execute the target task based on the execution plan information represented by the target latent space message. The task result of the target task is "The atmospheric pressure on the surface of Mars is extremely low, causing liquid water to be unable to exist stably." The input to the first communication intelligence module or the first AI model is task description information described in natural language, and its output is latent space messages. The input to the second communication intelligence module or the second AI model is the latent space messages output by the first communication intelligence module. Since the second communication intelligence module is user-oriented, its output is the task results described in natural language.

[0026] The implicit space communication method provided in this application is applied between two AI models, or between a first communication intelligence module and a second communication intelligence module based on AI models. The first AI model generates an implicit space message representing execution plan information based on the task description information of the task to be processed, and directly transmits the implicit space message representing the execution plan information to the second AI model. The second AI model directly processes the task based on the implicit space message representing the execution plan information. Therefore, communication intelligence modules based on AI models or AI models can communicate directly in the implicit space, effectively avoiding the unavoidable information loss when compressing implicit states into discrete, natural language. This achieves precise alignment of intentions between communication intelligence modules or AI models, improves task execution performance, significantly reduces communication latency and computational overhead, and increases communication efficiency between communication intelligence modules or AI models.

[0027] In this embodiment, the model structure of the first AI model is not limited. To facilitate understanding and description of the difference between latent space messages and natural language text in this embodiment, a general description of the model structure of the first AI model is provided. Optionally, the first AI model includes an input layer, multiple hidden layers, and an output layer, with the multiple hidden layers forming the latent space of the first AI model. The input layer's function is to convert the information input to the AI ​​model in natural language (e.g., task description information of the target task) into initial feature vectors suitable for processing in the latent space, i.e., to vectorize the information input to the AI ​​model in natural language. The multiple hidden layers are the true "intelligent core" and "thinking organ" of the AI ​​model, primarily responsible for gradually transforming the original, low-dimensional initial feature vectors into high-level, abstract feature representations. For example, through multi-level, non-linear complex calculations, the initial feature vectors can be transformed into meaningful abstract representations, completing a series of core cognitive tasks such as feature extraction, pattern recognition, knowledge storage, and logical reasoning in the process. It should be noted that the core cognitive tasks completed by the multiple hidden layers may differ depending on the model's function; this is not limited in this embodiment. The output layer is the interface through which the AI ​​model interacts with the external world. It is responsible for transforming the complex and abstract internal computation results completed by the hidden layer (i.e., the high-dimensional or low-dimensional, abstract vectors output by the hidden layer) into a meaningful final result that can be understood by the outside world (i.e., mapped to the natural language text required for the task).

[0028] In this embodiment, the specific implementation of the input layer, hidden layer, and output layer is not limited. The implementation of the input layer, hidden layer, and output layer can be flexibly designed according to the data type (such as image, text, audio, etc.), task type, and task requirements. Examples of the implementation methods for the input layer, hidden layer, and output layer are given below.

[0029] The input layer can be implemented using, but is not limited to, the following methods: Fully Connected Layer: Directly connects the input vector to the hidden layer, suitable for tabular data. Convolutional Layer: Used for image data to extract spatial features. Recurrent Neural Network (RNN): Suitable for sequence data, such as time series and text. Embedding Layer: Used to convert discrete categorical data (such as words, category IDs) into dense vectors. Positional Encoding: Provides positional information for sequence data in an encoder-decoder architecture. Feature Extractor: For example, uses a pre-trained model to extract features, then uses these features as input to the hidden layer. Data Preprocessing Layer: Such as normalization layers, standardization layers, and data augmentation layers (such as random cropping, flipping, etc.).

[0030] Hidden layers can be flexibly selected based on task requirements, but are not limited to the following implementation methods: Fully connected layers: Suitable for various types of data, especially data where there is no clear spatial or sequential relationship between features. Convolutional Neural Networks (CNNs): Used for grid data such as images and videos, extracting local features through convolutional kernels. Recurrent Neural Networks (RNNs) and their variants, such as Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRUs): Process sequential data, such as text and time series. Encoder-decoder networks: Based on self-attention mechanisms, suitable for sequential data. Graph Neural Networks (GNNs): Used for graph-structured data, such as Graph Convolutional Networks (GCNs) and Graph Attention Networks (GATs). Autoencoders: Used for dimensionality reduction and feature learning. Generative Adversarial Networks (GANs) generators and discriminators: Used for generation tasks. Attention mechanisms: Can be added on top of other layers to focus on important information. Normalization layers: such as batch normalization and layer normalization, are used to stabilize the training process.

[0031] Hidden layers can be flexibly selected based on task requirements, but are not limited to the following implementation methods: Fully connected layers and activation functions: For classification tasks, Softmax or Sigmoid activation functions can be used; for regression tasks, linear activation functions or ReLU, etc., can be used to ensure non-negative output. Convolutional layers: Used for pixel-level prediction, such as image segmentation (usually followed by a pixel-wise Softmax or Sigmoid). Recurrent neural network layers: Used for sequence generation tasks, such as machine translation and text generation. Output layer of encoder-decoder architecture: Connected to a linear layer and Softmax to generate the output distribution at each position in the sequence. Decoder of autoencoder: Reconstructs the input data; the output layer has the same shape as the input layer, and the activation function is selected according to the data range. CRF layer (Conditional Random Field): Used for sequence labeling tasks, jointly modeling the output sequence.

[0032] Based on the aforementioned first AI model, which includes an input layer, multiple hidden layers, and an output layer, the first AI model is invoked to generate a target latent space message in the latent space based on the task description information of the target task. This includes: inputting the task description information of the target task into the input layer for vectorization processing to obtain an initial feature vector; and performing nonlinear operations on the initial feature vector layer by layer through multiple hidden layers to obtain the hidden states of at least some of the hidden layers, which serve as the target latent space message.

[0033] Optionally, the first AI model can adopt an encoder-decoder architecture, an encoder-only architecture, or a decoder-only architecture; wherein each architecture includes an input layer and an output layer, and multiple hidden layers are distributed in the encoder and / or decoder of any architecture. That is, when the first AI model adopts an encoder-decoder architecture, the first AI model includes: an input layer, an encoder, a decoder, and an output layer; wherein the encoder includes multiple hidden layers, and the decoder includes multiple hidden layers; furthermore, in this architecture, the number of encoders and decoders can be one or more. When the first AI model adopts an encoder-only architecture, the first AI model includes: an input layer, an encoder, and an output layer; wherein the encoder includes multiple hidden layers; furthermore, in this architecture, the number of encoders can be one or more. When the first AI model adopts a decoder-only architecture, the first AI model includes: an input layer, a decoder, and an output layer; wherein the decoder includes multiple hidden layers; furthermore, in this architecture, the number of decoders can be one or more.

[0034] Based on the above model architecture, the above method of performing nonlinear operations on the initial feature vector layer by layer through multiple hidden layers to obtain the hidden states of at least some hidden layers as the target latent space message includes: if the first AI model adopts an encoder-decoder architecture, the initial feature vector is input into the encoder, encoded through multiple hidden layers in the encoder, and then decoded through multiple hidden layers in the decoder to obtain the hidden states of at least some hidden layers as the target latent space message; if the first AI model adopts an encoder-only architecture, the initial feature vector is input into the encoder, encoded through multiple hidden layers in the encoder to obtain the hidden states of at least some hidden layers as the target latent space message; if the first AI model adopts a decoder-only architecture, the initial feature vector is input into the decoder, decoded through multiple hidden layers in the decoder to obtain the hidden states of at least some hidden layers as the target latent space message.

[0035] Regardless of the model architecture described above, the target latent space message can originate from the hidden state sequence output by any one or more hidden layers in the first AI model. Each hidden layer's output sequence includes multiple hidden states, the number of which is related to the number of tokens included in the initial feature vector. The specific hidden state sequences from which the target latent space message originates can be determined based on task requirements and the model structure of the first AI model. Optionally, the target latent space message can originate from the hidden state sequence output by one or more hidden layers connected to the output layer. Further, optionally, in some model structures, the last hidden layer of the first AI model is connected to the output layer; in this case, the target latent space message can be the hidden state sequence output by the last hidden layer of the first AI model before outputting the natural language text. The target latent space message can be considered as the last hidden state sequence of the first AI model.

[0036] This document explains that the first AI model includes an output layer connected to its hidden layers. The output layer converts the hidden state sequence output by the connected hidden layers into natural language text. In other words, the hidden state sequence output by the hidden layers connected to the output layer in the first AI model is converted into natural language text after passing through the output layer. However, in this embodiment, within the first AI model, the hidden state sequence output by the hidden layers connected to the output layer no longer passes through the output layer but is directly output. That is, in this embodiment, during inference, the output layer in the first AI model is in a suspended state and no longer functions.

[0037] In the various embodiments of this application, the first AI model or the second AI model differs from traditional AI models in terms of model input, model output, and working principle. Instead, it possesses latent space communication capabilities. This capability allows the AI ​​model (or communication intelligence module) to directly transmit representations that naturally arise during computation, revealing its internal thought processes, rather than inferring thoughts from the linguistic space. This enables more direct, efficient, and information-dense collaboration. To further enable latent space communication capabilities in the AI ​​model or communication intelligence module, the embodiments of this application also provide the following model training methods to allow the AI ​​model to output, receive, and recognize latent space messages, providing conditions for communication based on these messages.

[0038] Figure 3 A flowchart illustrating a model training method provided in an embodiment of this application. See also... Figure 3 The method may include the following steps: 301. Obtain the first training sample, which includes: the first latent space message and context information that match the first sample task.

[0039] In practical applications, a second AI model can be trained to perform task processing based on latent space messages representing execution plan information. For ease of understanding and distinction, the training samples used to train the second AI model are referred to as first training samples, and there can be multiple first training samples. Steps 301 to 303 can be iteratively executed using multiple first training samples until the model training termination condition is met, resulting in a trained second AI model. The model training termination condition may include, for example, meeting the required number of training iterations or the convergence of the model parameters of the second AI model, but is not limited to these.

[0040] Specifically, the task involved in the first training sample is referred to as the first sample task. The task description information of the first sample task is input into the first AI model. The latent space message generated by the first AI model in the latent space based on the task description information of the first sample task is obtained as the first latent space message that matches the first sample task. The first latent space message represents the execution plan information of the first sample task.

[0041] When preparing the first training sample, the first latent space message and context information that match the first sample task are obtained. The context information may come from at least part of the task description information of the first sample task.

[0042] 302. Based on the first latent space message and context information, call the second AI model to perform task processing on the first sample task, so as to obtain at least one task result corresponding to the first training sample.

[0043] Optionally, during the preparation of the first training samples, a second latent space message that does not match the first sample task can also be obtained. For example, information can be removed or noise can be added to the first latent space message that matches the first sample task to obtain a second latent space message that does not match the first sample task. Alternatively, latent space messages that match other sample tasks can be used as the second latent space message that does not match the first sample task. These latent space messages that match other sample tasks are generated in the latent space by the first AI model based on the task description information of other sample tasks, which are different from the first sample task.

[0044] Optionally, in preparing the first training sample, the target language space message corresponding to the first sample task can also be obtained. The target language space message is used to represent the execution plan information of the first sample task.

[0045] Specifically, the task description information of the first sample task is input into the first AI model, and the execution plan information generated by the first AI model in the language space based on the task description information of the first sample task is obtained as the target language space message that matches the first sample task. The target language space message can be a piece of execution plan information described in natural language.

[0046] Optionally, based on the first latent space message and context information, a second AI model is invoked to perform task processing on the first sample task to obtain at least one task result, including performing at least one of the following task processing: Method 1: Based on the first latent space message and context information, the second AI model is invoked to perform task processing on the first sample task, resulting in a first task result. The first task result includes the predicted probability of the expected word at at least one time step or the predicted probability of at least one first word in the vocabulary at at least one time step. The predicted probability of the expected word at at least one time step is used as the first part of the first task result; the predicted probability of at least one first word in the vocabulary at at least one time step is used as the second part of the first task result.

[0047] In this context, the expected tokens are the tokens in the expected task result of the first sample task, used to supervise the accuracy of the second AI model's task execution. The expected task result can be understood as the standard, accurate task result of the first sample task. For example, if the task description information of the first sample task is "Why is there no liquid water on Mars?", the expected task result of the first sample task would be "The atmospheric pressure on the surface of Mars is extremely low, causing liquid water to be unable to exist stably." Each word in "The atmospheric pressure on the surface of Mars is extremely low, causing liquid water to be unable to exist stably" can be considered an expected token.

[0048] In practical applications, the second AI model processes the task step by step. Let S be the number of terms in the expected task result, and let t be the expected term in the expected task result. The value of t can be any positive integer from 1 to S. The second AI model will perform S time steps during task processing, and the t-th word... The predicted probability is denoted as , The context information required to input into the second AI model for executing the t-th time step. For the first hidden space message; at time step t, will and The input is used for task processing in the second AI model, and the output of the second AI model at time step t is... The predicted probability.

[0049] Suppose that the word sequence obtained by segmenting "task description information of the first sample task" is denoted as... , Each is a word element; during the execution of the first time step, It can be When predicting subsequent lexical units after the first lexical unit, the already predicted lexical units can be added as supplementary contextual information. In the middle, at this time, include And in prediction The previously predicted word groups.

[0050] The vocabulary includes multiple lexical units. These units can include expected lexical units from the expected task result of the first sample task, or other lexical units. The first lexical unit comes from the vocabulary and is the lexical unit predicted by the second AI model at each time step. For the t-th time step, the prediction probability of at least one first lexical unit in the vocabulary output at the t-th time step is denoted as... For example, at the first time step, the second AI model outputs a prediction probability of 0.92 for "fire," 0.03 for "earth," and 0.02 for "water." "Fire," "earth," and "water" are derived from lexical units in the vocabulary. Including the predicted probabilities of "fire", "earth", and "water".

[0051] Optionally, the first task result also includes the log odds for at least one first word in the vocabulary at each of at least one time step; the log odds for at least one first word in the vocabulary at each of at least one time step are considered as the third part of the first task result. The log odds for at least one first word in the vocabulary output at the t-th time step are denoted as... .

[0052] Method 2: Based on the second latent space message and context information, call the second AI model to perform task processing on the first sample task to obtain the second task result. The second task result includes the prediction probability of at least one second word element in the vocabulary corresponding to each time step.

[0053] Suppose the second latent space message that does not match the first sample task is denoted as... The second word element comes from the vocabulary list, and it represents the word elements predicted by the second AI model at each time step. For the t-th time step, the prediction probability of the second AI model for at least one second word element in the vocabulary list at the t-th time step is denoted as... .

[0054] Method 3: Based on the target language space message and context information corresponding to the first sample task, the second AI model is invoked to perform task processing for the first sample task, resulting in the third task result. The third task result includes the prediction probability for at least one third word element in the vocabulary at each of at least one time step. The prediction probability for at least one third word element in the vocabulary at each of at least one time step can be used as the first part of the third task result.

[0055] Assume the target language space message corresponding to the first sample task is denoted as P, where P represents the execution plan information of the first sample task. The third word element comes from the vocabulary, and the third word element is each word element predicted by the second AI model at each time step. For the t-th time step, the prediction probability of the second AI model at the t-th time step for at least one third word element in the vocabulary is denoted as... .

[0056] Optionally, the second part of the third task results includes: the log odds for at least one third word in the vocabulary at each time step.

[0057] For the t-th time step, the log probability of the output at the t-th time step for at least one third word in the vocabulary is denoted as . .

[0058] In practical applications, tasks can be processed using one or more of the three methods mentioned above, without any restrictions.

[0059] 303. Update the model parameters of the second AI model based on at least one task result corresponding to the first training sample until the model training termination condition is met.

[0060] In practical applications, the model parameters of the second AI model can be updated based on one or more of the results of the first, second, and third tasks until the model training termination condition is met.

[0061] Optionally, the task loss value can be determined based on the first part of the results from the first task and the cross-entropy loss function. Update the model parameters of the second AI model based on the task loss value. For example, calculate the task loss value of the first training sample according to formula (1). : ……(1) in, Let S be the predicted probability of the expected word at time step t, and the total number of time steps is S.

[0062] Specifically, the task loss value of the first training sample This enables the second AI model or a communication intelligence module based on the second AI model to accurately perform task processing, thereby improving the accuracy of task execution.

[0063] Optionally, the separation loss can be determined based on the second part of the results from the first task, the results from the second task, and the JS divergence function. Based on the separation loss, update the model parameters of the second AI model.

[0064] ……(2) in, The predicted probability for at least one first word element in the vocabulary is output at the t-th time step in the first task result. is the predicted probability of at least one second word in the vocabulary at the t-th time step output in the second task result; JS (Jensen-Shannon Divergence) is the JS divergence function, which is a measure of the similarity between two probability distributions.

[0065] Specifically, separation loss By employing a contrastive learning approach, it is possible to effectively distinguish between latent space information that matches the task at hand and latent space information that does not. This separation loss... This forces the second model to understand task-related information in the latent space information message, rather than simply matching surface statistical features.

[0066] Optionally, the plan alignment loss can be determined based on the second part of the results from the first task, the results from the third task, and the KL divergence function. Update the model parameters of the second AI model based on the plan alignment loss. For example, calculate the plan alignment loss according to formula (3). : ……(3) in, The predicted probability for at least one first word element in the vocabulary is output at the t-th time step in the first task result. This represents the predicted probability for at least one third word in the vocabulary, output at time step t in the results of the third task; KL() is the KL (Kullback-Leibler Divergence) divergence function, also known as relative entropy, used to measure the difference between two probability distributions. The larger its value, the greater the difference between the two probability distributions; the relative entropy is 0 when the two probability distributions are completely equal.

[0067] Optionally, the plan alignment loss can be determined based on the second part of the results from the first task, the results from the third task, and the KL divergence function. This includes: determining a first alignment loss value based on the second part of the results from the first task, the first part of the results from the third task, and the KL divergence function; determining a second alignment loss value based on the third part of the results from the first task, the second part of the results from the third task, and the cosine similarity function; and determining the planned alignment loss based on the first and second alignment loss values. For example, the plan alignment loss is calculated according to formula (4). : ……(4) Where α is a constant coefficient that can be set as needed, and β is a constant coefficient that can be set as needed; The log probability of at least one first word in the vocabulary, output at time step t in the first task result; The log probability of at least one third word in the vocabulary for the output at time step t in the third task result.

[0068] Specifically, the plan alignment loss is introduced. This allows the prediction results under the task execution plan conditions using latent space messages to be aligned with the prediction results under the corresponding natural language task execution plan conditions, effectively improving the task execution performance of the second AI model or the communication intelligence module based on the second AI model.

[0069] In practical applications, the task loss value can be based on the training samples. Separation loss Losses aligned with plan The total loss value of one or more training samples Update the model parameters of the second AI model based on the total loss value.

[0070] For example, a joint optimization objective containing three loss functions can be designed, and the total loss value of the training samples can be calculated according to formula (5). : ……(5) The technical solution provided in this application embodiment can train a second AI model to perform task processing based on the execution plan information represented by latent space messages, communicate directly in the latent space, and use the high information density of the hidden state as the latent space message. The information density of the latent space message is significantly higher than that of natural language, and more key task information can be carried in fewer positions. This fundamentally breaks through the language space bandwidth bottleneck and redundancy burden, reduces the expression bandwidth limitation and redundancy problem in language communication, and truly allows "models to communicate with each other through ideas rather than language".

[0071] Furthermore, in this embodiment, considering that the vectorization characteristics of latent space can encode information with much higher efficiency than linguistic space, the vectorization characteristics of latent space can be used to compress latent space messages transmitted during latent space communication. This allows the first AI model to generate and output shorter but more informative latent space messages, thereby improving communication efficiency and saving communication bandwidth and computing resources. Based on this, this embodiment also provides another model training method for training an AI model capable of compressing latent space messages. This allows the AI ​​model to carry key task information with fewer vector positions, thereby maintaining task execution performance while improving the compression ratio.

[0072] Figure 4 A flowchart illustrating another model training method provided in this application embodiment. See also... Figure 4 The method may include the following steps: 401. Obtain a second training sample based on the third latent space message, the fourth latent space message, and the sample task description information that match the second sample task, wherein the length of the third latent space message is less than the length of the fourth latent space message, and the second training sample includes: first input information that is at least related to the third latent space message, second input information that is at least related to the fourth latent space message, and third input information that is only related to the sample task description information.

[0073] In practical applications, model training can enable the first AI model to generate latent space messages representing task execution plan information based on task description information, and also to compress latent space messages, meaning the first AI model can output compressed latent space messages. For ease of understanding and distinction, the training samples used to train the first AI model are referred to as second training samples, and there can be multiple second training samples. Steps 401 to 403 can be iteratively executed using multiple second training samples until the model training termination condition is met, resulting in a trained first AI model. The model training termination condition may include, for example, meeting the required number of training iterations or the convergence of the model parameters of the first AI model, but is not limited to these.

[0074] Specifically, the task involved in the second training sample is called the second sample task, and the task description information of the second sample task is called the sample task description information.

[0075] The first AI model can generate latent space messages using an autoregressive approach according to formula (6): ……(6) Specifically, the sample task description information for the second sample task includes multiple word units, where the word vector sequence corresponding to the first i word units is denoted as... ;Will Input the first AI model, and the first AI model outputs the hidden state. ; and through the Generate vectors by performing projection operations ,Will and By concatenating or superimposing the sequences, we obtain the word vector sequence input at the next time step. ; Indicates will Input the first AI model for processing. Indicates splicing or overlaying operations, symbol This indicates a loop.

[0076] Specifically, by employing a continuous latent space autoregressive approach, the inference process of the first AI model avoids or reduces the overhead of going through the decoder, tokenizer (word segmentation module), and then encoding. While maintaining performance, it achieves extreme compression at the level of multiple words (e.g., 8 tokens), enabling latent space messages to carry an information density far exceeding that of natural language.

[0077] Specifically, the sample task description information of the second sample task can be input into the first AI model, and the first AI model generates a third hidden space message using an autoregressive method. The third hidden space message represents the execution plan information of the second sample task.

[0078] Specifically, the sample task description information of the second sample task can be input into the first AI model to obtain the fourth hidden space message output by the first AI model. The fourth hidden space message represents the execution plan information of the second sample task.

[0079] In practical applications, the length of the latent space message output by the first AI model can be effectively controlled by limiting its output length. For example, the output length of the first AI model can be limited by setting a maximum allowed number of output tokens. The larger the maximum allowed number of output tokens, the longer the latent space message output by the first AI model; the smaller the maximum allowed number of output tokens, the shorter the latent space message output by the first AI model.

[0080] In comparison, the length of the third latent space message is shorter than that of the fourth latent space message; the third latent space message is considered a compressed latent space message, while the fourth latent space message is considered an uncompressed latent space message. Compared to the fourth latent space message, the third latent space message can carry critical task information with fewer vector positions, thus maintaining task execution performance while improving the compression ratio.

[0081] In comparison, the first AI model performs fewer encoding and decoding operations when generating the third latent space message than when generating the fourth latent space message. For example, the first AI model performs no encoding or decoding operations when generating the third latent space message. The first AI model performs multiple encoding and decoding operations when generating the fourth latent space message. When preparing the second training sample, at least one of the following can be accurately selected: first input information at least related to the third latent space message, second input information at least related to the fourth latent space message, and third input information at least related to the sample task description information. Optionally, the first input information may be related only to the third latent space message, or the first input information may be related to the third latent space message and the sample task description information of the second sample task, but this is not a limitation.

[0082] Optionally, the second input information may be related only to the fourth latent space message, or the second input information may be related to the fourth latent space message and the sample task description information, but is not limited to these.

[0083] For example, the first input information is denoted as , The second input information is denoted as , The third input information can be recorded as: , .

[0084] in, The sample task description information represents the vectorized form of the second sample task; A specific identifier representing a vectorized form Specific identifier A marker pointing to the start of a hidden space message; A specific identifier representing a vectorized form Specific identifier A marker indicating the end of a hidden space message; This indicates a message in the third hidden space. The results obtained by performing various processes such as multi-head attention mechanisms or projection processing. messages in the fourth hidden space The results obtained by performing various processes such as multi-head attention mechanisms or projection processing.

[0085] Will , , , Perform operations such as concatenation or vector addition to obtain the first input information. .Will , , , Perform operations such as concatenation or vector addition to obtain the second input information. .Will As the third input information .

[0086] 402. Based on the first input information, the second input information, and the third input information, respectively, call the second AI model to perform task processing on the second sample task, so as to obtain at least one task result corresponding to the second training sample.

[0087] Specifically, tasks can be processed through one or more paths. Processing a task using the first input information can be called path A, processing a task using the second input information can be called path D, and processing a task using the third input information can be called path B.

[0088] For path A, the first input information (such as...) The input to path A is a second AI model (e.g., a task execution model), and the result of the fourth task is obtained by the second AI model performing task processing. Since the input information of path A is sufficient, the accuracy of the output result of the second AI model is usually high, or in other words, the uncertainty of the output result of the second AI model is low.

[0089] For path D, the second input information (such as...) Input the second AI model and obtain the fifth task result obtained by the second AI model in performing task processing. Since the input information of path D is sufficient, the accuracy of the output result of the second AI model is usually high, or in other words, the uncertainty of the output result of the second AI model is low.

[0090] For path B, the third input information (such as...) Input the second AI model and obtain the sixth task result obtained by the second AI model in performing task processing. Since path B only inputs... If the input information is insufficient, the accuracy of the output of the second AI model is usually low, or in other words, the uncertainty of the output of the second AI model is high.

[0091] Optionally, the first part of the results of the fourth task includes the prediction probability of the expected word element corresponding to at least one time step, and the second part of the results of the fourth task includes the prediction probability of at least one word element in the vocabulary corresponding to at least one time step.

[0092] Optionally, the first part of the results of the fifth task includes the prediction probability of the expected word element corresponding to at least one time step, and the second part of the results of the fifth task includes the prediction probability of at least one word element in the vocabulary corresponding to at least one time step.

[0093] Optionally, the first part of the results of the sixth task includes the prediction probability of the expected word element corresponding to at least one time step, and the second part of the results of the sixth task includes the prediction probability of at least one word element in the vocabulary corresponding to at least one time step.

[0094] Optionally, the information entropy associated with the target task result is determined based on the prediction probability of at least one word in the vocabulary at each time step in the target task result. The target task result can be any one of the fourth, fifth, or sixth task results.

[0095] Information entropy typically refers to a measure of a model's uncertainty about a given output, and is often used to evaluate a model's confidence level or prediction quality. Optionally, information entropy can be calculated according to formula (7). .

[0096] ……(7) in, This represents the word in the vocabulary output at time step t in the target task results. The predicted probability of each word element. This indicates the expression for all lexical elements. Perform summation.

[0097] 403. Update the model parameters of the first AI model based on at least one task result corresponding to the second training sample until the model training termination condition is met.

[0098] In practical applications, the model parameters of the first AI model can be updated based on one or more of the results of the fourth, fifth, and sixth tasks until the model training termination condition is met.

[0099] Optionally, based on the predicted probabilities of the expected words at at least one time step in the fourth task results and the cross-entropy loss function, the task loss value of the second training sample is determined; based on the task loss value of the second training sample, the model parameters of the first AI model are updated. Let the task loss value be denoted as... The task loss value is determined according to formula (8).

[0100] ……(8) The meaning of formula (8) is similar to that of formula (1) mentioned above. Please refer to the relevant introduction of formula (1) mentioned above.

[0101] Optionally, the uncertainty loss value of the second training sample is determined based on the results of the fifth task and its corresponding information entropy, the information entropy corresponding to the results of the sixth task, and the results of the fourth task; the model parameters of the first AI model are then updated based on the uncertainty loss value of the second training sample. Further, optionally, when determining the uncertainty loss value of the second training sample, weight coefficients can be generated based on the information entropy corresponding to the results of the fifth and sixth tasks, and the uncertainty loss value of the second training sample can be determined based on the results of the fourth and fifth tasks and the KL divergence function. The uncertainty loss value enables the compressed latent space information to reproduce the distributional behavior of the uncompressed latent space information at positions where uncertainty could have been significantly reduced, thereby maximizing the information carrying efficiency of the compressed latent space information.

[0102] Assume the uncertainty loss value is denoted as The uncertainty loss value is determined according to formula (9).

[0103] ……(9) in, The temperature coefficient is set as needed to adjust the smoothness of the probability distribution.

[0104] in, The weighting coefficients reflect the information value of the latent space message at time step t; where, they are determined according to formula (10). : ……(10) Wherein, the information entropy associated with the result of the sixth task determined according to formula (7) at time step t is expressed as: The information entropy associated with the fifth task result at time step t, as determined by formula (7), is expressed as follows: Here, max() represents the function that takes the maximum value.

[0105] in, Indicates input In the case of the first AI model outputting the predicted probability of at least one word element in the vocabulary at the t-th time step; in, Indicates input In the case of the first AI model outputting the predicted probability of at least one word element in the vocabulary at the t-th time step; in, , Here, softmax() is a normalization function; For input In the case of the first AI model outputting the original log odds for at least one word element in the vocabulary at the t-th time step; For input In the case of the first AI model outputting the original log odds for at least one word in the vocabulary at the t-th time step.

[0106] Optionally, the first step mean direction vector is determined based on the results of the fourth task, the second step mean direction vector is determined based on the results of the fifth task, and the geometric alignment loss value of the second training sample is determined based on the first step mean direction vector, the second step mean direction vector, and the cosine similarity function; the model parameters of the first AI model are updated based on the geometric alignment loss value of the second training sample.

[0107] In this application embodiment, the implementation of determining the first average direction vector based on the fourth task result and the second average direction vector based on the fifth task result is not limited. Optionally, determining the first average direction vector based on the fourth task result includes: averaging the prediction probabilities for at least one word in the vocabulary across all time steps in the fourth task result to obtain the first average direction vector; or, selecting a portion of time steps from the fourth task result and averaging the prediction probabilities for at least one word in the vocabulary across the selected time steps to obtain the first average direction vector. Correspondingly, determining the second average direction vector based on the fifth task result includes: averaging the prediction probabilities for at least one word in the vocabulary across all time steps in the fifth task result to obtain the second average direction vector; or, selecting a portion of time steps from the fifth task result and averaging the prediction probabilities for at least one word in the vocabulary across the selected time steps to obtain the second average direction vector.

[0108] Among them, the geometric alignment loss value ensures that the compressed latent space information (i.e., the compressed hidden state sequence) remains consistent with the uncompressed latent space information (i.e., the complete hidden state sequence) in the representation space of the second model, which helps to improve training stability and generalization.

[0109] Assume the formula for the step-average direction vector is as follows: ……(11) Specifically, for path A, The value is A. Represents the direction vector of the first step; represents the direction vector for path D. The value is D. This represents the direction vector of the second step; This represents the prediction probability for at least one word in the vocabulary at the k-th time step; the value of K can be any number of time steps.

[0110] Assume the geometric alignment loss value is denoted as Determined according to formula (12) .

[0111] ……(12) In practical applications, the task loss value can be based on the second training sample. Uncertainty loss value And geometric alignment loss value One or more determine the total loss value of the second training sample Based on the total loss value Update the model parameters of the first AI model.

[0112] For example, the total loss value of the training samples is calculated according to formula (12). : ……(13) In the model training process of this embodiment, the model parameters of the second AI model are frozen, and the second AI model is only used to evaluate the quality of the compressed latent space messages generated by the first AI model. The first AI model generates compressed latent space messages in an autoregressive manner (in the process of generating latent space messages, the hidden state output in the previous step is used as the input for the next step), and performs end-to-end optimization using the supervision signal from the second AI model whose model parameters are frozen.

[0113] The technical solution provided in this application embodiment can train a first AI model that can output execution plan information represented by latent space messages based on the task description information of the task to be processed. The vectorization characteristics of latent space messages enable them to encode information with much higher efficiency than language space. The trained first AI model can also generate shorter but more information-rich latent space messages, making the communication efficiency of transmitting latent space messages higher.

[0114] The technical solutions provided in the embodiments of this application are comprehensively analyzed and explained: 1. The embodiments of this application innovatively propose to break away from the "latent space communication paradigm" of natural language. That is, in a multi-communication intelligent module system, the communication intelligent modules or AI models can directly transmit the hidden states corresponding to the generated lexical units (such as the last hidden state) for communication. In other words, the "latent thoughts" of one communication intelligent module or AI model are transmitted to the other party in a continuous representation, rather than relying on natural language text (such as thought chain text). More key information of the task can be carried in fewer places, alleviating the expression bandwidth limitation and redundancy problem in language space communication, and achieving the goal of "communication intelligent modules or AI models communicating with thoughts rather than language".

[0115] 2. In the model training process of the second AI model, this application's embodiment takes "mental separation and alignment" based on contrastive learning and alignment regularization as the training objective. That is, it uses JS separation loss to distinguish "matching / non-matching" latent space communication, forcing the second AI model, as the receiving end, to truly understand the latent space message, rather than relying solely on statistical bias. At the same time, a planned alignment loss is introduced to prevent the model from exploiting loopholes and losing semantic effectiveness by simply increasing divergence with individual tokens.

[0116] 3. The embodiments of this application propose a "differentiable reasoning and compressed generation" mechanism for the first AI model in the latent space. That is, by generating latent space messages in an autoregressive manner in the latent space, the reasoning process does not go through the overhead of decoder to preprocessing and then to encoder. While ensuring the effect, higher compression is achieved, proving that latent space messages can carry information density far higher than that of language text tokens.

[0117] 4. The first AI model with compression capability provided in this application embodiment possesses the characteristic of "parallel thinking branches" in latent space communication, which outperforms a single linear thinking chain in terms of performance. That is, the probability distribution in latent space communication maintains multiple parallel paths, indicating that latent space communication is not only a "compressed version of language" but also a thinking mode of multi-hypothesis parallel reasoning. However, communication in the language space converges into a single path. Therefore, in comparison, latent space communication presents a stable vertical interval between each reasoning step, which is beneficial for providing more sufficient parallel alternatives for subsequent action decisions.

[0118] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, in practice, this electronic device includes a memory 51 and a processor 52.

[0119] Memory 51 is used to store computer programs and can be configured to store various other data to support operation on the computing platform. Examples of this data include instructions for any application or method operating on the computing platform, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0120] The processor 52, coupled to the memory 51, is used to execute the computer program in the memory 51 for: performing the steps in the implicit space communication method or model training method provided in the above embodiments of this application.

[0121] Furthermore, such as Figure 5 As shown, the electronic device also includes other components such as a communication component 53, a display 54, a power supply component 55, and an audio component 56. Figure 5 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 5 The components shown. Additionally... Figure 5 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the work node. In this embodiment, the work node can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device, or a server-side device such as a conventional server, cloud server, or server array. If the work node in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 5The components within the dashed box; if the working node in this embodiment is implemented as a server-side device such as a conventional server, cloud server, or server array, it may be omitted. Figure 5 The component within the dashed box.

[0122] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0123] The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.

[0124] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0125] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.

[0126] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0127] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Accordingly, this application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, enable the processor to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by a computer program or instructions.

[0128] In addition, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, so that the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device can be implemented as a means to implement the corresponding functions in the above method embodiments.

[0129] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0130] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for covert space communication, characterized in that, include: Obtain the text-based task description information of the target task; The first AI model is invoked to generate a target latent space message in the latent space based on the task description information of the target task. The target latent space message represents the execution plan information of the target task. The target latent space message is sent to the second AI model so that the second AI model can execute the target task based on the execution plan information represented by the target latent space message; The training method of the second AI model includes: acquiring a first training sample, the first training sample including: a first latent space message and context information matching the first sample task; based on the first latent space message and context information, calling the second AI model to perform task processing on the first sample task to obtain at least one task result corresponding to the first training sample; updating the model parameters of the second AI model according to at least one task result corresponding to the first training sample until the model training termination condition is met. The training method of the first AI model includes: obtaining a second training sample based on a third latent space message, a fourth latent space message, and sample task description information that match the second sample task, wherein the length of the third latent space message is less than the length of the fourth latent space message, and the second training sample includes: first input information at least related to the third latent space message, second input information at least related to the fourth latent space message, and third input information only related to the sample task description information; calling the second AI model to perform task processing on the second sample task according to the first input information, the second input information, and the third input information respectively, to obtain at least one task result corresponding to the second training sample; updating the model parameters of the first AI model according to at least one task result corresponding to the second training sample, until the model training termination condition is met.

2. The method according to claim 1, characterized in that, The first AI model includes an input layer, multiple hidden layers, and an output layer, wherein the multiple hidden layers form the latent space of the first AI model; The first artificial intelligence (AI) model is invoked to generate a target latent space message in the latent space based on the task description information of the target task, including: The task description information of the target task is input into the input layer for vectorization processing to obtain the initial feature vector. The initial feature vector is subjected to nonlinear operations layer by layer through the multiple hidden layers to obtain the hidden states of at least some of the hidden layers, which are used as the target latent space message.

3. The method according to claim 2, characterized in that, The first AI model adopts an encoder-decoder architecture, an encoder-only architecture, or a decoder-only architecture; wherein, any architecture includes the input layer and the output layer, and the multiple hidden layers are distributed in the encoder and / or decoder in any architecture; The initial feature vector is processed through the multiple hidden layers using nonlinear operations to obtain the hidden states of at least some of the hidden layers, which serve as the target latent space message, including: If the first AI model adopts an encoder-decoder architecture, the initial feature vector is input into the encoder, encoded by multiple hidden layers in the encoder, and then decoded by multiple hidden layers in the decoder to obtain the hidden state of at least some hidden layers, which serves as the target latent space message. If the first AI model adopts an encoder-only architecture, the initial feature vector is input into the encoder and encoded through multiple hidden layers in the encoder to obtain the hidden state of at least some hidden layers, which is used as the target latent space message. If the first AI model adopts a decoder-only architecture, the initial feature vector is input into the decoder and decoded through multiple hidden layers in the decoder to obtain the hidden states of at least some hidden layers, which are used as the target latent space message.

4. The method according to claim 1, characterized in that, The first training sample also includes: a second latent space message that does not match the first sample task, and / or a target language space message corresponding to the first sample task, wherein the target language space message is used to characterize the execution plan information of the first sample task; Accordingly, based on the first latent space message and context information, the second AI model is invoked to perform task processing on the first sample task to obtain at least one task result, including performing at least one of the following task processing: Based on the first latent space message and the context information, the second AI model is invoked to perform task processing on the first sample task to obtain the first task result. The first task result includes a first part of the result and a second part of the result. The first part of the result includes the prediction probability of the expected word element corresponding to at least one time step. The second part of the result includes the prediction probability of at least one first word element in the vocabulary corresponding to at least one time step. Based on the second latent space message and the context information, the second AI model is invoked to perform task processing on the first sample task to obtain the second task result. The second task result includes the prediction probability of at least one second word element in the vocabulary corresponding to each time step. Based on the target language space message corresponding to the first sample task and the context information, the second AI model is invoked to perform task processing on the first sample task to obtain the third task result. The third task result includes the prediction probability of at least one third word element in the vocabulary corresponding to each time step.

5. The method according to claim 4, characterized in that, Based on at least one task result corresponding to the first training sample, update the model parameters of the second AI model, including: The task loss value is determined based on the first part of the results in the first task and the cross-entropy loss function; Based on the second part of the results from the first task, the second task results, and the JS divergence function, determine the separation loss; Based on the second part of the results from the first task, the results from the third task, and the KL divergence function, determine the plan alignment loss; The model parameters of the second AI model are updated based on at least one of the task loss value, the separation loss, and the plan alignment loss.

6. The method according to claim 5, characterized in that, The third part of the first task result includes: the log odds of at least one first word element in the vocabulary for each of at least one time step; the first part of the third task result includes: the prediction probability of at least one third word element in the vocabulary for each of at least one time step; the second part of the third task result includes: the log odds of at least one third word element in the vocabulary for each of at least one time step. Based on the second part of the results from the first task, the third task results, and the KL divergence function, the plan alignment loss is determined, including: Based on the second part of the results from the first task, the first part of the results from the third task, and the KL divergence function, determine the first alignment loss value; The second alignment loss value is determined based on the third part of the results from the first task, the second part of the results from the third task, and the cosine similarity function. The planned alignment loss is determined based on the first alignment loss value and the second alignment loss value.

7. The method according to claim 6, characterized in that, The second AI model, invoked based on the first input information, the second input information, and the third input information, performs task processing on the second sample task to obtain at least one task result corresponding to the second training sample, including performing at least one of the following task processing: The second AI model is invoked based on the first input information to obtain the fourth task result output by the second AI model; The second AI model is invoked based on the second input information to obtain the fifth task result output by the second AI model; The second AI model is invoked based on the third input information to obtain the sixth task result output by the second AI model.

8. The method according to claim 7, characterized in that, Based on at least one task result corresponding to the second training sample, update the model parameters of the first AI model until the model training termination condition is met, including: Based on the predicted probability of the expected word element and the cross-entropy loss function corresponding to at least one time step in the fourth task results, the task loss value of the second training sample is determined. Based on the results of the fourth task, the results of the fifth task and their corresponding information entropy, and the information entropy corresponding to the results of the sixth task, determine the uncertainty loss value of the second training sample; The first step average direction vector is determined based on the results of the fourth task, and the second step average direction vector is determined based on the results of the fifth task; the geometric alignment loss value of the second training sample is determined based on the first step average direction vector, the second step average direction vector, and the cosine similarity function. The model parameters of the first AI model are updated based on at least one of the task loss value, uncertainty loss value, and geometric alignment loss value of the second training sample.

9. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer programs; The processor is coupled to the memory for executing the computer program to perform the steps of the method according to any one of claims 1-8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it causes the processor to perform the steps of the method according to any one of claims 1-8.

11. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, causes the processor to perform the steps of the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Network model reasoning method and device, network model data processing method and device, electronic equipment and medium

    CN119808949A

  • Model training method and device and text generation method and device

    CN120911405A