Task processing method and device, equipment, medium and product

By deploying a compressed Conformer-based speech feature extractor on user terminals, the method addresses data security and privacy concerns in AI systems by enabling secure, local processing of multi-modal data, ensuring efficient operation on resource-constrained devices.

CN120317359APending Publication Date: 2025-07-15IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510244265.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In artificial intelligence technology, how to run multimodal large models on resource-constrained terminal devices, while protecting user privacy and avoiding data breaches and malicious attacks.

Method used

By deploying a task processing model on the user terminal device, using the voice feature extraction module to perform K-fold downsampling of voice data, using the Conformer module to perform K-fold compression, reducing the model scale, and combining image and text data processing to generate reply content.

Benefits of technology

It realizes efficient operation of multimodal large models on resource-constrained devices, protects user privacy, reduces data leakage risks, reduces network dependence, and improves security and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120317359A_ABST
    Figure CN120317359A_ABST
Patent Text Reader

Abstract

The invention provides a task processing method and device, equipment, a medium and a product, and the method comprises the steps: receiving a user request which comprises input data, and the input data comprises at least one of image data, text data and voice data; processing the input data through a task processing model to generate reply content; wherein the task processing model is deployed on terminal equipment of a user, the task processing model comprises a voice feature extraction module which is used for carrying out K-time downsampling on voice data, the voice feature extraction module is obtained by carrying out K-time compression on a Conformer module, and K is a positive integer greater than 8. According to the method and the device, the multi-mode large model can be operated on the resource-limited terminal equipment, and the privacy of the user is effectively protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and particularly to a task processing method, apparatus, device, medium, and product. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, having a personal AI assistant is gradually becoming a reality. Such intelligent agents can assist users in completing various tasks, such as creating presentation slides (PPT), interactive chatting, and providing suggestions.

[0003] In the context of the current digital age, the amount of data has grown explosively, which provides rich resources for the development of AI while also introducing privacy risks. Specifically, in the process of processing user data, there may be a risk that the AI model leaks users' personal information and may even be exploited by malicious attackers. Therefore, how to ensure the security of this information has become an urgent problem to be solved. Summary of the Invention

[0004] Based on the above technical status quo, this application provides a task processing method, apparatus, device, medium, and product, which can enable a multi-modal large model to run on resource-constrained terminal devices and effectively protect user privacy.

[0005] To achieve the above technical objectives, this application specifically proposes the following technical solutions:

[0006] According to the first aspect of the embodiments of this application, a task processing method is provided, including: receiving a user request, where the user request includes input data, and the input data includes at least one of image data, text data, and voice data; processing the input data through a task processing model to generate a reply content; where the task processing model is deployed on the user's terminal device, and the task processing model includes a voice feature extraction module for performing K-fold downsampling on the voice data, and the voice feature extraction module is obtained by performing K-fold compression on the Conformer module, and K is a positive integer greater than 8.

[0007] According to the second aspect of the embodiments of this application, a task processing apparatus is provided, including: a receiving module for receiving a user request, where the user request includes input data, and the input data includes at least one of image data, text data, and voice data; a task processing module for processing the input data through a task processing model to generate a reply content; where the task processing model is deployed on the user's terminal device, and the task processing model includes a voice feature extraction module for performing K-fold downsampling on the voice data, and the voice feature extraction module is obtained by performing K-fold compression on the Conformer module, and K is a positive integer greater than 8.

[0008] According to a third aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor; the memory is connected to the processor and is used for storing a program; the processor is used for implementing the task processing method as described in the first aspect by running the program in the memory.

[0009] According to a fourth aspect of the embodiments of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is run by a processor, the task processing method as described in the first aspect is implemented.

[0010] According to a fifth aspect of the embodiments of the present application, a computer program product is provided, including computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute: the task processing method as described in the first aspect.

[0011] A task processing method, device, equipment, medium and product provided by the embodiments of the present application receive a user request including at least one modality of input data (such as image data, text data and voice data), and process the input data through a task processing model deployed on a user terminal device, and finally generate corresponding reply content; the task processing model includes a voice feature extraction module for performing K-fold downsampling on voice data, where K is a positive integer greater than 8, and the voice feature extraction module is obtained by performing K-fold compression on a Conformer module. This compression method can reduce the scale of the voice feature extraction module and enable it to operate efficiently on resource-constrained terminal devices. Since all processing is completed on the user's terminal device without uploading the user's data to the cloud, it can not only reduce the dependence on the external network, but also fundamentally protect the user's privacy and security. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0013] Figure 1 It is a flowchart of a task processing method provided by the embodiments of the present application;

[0014] Figure 2 It is a schematic structural diagram of a task processing model provided by the embodiments of the present application;

[0015] Figure 3 It is a schematic structural diagram of a voice feature extraction module provided by the embodiments of the present application;

[0016] Figure 4 It is a schematic flow diagram of K-fold downsampling provided by an embodiment of the present application;

[0017] Figure 5 It is a schematic structural diagram of a Conformer module provided by an embodiment of the present application;

[0018] Figure 6 It is a schematic flow diagram of the supervised fine-tuning process of a task processing model provided by an embodiment of the present application;

[0019] Figure 7 It is a schematic structural diagram of a task processing device provided by an embodiment of the present application;

[0020] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0021] The technical solutions provided by the embodiments of the present application can be exemplarily applied to hardware devices such as processors, electronic devices, and servers (including cloud servers), or packaged into software programs to be run. When the hardware device executes the processing process of the technical solutions of the embodiments of the present application, or when the above software program is run, it can realize the automatic splitting of the target task and automatically call the application program interfaces required for the task, so as to complete the purpose of the target task. The embodiments of the present application only exemplarily introduce the specific processing process of the technical solutions of the present application, and do not limit the specific implementation form of the technical solutions of the present application. Any technical implementation form that can execute the processing process of the technical solutions of the present application can be adopted by the embodiments of the present application.

[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.

[0023] Before introducing the solutions of the present application, the related technologies will be introduced first:

[0024] The continuous growth of the data volume provides rich resources for the development of AI, but at the same time brings challenges. On the one hand, it is necessary to efficiently process and analyze a large amount of data to extract valuable information and provide more accurate decision-making basis for AI agents; on the other hand, the security and privacy protection of data have become crucial issues.

[0025] In the context of the widespread application of AI, numerous models have collected and processed a large amount of data containing users' personal information to provide customized services. Therefore, in the AI era, ensuring personal privacy and data security has become a key issue that urgently needs to be solved. To enhance the understanding and decision-making abilities of AI agents, it is necessary to continuously explore new technologies and methods. For example, through deep learning algorithms, AI agents can learn knowledge and patterns from large-scale data, improving their understanding and decision-making abilities; at the same time, the progress of natural language processing technology enables AI agents to better understand users' language and intentions, realizing more natural human-computer interaction. However, the development of these technologies is also accompanied by new privacy risks, such as the potential for data leakage or being maliciously exploited.

[0026] In view of this, when developing AI agents, great attention must be paid to data security and privacy protection. In recent years, large models have developed rapidly due to the improvement of computing power and model capacity and have been widely applied in multiple fields. In addition to common intelligent customer service, intelligent recommendation, and sentiment analysis, large models have gradually penetrated into vertical fields such as education, healthcare, and finance. Through the path of "big data + big computing power + strong algorithms", the generality and generalization ability of the model have been enhanced, promoting the transformation of AI from small-scale customized models to general large model pre-training.

[0027] In addition, the emergence of methods such as optimized algorithms, parallel computing technology, and model compression has further improved the training speed, efficiency, and real-time performance of model deployment.

[0028] Cross-modal large models, as a research feature, can process various modal data such as text, images, and speech, realizing information fusion and interaction.

[0029] In the process of building a large model, it is necessary to collect a large amount of text data from various sources such as books, news, papers, and web pages, and through cleaning and preprocessing, remove noise and error information to provide high-quality input for model training. High-performance graphics processing unit (GPU) clusters or dedicated AI chips, as the key support, accelerate the training and inference processes of the model. The distributed computing framework distributes large-scale computing tasks to multiple nodes, improving the computing efficiency.

[0030] The Transformer architecture is one of the commonly used model architectures. Its self-attention mechanism can effectively capture long-distance dependencies, enhancing the ability of language understanding and generation. The huge number of parameters enables large models to learn richer language patterns and knowledge. During the training process, the methods of unsupervised learning and supervised learning are usually combined: the former is used for pre-training of large-scale text data, and the latter is used for fine-tuning specific tasks, such as text classification, machine translation, and question-answering systems, so as to improve the performance of the model on specific tasks.

[0031] A complete AI agent consists of three parts: the control end, the perception end, and the action end, which are responsible for storing memory and knowledge, processing information and making decisions, and performing operations to adapt to or change the environment. Cloud-based large models provide strong support for personal assistants, but there are problems such as the risk of user privacy leakage and high deployment costs. In contrast, the end-side multimodal large model can run efficiently on resource-constrained devices while providing better privacy protection thanks to its innovative model compression technology. This not only reduces deployment costs, but also shows unique advantages in application scenarios that focus on privacy and security. Therefore, it is particularly necessary to develop a multimodal large model solution that is suitable for the end side and takes into account privacy protection.

[0032] In view of this, the embodiments of the present application are dedicated to providing a task processing method, device, equipment, medium and product, which compresses a large model and deploys it on the user's terminal device, so that task processing can be completed locally without uploading the user's personal data to the cloud, avoiding data leakage and protecting user privacy. The following embodiments are described in detail one by one.

[0033] Exemplary method

[0034] Figure 1 A flowchart of a task processing method provided in an embodiment of the present application. Figure 1 As shown, the task processing method provided in this embodiment includes step S101 and step S102:

[0035] S101. Receive a user request, where the user request includes input data, and the input data includes at least one of image data, text data, and voice data.

[0036] The execution subject of the method of this embodiment may be a user's terminal device. In step S101, a request from a user is first received. The user request not only includes the user's specific needs or questions expressed in text form or voice form, but also includes at least one type of data, which may be any combination of image data, text data, and voice data.

[0037] In some scenarios, when user requests involve visual information, such as identifying objects in photos, analyzing medical images, or evaluating works of art, users may upload one or more images and ask specific questions or requests through text data. In this case, image data and text data are used as input data together. For example, a user may upload a photo of a work of art and enter a question in text form asking about the creation time or style characteristics of the work.

[0038] In some scenarios, for requests that need to be processed based on text information, such as translating a foreign language text, answering specific questions, or writing a report, the user will provide the corresponding text data. At this time, the text data is the input data. For example, the user may submit an article that needs to be translated into English, or request an analysis report to be written based on a given topic.

[0039] In some specific application scenarios, such as using voice assistant services, telephone customer service automation, or controlling smart devices through voice commands, the user may submit requests in voice form. At this time, the voice data is the input data.

[0040] It should be noted that the user can flexibly select the form of the input data according to their own needs. The input data can be any one of image data, text data, and voice data, or a combination of any two, or even a combination of all three. This flexibility enables the user to generate multimodal task processing requests according to specific task requirements.

[0041] S102. Process the input data through a task processing model to generate a response content; wherein, the task processing model is deployed on the user's terminal device, and the task processing model includes a voice feature extraction module for performing K-fold downsampling on the voice data. The voice feature extraction module is obtained by performing K-fold compression on the Conformer module, and K is a positive integer greater than 8.

[0042] In step S102, the input data can be processed by using the task processing model deployed on the user's terminal device to generate the corresponding response content. This task processing model can be used to process multimodal data, including voice, image, and text data. The following will introduce the specific implementation process of step S102 in detail with reference to the accompanying drawings:

[0043] Figure 2 This is a schematic structural diagram of the task processing model provided by the embodiment of the present application. As Figure 2 shown, the task processing model includes an image feature extraction module 21, a voice feature extraction module 22, a multilayer perceptron (MLP) 23, a pre-trained large language model (LLM) 24, an image generation module 25, and a voice generation module 26.

[0044] Among them, the image feature extraction module 21 can be an architecture of Vision Transformer (ViT), which is used to divide the image data into multiple image patches and input the multiple image patches into the Transformer model for image feature extraction. ViT can capture the global information in the image, so as to better understand the image content.

[0045] A voice feature extraction module 22 for extracting voice features from the voice data in the input data.

[0046] A multi-layer perceptron 23 for further processing the features obtained from the voice feature extraction module and the image feature extraction module, and extracting deeper voice features and image features through a deep network.

[0047] A pre-trained large language model 24, which can adopt a 10B-parameter large model of a causal decoder, for combining the text data in the input data and the deep voice features and deep image features obtained through the MLP, deeply understanding the overall input data, and generating corresponding reply content.

[0048] An image generation module 25, which can adopt a painting generation tool, such as a Stable Diffusion structure, for generating the required image according to the input data when the user requests to generate a new image, such as adjusting the style of an existing image. For example, when the user requests to adjust the style of an image, the image generation module 25 will generate an image with the adjusted style according to the user's requirements.

[0049] A voice generation module 26, which can adopt an audio generation model, such as an AudioLDM (Audio Language Diffusion Model) structure, for generating a corresponding voice output according to the input data when the user requests to generate a new segment of voice data, such as translating the voice of one language into another language. For example, the user can request to translate a segment of voice in the original language into the target language and generate voice data in the target language. At this time, the voice generation module will generate the corresponding target language voice based on the translated text.

[0050] Based on the above architecture design of the task processing model, step S102 includes: extracting image features from the image data through the image feature extraction module 21 to obtain image features; extracting voice features from the voice data through the voice feature extraction module 22 to obtain voice features; extracting deep voice features according to the voice features and deep image features according to the image features through the MLP; generating reply content through the pre-trained LLM according to at least one of the text data, deep image features, and deep voice features.

[0051] Among them, the reply content includes text content, new image data, and new voice data. When the reply content includes new image data, it is necessary to generate new image data that meets the user's needs through the image generation module 25 according to the image data and the user's instructions. When the reply content includes new voice data, it is necessary to generate new voice data that meets the user's needs through the voice generation module 26 according to the voice data and the user's instructions.

[0052] It should be noted that the voice feature extraction module 22 is implemented using the Conformer structure. Given that the scale of the original Conformer structure is relatively large, in order to adapt to resource-limited devices such as personal terminals, it needs to be optimized and compressed for actual deployment applications. The following will introduce the streamlining process of the voice feature extraction module in detail with reference to the accompanying drawings:

[0053] Figure 3 This is a schematic structural diagram of the voice feature extraction module provided by the embodiment of the present application. As Figure 3 shown, the voice feature extraction module 22 includes: a feature enhancement (SpecAug) layer 221, a feature extraction (Convolution Subsampling) layer 222, a linear transformation (Linear) layer 223, a regularization (Dropout) layer 224, and a Conformer module 225.

[0054] Among them, the SpecAug layer 221 is used to perform enhancement processing on the initial voice features corresponding to the voice data to obtain enhanced voice features.

[0055] The Convolution Subsampling layer 222, also known as the pre-downsampling layer, is used to perform downsampling on the enhanced voice features using a convolutional layer to reduce the time dimension of the features.

[0056] The Linear layer 223 is used to map the downsampled voice features to a higher-dimensional space so that subsequent modules can better capture features.

[0057] The Dropout layer 224 is used to adopt a regularization technique to randomly discard a part of the neurons of the voice features mapped to a high dimension to prevent overfitting and obtain regularized voice features.

[0058] The Conformer module 225 is used to perform deep feature extraction based on the regularized voice features to obtain the final voice features.

[0059] Among them, the Conformer module 225 includes a plurality of Conformer Blocks arranged in a stack. Each Conformer Block includes a multi-head self-attention mechanism (Multi-Head Self Attention), a convolution module (ConvolutionModule), a feed-forward module (Feed Forward Module), and a normalization layer (Layernorm), and can capture long-distance dependencies and improve the ability to capture local features.

[0060] In some embodiments, a post-downsampling layer is embedded in multiple Conformer Blocks, and its function is to perform K-fold downsampling processing on the received speech features. Wherein, K is a positive integer greater than 8, such as 16, 32, or 64, etc. Next, the case where K is 16 will be introduced in detail:

[0061] As Figure 4 shown, when step S102 performs K-fold downsampling on the received speech features through the post-downsampling layer and generates a response content based on the downsampled speech features, it specifically includes the following steps S401 to S403:

[0062] S401. Convolutionally downsample the initial speech features corresponding to the speech data through a pre-downsampling layer using a first convolution stride to obtain speech features with a first sequence length. The first convolution stride is 4 times the unit convolution stride.

[0063] Wherein, the unit convolution stride means the stride is 1.

[0064] Continue to refer to Figure 3 , assuming that for a piece of speech data with a duration of 2 seconds, the time resolution of the initial speech features is 10ms rate (that is, one sample is taken every 10ms), which means that there are 200 data points in this 2-second speech data. Next, when these initial speech features are downsampled through the pre-downsampling layer using the first convolution stride (such as taking one sample every 40ms), the time resolution is adjusted from 10ms to 40ms. At this time, a new data point is generated every 40ms. Therefore, for the above 2-second speech data, its sequence length can be reduced to 50 data points.

[0065] S402. Convolutionally downsample the speech features with the first sequence length through the post-downsampling layer using a second convolution stride at least twice to obtain speech features with a second sequence length. The second convolution stride is 2 times the unit convolution stride.

[0066] Continue to refer to Figure 3 , in order to perform data processing and feature extraction more efficiently, the speech features can also pass through the post-downsampling layer to further reduce the time resolution to 160ms. At this time, for the originally 2-second long speech segment, its sequence length can be further reduced to about 12 data points.

[0067] When implementing step S402, in order to more finely control the downsampling process, the post-downsampling layer can be subdivided into a first downsampling layer and a second downsampling layer, both of which use 2 times the unit convolution stride, that is, a convolution stride of 2 is designed. This kind of hierarchical design can gradually reduce the dimension of the data while maintaining the key information of the speech features, thereby improving the processing efficiency.

[0068] Specifically, in a Conformer module containing multiple Conformer Blocks, the first downsampling layer and the second downsampling layer can be respectively set at the front half of the module. For example, assuming that a Conformer module contains 16 Conformer Blocks, the first downsampling layer can be set between the 4th Conformer Block and the 5th Conformer Block, and the second downsampling layer can be set between the 6th Conformer Block and the 7th Conformer Block. This layout ensures that each downsampling layer can effectively reduce the dimension when the feature representation reaches a certain depth, which not only ensures the richness of the feature expression, but also realizes the effective use of computing resources.

[0069] Of course, in other examples, the first downsampling layer may also be placed between the 4th and 5th Conformer Blocks, but the second downsampling layer may be arranged between the 5th and 6th Conformer Blocks. Alternatively, the first downsampling layer may also be placed between the 5th and 6th Conformer Blocks, but the second downsampling layer may be arranged between the 6th and 7th Conformer Blocks.

[0070] In some embodiments, N can be set to 1, and the second sampling rate is 2*1 times the original sampling rate (ie, 2 times). Based on this, step S402 can include the following steps a1 and a2:

[0071] Step a1: Perform convolution downsampling on the speech features of the first sequence length using the second convolution step size through the first downsampling layer to obtain speech features of 1 / 8 times the original sequence length.

[0072] In step a1, when the speech feature with the first sequence length is processed by the first downsampling layer, a downsampling operation can be performed on it with a unit convolution step size of 2, thereby reducing the amount of input speech feature data by half while retaining key information and feature expressions as much as possible. Through this downsampling method, the sequence length of the speech feature can be reduced from the original length to 1 / 8 times.

[0073] In order to implement the downsampling process in step a1, the first downsampling layer may adopt a convolution module with a downsampling function, such as a convolution layer with a stride-2 convolution of 2 times the unit convolution step.

[0074] Specifically, when the speech features of the original sequence length enter the first downsampling layer, this layer performs downsampling by applying a convolutional operation with a stride set to 2. This means that for each output unit, the convolutional kernel only moves two positions instead of the usual single position. In this way, the spatial resolution of the output feature map can be halved.

[0075] Step a2: Use the second downsampling layer to perform convolutional downsampling on the speech features of 1 / 8 times the original sequence length with a second convolutional stride to obtain speech features of 1 / 16 times the original sequence length.

[0076] In step a2, the second downsampling layer is used to continue processing the speech features after the first downsampling (whose sequence length has been reduced to 1 / 8 of the original length). Here, a convolutional stride of 2 times can also be used for further downsampling operations to halve the sequence length of the speech features again based on the already compressed 1 / 8 times, and finally obtain speech features with a sequence length of 1 / 16 times the original length.

[0077] To implement the downsampling process in step a2, the second downsampling layer can also use a convolutional module with a downsampling function, such as a convolutional layer with a stride of 2 (stride-2 convolution).

[0078] Specifically, when the speech features of 1 / 8 times the original sequence length enter the first downsampling layer, this layer performs downsampling by applying a convolutional operation with a stride set to 2. Similarly, for each output unit, the convolutional kernel only moves two positions, thereby further halving the spatial resolution of the output feature map.

[0079] Continue to refer to Figure 4 , after step S402, the specific implementation manner of step S102 in the method of this embodiment may further include the following step S403:

[0080] S403: Generate a response content based on the speech features of the second sequence length.

[0081] In step S403, the response content is generated based on the speech features of the second sequence length (i.e., 1 / 16 times the original sequence length) obtained after two downsamplings.

[0082] Specifically, the compressed speech features, the image features extracted in the previous steps, and the obtained text data are used as input data and processed through a pre-trained deep learning model (such as a large language model LLM) to generate the corresponding response content.

[0083] In the above embodiments, it is introduced that the speech features of the first sequence length are downsampled twice by the first downsampling layer and the second downsampling layer using the second convolution stride, so as to obtain the speech features of the second sequence length. However, in some embodiments, the post-downsampling layer may include a third downsampling layer. Then, the input data is processed by the task processing model to generate a response content, including: performing convolutional downsampling on the initial speech features corresponding to the speech data by the pre-downsampling layer using the first convolution stride to obtain the speech features of the first sequence length, where the first convolution stride is 4 times the unit convolution stride; performing convolutional downsampling on the speech features of the first sequence length by the third downsampling layer using the third convolution stride to obtain the speech features of the second sequence length, where the third convolution stride is 4 times the unit convolution stride; generating a response content according to the speech features of the second sequence length.

[0084] In this embodiment, first, the pre-downsampling layer is used to perform preliminary downsampling on the initial speech features corresponding to the original speech data, and the first convolution stride is used to reduce the data volume to 1 / 4 to obtain the speech features of the first sequence length.

[0085] Next, the speech features of the first sequence length are downsampled by the third downsampling layer using the third convolution stride. Specifically, the third downsampling layer may adopt a convolutional module with a downsampling function, such as a convolutional layer with a stride of 4 (stride-4 convolution). When the input is the speech features of the first sequence length (for example, 1 / 4 times the original sequence length after one or more downsamplings), this layer performs downsampling by applying a convolutional operation with a stride set to 4. This means that for each output unit, the convolutional kernel moves four positions, resulting in a 4-fold reduction in the spatial resolution of the output feature map. Therefore, combined with the previous downsampling steps, an overall downsampling effect of 16 times is finally achieved, and the speech features of the second sequence length are obtained.

[0086] Finally, based on the speech features of the second sequence length obtained after the above processing, a corresponding response content can be generated.

[0087] To further improve the speech recognition effect of the model, progressive modeling can be performed using modeling units of multiple granularities, so as to more accurately extract speech features. Specifically, it may include: performing phoneme modeling on the initial speech features of the first sequence length corresponding to the input data through multiple Conformer Blocks to obtain phoneme features; performing initial consonant and final modeling based on the phoneme features to obtain initial consonant and final features; performing word segmentation modeling based on the initial consonant and final features to obtain word segmentation features; generating a response content according to the word segmentation features.

[0088] In this embodiment, first, multiple Conformer Blocks are used to perform phoneme-level modeling on the initial speech features of the first sequence length corresponding to the input data to obtain phoneme features, thereby obtaining the most basic sound units in the speech and providing a basis for subsequent complex modeling.

[0089] Next, based on the obtained phoneme features, initial consonant and vowel modeling can be further performed to obtain initial consonant and vowel features. By modeling the unique initial consonant and vowel structures in languages such as Chinese, it helps to improve the ability to understand specific language characteristics.

[0090] Furthermore, word segmentation modeling can be performed according to the initial consonant and vowel features to generate word segmentation features. Specifically, techniques such as Byte-Pair Encoding (BPE) can be used to split a continuous character sequence into meaningful lexical units, thereby helping to accurately understand natural language.

[0091] Finally, based on the word segmentation features obtained in the above steps, the corresponding response content can be generated.

[0092] Next, the progressive modeling method will be introduced in detail in combination with the structure of the Conformer module:

[0093] Figure 5 FIG. is a schematic structural diagram of the Conformer module provided in the embodiment of the present application. As Figure 5 shown, the multiple Conformer Blocks stacked in the Conformer module can be divided into a first group of Conformer Blocks, a second group of Conformer Blocks, and a third group of Conformer Blocks.

[0094] Taking the Conformer module including 16 Conformer Blocks as an example, the first group of Conformer Blocks includes the 1st to 4th Conformer Blocks, which are used to perform phoneme modeling on the initial speech features of the first sequence length corresponding to the input data to obtain phoneme features.

[0095] The second group of Conformer Blocks includes the 5th to 12th Conformer Blocks, which are used to perform initial consonant and vowel modeling according to the phoneme features to obtain initial consonant and vowel features.

[0096] The third group of Conformer Blocks includes the 13th to 16th Conformer Blocks, which are used to perform word segmentation modeling according to the initial consonant and vowel features to obtain word segmentation features.

[0097] Among them, the first group of Conformer Blocks, the second group of Conformer Blocks, and the third group of Conformer Blocks can be regarded as the shallow layer, middle layer, and high layer of the Conformer module respectively. The shallow layer focuses on the recognition of the most basic voice elements. The middle layer deeply analyzes the specific pronunciation characteristics of the language. The high layer focuses on converting speech into understandable language units and optimizing the quality of the final output speech features.

[0098] In this embodiment, through this hierarchical feature extraction method, that is, the step-by-step refinement process from phonemes to initials and finals to word segmentation, the Conformer module can significantly improve the effect of speech understanding in low-resource scenarios.

[0099] The inference application process of the task processing model was introduced in the above embodiment. However, before the model is applied in inference, it needs to be trained with multi-modal sample data to learn how to accurately generate response content according to multi-modal data and achieve multi-modal task processing.

[0100] In this embodiment, the image feature extraction module, the speech feature extraction module, the pre-trained LLM, the image generation module, and the speech generation module are all pre-trained. Among them, the image feature extraction module, the image generation module, and the speech generation module can directly adopt open-source pre-trained models. The speech feature extraction module can be pre-trained using the wav2vec2.0 strategy and contrastive learning. Contrastive learning can enhance the similarity of speech segment representations. The pre-training process of the pre-trained LLM will be introduced in the following embodiments.

[0101] Based on pre-training, supervised fine-tuning can further optimize the model in specific user requirements and application scenarios. Supervised fine-tuning is directionally adjusted through user data (such as the combination of image samples, speech samples, and text samples), so that the model can better understand user instructions and make reasonable feedback. The supervised fine-tuning process of the model will be introduced below:

[0102] Figure 6 It is a schematic flowchart of the supervised fine-tuning process of the task processing model provided by the embodiment of the present application. As Figure 6 shown, the supervised fine-tuning process of the task processing model includes the following steps S601 to step S604:

[0103] S601. Extract features from the image sample through the image feature extraction module to obtain an image feature sample.

[0104] In this embodiment, a training sample in the supervised fine-tuning training process can include any one of an image sample, a speech sample, and a text sample. For example, a training sample includes: when the user asks "What day of the week is today?", the marked response content is "Today is Monday".

[0105] It should be noted that in a multi-modal scenario, users may input in any combination of pictures, voices, and texts. For example, they may send a photo of an animal and ask the agent about the type of the animal. The task processing model needs to comprehensively generate accurate responses based on picture, voice, and text information.

[0106] In some embodiments, to enhance the adaptability of the model, in addition to the manually constructed high-quality supervised data, these data can be used as examples, and more similar data can be generated as supplements through the open-source GPT model.

[0107] In step S601, for the image samples in the training samples, the image samples are input into the image feature extraction module. The image samples are segmented into multiple image patches by the image feature extraction module, and the multiple image patches are input into the Transformer model for image feature extraction to obtain image feature samples.

[0108] S602. Feature extraction is performed on the voice samples through the voice feature extraction module to obtain voice feature samples.

[0109] In step S602, for the voice samples in the training samples, the voice samples are input into the voice feature extraction module, and the voice feature extraction module extracts voice feature samples from the voice samples.

[0110] Specifically, step S602 includes: downsampling the initial voice feature samples corresponding to the voice samples through the pre-downsampling layer at the first sampling rate to obtain voice feature samples with the first sequence length, and downsampling the voice features with the first sequence length through the post-downsampling layer at the second sampling rate at least twice to obtain voice features with the second sequence length.

[0111] Among them, downsampling the voice features with the first sequence length through the post-downsampling layer at the second sampling rate at least twice to obtain voice features with the second sequence length can be implemented by at least one of the following implementation methods:

[0112] In some implementation methods, it can be implemented by performing two downsamplings through the first downsampling layer and the second downsampling layer respectively at 2*1 times the original sampling rate. Specifically, it includes: downsampling the voice feature samples with the first sequence length through the first downsampling layer at 2*1 times the original sampling rate to obtain voice feature samples with 1 / 8 times the original sequence length; downsampling the voice feature samples with 1 / 8 times the original sequence length through the second downsampling layer at 2*1 times the original sampling rate to obtain voice feature samples with 1 / 16 times the original sequence length.

[0113] In some implementations, it can also be achieved by performing downsampling once through a third downsampling layer at 4 times the original sampling rate. Specifically, it includes: downsampling the speech features of the first sequence length through the third downsampling layer at the third sampling rate to obtain the speech features of the second sequence length, where the third sampling rate is 4 times the original sampling rate.

[0114] S603. Generate a predicted response content by means of a pre-trained LLM model based on at least one of a text sample, an image feature sample, and a speech feature sample.

[0115] In step S603, the pre-trained LLM needs to be pre-trained. The pre-training process includes: obtaining a large number of diverse text data as the pre-training dataset, using the language modeling loss and the denoising auto-encoding loss as loss functions, and combining the AdamW optimizer to update the model parameters.

[0116] Among them, the language modeling loss is used to evaluate the model's ability to predict the next word, prompting the model to learn to understand and generate natural and fluent language. The denoising auto-encoding loss enhances the model's semantic understanding ability and robustness by adding noise to the input text and attempting to recover the original text.

[0117] To further improve the training effect, a periodic learning rate adjustment strategy can also be used to dynamically adjust the learning rate, enabling the model to flexibly adapt to data changes at different stages of training, thereby improving the final performance.

[0118] Through the above pre-training process, basic reasoning and information understanding capabilities can be endowed to the LLM, providing a basis for subsequent fine-tuning training and reinforcement learning.

[0119] After obtaining the pre-trained LLM, the model can be further trained by inputting at least one of a text sample, an image feature sample, and a speech feature sample, enabling it to understand the interaction between text and image or between text and speech, and predict the corresponding response content.

[0120] S604. Determine the comprehensive loss based on the difference between the predicted response content and the annotated response content, and converge the task processing model to be trained based on the comprehensive loss to obtain the trained task processing model.

[0121] Among them, the labeled response content can be manually labeled or real response content generated by other models, which is used to guide the fine-tuning training process of the model. For example, when the user asks "What day of the week is today?", the model needs to recognize its intention, query the calendar, and predict the response content, such as "Today is Monday". In a multi-modal scenario, the user may input through pictures, voices, and texts simultaneously. For example, when sending a photo of an animal and asking the agent about the type of the animal, it is necessary to comprehensively consider the information of pictures, voices, and texts to accurately predict the response content.

[0122] Specifically, there are various implementation methods for determining the comprehensive loss according to the difference between the predicted response content and the labeled response content, which can be specifically as follows:

[0123] In some implementation methods, determining the comprehensive loss according to the difference between the predicted response content and the labeled response content in step S604 includes: determining the response content prediction loss according to the difference between the predicted response content and the labeled response content; determining the image generation prediction loss according to the difference between the predicted image generated by the image generation module and the real image; and determining the voice generation prediction loss according to the difference between the predicted voice generated by the voice generation module and the real voice; determining the comprehensive loss according to the sum of the response content prediction loss, the image generation prediction loss, and the voice generation prediction loss.

[0124] In some implementation methods, determining the comprehensive loss according to the difference between the predicted response content and the labeled response content in step S604 includes the following steps a1 to a5:

[0125] Step a1: Determine the response content prediction loss according to the difference between the predicted response content and the labeled response content.

[0126] In step a1, the response content prediction loss is determined by comparing the predicted response content with the real response content that is manually or automatically labeled and calculating the difference between the two. This process can use the cross-entropy loss function or other loss functions suitable for generative tasks for quantification.

[0127] The response content prediction loss aims to evaluate the model's ability to understand and respond to input data, ensuring that the responses it generates are both accurate and relevant. For example, in a dialogue system, if the user asks about the weather, the model should be able to accurately extract relevant information and generate an appropriate answer. By performing this evaluation on a large number of training samples, the model parameters can be effectively adjusted to enable it to gradually learn how to understand the context more accurately and generate high-quality responses.

[0128] Step a2: Determine the phoneme prediction loss according to the difference between the predicted phoneme corresponding to the phoneme feature sample and the real phoneme.

[0129] Phoneme is the most basic sound unit in speech. In step a2, the first group of Conformer Blocks in the plurality of Conformer Blocks may be used to perform phoneme modeling on the initial speech feature samples of the speech sample to obtain phoneme feature samples.

[0130] Afterwards, the predicted phonemes are compared with the real phonemes and the phoneme prediction loss is calculated. This not only helps improve the model's ability to recognize the basic units of speech signals, but also provides a basis for subsequent higher-level language processing. For example, for a speech clip containing "hello", the model should be able to accurately decompose it into the corresponding phonemes [h], [e], [l], [o]. The phoneme prediction loss helps the model learn how to capture these basic sound units more accurately, thereby improving the accuracy of overall speech recognition.

[0131] Step a3: Determine the prediction loss of the finals according to the difference between the predicted finals and the real finals corresponding to the final feature samples.

[0132] The initial speech feature samples of the speech sample are modeled by the second group of Conformer Blocks in the plurality of ConformerBlocks to obtain the feature samples of the initial speech feature samples of the speech sample.

[0133] Next, the predicted initials and finals are compared with the actual initials and finals, and the initial and final prediction loss is calculated. For example, in Chinese, "hello" is decomposed into initials [n], [n] and finals [iɑu], [ou]. In this way, the model can better understand and process the pronunciation rules in different languages, thereby improving the accuracy and robustness of speech recognition.

[0134] Step a4: Determine the word segmentation prediction loss based on the difference between the predicted word segmentation and the true word segmentation corresponding to the word segmentation feature sample.

[0135] Word segmentation refers to the process of segmenting a continuous character sequence into meaningful vocabulary units. In step a4, the initial speech feature sample of the speech sample is subjected to word segmentation modeling by a third group of Conformer Blocks among the plurality of Conformer Blocks, thereby obtaining a word segmentation feature sample.

[0136] After that, the predicted word segmentation results are compared with the actual word segmentation results to calculate the word segmentation prediction loss. For example, for the sentence "I like cats", the correct word segmentation should be "I / like / cats". By calculating the word segmentation prediction loss, the model can learn how to perform word segmentation more accurately.

[0137] Step a5: Determine the comprehensive loss based on the response content prediction loss, phoneme prediction loss, initial consonant and final vowel prediction loss, and word segmentation prediction loss.

[0138] To comprehensively evaluate the overall performance of the model, it is necessary to integrate all the individual loss functions obtained in steps a1 to a4 to form a comprehensive loss. Specifically, the response content prediction loss, phoneme prediction loss, initial consonant and final vowel prediction loss, and word segmentation prediction loss can be accumulated to obtain the comprehensive loss. By minimizing the comprehensive loss, the model parameters can be effectively adjusted to gradually approach the optimal solution, and finally a task processing model that is fully trained and has excellent performance can be obtained.

[0139] In some implementation manners, determining the comprehensive loss according to the difference between the predicted response content and the annotated response content in step S604 includes the following steps b1 to b3:

[0140] Step b1: Determine the response content prediction loss according to the difference between the predicted response content and the annotated response content.

[0141] For the specific implementation process of step b1, reference can be made to the specific implementation process of the aforementioned step a1, which will not be elaborated here.

[0142] Step b2: Determine the first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss.

[0143] Since the speech feature extraction module uses 16-fold downsampling compression, which will cause the loss of acoustic information and semantic information, therefore, interaction losses can be applied before and after each of the first downsampling layer and the second downsampling layer respectively to supplement the acoustic information and semantic information lost during downsampling. Among them, the interaction losses include: the first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss, and the interaction losses are used to supplement the acoustic information and semantic information lost during downsampling. Specifically, determining the first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss includes the following steps b21 to b24:

[0144] Step b21: Determine the first predicted downsampled character corresponding to the phoneme feature sample, and determine the first predicted downsampled character loss according to the difference between the first predicted downsampled character and the real character.

[0145] Continue to refer to Figure 5, in step b21, first, it is necessary to determine the first predicted downsampled character corresponding to the phoneme feature sample. Specifically, this can be accomplished through a linear layer (not shown in the figure), which performs character prediction based on the phoneme feature sample obtained from the first group of ConformerBlocks. Subsequently, the first predicted downsampled character loss is calculated according to the difference between the predicted first predicted downsampled character and the true character.

[0146] Specifically, in the supervised fine-tuning training phase, after extracting the phoneme feature sample through the first group of Conformer Blocks, the above-mentioned linear layer can be used for character-level prediction. Then, the predicted character is compared with the true character, and the loss value is determined based on this difference. This loss calculation can adopt the CTC (Connectionist Temporal Classification) loss function to effectively measure the gap between the prediction result and the true label.

[0147] Step b22: Determine the second predicted downsampled character corresponding to the phoneme feature sample with 1 / 8 of the original sequence length, and determine the second predicted downsampled character loss according to the difference between the second predicted downsampled character and the true character; the phoneme feature sample with 1 / 8 of the original sequence length is obtained by downsampling the phoneme feature sample through the first downsampling layer using the second convolution stride.

[0148] As mentioned above, by downsampling the speech sample with the original sequence length through the pre-downsampling layer, an initial speech feature sample with 1 / 4 of the original sequence length can be obtained.

[0149] Next, continue to refer to Figure 5 , through the first group of Conformer Blocks, phoneme features with the same sequence length can be extracted according to the initial speech feature sample with 1 / 4 of the original sequence length, and a phoneme feature sample with 1 / 4 of the original sequence length can be obtained; then, through the first downsampling layer, the phoneme feature sample with 1 / 4 of the original sequence length is downsampled to obtain a phoneme feature sample with 1 / 8 of the original sequence length.

[0150] In step b22, a linear layer can be used to predict the corresponding character (i.e., the second predicted downsampled character) according to these phoneme feature samples with 1 / 8 of the original sequence length. Then, the second predicted downsampled character loss is calculated according to the difference between the predicted second predicted downsampled character and the true character.

[0151] This stage uses the CTC (Connectionist Temporal Classification) loss function for calculation to accurately evaluate the gap between the prediction result and the true label.

[0152] Step b23: Determine the third predicted downsampled character corresponding to the phoneme and vowel feature sample with 1 / 8 of the original sequence length, and determine the loss of the third predicted downsampled character based on the difference between the third predicted downsampled character and the true character; the phoneme and vowel feature sample with 1 / 8 of the original sequence length is obtained by performing phoneme and vowel modeling on the phoneme feature sample with 1 / 8 of the original sequence length through at least one spaced Conformer Block.

[0153] Continue to refer to Figure 5 , after obtaining the phoneme feature sample with 1 / 8 of the original sequence length by downsampling the phoneme feature sample with 1 / 4 of the original sequence length through the first downsampling layer, at least one spaced Conformer Block located between the first downsampling layer and the second downsampling layer can be used to continue performing phoneme and vowel modeling on the phoneme feature sample with 1 / 8 of the original sequence length to obtain the phoneme and vowel feature sample with 1 / 8 of the original sequence length.

[0154] Next, a linear layer can be used to predict the corresponding characters (i.e., the third predicted downsampled characters) based on these phoneme and vowel feature samples with 1 / 8 of the original sequence length. Then, calculate the loss of the third predicted downsampled character based on the difference between the predicted third predicted downsampled character and the true character.

[0155] This stage uses the CTC (Connectionist Temporal Classification) loss function for calculation to accurately evaluate the gap between the prediction result and the true label.

[0156] Step b24: Determine the fourth predicted downsampled character corresponding to the phoneme and vowel feature sample with 1 / 16 of the original sequence length, and determine the loss of the fourth predicted downsampled character based on the difference between the fourth predicted downsampled character and the true character; the phoneme and vowel feature sample with 1 / 16 of the original sequence length is obtained by downsampling the phoneme and vowel feature sample with 1 / 8 of the original sequence length through the second downsampling layer using the second convolution stride.

[0157] Continue to refer to Figure 5 , after obtaining the phoneme and vowel feature sample with 1 / 8 of the original sequence length by continuing to perform phoneme and vowel modeling on the phoneme feature sample with 1 / 8 of the original sequence length through at least one spaced Conformer Block, the second downsampling layer can be used to downsample the phoneme and vowel feature sample with 1 / 8 of the original sequence length to obtain the phoneme and vowel feature sample with 1 / 16 of the original sequence length.

[0158] Next, a linear layer can be used to predict the corresponding characters (i.e., the fourth predicted downsampled characters) based on these initial consonant and vowel feature samples with a length of 1 / 16 times the original sequence length. Then, the fourth predicted downsampled character loss is calculated based on the difference between the predicted fourth predicted downsampled characters and the true characters.

[0159] This stage uses the CTC (Connectionist Temporal Classification) loss function for calculation to accurately evaluate the gap between the prediction result and the true label.

[0160] Step b3: Determine the comprehensive loss based on the response content prediction loss, the first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss.

[0161] Specifically, the response content prediction loss, the phoneme prediction loss, the initial consonant and vowel prediction loss, and the word segmentation prediction loss can be accumulated to obtain the comprehensive loss.

[0162] After obtaining the comprehensive loss through steps b1 to b3, the model parameters can be effectively adjusted by minimizing the comprehensive loss, making it gradually approach the optimal solution, and finally obtaining a well-trained task processing model with excellent performance.

[0163] In some implementation manners, determining the comprehensive loss according to the difference between the predicted response content and the annotated response content in step S604 includes the following steps c1 to c3:

[0164] Step c1: Determine the response content prediction loss according to the difference between the predicted response content and the annotated response content.

[0165] Regarding the specific implementation process of step c1, reference can be made to the specific implementation process of the foregoing step a1, which will not be elaborated here.

[0166] Step c2: Determine the first predicted upsampled character loss, the second predicted upsampled character loss, the third predicted upsampled character loss, and the fourth predicted upsampled character loss. The first predicted upsampled character loss, the second predicted upsampled character loss, the third predicted upsampled character loss, and the fourth predicted upsampled character loss are used to supplement the acoustic information and semantic information of the upsampling loss.

[0167] In this embodiment, the speech feature extraction module uses 16-fold downsampling compression. Therefore, it is necessary to restore the original sequence length through upsampling after downsampling. Continuing to refer to Figure 5 Furthermore, upsampling layers can be inserted into multiple Conformer Blocks. The upsampling layer includes a first upsampling layer and a second upsampling layer.

[0168] Specifically, in a Conformer module that includes multiple Conformer Blocks, the first upsampling layer and the second upsampling layer can be respectively arranged at the latter half position of the module. For example, assuming that a Conformer module includes 16 Conformer Blocks, the first upsampling layer can be arranged between the 10th Conformer Block and the 11th Conformer Block, and the second downsampling layer can be arranged between the 12th Conformer Block and the 13th Conformer Block.

[0169] Of course, in some other examples, the first upsampling layer can also be placed between the 10th and the 11th Conformer Blocks, but the second upsampling layer can be arranged between the 11th and the 12th Conformer Blocks. Or, the first upsampling layer can be placed between the 11th and the 12th Conformer Blocks, but the second upsampling layer can be arranged between the 12th and the 13th Conformer Blocks.

[0170] Continue to refer to Figure 5 , based on this upsampling layer, the method of this embodiment further includes: upsampling the speech feature samples with 1 / 16 times the original sequence length through the first upsampling layer, and speech feature samples with 1 / 8 times the original sequence length can be obtained; upsampling the speech feature samples with 1 / 8 times the original sequence length through the second upsampling layer, and speech feature samples with 1 / 4 times the original sequence length can be obtained.

[0171] After that, the speech feature samples with 1 / 4 times the original sequence length can also be upsampled through a third upsampling layer corresponding to the preposed upsampling layer to obtain speech feature samples with the original sequence length.

[0172] However, during the upsampling process, information loss may also occur. Therefore, an interaction loss can be applied respectively before and after the first upsampling layer and the second upsampling layer to supplement the acoustic information and semantic information lost during upsampling. The interaction loss includes: the first predicted upsampling character loss, the second predicted upsampling character loss, the third predicted upsampling character loss, and the fourth predicted upsampling character loss. The interaction loss is used to supplement the acoustic information and semantic information lost during upsampling. Specifically, determining the first predicted upsampling character loss, the second predicted upsampling character loss, the third predicted upsampling character loss, and the fourth predicted upsampling character loss includes the following steps c21 to step c24:

[0173] Step c21: Determine the first predicted upsampled character corresponding to the initial consonant and vowel feature sample with a length of 1 / 16 times the original sequence length, and determine the loss of the first predicted upsampled character based on the difference between the first predicted upsampled character and the true character.

[0174] Continue to refer to Figure 5 , in step c21, it is first necessary to determine the first predicted upsampled character corresponding to the initial consonant and vowel feature sample with a length of 1 / 16 times the original sequence length. Specifically, this can be accomplished through a linear layer that performs character prediction based on the initial consonant and vowel feature sample with a length of 1 / 16 times the original sequence length obtained from the Conformer Block located before the first upsampling layer. Subsequently, the loss of the first predicted upsampled character is calculated based on the difference between the predicted first predicted upsampled character and the true character.

[0175] Specifically, in the supervised fine-tuning training phase, after extracting the initial consonant and vowel feature sample through the Conformer Block located before the first upsampling layer, the above-mentioned linear layer can be used for character-level prediction. Then, the predicted character is compared with the true character, and the loss value is determined based on this difference. This loss calculation can use the CTC (Connectionist Temporal Classification) loss function to effectively measure the gap between the prediction result and the true label.

[0176] Step c22: Determine the second predicted upsampled character corresponding to the initial consonant and vowel feature sample with a length of 1 / 8 times the original sequence length, and determine the loss of the second predicted upsampled character based on the difference between the second predicted upsampled character and the true character; the initial consonant and vowel feature sample with a length of 1 / 8 times the original sequence length is obtained by upsampling the initial consonant and vowel feature sample with a length of 1 / 16 times the original sequence length through the first upsampling layer.

[0177] Continue to refer to Figure 5 , by upsampling the initial consonant and vowel with a length of 1 / 16 times the original sequence length through the first upsampling layer, an initial consonant and vowel feature sample with a length of 1 / 8 times the original sequence length can be obtained.

[0178] In step c22, a linear layer can be used to predict the corresponding character (i.e., the second predicted upsampled character) based on these initial consonant and vowel feature samples with a length of 1 / 8 times the original sequence length. Then, the loss of the second predicted upsampled character is calculated based on the difference between the predicted second predicted upsampled character and the true character.

[0179] This stage uses the CTC (Connectionist Temporal Classification) loss function for calculation to accurately evaluate the gap between the prediction result and the true label.

[0180] Step c23: Determine the third predicted upsampled character corresponding to the initial consonant and vowel feature samples with a length of 1 / 8 of the original sequence length, and determine the loss of the third predicted upsampled character according to the difference between the third predicted upsampled character and the true character; the initial consonant and vowel feature samples with a length of 1 / 8 of the original sequence length are obtained by performing initial consonant and vowel modeling on the phoneme feature samples with a length of 1 / 8 of the original sequence length through at least one spaced Conformer Block.

[0181] Continue to refer to Figure 5 , after downsampling the initial consonant and vowel feature samples with a length of 1 / 16 of the original sequence length through the first upsampling layer to obtain the initial consonant and vowel feature samples with a length of 1 / 8 of the original sequence length, at least one spaced Conformer Block located between the first upsampling layer and the second upsampling layer can be used to continue performing initial consonant and vowel modeling on the initial consonant and vowel feature samples with a length of 1 / 8 of the original sequence length, so as to obtain the deep initial consonant and vowel feature samples with a length of 1 / 8 of the original sequence length.

[0182] Next, a linear layer can be used to predict the corresponding characters (i.e., the third predicted upsampled characters) according to these deep initial consonant and vowel feature samples with a length of 1 / 8 of the original sequence length. Then, calculate the loss of the third predicted upsampled character according to the difference between the predicted third predicted upsampled character and the true character.

[0183] In this stage, the CTC (Connectionist Temporal Classification) loss function is used for calculation to accurately evaluate the gap between the prediction result and the true label.

[0184] Step c24: Determine the fourth predicted upsampled character corresponding to the initial consonant and vowel feature samples with a length of 1 / 4 of the original sequence length, and determine the loss of the fourth predicted upsampled character according to the difference between the fourth predicted upsampled character and the true character; the initial consonant and vowel feature samples with a length of 1 / 4 of the original sequence length are obtained by upsampling the deep initial consonant and vowel feature samples with a length of 1 / 8 of the original sequence length through the second upsampling layer.

[0185] Continue to refer to Figure 5 , upsampling the deep initial consonant and vowel feature samples with a length of 1 / 8 of the original sequence length through the second upsampling layer to obtain the initial consonant and vowel feature samples with a length of 1 / 4 of the original sequence length.

[0186] Next, a linear layer can be used to predict the corresponding characters (i.e., the fourth predicted upsampled characters) according to these initial consonant and vowel feature samples with a length of 1 / 4 of the original sequence length. Then, calculate the loss of the fourth predicted upsampled character according to the difference between the predicted fourth predicted upsampled character and the true character.

[0187] In this stage, the CTC (Connectionist Temporal Classification) loss function is used for calculation to accurately evaluate the gap between the prediction result and the true label.

[0188] Step c3: Determine the comprehensive loss based on the reply content prediction loss, the first predicted upsampled character loss, the second predicted upsampled character loss, the third predicted upsampled character loss, and the fourth predicted upsampled character loss.

[0189] Specifically, the reply content prediction loss, the phoneme prediction loss, the initial and final consonant prediction loss, and the word segmentation prediction loss can be accumulated to obtain the comprehensive loss.

[0190] After obtaining the comprehensive loss through steps c1 to c3, the model parameters can be effectively adjusted by minimizing the comprehensive loss, making it gradually approach the optimal solution, and finally obtaining a well-trained task processing model with excellent performance.

[0191] In some embodiments, the reply content prediction loss can also be combined with any at least one of the image generation prediction loss, the speech generation prediction loss, the first predicted upsampled character loss, the second predicted upsampled character loss, the third predicted upsampled character loss, the fourth predicted upsampled character, the phoneme prediction loss, the initial and final consonant prediction loss, the word segmentation prediction loss, the first predicted upsampled character loss, the second predicted upsampled character loss, the third predicted upsampled character loss, and the fourth predicted upsampled character loss to determine the comprehensive loss.

[0192] During this supervised fine-tuning training process, all the model parameters of the MLP can be opened, while only some parameters of other modules (such as the image feature extraction module, the speech feature extraction module, the pre-trained LLM, the image generation module, and the speech generation module) are opened. That is to say, when adjusting the model parameters according to the comprehensive loss, all the model parameters of the MLP are adjusted, while some model parameters of other modules are adjusted. Specifically, when adjusting some model parameters of other modules, the LoRA fine-tuning method can be used to achieve this.

[0193] After completing the model training of the above embodiments, the model can be deployed to the user's terminal device as the user's AI personal assistant to assist the user in processing various tasks. To further optimize the model performance, during the actual use process, interaction data between the user and the model can also be collected, such as the user's rating feedback on the model's reply content. This interaction data can be used to perform reinforcement learning on the model regularly to make its behavior more in line with the user's expectations.

[0194] Specifically, the reinforcement learning part can adopt the classical Proximal Policy Optimization (PPO) algorithm. PPO ensures the stability during policy adjustment by restricting the update amplitude of the old and new policies. This algorithm improves on Trust Region Policy Optimization (TRPO) by using a clipping function to control the policy update amplitude, thus avoiding training instability caused by excessive policy changes during the training process. The core of this algorithm lies in optimizing an objective based on the advantage function, aiming to improve the expected return of the content generated by the model. A clipping mechanism is used to constrain the difference between the old and new policies, thereby preventing unnecessary large deviations in the output results.

[0195] Therefore, the reinforcement learning algorithm based on the PPO algorithm can not only significantly improve the user's interaction experience but also enhance the model's content creation ability. For example, the model can update its parameters periodically on a monthly basis to continuously adapt to and meet the changing user needs. This iterative improvement method can ensure that the model can effectively adapt to the evolving expectations and preferences of users in the long term.

[0196] In summary, by performing K-fold downsampling on the encoder of the speech modality in the embodiments of the present application, the inference cost of the model can be significantly reduced, and the computational amount and memory occupancy can be decreased; and by adopting the training method of multi-granularity modeling units based on progressive learning, the model can learn different levels of features and knowledge at different stages, model the data from different granularities and perspectives, and improve the expression ability and generalization ability of the model. Therefore, it is possible to achieve the efficient operation of the end-side multi-modal large model on resource-constrained devices while maintaining an effect close to that of the cloud model, and effectively protecting user privacy.

[0197] In addition, by continuously optimizing the model compression algorithm and training method, it is possible to further reduce the size and computational requirements of the model on the premise of ensuring the model performance, enabling the end-side multi-modal large model to better adapt to the resource limitations of various terminal devices.

[0198] Compared with deploying the large model in the cloud, the end-side multi-modal large model can run directly on the user's device without uploading data to the cloud, greatly reducing the risk of data leakage. While reducing the dependence on the external network through model compression, measures such as data encryption and access control can also be taken to enhance privacy protection and ensure that user data is fully protected throughout the process. Running on the end side can reduce the risk of network attacks and provide stable services even in the case of unstable or disconnected networks, enhancing the security and reliability of the application.

[0199] Exemplary device

[0200] Corresponding to the above task processing method, the embodiments of the present application also provide a task processing device. Figure 7It is a schematic structural diagram of a task processing device provided by an embodiment of the present application. As Figure 7 shown, the task processing device provided by the embodiment of the present application includes: a receiving unit 701 and a task processing unit 702; wherein, the receiving unit 701 is configured to receive a user request, and the user request includes input data, and the input data includes at least one of image data, text data, and voice data; the task processing unit 702 is configured to process the input data through a task processing model to generate a reply content; wherein, the task processing model is deployed on the user's terminal device, and the task processing model includes a voice feature extraction module for performing K-fold downsampling on voice data, and the voice feature extraction module is obtained by performing K-fold compression on the Conformer module, and K is a positive integer greater than 8.

[0201] In some embodiments, the voice feature extraction module includes a pre-downsampling layer and a Conformer module, and the Conformer module includes a plurality of stacked Conformer Blocks, and a post-downsampling layer is inserted in the plurality of Conformer Blocks; wherein, the task processing unit 702 processes the input data through the task processing model to generate a reply content, including: downsampling the initial voice features corresponding to the voice data through the pre-downsampling layer at a first sampling rate to obtain voice features with a first sequence length; the first sampling rate is 4 times the original sampling rate; downsampling the voice features with the first sequence length through the post-downsampling layer at a second sampling rate for at least two times to obtain voice features with a second sequence length; the second sampling rate is 2*N times the original sampling rate, and N is a positive integer; generating the reply content according to the voice features with the second sequence length.

[0202] In some embodiments, the post-downsampling layer includes a first downsampling layer and a second downsampling layer, and the second sampling rate is 2*1 times the original sampling rate; wherein, the task processing unit 702 downsamples the voice features with the first sequence length through the post-downsampling layer at the second sampling rate to obtain voice features with a second sequence length, including: downsampling the voice features with the first sequence length through the first downsampling layer at 2*1 times the original sampling rate to obtain voice features with 1 / 8 times the original sequence length; downsampling the voice features with 1 / 8 times the original sequence length through the second downsampling layer at 2*1 times the original sampling rate to obtain voice features with 1 / 16 times the original sequence length.

[0203] In some embodiments, the speech feature extraction module includes a pre-downsampling layer and a Conformer module. The Conformer module includes a plurality of stacked Conformer Blocks, and a post-downsampling layer is inserted into the plurality of Conformer Blocks. The post-downsampling layer includes a third downsampling layer. Among them, the task processing unit 702 processes the input data through a task processing model to generate a response content, including: downsampling the initial speech features corresponding to the speech data by a first sampling rate through the pre-downsampling layer to obtain speech features with a first sequence length; the first sampling rate is 4 times the original sampling rate; downsampling the speech features with the first sequence length by a third sampling rate through the third downsampling layer to obtain speech features with a second sequence length; the third sampling rate is 4 times the original sampling rate; generating the response content according to the speech features with the second sequence length.

[0204] In some embodiments, the task processing unit 702 processes the input data through a task processing model to generate a response content, including: performing phoneme modeling on the initial speech features with the first sequence length corresponding to the input data through the plurality of Conformer Blocks to obtain phoneme features; performing initial-final modeling on the phoneme features to obtain initial-final features; performing word segmentation modeling on the initial-final features to obtain word segmentation features; generating the response content according to the word segmentation features.

[0205] In some embodiments, the task processing model is trained through multimodal sample data, and the multimodal sample data includes at least one of image samples, speech samples, and text samples. The task processing model further includes an image feature extraction module and a pre-trained LLM model. Among them, the task processing model is trained by the following steps: extracting image features from the image samples through the image feature extraction module to obtain image feature samples; extracting speech features from the speech samples through the speech feature extraction module to obtain speech feature samples; generating a predicted response content by the pre-trained LLM model according to at least one of the text samples, the image feature samples, and the speech feature samples; determining a comprehensive loss according to the difference between the predicted response content and the labeled response content, and converging the task processing model to be trained based on the comprehensive loss to obtain a trained task processing model.

[0206] In some embodiments, the speech feature extraction module includes a Conformer module, and the Conformer module includes a plurality of stacked Conformer Blocks; wherein, determining the comprehensive loss according to the difference between the predicted response content and the annotated response content includes: determining the response content prediction loss according to the difference between the predicted response content and the annotated response content; determining the phoneme prediction loss according to the difference between the predicted phoneme and the true phoneme corresponding to the phoneme feature sample; determining the initial consonant and vowel prediction loss according to the difference between the predicted initial consonant and vowel and the true initial consonant and vowel corresponding to the initial consonant and vowel feature sample; determining the word segmentation prediction loss according to the difference between the predicted word segmentation and the true word segmentation corresponding to the word segmentation feature sample; determining the comprehensive loss based on the response content prediction loss, the phoneme prediction loss, the initial consonant and vowel prediction loss, and the word segmentation prediction loss; wherein, the phoneme feature sample, the initial consonant and vowel feature sample, and the word segmentation feature sample are obtained by performing phoneme modeling, initial consonant and vowel modeling, and word segmentation modeling in sequence on the initial speech feature sample corresponding to the speech sample through the plurality of Conformer Blocks.

[0207] In some embodiments, the speech feature extraction module includes a Conformer module, the Conformer module includes a plurality of stacked Conformer Blocks, a post-downsampling layer is inserted in the plurality of Conformer Blocks, the post-downsampling layer includes a first downsampling layer and a second downsampling layer, and at least one interval Conformer Block is arranged between the first downsampling layer and the second downsampling layer; wherein, determining the comprehensive loss according to the difference between the predicted response content and the annotated response content includes: determining the response content prediction loss according to the difference between the predicted response content and the annotated response content; determining the first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss, and the first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss are used to supplement the loss acoustic information and loss semantic information during downsampling; determining the comprehensive loss based on the response content prediction loss, the first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss.

[0208] In some embodiments, determining the first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss includes: determining a first predicted downsampled character corresponding to a phoneme feature sample, and determining the first predicted downsampled character loss according to the difference between the first predicted downsampled character and the true character; determining a second predicted downsampled character corresponding to a phoneme feature sample with 1 / 8 of the original sequence length, and determining the second predicted downsampled character loss according to the difference between the second predicted downsampled character and the true character; the phoneme feature sample with 1 / 8 of the original sequence length is obtained by downsampling the phoneme feature sample by a factor of 2*1 using the first downsampling layer; determining a third predicted downsampled character corresponding to the initial and final consonant feature sample with 1 / 8 of the original sequence length, and determining the third predicted downsampled character loss according to the difference between the third predicted downsampled character and the true character; the initial and final consonant feature sample with 1 / 8 of the original sequence length is obtained by performing initial and final consonant modeling on the phoneme feature sample with 1 / 8 of the original sequence length using the at least one spaced Conformer Block; determining a fourth predicted downsampled character corresponding to the initial and final consonant feature sample with 1 / 16 of the original sequence length, and determining the fourth predicted downsampled character loss according to the difference between the fourth predicted downsampled character and the true character; the initial and final consonant feature sample with 1 / 16 of the original sequence length is obtained by downsampling the initial and final consonant feature sample with 1 / 8 of the original sequence length by a factor of 2*1 using the second downsampling layer.

[0209] In some embodiments, the speech feature extraction module includes a Conformer module, the Conformer module includes a plurality of stacked Conformer Blocks, an upsampling layer is inserted in the plurality of Conformer Blocks, and the upsampling layer includes a first upsampling layer and a second upsampling layer; wherein, determining the comprehensive loss according to the difference between the predicted response content and the labeled response content includes: determining a response content prediction loss according to the difference between the predicted response content and the labeled response content; determining a first predicted upsampled character loss, a second predicted upsampled character loss, a third predicted upsampled character loss, and a fourth predicted upsampled character loss, and the first predicted upsampled character loss, the second predicted upsampled character loss, the third predicted upsampled character loss, and the fourth predicted upsampled character loss are used to supplement the acoustic information and semantic information of the upsampling loss; determining the comprehensive loss based on the response content prediction loss, the first predicted upsampled character loss, the second predicted upsampled character loss, the third predicted upsampled character loss, and the fourth predicted upsampled character loss.

[0210] The task processing device provided in this embodiment belongs to the same inventive concept as the task processing method provided in the foregoing embodiments of the present application, and can execute the task processing methods provided in any of the foregoing embodiments of the present application, and has corresponding functional modules and beneficial effects for executing the task processing method. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the task processing method provided in the foregoing embodiments of the present application, which will not be elaborated here.

[0211] The functions implemented by the above receiving unit 701 and task processing unit 702 can be implemented by the same or different processors respectively, which is not limited in the embodiments of the present application.

[0212] It should be understood that the units in the above device can be implemented in the form of a processor invoking software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor invokes the instructions stored in the memory to implement any of the above methods or the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory inside the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of a hardware circuit, and the functions of some or all of the units can be realized by designing the hardware circuit. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured through a configuration file, so as to realize the functions of some or all of the above units. All units of the above device can be all implemented in the form of a processor invoking software, or all implemented in the form of a hardware circuit, or some implemented in the form of a processor invoking software, and the remaining part implemented in the form of a hardware circuit.

[0213] In the embodiments of the present application, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and running capabilities, such as a CPU, a microprocessor, a GPU, or a DSP, etc.; in another implementation, the processor can realize certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or can be reconstructed. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA, etc. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a kind of ASIC, such as an NPU, a TPU, a DPU, etc.

[0214] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method. For example: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0215] In addition, each unit in the above device can be integrated in whole or in part, or can be independently implemented. In one implementation, these units are integrated together and implemented in the form of an SOC. The SOC can include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The types of the at least one processor can be different. For example, it can include a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0216] Exemplary electronic device

[0217] An embodiment of the present application provides an electronic device. Refer to Figure 8 As shown, the electronic device includes:

[0218] A memory 200 and a processor 210;

[0219] Wherein, the memory 200 is connected to the processor 210 and is used for storing programs;

[0220] The processor 210 is used for implementing the task processing method disclosed in any of the above embodiments by running the programs stored in the memory 200.

[0221] Specifically, the above electronic device may further include: a bus, a communication interface 220, an input device 230, and an output device 240.

[0222] The processor 210, the memory 200, the communication interface 220, the input device 230, and the output device 240 are interconnected through the bus. Among them:

[0223] The bus may include a path for transmitting information between various components of the computer system.

[0224] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0225] The processor 210 may include a main processor, and may also include a baseband chip, a modem, etc.

[0226] The memory 200 stores a program for implementing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, etc.

[0227] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.

[0228] The output device 240 may include a device for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.

[0229] The communication interface 220 may include a device of any transceiver type for communicating with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0230] The processor 210 executes the program stored in the memory 200 and calls other devices, and can be used to implement each step of any task processing method provided in the above embodiments of the present application.

[0231] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored on a memory through the data interface to execute the task processing method introduced in any of the above embodiments. The specific processing process and its beneficial effects can be referred to the embodiment introduction of the above task processing method.

[0232] Exemplary computer program product and storage medium

[0233] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the task processing method according to various embodiments of the present application described in any of the above embodiments of this specification.

[0234] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0235] In addition, an embodiment of the present application can also be a storage medium on which a computer program is stored. The computer program is executed by a processor to perform the steps in the task processing method according to various embodiments of the present application described in any of the above embodiments of the present specification, and specifically can implement the steps of the above task processing method.

[0236] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0237] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0238] The steps in the methods of the various embodiments of the present application can be adjusted, combined, and deleted according to actual needs. The technical features recorded in the various embodiments can be replaced or combined.

[0239] The modules and sub-modules in the devices and terminals in the various embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0240] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in electrical, mechanical, or other forms.

[0241] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place or distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0242] In addition, each functional module or sub-module in various embodiments of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware or in the form of software functional modules or sub-modules.

[0243] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0244] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0245] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0246] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A task processing method, characterized in that, Including: Receiving a user request, where the user request includes input data, and the input data includes at least one of image data, text data, and voice data; Processing the input data through a task processing model to generate a reply content; Wherein, the task processing model is deployed on the user's terminal device, and the task processing model includes a voice feature extraction module for performing K-fold downsampling on voice data, and the voice feature extraction module is obtained by performing K-fold compression on the Conformer module, and K is a positive integer greater than 8.

2. The method according to claim 1, wherein The voice feature extraction module includes a pre-downsampling layer and a Conformer module, and the Conformer module includes a plurality of stacked Conformer Blocks, and a post-downsampling layer is inserted into the plurality of Conformer Blocks; Wherein, processing the input data through the task processing model to generate a reply content includes: Performing convolutional downsampling on the initial voice features corresponding to the voice data through the pre-downsampling layer using a first convolutional stride to obtain voice features with a first sequence length; the first convolutional stride is 4 times the unit convolutional stride; Performing convolutional downsampling on the voice features with the first sequence length through the post-downsampling layer using a second convolutional stride at least twice to obtain voice features with a second sequence length; the second convolutional stride is 2 times the unit convolutional stride; Generating the reply content according to the voice features with the second sequence length.

3. The method according to claim 2, characterized in that The post-downsampling layer includes a first downsampling layer and a second downsampling layer; Wherein, performing downsampling on the voice features with the first sequence length through the post-downsampling layer using a second sampling rate to obtain voice features with a second sequence length includes: Performing convolutional downsampling on the voice features with the first sequence length through the first downsampling layer using the second convolutional stride to obtain voice features with 1 / 8 times the original sequence length; Performing convolutional downsampling on the voice features with 1 / 8 times the original sequence length through the second downsampling layer using the second convolutional stride to obtain voice features with 1 / 16 times the original sequence length.

4. The method according to claim 1, characterized in that, The voice feature extraction module includes a pre-downsampling layer and a Conformer module, the Conformer module includes a plurality of stacked Conformer Blocks, a post-downsampling layer is inserted into the plurality of Conformer Blocks, and the post-downsampling layer includes a third downsampling layer; Wherein, processing the input data through the task processing model to generate a reply content includes: Performing convolutional downsampling on the initial voice features corresponding to the voice data through the pre-downsampling layer using a first convolutional stride to obtain voice features with a first sequence length; the first convolutional stride is 4 times the unit convolutional stride; Performing convolutional downsampling on the voice features with the first sequence length through the third downsampling layer using a third convolutional stride to obtain voice features with a second sequence length; the third convolutional stride is 4 times the unit convolutional stride; Generate the reply content according to the speech features of the second sequence length.

5. The method according to any one of claims 2-4, characterized in that, The task processing model processes the input data to generate a reply content, including: Phoneme modeling is performed on the initial speech features of the first sequence length corresponding to the input data through the multiple Conformer Blocks to obtain phoneme features; Initial consonant and vowel modeling is performed based on the phoneme features to obtain initial consonant and vowel features; Word segmentation modeling is performed based on the initial consonant and vowel features to obtain word segmentation features; Generate the reply content based on the word segmentation features.

6. The method according to any one of claims 1-4, characterized in that, The task processing model is trained through multi-modal sample data, and the multi-modal sample data includes at least one of image samples, speech samples, and text samples. The task processing model also includes an image feature extraction module and a pre-trained LLM model; Among them, the task processing model is trained by the following steps: The image feature extraction module extracts features from the image samples to obtain image feature samples; The speech feature extraction module extracts features from the speech samples to obtain speech feature samples; The pre-trained LLM model generates a predicted reply content based on at least one of the text samples, the image feature samples, and the speech feature samples; Determine the comprehensive loss according to the difference between the predicted reply content and the annotated reply content, and converge the task processing model to be trained based on the comprehensive loss to obtain a trained task processing model.

7. The method according to claim 6, wherein The speech feature extraction module includes a Conformer module, and the Conformer module includes a plurality of Conformer Blocks arranged in a stack; Among them, determining the comprehensive loss according to the difference between the predicted reply content and the annotated reply content includes: Determine the reply content prediction loss according to the difference between the predicted reply content and the annotated reply content; Determine the phoneme prediction loss according to the difference between the predicted phoneme corresponding to the phoneme feature sample and the true phoneme; Determine the initial consonant and vowel prediction loss according to the difference between the predicted initial consonant and vowel corresponding to the initial consonant and vowel feature sample and the true initial consonant and vowel; Determine the word segmentation prediction loss according to the difference between the predicted word segmentation corresponding to the word segmentation feature sample and the true word segmentation; Based on the reply content prediction loss, the phoneme prediction loss, the initial consonant and vowel prediction loss, and the word segmentation prediction loss, determine the comprehensive loss; Among them, the phoneme feature samples, the initial consonant and vowel feature samples, and the word segmentation feature samples are obtained by performing phoneme modeling, initial consonant and vowel modeling, and word segmentation modeling on the initial speech feature samples corresponding to the speech samples through the multiple Conformer Blocks in sequence.

8. The method according to claim 6, wherein The speech feature extraction module includes a Conformer module, and the Conformer module includes a plurality of Conformer Blocks arranged in a stack. A post-downsampling layer is inserted into the plurality of Conformer Blocks. The post-downsampling layer includes a first downsampling layer and a second downsampling layer, and at least one spaced Conformer Block is arranged between the first downsampling layer and the second downsampling layer; Among them, determining the comprehensive loss according to the difference between the predicted response content and the labeled response content includes: Determining the response content prediction loss according to the difference between the predicted response content and the labeled response content; Determining a first predicted downsampled character loss, a second predicted downsampled character loss, a third predicted downsampled character loss, and a fourth predicted downsampled character loss. The first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss are used to supplement the loss acoustic information and loss semantic information during downsampling; Determining the comprehensive loss based on the response content prediction loss, the first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss.

9. The method according to claim 8, wherein The determining the first predicted downsampled character loss, the second predicted downsampled character loss, the third predicted downsampled character loss, and the fourth predicted downsampled character loss includes: Determining a first predicted downsampled character corresponding to the phoneme feature sample, and determining the first predicted downsampled character loss according to the difference between the first predicted downsampled character and the true character; Determining a second predicted downsampled character corresponding to the phoneme feature sample with 1 / 8 of the original sequence length, and determining the second predicted downsampled character loss according to the difference between the second predicted downsampled character and the true character; the phoneme feature sample with 1 / 8 of the original sequence length is obtained by performing convolutional downsampling on the phoneme feature sample by the first downsampling layer using a second convolutional stride; Determining a third predicted downsampled character corresponding to the initial-final feature sample with 1 / 8 of the original sequence length, and determining the third predicted downsampled character loss according to the difference between the third predicted downsampled character and the true character; the initial-final feature sample with 1 / 8 of the original sequence length is obtained by performing initial-final modeling on the phoneme feature sample with 1 / 8 of the original sequence length by the at least one spaced Conformer Block; Determining a fourth predicted downsampled character corresponding to the initial-final feature sample with 1 / 16 of the original sequence length, and determining the fourth predicted downsampled character loss according to the difference between the fourth predicted downsampled character and the true character; the initial-final feature sample with 1 / 16 of the original sequence length is obtained by performing downsampling on the initial-final feature sample with 1 / 8 of the original sequence length by the second downsampling layer using the second convolutional stride, and the second convolutional stride is 2 times the unit convolutional stride.

10. The method according to claim 6, wherein The speech feature extraction module includes a Conformer module, and the Conformer module includes a plurality of stacked Conformer Blocks. An upsampling layer is inserted into the plurality of Conformer Blocks, and the upsampling layer includes a first upsampling layer and a second upsampling layer; Among them, determining the comprehensive loss according to the difference between the predicted response content and the labeled response content includes: Determining the response content prediction loss according to the difference between the predicted response content and the labeled response content; Determining a first predicted upsampling character loss, a second predicted upsampling character loss, a third predicted upsampling character loss, and a fourth predicted upsampling character loss, where the first predicted upsampling character loss, the second predicted upsampling character loss, the third predicted upsampling character loss, and the fourth predicted upsampling character loss are used to supplement the acoustic information and semantic information of the upsampling loss; Determining the comprehensive loss based on the response content prediction loss, the first predicted upsampling character loss, the second predicted upsampling character loss, the third predicted upsampling character loss, and the fourth predicted upsampling character loss.

11. A task processing device, characterized in that, Including: A receiving module, configured to receive a user request, where the user request includes input data, and the input data includes at least one of image data, text data, and voice data; A task processing module, configured to process the input data through a task processing model to generate response content; Among them, the task processing model is deployed on the user's terminal device, and the task processing model includes a speech feature extraction module, configured to perform K-fold downsampling on the speech data. The speech feature extraction module is obtained by compressing the Conformer module by K times, and K is a positive integer greater than 8.

12. An electronic device, characterized in that, Including a memory and a processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the method according to any one of claims 1 to 10 by running the program in the memory.

13. A storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is run by the processor, the method according to any one of claims 1 to 10 is implemented.

14. A computer program product, characterized in that, Including computer program instructions, and when the computer program instructions are run by the processor, the processor is caused to implement the method according to any one of claims 1 to 10.