Information processing method and related device

By acquiring discrete and continuous feature information, and combining a feature extraction module with a generative machine learning model, the problem that generative models cannot directly generate audio is solved, achieving high-quality audio generation and image understanding, and improving the user experience.

WO2026051936A1PCT designated stage Publication Date: 2026-03-12HUAWEI TECH CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing generative machine learning models can only generate discrete feature information and cannot directly output audio. They also cannot effectively understand continuous audio feature information, resulting in the inability to generate audio end-to-end.

Method used

By acquiring discrete and continuous feature information of the input content, using discrete and continuous feature extraction modules to extract feature information, and combining it with a generative machine learning model, feature fusion of audio and image is achieved to generate high-quality audio content.

Benefits of technology

It achieves end-to-end audio generation capabilities and improves the understanding of image information, resulting in higher quality and more natural content, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025118594_12032026_PF_FP_ABST
    Figure CN2025118594_12032026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an information processing method and a related device, applied to the field of artificial intelligence. The method comprises: acquiring initial feature information of input content; and inputting the initial feature information into a first machine learning model to obtain generated content. The generated content is generated after the first machine learning model executes a task. If the input content comprises an audio, initial feature information of the audio is discrete feature information, and thus the first machine learning model uses the discrete feature information to understand the audio, so that the first machine learning model has the capability to directly generate the audio. When the task comprises generating an audio, the generated content comprises an audio. If the input content comprises an image, initial feature information of the image is continuous feature information. By combining the discrete feature information of the audio with the continuous feature information of the image, end-to-end audio generation can be realized, and the image can be well understood, thereby improving the quality of the generated content.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method and related device

[0001] The present application claims priority from the Chinese patent application No. 202411240376.X filed on September 4, 2024, and entitled "Information processing method and related device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the field of artificial intelligence, and in particular to an information processing method and related device. BACKGROUND

[0003] Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is the design principle and implementation method of various intelligent machines, enabling machines to have perception, reasoning and decision-making functions.

[0004] As a cutting-edge technology in the field of artificial intelligence, generative machine learning models have a very wide range of application scenarios. However, the data type of the generated content directly output by the current generative machine learning model is only text. In order to interact with the user in the form of audio, the text generated by the machine learning model needs to be converted into audio by calling an externally hung text-to-speech (TTS) module. A scheme for directly generating audio through a machine learning model is urgently needed. SUMMARY

[0005] The present application provides an information processing method and related device, which not only realizes the ability to generate audio end-to-end, but also has a good understanding of images in the input content, which is beneficial to improve the quality of the generated content.

[0006] The present application provides the following technical solutions:

[0007] In a first aspect, the present application provides an information processing method, which can be used in the field of artificial intelligence. The method comprises: a first device obtaining at least one initial feature information of input content, inputting the at least one initial feature information of the input content into a first machine learning model, and obtaining generated content corresponding to the input content. If the input content includes a first audio, the initial feature information of the first audio is discrete feature information, and the initial feature information of the first audio is obtained by performing feature extraction on the first audio. If the input content includes an image, the initial feature information of the image is continuous feature information, and the initial feature information of the first image is obtained by performing feature extraction on the first image.

[0008] The generated content corresponding to the input content is generated after the first machine learning model performs a task. In the case where the first machine learning model performs a task including generating an audio, the generated content includes a second audio corresponding to the input content. Exemplarily, the first device inputs at least one initial feature information of the input content and prompt information into the first machine learning model to obtain generated content corresponding to the input content. The prompt information can indicate a task that the first machine learning model needs to perform.

[0009] Further, in one case, the first device inputs at least one initial feature information of the input content into the first machine learning model to obtain generated content corresponding to the input content, which comprises: the first device inputs at least one initial feature information of the input content into the first machine learning model, and generates the generated content corresponding to the input content through the generative first machine learning model. In another case, the first device inputs at least one initial feature information of the input content into the first machine learning model to obtain generated content corresponding to the input content, which comprises: the first device sends the execution device of the first machine learning model, receives the generated content corresponding to the input content sent by the execution device of the first machine learning model, and the generated content corresponding to the input content is generated by the execution device through the first machine learning model.

[0010] The technical personnel find that, since the current generative machine learning model can only generate discrete feature information, and the initial feature information of the audio often adopts continuous feature information, the machine learning model can only adopt the way of continuous feature information to understand the audio, resulting in the inability to restore the audio based on the generated discrete feature information. In the implementation mode, the initial feature information of the audio obtained is discrete feature information, so the first machine learning model can adopt discrete feature information to understand the audio, so that the first machine learning model has the ability to directly generate audio. In addition, the discrete feature information of the audio is combined with the continuous feature information of the image, and the continuous feature information of the image can retain rich information in the image, so not only the ability of end-to-end generated audio is realized, but also the image in the input content is well understood, which is conducive to improving the quality of the generated content.

[0011] In a possible implementation mode, the values in the first value space of the feature values in the discrete feature information are discrete, and the first value space refers to a set composed of all values that the feature values in the discrete feature information can adopt, and the number of values in the first value space is limited. The second value space of the feature values in the continuous feature information is continuous, and the second value space refers to a set composed of all values that the feature values in the discrete feature information can adopt, and the number of values in the second value space is not limited, in other words, the number of values in the second value space can be infinite.

[0012] Exemplarily, the initial feature information of the first audio can include at least one first feature value, and any one of the at least one first feature value is included in a limited value set (that is, a first value interval). The initial feature information of the first image can include at least one second feature value, and any one of the at least one second feature value can take any real value in a value range, that is, the second value space is any real value in the value range; or, any one of the at least one second feature value can take any real value, that is, the second value space is any real value.

[0013] In the embodiment of the application, the concepts of continuous feature information and discrete feature information are further clarified, thereby reducing the implementation difficulty of the scheme and improving the realizability of the scheme.

[0014] In a possible implementation, the first device obtains at least one initial feature information of the input content, including: if the input content includes first audio, the first device can input the first audio in the input content into a first feature extraction module, and obtain initial feature information of the first audio by performing discretization processing on the first audio through the first feature extraction module; for example, the initial feature information of the first audio includes at least one token. The first feature extraction module is a discrete feature extraction module, and optionally, the first feature extraction module can be a discrete audio encoder.

[0015] If the input content includes first image, the first device can input the first image in the input content into a second feature extraction module to obtain feature information of the first image generated by the second feature extraction module; for example, if the feature information of the first image generated by the second feature extraction module includes at least one feature map, the first device can further input the feature information of the first image generated by the second feature extraction module into a conversion module to obtain initial feature information of the first image generated by the conversion module; for example, the initial feature information of the first audio includes at least one token. The second feature extraction module is a continuous feature extraction module, and optionally, the second feature extraction module can be a continuous vision encoder; the conversion module is configured to convert the at least one feature map generated by the second feature extraction module into feature information in the form of token.

[0016] In a possible implementation, the second audio is generated by the first machine learning model based on the first feature information; for example, the first machine learning model can include a feature processing module corresponding to the first feature extraction module, and the feature processing module is configured to obtain the second audio based on the first feature information generated by the first machine learning model, where the first feature information can be discrete feature information, for example, the feature processing module can also be understood as an audio decoder.

[0017] In a possible implementation, the second audio is generated by the first machine learning model based on the first feature information and the second feature information, where the first feature information indicates semantic content of the second audio, and the second feature information indicates emotional style of the second audio, for example, the emotional style of the second audio can be any one of the following: gentle, happy, neutral, surprised, scared, or other emotional styles, etc.

[0018] Optionally, the first feature information also indicates a pitch of the second audio. For example, if the task required to be performed by the first machine learning model includes determining a pitch of the generated audio, the first feature information can also indicate the pitch of the second audio. For example, the pitch of the second audio can be any one of the following: low, normal, high, or other pitches.

[0019] Optionally, the second feature information also indicates a speech rate of the second audio. For example, if the task required to be performed by the first machine learning model includes determining a speech rate of the generated audio, the first feature information can also indicate the speech rate of the second audio. For example, the speech rate of the second audio can be any one of the following: slow, normal, fast, or other speech rates.

[0020] In the embodiments of the present application, the emotion is injected into the second audio, which is beneficial to improve the naturalness of the output second audio and improve the emotional expression effect of the second audio, so as to improve the user experience of the present scheme. In addition, the second feature information specially used to indicate the emotional style of the second audio is introduced, and the first machine learning model can generate the second audio in combination with the semantic content of the second audio and the emotional style of the second audio, which decouples the semantic content and the emotional style of the second audio, is beneficial to reduce the difficulty of the first machine learning model in generating the second audio, and is beneficial to obtain a more natural and humanized second audio.

[0021] In a possible implementation manner, the second feature information is obtained by a second machine learning model based on an emotion category corresponding to the input content, and the emotion category is generated by the first machine learning model. The prompt information can indicate that the task required to be performed by the generative first machine learning model includes identifying the emotion category of the input content, so that the first machine learning model can generate the emotion category corresponding to the input content. The emotion category corresponding to the input content can be at least one of the following: neutral, happy, sad, angry, surprised, disgusted, scared, anxious, or other emotion categories. Optionally, the first device inputs at least one initial feature information of the input content into the first machine learning model to obtain generated content corresponding to the input content, including: after the first device inputs at least one initial feature information of the input content and the prompt information into the first machine learning model, the first feature information and the emotion category of the input content are generated by the first machine learning model; the first device inputs the emotion category of the input content into the second machine learning model to obtain the second feature information generated by the second machine learning model; and the first device can generate the second audio based on the first feature information and the second feature information through a feature processing module in the first machine learning model.

[0022] In the embodiments of the present application, the first machine learning model is used to generate the mood category of the input content, and then the second machine learning model is used to generate the second feature information for representing the emotional style of the second audio. The second machine learning model is independent, which is conducive to generating second feature information with better quality, and is conducive to making the emotional style of the second audio more suitable for the mood category of the input content, so as to further improve the user experience of the present solution. For example, when the input content expresses an anxious mood, the second audio can be of a gentle emotional style, and the intelligent assistant can respond in a gentle tone, thereby achieving a calming effect. For another example, when the input content expresses a neutral mood, the second audio can be of a neutral emotional style, and the intelligent assistant can respond in a neutral tone. For another example, the emotional style of the second audio can be determined according to the mood type expressed by the student through the input content, and then the educational assistant can explain the teaching content in the aforementioned emotional style. For another example, for a long-term care patient, the medical consultation system can adopt a gentle emotional style, and the like, to provide a more personalized second audio, thereby further improving the user experience of the present solution.

[0023] In a possible implementation manner, the first training stage of the first machine learning model can also be referred to as a pre-training stage of the first machine learning model. In the first training stage of the first machine learning model, the training task of the first machine learning model includes generating a text description corresponding to an input training sample. The data types of the training sample include images, audio, or text, and different data types of training samples are cross-used. In other words, the first training stage of the first machine learning model can include multiple rounds of training of the first machine learning model, and different data types of first training samples can be used in adjacent rounds of training.

[0024] In the embodiments of the present application, in the first training stage of the first machine learning model, different data types of training samples are cross-used to train the first machine learning model, so that the first machine learning model cross-understands different data types of training samples. The confusion degree of the first training stage is improved, which is conducive to increasing the difficulty of the first training stage of the first machine learning model, thereby being conducive to improving the understanding ability of the trained first machine learning model for various data types of data, and being conducive to making the first machine learning model generate better generated content on the premise of fully understanding the input content.

[0025] In a possible implementation manner, the second training stage of the first machine learning model can also be referred to as a post-training stage of the first machine learning model. In the second training stage of the first machine learning model, in each inference process of the first machine learning model, the training task performed by the first machine learning model includes generating at least two of images, audio, and text. Optionally, the training task performed by the first machine learning model includes generating images, audio, and text.

[0026] In the second training stage of the first machine learning model in the embodiments of the present application, the training task of each inference process of the first machine learning model includes generating at least two of images, audio and text, that is, generating data of at least two data types, which improves the difficulty of the second training stage, is conducive to stimulating the generation ability of the first machine learning model for data of multiple data types, and is conducive to improving the quality of newly generated images, audio and text obtained by the trained first machine learning model.

[0027] In a second aspect, the present application provides an information processing device, which can be used in the field of artificial intelligence, and the device comprises: an acquisition module configured to acquire at least one initial feature information of input content, wherein if the input content comprises a first audio, the initial feature information of the first audio is discrete feature information, and if the input content comprises an image, the initial feature information of the image is continuous feature information; and a processing module configured to input the at least one initial feature information into a first machine learning model to obtain generated content corresponding to the input content, wherein the generated content is generated after the first machine learning model performs a task, and in the case that the task comprises generating audio, the generated content comprises second audio corresponding to the input content.

[0028] In the second aspect of the present application, the information processing device is further configured to perform the steps performed by the first device in the first aspect and various possible implementation manners of the first aspect. The specific implementation manners of the steps, the meanings of the terms and the beneficial effects brought by the second aspect can be referred to the first aspect, and will not be described here.

[0029] In a third aspect, the present application provides a device comprising a processor and a memory, wherein the processor is coupled to the memory, the memory is configured to store a program, and the processor is configured to execute the program in the memory, so that the first device executes the method of the first aspect described above.

[0030] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and when the computer program runs on a computer, the computer program causes the computer to execute the method of the first aspect described above.

[0031] In a fifth aspect, the present application provides a computer program product, wherein the computer program product comprises a program, and when the program runs on a computer, the program causes the computer to execute the method of the first aspect described above.

[0032] In a sixth aspect, the present application provides a chip system, which comprises a processor for supporting the implementation of the functions involved in the above aspects, such as sending or processing the data and / or information involved in the above methods. In a possible design, the chip system further comprises a memory, and the memory is configured to store necessary program instructions and data of the terminal device or the communication device. The chip system can be composed of a chip, or can comprise a chip and other discrete devices.

[0033] The third aspect to the sixth aspect of the present application correspond to the first aspect or various possible manners of the first aspect, and have corresponding beneficial effects. BRIEF DESCRIPTION OF DRAWINGS

[0034] FIG. 1 is a structural schematic diagram of an artificial intelligence subject framework provided by the present application;

[0035] FIG. 2a is a system architecture diagram of an information processing system provided by an embodiment of the present application;

[0036] FIG. 2b is another system architecture diagram of an information processing system provided by an embodiment of the present application;

[0037] FIG. 3 is a flow schematic diagram of an information processing method provided by an embodiment of the present application;

[0038] FIG. 4 is a schematic diagram of generated content obtained through a first machine learning model provided by an embodiment of the present application;

[0039] FIG. 5 is another schematic diagram of generated content obtained through a first machine learning model provided by an embodiment of the present application;

[0040] FIG. 6 is a flow schematic diagram of a model training method provided by an embodiment of the present application;

[0041] FIG. 7 is a schematic diagram of generating a text description corresponding to an input first training sample provided by an embodiment of the present application;

[0042] FIG. 8 is a schematic diagram of input content and prompt information provided by an embodiment of the present application;

[0043] FIG. 9 is a structural schematic diagram of an information processing apparatus provided by an embodiment of the present application;

[0044] FIG. 10 is a structural schematic diagram of an apparatus provided by an embodiment of the present application;

[0045] FIG. 11 is a structural schematic diagram of a chip provided by an embodiment of the present application. DETAILED DESCRIPTION

[0046] The embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art can understand that the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems as the technology develops and new scenarios appear.

[0047] The terms "first", "second", and the like in the description and claims of the present application and above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the terms used in this way can be interchanged, as appropriate, and are merely used to distinguish between objects with the same attributes in the description of the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device that includes a series of units does not necessarily have to be limited to those units, but can include other units not clearly listed or inherent to the process, method, product or device.

[0048] In the embodiments of the present application, "sending" and "receiving" represent the direction of signal transmission. For example, "sending information to XX device" can be understood as the destination of the information being XX device, which can include direct transmission through the air interface, or indirect transmission through the air interface by other units or modules. "Receiving information from YY device" can be understood as the source of the information being YY device, which can include direct reception from YY device through the air interface, or indirect reception from YY device through the air interface from other units or modules. "Sending" can also be understood as "output" of the chip interface, and "receiving" can also be understood as "input" of the chip interface. In other words, sending and receiving can be between devices, or within a device, such as between components, modules, chips, software modules or hardware modules within a device through a bus, wire or interface. It can be understood that the information between the source and the destination of the information transmission may be processed as necessary, such as encoding, modulation, etc., but the destination can understand the valid information from the source. Similar expressions in the present application can be similarly understood, and will not be repeated.

[0049] In the embodiments of the present application, the indication can include direct indication and indirect indication, and can also include explicit indication and implicit indication. The information indicated by certain information (indication information described below) is referred to as to-be-indicated information. In the implementation process, there are many ways to indicate the to-be-indicated information, for example, but not limited to, the to-be-indicated information can be directly indicated, such as the to-be-indicated information itself or an index of the to-be-indicated information. The to-be-indicated information can also be indirectly indicated by indicating other information, where the other information and the to-be-indicated information have an association relationship. The to-be-indicated information can also be indicated only by a part of the to-be-indicated information, and the other part of the to-be-indicated information is known or agreed in advance. For example, the arrangement order of each information agreed in advance (for example, protocol predefined) can be used to indicate a specific information, thereby reducing the indication overhead to a certain extent. The specific manner of indication is not limited in the present application. It can be understood that the indication information can be used to indicate the to-be-indicated information for the sender of the indication information, and the indication information can be used to determine the to-be-indicated information for the receiver of the indication information.

[0050] First, the overall workflow of the artificial intelligence system is described. Please refer to FIG. 1, which is a structural schematic diagram of an artificial intelligence main body framework provided by the present application. The artificial intelligence main body framework is described below from two dimensions of “intelligent information chain” (horizontal axis) and “IT value chain” (vertical axis). The “intelligent information chain” reflects a process from data acquisition to processing. For example, it can be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes a condensation process of “data-information-knowledge-wisdom”. The “IT value chain” reflects the value brought by artificial intelligence to the information technology industry from the bottom infrastructure of human intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.

[0051] (1) Infrastructure

[0052] The infrastructure provides computing power support for the artificial intelligence system, realizes communication with the external world, and realizes support through the underlying platform. Communication with the outside world is realized through sensors; computing power is provided by an intelligent chip, which can specifically adopt a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA) hardware acceleration chip; the underlying platform includes a distributed computing framework and related platform guarantees and supports such as a network, which can include cloud storage and computing, an interconnection network, etc. For example, the sensor and external communication obtain data, which are provided to the intelligent chip in the distributed computing system provided by the underlying platform for calculation.

[0053] (2) Data

[0054] The data of the upper layer of the infrastructure is used to represent the data source in the field of artificial intelligence. The data relates to graphics, images, voice, text, and also relates to the Internet of Things data of traditional devices, including the business data of existing systems and sensing data such as force, displacement, liquid level, temperature, and humidity.

[0055] (3) Data processing

[0056] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0057] Among them, machine learning and deep learning can model, extract, preprocess, train, etc. symbolic and formalized intelligent information of data.

[0058] Reasoning refers to the process of simulating human intelligent reasoning methods in a computer or intelligent system, using formalized information to perform machine thinking and solve problems according to reasoning control strategies, and the typical function is search and matching.

[0059] Decision-making refers to the process of decision-making after intelligent information is reasoned, which usually provides functions such as classification, sorting, and prediction.

[0060] (4) General capabilities

[0061] After the data is processed by the above-mentioned data processing, some general capabilities can be formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0062] (5) Intelligent product and industry application

[0063] Intelligent product and industry application refers to the product and application of artificial intelligence system in various fields, which is the packaging of the overall solution of artificial intelligence, and realizes the application of intelligent information decision productization. Its application fields mainly include: intelligent terminal, smart home, intelligent medical care, intelligent security, intelligent driving, intelligent transportation, smart city, intelligent manufacturing, etc.

[0064] The method provided in the present application can be applied in various application fields of artificial intelligence technology. For example, it can be applied in various application scenarios using a generative first machine learning model. The first machine learning model can be understood as obtaining new generated content based on input content. For example, the generated content can include images, videos, texts, voices or other content, etc. Optionally, the data types supported by the first machine learning model for processing can include at least one of the following: text, image, video, voice or other types of data. In other words, the first machine learning model can support processing of multi-modal data. Alternatively, the first machine learning model can only support processing of single-modal data. The specific determination can be combined with the actual application scenario.

[0065] Since the intelligent terminal, intelligent medical care, smart home and other fields may use the above-mentioned generative first machine learning model, the application scenarios in the application fields of the method provided in the present application are exemplified as follows.

[0066] Application field 1: intelligent terminal field

[0067] The intelligent terminal can be a mobile phone, computer, tablet, robot, wearable device or other types of terminal device, etc. which is not exhaustive here. For example, there can be an intelligent assistant in the intelligent terminal field. The intelligent assistant can be an application program (APP) that provides assistant functions in the terminal device, a robot or other product form with assistant functions, etc. Voice, text, pictures and / or videos, etc. can be input into the intelligent assistant. The intelligent assistant can obtain generated content corresponding to the aforementioned input content through the generative first machine learning model, and then output the generated content to the user. Optionally, if the picture is a picture of a gesture, it can also be understood that the user communicates with the intelligent assistant through gestures. Optionally, in some scenarios, the intelligent assistant may need to answer complex questions by adopting the form of audio.

[0068] For example, in the field of intelligent terminals, there can be an education assistant that can be manifested as an online learning platform, an application program providing teaching assistance functions, a terminal device (such as a tablet specially used for education learning) or other product forms providing teaching assistance functions, etc., and voice, text or pictures, etc. can be input to the education assistant, the education assistant can obtain generated content corresponding to the aforementioned input content through a first machine learning model, and the generated content can include teaching content. Optionally, in some scenarios, the education assistant can need to combine audio-form instructions to more comprehensively explain the teaching content.

[0069] For example, in the field of intelligent terminals, it can be necessary to automatically generate subtitles corresponding to the video, and the video can be input into the first machine learning model to generate subtitles matching each video frame through the first machine learning model.

[0070] For example, in the scenario of using intelligent terminals for online meetings, an intelligent meeting system can provide the functions of automatically generating meeting records and extracting key content in the meeting, that is, the video generated during the online meeting can be input into the first machine learning model to obtain the meeting records and the key content in the meeting generated by the first machine learning model corresponding to the aforementioned video; optionally, the intelligent meeting system can also use the form of voice to feed back the key content in the meeting.

[0071] Application field 2: intelligent medical field

[0072] For example, in the field of intelligent medical treatment, a medical consultation system can be deployed, which can be manifested as an online medical consultation system, an application program providing medical consultation functions, an intelligent device providing medical consultation functions or other product forms, etc.; multi-modal data such as images, texts and / or audios can be input into the medical consultation system, for example, the image can be an X-ray film, a nuclear magnetic resonance image (MRI) or other images used in the diagnosis and treatment process, for example, the text can be a medical record, the audio can be a voice-form medical history description, etc., the medical consultation system can include a generative first machine learning model, and the medical consultation system can generate generated content corresponding to the aforementioned input content through the first machine learning model, and the generated content can include a diagnosis report, which can provide more comprehensive diagnostic support for patients and / or doctors. Optionally, the medical consultation system can also use the form of audio to explain the diagnosis report.

[0073] Application field 3: intelligent home field

[0074] For example, in the field of smart home, after the smart sound is started, the smart sound can continuously acquire audio in the surrounding environment, understand the semantic content of the audio through the first machine learning model, obtain generated content corresponding to the input audio, and then the smart sound can reply in the form of voice.

[0075] It should be noted that the method provided in the present application can also be applied to other application scenarios, such as application scenarios in the field of intelligent driving. The above examples of various application scenarios of the present application are only for easy understanding of the present application, and the application scenarios of the present application are not exhausted.

[0076] Since the current generative first machine learning model can only output text as the data type of the generated data, but in some application scenarios, it is necessary to output in the form of audio. In order to be able to output in the form of audio, after obtaining the text form of the generated content generated by the first machine learning model, the text generated by the machine learning model can be converted into audio by calling an externally hung text-to-speech (TTS) module. However, the externally hung TTS module may not be compatible with the first machine learning model. In order to enable the generative first machine learning model to have the ability to directly generate audio, the present application discloses: obtaining at least one initial feature information of the input content. In the case that the input content includes a first audio, the initial feature information of the first audio is discrete feature information. In the case that the input content includes an image, the initial feature information of the image is continuous feature information. Inputting the at least one initial feature information and prompt information into the first machine learning model to obtain generated content generated by the first machine learning model. The prompt information indicates the task that the first machine learning model needs to perform. In the case that the task that the first machine learning model needs to perform includes generating audio, the generated content includes a second audio corresponding to the input content. The skilled person found that since the current generative machine learning model can only generate discrete feature information, and the initial feature information of the audio often uses continuous feature information, the machine learning model can only use the continuous feature information to understand the audio, resulting in the inability to restore the audio based on the generated discrete feature information. In the present application, the initial feature information of the audio is discrete feature information, so the first machine learning model can use discrete feature information to understand the audio, so that the first machine learning model has the ability to directly generate audio. In addition, the discrete feature information of the audio is combined with the continuous feature information of the image. The continuous feature information of the image can retain the rich information in the image, so as to not only realize the end-to-end generated audio capability, but also have a good understanding of the image in the input content, which is conducive to improving the quality of the generated content.

[0077] Before the method provided in the present application is described in detail, the information processing system provided in the embodiments of the present application is introduced by means of FIG. 2a and FIG. 2b. Please refer to FIG. 2a first, which is a system architecture diagram of the information processing system provided in the embodiments of the present application. In FIG. 2a, the information processing system 200 includes a training device 210, a database 220, an execution device 230 and a data storage system 240. The execution device 230 includes a computing module 231.

[0078] In the database 220, a training data set is stored. In the training stage of the generative first machine learning model 201, the training device 210 performs iterative training on the first machine learning model 201 by using the training data set, so as to obtain the first machine learning model 201 after training. The first machine learning model 201 can be a neural network or a non-neural network model. Optionally, the first machine learning model in the present application can be a deep learning model. For example, the first machine learning model can be a neural network based on an attention mechanism, a residual neural network, a fully connected neural network or a convolutional neural network, etc. For example, the first machine learning model can be a Pangu model or other generative machine learning model, etc.

[0079] The first machine learning model 201 after training can be deployed in the computing module 231 of the execution device 230. For example, in the application stage of the first machine learning model 201, after obtaining at least one initial feature information of the input content, the execution device 230 can input the at least one initial feature information of the input content and the prompt information into the first machine learning model 201 after training, and generate the generated content corresponding to the input content by using the first machine learning model 201 after training.

[0080] Optionally, as shown in FIG. 2a, the execution device 230 and the client device can be integrated in the same device, that is, the user can directly interact with the execution device 230. For example, the execution device 230 can be a mobile phone, a computer, a tablet, a vehicle, a virtual reality (VR) device, a robot, or other types of devices, and the like. For example, the execution device 230 can be a module in a host central processing unit (CPU) of the client device for processing data using the first machine learning model. The execution device 230 can also be a graphics processing unit (GPU), a neural network processing unit (NPU), or a tensor processing unit (TPU) in the client device, which is mounted on the host CPU as a co-processor and is assigned tasks by the host CPU.

[0081] The execution device 230 can call data, code, and the like in the data storage system 240, or store data, instructions, and the like in the data storage system 240. The data storage system 240 can be located in the execution device 230, or can be an external storage system 240 relative to the execution device 230.

[0082] Please continue to refer to FIG. 2b, which is another system architecture diagram of the information processing system provided by the embodiments of the present application. The system architecture shown in FIG. 2b can be understood in combination with the above description of FIG. 2a, and the difference is that the execution device 230 and the client device 250 in the information processing system 200 shown in FIG. 2b are separate devices. The execution device 230 is configured with an input / output (I / O) interface, and the execution device 230 can interact with the client device 250 through the I / O interface. For example, the client device 250 can send at least one initial feature information of the input content and the prompt information to the execution device 230 through the I / O interface. After the execution device 230 obtains the generated content corresponding to the input content through the first machine learning model 201, the execution device 230 can send the generated content to the client device 250 through the I / O interface.

[0083] It should be noted that FIGS. 2a and 2b are only schematic diagrams of the architecture of the information processing system provided by the embodiments of the present application, and the positional relationship between the devices, components, modules, and the like shown in the figures does not constitute any limitation. For example, in other embodiments of the present application, the training device 210 and the execution device 230 can also be integrated in the same device, and the specific architecture of the information processing system can be determined according to the actual application scenario.

[0084] In combination with the foregoing description, the application stage and the training stage of the first machine learning model 201 are described below, respectively.

[0085] I. Application stage

[0086] Specifically, refer to FIG. 3, which is a flowchart of an information processing method provided by an embodiment of the present application. The information processing method provided by the embodiment of the present application can include the following steps.

[0087] 301. Obtain at least one initial feature information of the input content. If the input content includes a first audio, the initial feature information of the first audio is discrete feature information. If the input content includes a first image, the initial feature information of the first image is continuous feature information.

[0088] For example, the initial feature information of the first audio is obtained by feature extraction of the first audio. In the present application, the discrete feature information can also be referred to as "discretized feature information", "discrete feature representation", "discretized feature representation", or "discrete representation", etc. The values in the first value space of the feature values in the discrete feature information are discrete. The first value space refers to a set of all values that the feature values can take. The number of values in the first value space is finite.

[0089] The audio is continuous data, and the initial feature information of the first audio is discrete feature information. This can be understood as follows: in the process of obtaining the initial feature information of the first audio, the continuous data (i.e., the first audio) is mapped into a plurality of discrete value intervals, and each discrete value interval includes a limited number of values. For example, the initial feature information of the first audio can include at least one first feature value, and any one of the at least one first feature value is included in a limited value set (i.e., a first value interval). The limited value set includes the values in the plurality of discrete value intervals. Since the initial feature information of the audio in the present application is discretized feature information, it can be understood that the initial feature information of any audio can be expressed by a limited number of values.

[0090] The initial feature information of the first image is obtained by feature extraction of the first image. In the present application, the continuous feature information can also be referred to as "continuous feature information", "continuous feature representation", "continuous feature representation", or "continuous representation", etc. The second value space of the feature values in the continuous feature information is continuous. The second value space refers to a set of all values that the feature values can take. The number of values in the second value space is not limited, in other words, the number of values in the second value space can be infinite.

[0091] The first image is continuous data, and the initial feature information of the first image is continuous feature information. It can be understood that the initial feature information of the first image is represented in a continuous form. For example, the initial feature information of the first image can include at least one second feature value. Any one of the at least one second feature value can take any real value in a certain value range, that is, the second value space is any real value in the value range. Alternatively, any one of the at least one second feature value can take any real value, that is, the second value space is any real value.

[0092] For example, the first value space of the feature value in the discrete feature information includes 0, 1, 2, 3, and 4, and the second value space of the feature value in the continuous feature information is any real number in 0-1. Through the foregoing comparison, it can be seen that the number of values that can be taken by the feature value in the discrete feature information is limited, while the number of values that can be taken by the feature value in the continuous feature information is infinite. It should be understood that the example is only for the convenience of understanding the difference between the discrete feature information and the continuous feature information, and is not used to limit the present application.

[0093] In the embodiments of the present application, the concepts of continuous feature information and discrete feature information are further clarified, thereby reducing the implementation difficulty of the present application and improving the realizability of the present application.

[0094] For example, if the input content includes a first text, the initial feature information of the first text is discrete feature information. For example, the initial feature information of the first text can include at least one token corresponding to the first text. For example, each token in the present application can be represented as a vector.

[0095] The initial feature information of the first audio is obtained by feature extraction of the first audio. If the input content includes a first video, the first video can be understood as including an image part and an audio part. The image part includes multiple frames of images (which can also be referred to as video frames) in the first video, and the audio part includes audio in the first video. The initial feature information of each frame of image in the multiple frames of images included in the first video can be obtained, and the initial feature information of all images in the multiple frames of images is fused to obtain the initial feature information of the image part in the first video. For example, the fusion operation includes splicing and compression. The initial feature information of the audio included in the first video is also obtained, so as to obtain the initial feature information of the first video.

[0096] Exemplarily, in an implementation, a first device can be deployed with a first feature extraction module corresponding to audio and a second feature extraction module corresponding to a first image; optionally, the first device can be further deployed with a conversion module corresponding to the second feature extraction module, for example, the conversion module can also be referred to as a mapping module or a projector. Optionally, the first device can be further deployed with a third feature extraction module corresponding to text;

[0097] The first feature extraction module is a discrete feature extraction module, and the first feature extraction module is configured to perform feature extraction on the audio. For example, the first feature extraction module can be a fully connected neural network, an attention mechanism-based neural network, a residual neural network, or a support vector machine, and the like, which will not be listed here. Optionally, the first feature extraction module can be a discrete audio encoder. Exemplarily, the feature information generated by the first feature extraction module (for example, initial feature information of the first audio) can include at least one token corresponding to the audio.

[0098] The second feature extraction module is a continuous feature extraction module, and the second feature extraction module is configured to perform feature extraction on the first image. For example, the second feature extraction module can be a convolutional neural network, a fully connected neural network, an attention mechanism-based neural network, a residual neural network, or a multi-layer perceptron (MLP), and the like, which will not be listed here. Optionally, the second feature extraction module can be a continuous vision encoder, and the function of the vision encoder is to encode the original first image data into initial feature information of the first image which is more compact and easy for the generative first machine learning model to understand, and the initial feature information of the first image retains rich information in the first image.

[0099] Exemplarily, in one case, the feature information of the first image generated by the second feature extraction module can include at least one feature corresponding to the first image, and each of the at least one feature can be expressed as a matrix. Optionally, the at least one feature generated by the second feature extraction module can be converted into token-form feature information by means of the conversion module. Exemplarily, the foregoing conversion operation can be understood as converting the matrix-form feature information into vector-form feature information, thereby obtaining the initial feature information of the first image, and the initial feature information of the first image can include a plurality of tokens corresponding to the first image. In another case, the feature information of the first image generated by the second feature extraction module can include at least one token corresponding to the first image, and then the second feature extraction module does not need to be connected with a conversion module.

[0100] The third feature extraction module is a discrete feature extraction module, and the third feature extraction module is configured to perform feature extraction on the text. For example, the third feature extraction module can be a recurrent neural network, a fully connected neural network, a neural network based on an attention mechanism, a residual neural network, or a support vector machine, and the like, which are not exhaustively listed herein. Optionally, the third feature extraction module can be a discrete text encoder. For example, the feature information (e.g., initial feature information of the first text) generated by the third feature extraction module can include at least one token corresponding to the text.

[0101] In step 301, after obtaining the input content, if the input content includes the first audio, the first device can input the first audio in the input content to the first feature extraction module, and perform discrete processing on the first audio by the first feature extraction module to obtain initial feature information of the first audio. For example, the discrete processing process can be understood as discretizing the continuous audio into at least one semantic unit (speech unit). Optionally, the discrete process can be based on the semantic content of the audio. In this way, the semantic information of the audio can be largely retained, and the generative first machine learning model can relatively easily understand the first audio based on the initial feature information of the first audio. For example, if the first audio is a 5-second audio, the semantic content of the first audio is “Please ask XXX subway station how to go”, the first audio can be discretized into 3 audio segments, i.e., the first audio is discretized into 3 semantic units, wherein the content in the first audio segment is “please”, the content in the second audio segment is “XXXX subway station”, and the content in the third audio segment is “how to go”. Then, the token corresponding to each audio segment can be obtained, and the initial feature information of the first audio is obtained. It should be understood that the above example is only for the convenience of understanding the present application, and is not used to limit the present application.

[0102] Alternatively, the above discrete processing process can also be based on time. For example, a segmentation operation is performed every preset time length, and the preset time length can be 1 second, 2 seconds, or other time lengths. For example, if the first audio is a 5-second audio and the preset time length is 1 second, the first audio can be segmented into 5 audio segments, and the token corresponding to each audio segment can be obtained, and the initial feature information of the first audio is obtained. The above discrete processing methods are not exhaustively listed herein.

[0103] If the input content includes the first image, the first device can input the first image in the input content into the second feature extraction module to obtain the feature information of the first image generated by the second feature extraction module; for example, if the feature information of the first image generated by the second feature extraction module contains at least one feature map, the first device can further input the feature information of the first image generated by the second feature extraction module into the conversion module to obtain the initial feature information of the first image generated by the conversion module. Alternatively, if the feature information of the first image generated by the second feature extraction module contains at least one token, the feature information of the first image generated by the second feature extraction module can also be used as the initial feature information of the first image.

[0104] If the input content includes the first text, the first text can be directly regarded as the initial feature information of the first text; or the first text in the input content can be input into the third feature extraction module to generate the initial feature information of the first text.

[0105] If the input content includes the first video, the initial feature information of the audio in the first video can be obtained through the first feature extraction module, the initial feature information of each image in the plurality of images (also referred to as video frames) included in the first video can be obtained through the second feature extraction module, and then the initial feature information of all images in the plurality of images is fused to obtain the initial feature information of the image part in the first video. The specific implementation mode of the foregoing steps can be referred to the description of the acquisition process of the initial feature information of the first audio and the initial feature information of the first image, which will not be described here. The initial feature information of the first video includes the initial feature information of the audio in the first video and the initial feature information of the image part in the first video.

[0106] In another implementation mode, the fourth feature extraction module with adaptive capability can be deployed in the first device. The fourth feature extraction module with adaptive capability can adaptively generate corresponding feature information according to the data type of the input data. The fourth feature extraction module can also be referred to as a hybrid encoder. Step 301 can include: after obtaining the input content, if the input content includes the first audio, the first audio can be input into the fourth feature extraction module, and the first audio is discretized by the fourth feature extraction module to obtain the initial feature information of the first audio. The initial feature information of the first audio is discrete feature information.

[0107] If the input content includes the first image, the first image can be input into the fourth feature extraction module to obtain the initial feature information of the first image generated by the fourth feature extraction module. The initial feature information of the first image is continuous feature information.

[0108] If the input content includes the first text, the first text can be directly regarded as initial feature information of the first text; or the first text can be input into the fourth feature extraction module to obtain initial feature information of the first text generated by the fourth feature extraction module, and the initial feature information of the first text is discrete feature information.

[0109] If the input content includes the first video, a plurality of images included in the first video can be input into the fourth feature extraction module to obtain initial feature information of each image in the first video generated by the fourth feature extraction module, the initial feature information of each image in the first video is continuous feature information, and then the initial feature information of all images in the plurality of images is fused to obtain initial feature information of the image part in the first video; the audio included in the first video is input into the fourth feature extraction module to obtain initial feature information of the audio in the first video generated by the fourth feature extraction module, and the initial feature information of the audio in the first video is discrete feature information, so that the initial feature information of the first video is obtained.

[0110] 302、input at least one initial feature information into the first machine learning model to obtain generated content corresponding to the input content, wherein the generated content is generated by the first machine learning model after performing a task, and in the case that the task includes generating audio, the generated content includes second audio corresponding to the input content.

[0111] Exemplarily, before performing step 302, the first device can also obtain prompt information, and the prompt information can indicate a task that the generative first machine learning model needs to perform, for example, the task can include at least one of the following: content description of the input image, speech recognition of the input audio, generation of text corresponding to the input text, generation of audio corresponding to the input content, generation of images corresponding to the input content, or other tasks, etc., which are not exhaustive here.

[0112] Optionally, the task can also include at least one of the following: identifying the emotional category of the input content, determining the speech rate of the generated audio, determining the pitch of the generated audio, determining the emotional style of the generated audio, or other tasks, etc., which can be determined in combination with actual application scenarios.

[0113] Optionally, the prompt information can further indicate at least one of the following: at least one emotion category, at least one speech speed, at least one pitch, or at least one emotional style, etc. The at least one emotion category can assist the generative first machine learning model in identifying the emotion category of the input content, for example, the at least one emotion category can include at least one of the following: neutral, happy, sad, angry, surprised, disgusted, fearful, anxious, or other emotion categories, which are not exhaustively listed here.

[0114] The at least one speech speed can assist the generative first machine learning model in determining the speech speed of the audio in the generated content, for example, the at least one speech speed can include at least one of the following: slow, normal, fast, or other speech speeds, etc. The at least one pitch can assist the generative first machine learning model in determining the pitch of the audio in the generated content, for example, the at least one pitch can include at least one of the following: low, normal, high, or other pitches, etc. The at least one emotional style can assist the generative first machine learning model in determining the emotional style of the audio in the generated content, for example, the at least one emotional style can include at least one of the following: gentle, happy, neutral, surprised, fearful, or other emotional styles, etc. It should be noted that the examples provided herein are only for the convenience of understanding the present solution, and the specific implementation can be determined in combination with the actual application scenario.

[0115] Optionally, the prompt information can further indicate the data format of the generated content output by the first machine learning model, for example, the prompt information indicates that the generated content needs to adopt the format of a json file; optionally, the prompt information can further include images, etc. The specific information included in the prompt information can be determined in combination with the specific application scenario.

[0116] Exemplarily, the first device inputs at least one initial feature information of the input content and the prompt information into the first machine learning model to obtain the generated content corresponding to the input content. Further, in one case, if the first device is the execution device in the schematic diagram shown in FIG. 2a, or the first device is the execution device in the schematic diagram shown in FIG. 2b, step 302 can include that the first device can input at least one initial feature information of the input content (optionally, the prompt information is also included) into the first machine learning model, and generate the generated content corresponding to the input content through the generative first machine learning model. In another case, if the first device is the client device in the schematic diagram shown in FIG. 2b, step 302 can include that the first device sends at least one initial feature information of the input content (optionally, the prompt information is also included) to the execution device of the first machine learning model, and receives the generated content corresponding to the input content sent by the execution device of the first machine learning model; the generated content corresponding to the input content is generated by the execution device through the first machine learning model.

[0117] Exemplarily, in the case that the task to be performed by the first machine learning model includes generating a second text corresponding to the input content, the generated content includes the second text corresponding to the input content. In the case that the task to be performed by the first machine learning model includes generating a second image corresponding to the input content, the first machine learning model can only generate feature information of the second image corresponding to the input content, and the feature information of the second image is discrete feature information, and the second image corresponding to the input content can be obtained based on the feature information of the second image by calling other modules other than the first machine learning model.

[0118] Since the first machine learning model can only generate discrete feature information, and the initial feature information of the first audio in the present application is discrete feature information, in the case that the task to be performed by the first machine learning model includes generating a second audio corresponding to the input content, the generated content includes the second audio corresponding to the input content. Exemplarily, the first machine learning model can include a feature processing module corresponding to the first feature extraction module, and the feature processing module is used to obtain the second audio based on the first feature information generated by the first machine learning model, wherein the first feature information can be discrete feature information, for example, the feature processing module can also be understood as an audio decoder, and the feature processing module can include a recurrent neural network layer, a fully connected neural network layer, a residual neural network layer, a neural network layer based on an attention mechanism, an MLP or other types of neural network layers, etc., which can be rated in combination with actual application scenarios.

[0119] Exemplarily, the first feature information can indicate semantic content of the second audio; optionally, the first feature information further indicates a pitch of the second audio, for example, if the task required to be performed by the first machine learning model includes determining a pitch of the generated audio, the first feature information can further indicate the pitch of the second audio; optionally, the second feature information further indicates a speech rate of the second audio, for example, if the task required to be performed by the first machine learning model includes determining a speech rate of the generated audio, the first feature information can further indicate the speech rate of the second audio.

[0120] For a more intuitive understanding of the scheme, please refer to FIG. 4, which is a schematic diagram of obtaining generated content through a first machine learning model provided by an embodiment of the present application. In FIG. 4, taking an example in which the input content includes a first image, a first text and a first audio, and the generated content includes a second text and a second audio, as shown in the figure, the first image is input into a continuous visual encoder to obtain feature information of the first image generated by the visual encoder, the feature information of the first image is input into a connector to obtain initial feature information of the first image generated by the connector, and the initial feature information of the first image is continuous feature information. The first audio is input into a discrete audio encoder to obtain initial feature information of the first audio generated by the audio encoder, and the initial feature information of the first audio is discrete feature information.

[0121] The initial feature information of the first image, the first text (which can also be referred to as initial feature information of the first text) and the initial feature information of the first audio are input into a generative first machine learning model, and are processed through multiple neural network layers in the first machine learning model to generate first feature information and feature information corresponding to the second text, both of which are discrete feature information; the second audio is generated by an audio decoder in the first machine learning model based on the first feature information; and the second text is obtained based on the aforementioned feature information corresponding to the second text. It should be understood that the example in FIG. 4 is only for the convenience of understanding the scheme and is not used to limit the scheme.

[0122] Optionally, the second audio is generated by the first machine learning model based on the first feature information and the second feature information, and optionally, the second audio is generated by an audio decoder in the first machine learning model based on the first feature information and the second feature information. The second feature information indicates an emotional style of the second audio, for example, the second feature information can be discrete feature information, for example, the second feature information can include at least one vector. Optionally, the size of the second feature information is the same as the size of the first feature information, for example, the number of vectors included in the second feature information and the first feature information is the same, and the number of vector values in each vector is also the same.

[0123] In the embodiments of the present application, emotions are injected into the second audio, which is beneficial to improve the naturalness of the output second audio and improve the emotional expression effect of the second audio, so as to improve the user experience of the present scheme; and the second feature information specially used to indicate the emotional style of the second audio is additionally introduced, and the first machine learning model can generate the second audio in combination with the semantic content of the second audio and the emotional style of the second audio, which decouples the semantic content and the emotional style of the second audio, is beneficial to reduce the difficulty of the first machine learning model in generating the second audio, and is beneficial to obtain a more natural and humanized second audio.

[0124] Further, in one case, the second feature information is obtained by a second machine learning model based on an emotion category corresponding to the input content, and the emotion category is generated by the first machine learning model; the prompt information can indicate that the task to be performed by the generated first machine learning model includes identifying the emotion category of the input content, so that the first machine learning model can generate the emotion category corresponding to the input content. For example, the second machine learning model can be a fully connected neural network, a convolutional neural network, a residual neural network, a recurrent neural network, an MLP, a support vector machine or other types of machine learning models, etc., which can be determined in combination with actual application scenarios.

[0125] Exemplarily, the execution device deploying the first machine learning model inputs at least one initial feature information of the input content and the prompt information into the first machine learning model, and generates the first feature information and the emotion category of the input content through the first machine learning model; the execution device inputs the emotion category of the input content into the second machine learning model to obtain the second feature information generated by the second machine learning model; the execution device can generate the second audio through the feature processing module in the first machine learning model based on the first feature information and the second feature information; it should be noted that if the first device is an execution device deploying the first machine learning model, the foregoing steps can also be understood as being executed by the first device, for example, if the first device is the execution device in FIG. 2a or FIG. 2b, the foregoing steps can also be understood as being executed by the first device.

[0126] In order to more intuitively understand the present scheme, please refer to FIG. 5, which is another schematic diagram of obtaining generated content through the first machine learning model provided by the embodiments of the present application. In FIG. 5, the input content includes a first image, a first text and a first audio, and the generated content includes a second audio as an example. The initial feature information of the first image, the first text and the initial feature information of the first audio are obtained, and the foregoing steps can be understood in combination with the description of FIG. 4 above, which will not be repeated here.

[0127] The initial feature information of the first image, the first text, and the initial feature information of the first audio are input into the generative first machine learning model, and the first feature information and the emotion category of the input content can be generated by processing through multiple neural network layers in the first machine learning model. The emotion category of the input content is input into the second machine learning model to obtain the second feature information generated by the second machine learning model. The second audio is generated by the audio decoder in the first machine learning model based on the first feature information and the second feature information. It should be understood that the examples in FIG. 5 are only for the convenience of understanding the present scheme and do not limit the present scheme.

[0128] In the embodiments of the present application, the emotion category of the input content is first generated by the first machine learning model, and then the second feature information for representing the emotional style of the second audio is generated by the independent second machine learning model. This is beneficial to generating second feature information of better quality by the independent second machine learning model, and is beneficial to making the emotional style of the second audio more suitable for the emotion category of the input content, so as to further improve the user experience of the present scheme. For example, when the input content expresses an anxious emotion, the second audio can be of a gentle emotional style, and the intelligent assistant can respond in a gentle tone, thereby achieving a calming effect. For another example, when the input content expresses a neutral emotion, the second audio can be of a neutral emotional style, and the intelligent assistant can respond in a neutral tone. For another example, the emotional style of the second audio can be determined according to the emotion type expressed by the student through the input content, and then the educational assistant will explain the teaching content in the aforementioned emotional style. For another example, for long-term care patients, the medical consultation system can adopt a gentle emotional style, and the like, to provide more personalized second audio, so as to further improve the user experience of the present scheme.

[0129] In another case, the second feature information is generated by the first machine learning model. For example, the prompt information can indicate that the task to be performed by the generative first machine learning model includes determining the emotional style of the generated audio. Illustratively, after the execution device deploying the first machine learning model inputs at least one initial feature information of the input content and the prompt information into the first machine learning model, the first feature information and the second feature information are generated by the first machine learning model, and then the second audio can be generated based on the first feature information and the second feature information by the feature processing module in the first machine learning model. It should be noted that if the first device is an execution device deploying the first machine learning model, the foregoing steps can also be understood as being performed by the first device. For example, if the first device is the execution device in FIG. 2a or FIG. 2b, the foregoing steps can also be understood as being performed by the first device.

[0130] Since the current generative machine learning model can only generate discrete feature information, and the initial feature information of the audio often adopts continuous feature information, the machine learning model can only adopt the way of continuous feature information to understand the audio, resulting in the inability to restore the audio based on the generated discrete feature information. The initial feature information of the audio obtained in the present application is discrete feature information, so the first machine learning model can adopt discrete feature information to understand the audio, so that the first machine learning model has the ability to directly generate audio. In addition, the discrete feature information of the audio is combined with the continuous feature information of the image. The continuous feature information of the image can retain rich information in the image. Therefore, not only the end-to-end audio generation capability is realized, but also the image in the input content can be well understood, which is conducive to improving the quality of the generated content.

[0131] II. Training phase

[0132] Specifically, please refer to FIG. 6, which is a flowchart of a model training method provided by an embodiment of the present application. The model training method provided by the embodiment of the present application can include:

[0133] 601. Obtain at least one initial feature information of a first training sample. If the first training sample includes a first audio, the initial feature information of the first audio is discrete feature information. If the first training sample includes a first image, the initial feature information of the first image is continuous feature information.

[0134] 602. In the first training phase of the first machine learning model, input at least one initial feature information of the first training sample into the first machine learning model, and generate a first predicted generated content through the first machine learning model.

[0135] The first training stage of the first machine learning model can also be referred to as a pre-training stage of the first machine learning model. In the first training stage of the first machine learning model, the task to be performed by the first machine learning model includes generating a text description corresponding to an input first training sample, and the data type of the first training sample includes an image, audio, or text. For a better understanding of the solution, refer to FIG. 7, which is a schematic diagram of generating a text description corresponding to an input first training sample according to an embodiment of the present application. FIG. 7 includes three sub-diagrams on the left, in the middle, and on the right. The left sub-diagram of FIG. 7 shows an image and a text description corresponding to the image. The middle sub-diagram of FIG. 7 shows audio and a text description corresponding to the audio. The right sub-diagram of FIG. 7 shows a question in text form and an answer in text form corresponding to the question (i.e., a text description corresponding to text). In FIG. 7, the generated text description is in Chinese. The first machine learning model can also generate a text description in other languages, such as English, German, Spanish, etc. The specific implementation can be determined according to the actual application scenario. The examples in FIG. 7 are for the convenience of understanding the solution and do not limit the solution.

[0136] The specific implementation of the training device performing steps 601 and 602 can refer to the description in the corresponding embodiment of FIG. 3. The difference is that the “input content” in the corresponding embodiment of FIG. 3 is replaced by the “first training sample”. The description is not repeated here.

[0137] 603, training the first machine learning model based on the first predicted generated content.

[0138] Optionally, in the first training stage of the first machine learning model, supervised learning can be used. For example, after obtaining the first predicted generated content, the training device can determine the function value of the first loss function based on the first expected generated content corresponding to the first training sample and the first predicted generated content. Based on the function value of the first loss function, the weight parameters of the first machine learning model are updated using a backpropagation algorithm. Optionally, the weight parameters of the feature extraction module used when generating the at least one initial feature information of the first training sample are also updated. Optionally, the weight parameters of the second machine learning model are also updated to implement one round of training of the first machine learning model.

[0139] For example, the feature extraction module used when generating the at least one initial feature information of the first training sample can include a first feature extraction module and a second feature extraction module. Optionally, it also includes a conversion module. Optionally, it also includes a third feature extraction module. Alternatively, the feature extraction module used when generating the at least one initial feature information of the first training sample can include a fourth feature extraction module.

[0140] Exemplarily, the first loss function indicates a similarity between the first expected generated content corresponding to the first training sample and the first predicted generated content, and a target of training by using the first loss function includes improving the similarity between the first expected generated content and the first predicted generated content; the first expected generated content corresponding to the first training sample can also be referred to as a true value corresponding to the first training sample or a first correct generated content corresponding to the first training sample, and the like. Exemplarily, the first expected generated content and the first predicted generated content can both be texts.

[0141] The training device repeatedly performs steps 601 to 603 multiple times until a first convergence condition is met, thereby completing the first training phase of the first machine learning model, and the first convergence condition includes that a convergence condition of the first loss function is met and / or the number of times of performing steps 601 to 603 reaches a first preset number. Optionally, in the first training phase of the first machine learning model, the first training samples of different data types are cross-used; in other words, the first training phase of the first machine learning model can include multiple rounds of training of the first machine learning model, and different data types of the first training samples can be used in adjacent rounds of training.

[0142] Exemplarily, the round 1, the round 2, the round 3, the round 4 and the round 5 are five continuous rounds of training of the first machine learning model, the data type of the first training sample used in the round 1 is an image, the data type of the first training sample used in the round 2 is a text, the data type of the first training sample used in the round 3 is an audio, the data type of the first training sample used in the round 4 is a text, and the data type of the first training sample used in the round 4 is an image. It should be understood that the examples herein are only for the convenience of understanding the present scheme and are not used to limit the present scheme.

[0143] In the first training phase of the first machine learning model, different data types of training samples are cross-used to train the first machine learning model, so that the first machine learning model cross-understands different data types of training samples, improves the confusion degree of the first training phase, is conducive to increasing the difficulty of the first training phase of the first machine learning model, thereby being conducive to improving the understanding ability of the trained first machine learning model to various data types of data, and on the premise of fully understanding the input content, being conducive to making the first machine learning model generate better generated content.

[0144] Or, in the first training stage of the first machine learning model, the first training samples of the same data type can be used continuously. For example, the first training stage of the first machine learning model includes 100 rounds of training, i.e., round 1 to round 100. In round 1 to round 34, the first training samples used are all of the text type, in round 35 to round 67, the first training samples used are all of the image type, and in round 68 to round 100, the first training samples used are all of the audio type. It should be understood that the examples herein are only for the convenience of understanding the present scheme and do not limit the present scheme.

[0145] 604、obtaining at least one initial feature information of the second training sample, wherein if the second training sample includes the first audio, the initial feature information of the first audio is discrete feature information, and if the second training sample includes the first image, the initial feature information of the first image is continuous feature information.

[0146] 605、in the second training stage of the first machine learning model, inputting at least one initial feature information of the second training sample into the first machine learning model, and generating the second predicted generated content through the first machine learning model.

[0147] Wherein, the second training stage of the first machine learning model can also be referred to as the post-training stage of the first machine learning model. Optionally, in the second training stage of the first machine learning model, in each inference process of the first machine learning model, the training task performed by the first machine learning model includes generating at least two of the image, the audio and the text, i.e., each training task performed by the first machine learning model is to generate multi-modal data. Optionally, the training task performed by the first machine learning model includes generating the image, the audio and the text, i.e., each training task performed by the first machine learning model is to generate full-modal data.

[0148] To understand the scheme more intuitively, refer to FIG. 8, which is a schematic diagram of input content and prompt information provided by an embodiment of the present application. As shown in FIG. 8, the input content is a first audio, and the prompt information indicates that the first machine learning model needs to perform tasks including: recognizing text in the first audio, generating a second audio corresponding to the recognized text, recognizing an emotion category of the first audio, determining a speech rate of the second audio, and determining a pitch of the second audio; the prompt information further indicates that the emotion category of the first audio can be selected from neutral, happy, sad, angry, surprised, disgusted, and fearful, the pitch of the second audio can be selected from low, normal, and high, and the speech rate of the second audio can be selected from slow, normal, and fast. The prompt information further indicates that the text in the first audio and the second audio are output in a file format of json. It should be understood that the example in FIG. 8 is only for the convenience of understanding the scheme and does not limit the scheme.

[0149] In the second training phase of the first machine learning model, the training task of each inference process of the first machine learning model includes generating at least two of images, audio, and text, i.e., generating data of at least two data types, which increases the difficulty of the second training phase, is conducive to stimulating the generation ability of the first machine learning model for data of multiple data types, and is conducive to improving the quality of newly generated images, audio, and text obtained by the trained first machine learning model.

[0150] Alternatively, in the second training phase of the first machine learning model, in each inference process of the first machine learning model, the training task performed by the first machine learning model only includes generating single-modal data, i.e., each training task performed by the first machine learning model is only generating images, audio, or text. Alternatively, in the second training phase of the first machine learning model, in each inference process of the first machine learning model, the training task performed by the first machine learning model only includes generating double-modal data, i.e., each training task performed by the first machine learning model is generating two of images, audio, or text.

[0151] The specific implementation of the training device performing steps 604 and 605 can refer to the description in the corresponding embodiment of FIG. 3 described above, with the difference being that the “input content” in the corresponding embodiment of FIG. 3 is replaced by “second training sample”. No repeated description is given here.

[0152] 606, training the first machine learning model based on the generated content of the second prediction.

[0153] Optionally, in the second training stage of the first machine learning model, supervised learning can be adopted. After the training device obtains the second predicted generated content, the function value of the second loss function can be determined based on the second expected generated content corresponding to the first training sample and the second predicted generated content. Then, the weight parameters of the first machine learning model are updated based on the function value of the first loss function by using the back propagation algorithm. Optionally, the weight parameters of the feature extraction module used when generating the at least one initial feature information of the first training sample are also updated. Optionally, the weight parameters of the second machine learning model are also updated to realize one training of the first machine learning model.

[0154] The specific implementation of the above steps can refer to the description of step 603, and the difference is that the "first expected generated content" is replaced by the "second expected generated content", the "first predicted generated content" is replaced by the "second predicted generated content", and the "first loss function" is replaced by the "second loss function". Here, the description is not repeated.

[0155] In order to have a more intuitive understanding of the beneficial effects brought by the method provided in the present application, experiments were carried out on the MME, OCRbench and Librispeech data sets respectively, wherein MME is an image data set, OCRbench is a text data set, and Librispeech is an audio data set. The task performed by the machine learning model when testing is to generate a text description corresponding to the input content. The experimental results are shown in Tables 1 and 2.

[0156] Table 1

[0157] Among them, VITA represents using an existing generative machine learning model to generate a text description corresponding to the input content, and the present application represents using the method provided in the present application to obtain a text description corresponding to the input content. The higher the index value obtained when the experiment is performed on MME and OCRbench, the better. As can be seen from Table 1 above, the understanding ability of the method provided in the present application for images and texts is stronger.

[0158] Table 2

[0159] Among them, the lower the index value obtained when the experiment is performed on Librispeech, the better. As can be seen from Table 2 above, the understanding ability of the method provided in the present application for audio is stronger.

[0160] On the basis of the embodiments corresponding to FIG. 1 to FIG. 8, in order to better implement the above-mentioned scheme of the embodiments of the present application, the following also provides a related device for implementing the above-mentioned scheme. Referring to FIG. 9, FIG. 9 is a structure schematic diagram of an information processing apparatus provided by the embodiments of the present application, the information processing apparatus 900 comprises: an acquisition module 901, configured to acquire at least one initial feature information of input content, wherein, if the input content comprises a first audio, the initial feature information of the first audio is discrete feature information, and if the input content comprises an image, the initial feature information of the image is continuous feature information; a processing module 902, configured to input the at least one initial feature information into a first machine learning model to obtain generated content corresponding to the input content, wherein, the generated content is generated after the first machine learning model performs a task, and in the case that the task comprises generating an audio, the generated content comprises a second audio corresponding to the input content.

[0161] Optionally, the values in the first value space of the feature values in the discrete feature information are discrete, and the number of values in the first value space is limited, and the second value space of the feature values in the continuous feature information is continuous.

[0162] Optionally, the second audio is generated by the first machine learning model based on first feature information and second feature information, the first feature information indicates semantic content of the second audio, and the second feature information indicates emotional style of the second audio.

[0163] Optionally, the second feature information is obtained by a second machine learning model based on an emotion category corresponding to the input content, and the emotion category is generated by the first machine learning model.

[0164] Optionally, in a first training stage of the first machine learning model, a training task of the first machine learning model comprises generating a text description corresponding to an input training sample, a data type of the training sample comprises an image, an audio or a text, and training samples of different data types are cross-used.

[0165] Optionally, in a second training stage of the first machine learning model, in each inference process of the first machine learning model, a training task performed by the first machine learning model comprises generating at least two of an image, an audio and a text.

[0166] It should be noted that the information interaction and execution process between the modules / units in the information processing apparatus 900 are based on the same concept as the method embodiments corresponding to FIG. 1 to FIG. 8 of the present application, and the specific content can be referred to the description in the method embodiments of the present application.

[0167] The embodiment of the present application further provides a device, as shown in Figure 10, which is a structural schematic diagram of the device provided by the embodiment of the present application. Optionally, the device 1000 executes the functions of the first device in each method embodiment corresponding to Figures 1 to 8 and / or executes the functions of the device.

[0168] The device 1000 comprises a memory 1002 and at least one processor 1001. Optionally, the processor 1001 realizes the method in the above embodiment by reading the instructions saved in the memory 1002, or the processor 1001 can also realize the method in the above embodiment by reading the instructions saved in the memory 1002. In the case that the processor 1001 realizes the method in the above embodiment by reading the instructions saved in the memory 1002, the memory 1002 saves the instructions for realizing the method provided by the above embodiment of the present application.

[0169] Optionally, the at least one processor 1001 is one or more CPUs, or a single core CPU, or a multi-core CPU. The memory 1002 comprises, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a flash memory, or an optical memory, etc. The memory 1002 saves the instructions of an operating system. After the program instructions stored in the memory 1002 are read by the at least one processor 1001, the device 1000 performs the corresponding operations in the above embodiment.

[0170] Optionally, the device 1000 further comprises a network interface 1003, which can be a wired interface or a wireless interface. The network interface 1003 is configured to receive and send data in each method embodiment corresponding to Figures 1 to 8.

[0171] It should be understood that the network interface 1003 has the functions of receiving data and sending data. The function of receiving data and the function of sending data can be integrated in the same transceiver interface, or the function of receiving data and the function of sending data can be realized in different interfaces respectively, which is not limited here. In other words, the network interface 1003 can comprise one or more interfaces for realizing the function of receiving data and the function of sending data.

[0172] After the processor 1001 reads the program instructions in the memory 1002, the device 1000 can perform other functions, which can be referred to the description in each method embodiment.

[0173] Optionally, the device 1000 further comprises a bus 1004. The above processor 1001 and memory 1002 are usually connected with each other through the bus 1004, or can be connected with each other in other manners.

[0174] The device 1000 provided in the embodiments of the present application is configured to execute the methods performed by the execution device, the second device or the training device in the various method embodiments described above, and achieve the corresponding beneficial effects. The specific implementation modes of the device 1000 shown in FIG. 10 can be referred to the descriptions in the various method embodiments described above, and will not be described here in detail.

[0175] In the embodiments of the present application, a computer readable storage medium is also provided, which stores a program, and when the program is executed on a computer, the computer is caused to execute the steps performed by the first device and / or the execution device in the methods described in the embodiments of FIG. 1 to FIG. 8.

[0176] In the embodiments of the present application, a computer program product is also provided, which includes a program, and when the program is executed on a computer, the computer is caused to execute the steps performed by the first device and / or the execution device in the methods described in the embodiments of FIG. 1 to FIG. 8.

[0177] In the embodiments of the present application, a circuit system is also provided, which includes a processing circuit configured to execute the steps performed by the first device and / or the execution device in the methods described in the embodiments of FIG. 1 to FIG. 8.

[0178] The first device and the execution device provided in the embodiments of the present application can be a chip, which includes a processing unit, for example, a processor. Optionally, the chip also includes a communication unit, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute computer execution instructions stored in a storage unit, so that the chip executes the methods described in the embodiments of FIG. 1 to FIG. 8. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0179] Exemplarily, refer to FIG. 11, which is a structural schematic diagram of a chip provided in the embodiments of the present application. The chip can be a neural network processor NPU 110, which is mounted on a host CPU (Host CPU) as a coprocessor and is assigned tasks by the Host CPU. The core part of the NPU is an operation circuit 1103, which extracts matrix data in a memory and performs multiplication operation under the control of a controller 1104.

[0180] In some implementations, the arithmetic circuit 1103 includes a plurality of processing units (PEs) inside. In some implementations, the arithmetic circuit 1103 is a two-dimensional systolic array. The arithmetic circuit 1103 can also be a one-dimensional systolic array or other electronic circuit capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1103 is a general-purpose matrix processor.

[0181] For example, assume that there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit takes the data of the matrix B corresponding from the weight memory 1102 and caches it on each PE of the arithmetic circuit. The arithmetic circuit takes the data of the matrix A from the input memory 1101 and performs matrix operation with the matrix B to obtain a partial result or a final result of the matrix, which is saved in the accumulator 1108.

[0182] The unified memory 1106 is used to store input data and output data. The weight data is transferred to the weight memory 1102 through the direct memory access controller (DMAC) 1105. The input data is also transferred to the unified memory 1106 through the DMAC.

[0183] The bus interface unit (BIU) 1110 is used for the interaction between the AXI bus and the DMAC and the instruction fetch buffer (IFB) 1109.

[0184] The bus interface unit (BIU) 1110 is used for the instruction fetch buffer 1109 to obtain instructions from the external memory, and is also used for the direct memory access controller 1105 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0185] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1106, or to transfer the weight data to the weight memory 1102, or to transfer the input data to the input memory 1101.

[0186] The vector calculation unit 1107 includes a plurality of arithmetic processing units, which further process the output of the arithmetic circuit as needed, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / full connection layer network calculation in neural networks, such as batch normalization, pixel-level summation, upsampling of feature planes, etc.

[0187] In some implementations, the vector computation unit 1107 can store the processed output vector to the unified memory 1106. For example, the vector computation unit 1107 can apply a linear function and / or a non-linear function to the output of the arithmetic circuit 1103, such as linear interpolation on the feature planes extracted by a convolution layer, and / or accumulate the vector of values to generate activation values. In some implementations, the vector computation unit 1107 generates normalized values, pixel-wise summed values, or both. In some implementations, the processed output vector can be used as activation input to the arithmetic circuit 1103, such as for use in a subsequent layer in a neural network.

[0188] The controller 1104 is connected to an instruction fetch buffer 1109 for storing instructions used by the controller 1104;

[0189] The unified memory 1106, the input memory 1101, the weight memory 1102, and the instruction fetch buffer 1109 are on-chip memories. Off-chip memories are private to the NPU hardware architecture.

[0190] In the above description, the operations of the layers of the first machine learning model and / or the second machine learning model can be performed by the arithmetic circuit 1103 or the vector computation unit 1107.

[0191] In the above description, the processor can be a general central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the method of the first aspect.

[0192] It should be noted that the apparatus embodiments described above are merely exemplary, and the units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the connection relationship between the modules in the apparatus embodiment provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.

[0193] Those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware, and of course can also be implemented by means of dedicated hardware including special integrated circuit, special CLU, special memory, special component, etc. Generally, any function completed by computer program can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the application, the software program implementation is a better embodiment. Based on such understanding, the technical solution of the application or the part of the application which makes contribution to the prior art can be embodied in the form of software product, which is stored in readable storage medium, such as floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a plurality of instructions for making a computer device (which can be personal computer, server or network device, etc.) execute the method described in various embodiments of the application.

[0194] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof, entirely or partially. When implemented by software, the implementation can be in the form of computer program product entirely or partially.

[0195] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the flow or function described in the embodiments of the application is entirely or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as server, data center, etc. integrated with one or more available media. The available medium can be magnetic medium (such as floppy disk, hard disk, magnetic tape), optical medium (such as DVD) or semiconductor medium (such as solid state disk (SSD)) etc.

Claims

1. An information processing method characterized by comprising: The method comprises: obtaining at least one initial feature information of input content, wherein if the input content comprises a first audio, the initial feature information of the first audio is discrete feature information, and if the input content comprises an image, the initial feature information of the image is continuous feature information; inputting the at least one initial feature information into a first machine learning model to obtain generated content corresponding to the input content, wherein the generated content is generated after the first machine learning model performs a task, and in the case that the task comprises generating an audio, the generated content comprises a second audio corresponding to the input content.

2. The method of claim 1, wherein, In the discrete feature information, the values in a first value space of feature values are discrete, and the number of values in the first value space is finite. In the continuous feature information, a second value space of feature values is continuous.

3. The method according to claim 1 or 2, characterized in that, The second audio is generated by the first machine learning model based on first feature information and second feature information, the first feature information indicates semantic content of the second audio, and the second feature information indicates emotional style of the second audio.

4. The method of claim 3, wherein, The second feature information is obtained by a second machine learning model based on an emotion category corresponding to the input content, and the emotion category is generated by the first machine learning model.

5. The method according to claim 1 or 2, characterized in that, In a first training phase of the first machine learning model, the training task of the first machine learning model comprises generating a text description corresponding to an input training sample, the data type of the training sample comprises an image, an audio or a text, and different data types of training samples are cross-used.

6. The method of claim 1 or 2, wherein, In a second training phase of the first machine learning model, in each inference process of the first machine learning model, the training task performed by the first machine learning model comprises generating at least two of an image, an audio and a text.

7. An information processing apparatus, characterized by comprising: The device comprises: an acquisition module configured to obtain at least one initial feature information of input content, wherein if the input content comprises a first audio, the initial feature information of the first audio is discrete feature information, and if the input content comprises an image, the initial feature information of the image is continuous feature information; a processing module configured to input the at least one initial feature information into a first machine learning model to obtain generated content corresponding to the input content, wherein the generated content is generated after the first machine learning model performs a task, and in the case that the task comprises generating an audio, the generated content comprises a second audio corresponding to the input content.

8. The apparatus of claim 7, wherein, In the discrete feature information, the values in a first value space of feature values are discrete, and the number of values in the first value space is finite. In the continuous feature information, a second value space of feature values is continuous.

9. The apparatus of claim 7 or 8, wherein, The second audio is generated by the first machine learning model based on first feature information and second feature information, the first feature information indicates semantic content of the second audio, and the second feature information indicates emotional style of the second audio.

10. The apparatus of claim 9, wherein, The second feature information is obtained by a second machine learning model based on an emotion category corresponding to the input content, and the emotion category is generated by the first machine learning model.

11. The apparatus of claim 7 or 8, wherein, In the first training stage of the first machine learning model, the training task of the first machine learning model includes generating a text description corresponding to an input training sample, and the data type of the training sample includes an image, audio or text, and different data types of training samples are cross-used.

12. The apparatus of claim 7 or 8, wherein, In the second training stage of the first machine learning model, in each inference process of the first machine learning model, the training task performed by the first machine learning model includes generating at least two of an image, audio and text.

13. An apparatus, comprising: The processor is coupled to the memory, and the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method in any one of claims 1 to 6 is implemented.

14. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a program, and when the program runs on the computer, the computer executes the method in any one of claims 1 to 6.

15. A computer program product, characterised in that, The computer program product includes a program, and when the program runs on the computer, the computer executes the method in any one of claims 1 to 6.

16. A chip, characterized by The chip includes a processor configured to perform the steps in the method in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Semantic recognition method and system

    CN114550741A

  • Audio understanding and generating method based on large-scale audio representation language model

    CN116741153A

  • Emotional speech synthesis method and device based on AI large model

    CN117174073A

  • Speech processing method, pre-training language model training method and speech recognition method

    CN117975943A

  • Emotion recognition interaction method and system based on large model and emotion wheel disc

    CN118053453A