Information interaction method and device, equipment, medium and product

By displaying and controlling voice files in the information interaction interface, the issues of flexibility and reliability in user interaction with large models are resolved, achieving a more realistic and flexible information interaction experience.

CN122431570APending Publication Date: 2026-07-21TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-01-20
Publication Date
2026-07-21

Smart Images

  • Figure CN122431570A_ABST
    Figure CN122431570A_ABST
Patent Text Reader

Abstract

The application provides an information interaction method, device, equipment, medium and product. The method comprises: displaying an information interaction interface, the information interaction interface comprising an information input area; in response to an information interaction operation in the information interaction interface, displaying a first voice file in the information input area; in response to an output operation on the first voice file displayed in the information input area, displaying a second voice file in the information interaction interface, the second voice file being composed of the first voice file and reply data for replying to the voice data. In the application, the voice data input by the user is displayed as a voice file in the information input area, ensuring the data reliability and flexibility in the information interaction process, thereby improving the information interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to an information interaction method, an information interaction device, a computer equipment, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the rapid development of artificial intelligence technology and the increasing sophistication of large language models, the information interaction function between users and large models has permeated all aspects of daily life.

[0003] Currently, the dialogue and interaction between users and large models often involves users inputting text information into a dialog box (i.e., the information interaction interface), and the model's response is usually displayed in text form on the information interaction interface. This information interaction method is relatively simple and not flexible enough. Summary of the Invention

[0004] This application provides an information interaction method, apparatus, device, medium, and product that supports displaying user-inputted voice data as a voice file in an information input area, ensuring data reliability and flexibility during the information interaction process, thereby improving the information interaction experience.

[0005] On the one hand, this application provides an information exchange method, which includes:

[0006] The information interaction interface is displayed and used to interact with the interaction object. The information interaction interface includes an information input area.

[0007] In response to an information interaction operation in the information interaction interface, a first voice file is displayed in the information input area; the first voice file is generated based on the voice data input by the first object;

[0008] In response to the output operation of the first voice file displayed in the information input area, a second voice file is displayed in the information interaction interface; the second voice file consists of the first voice file and response data that responds to the voice data.

[0009] On one hand, this application provides an information interaction device, which includes:

[0010] The display unit is used to display the information interaction interface, which is used to interact with the interaction object and includes an information input area.

[0011] The processing unit is configured to respond to an information interaction operation in the information interaction interface and display a first voice file in the information input area; the first voice file is generated based on the voice data input by the first object;

[0012] The processing unit is also configured to, in response to an output operation of a first voice file displayed in the information input area, display a second voice file in the information interaction interface; the second voice file consists of the first voice file and response data that responds to the voice data.

[0013] On one hand, embodiments of this application provide a computer device, which includes a processor and a memory; the memory stores a computer program; when the computer program is executed by the processor, it performs the aforementioned information interaction method.

[0014] On the one hand, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the aforementioned information interaction method.

[0015] On the one hand, embodiments of this application provide a computer program product, which includes a computer program. When the computer program is executed by a processor, it performs the aforementioned information interaction method.

[0016] In this embodiment, an information interaction interface is displayed for information interaction with an interactive object, and the interface includes an information input area. In response to an information interaction operation in the interface, a first voice file is displayed in the input area. The first voice file is generated based on voice data input by a first object. In response to an output operation on the first voice file displayed in the input area, a second voice file can be output in the interface. This second voice file consists of the first voice file and response data replying to the voice data. It can be seen that this application supports displaying voice files input by a user (such as a first object) in the input box (i.e., the information input area). Compared to the automatic output method after the user inputs voice data, displaying the user-recorded voice file in the input box provides a preview space for the user, thereby improving the reliability of the user's input data. Furthermore, it supports outputting voice data in file format. Compared to directly outputting voice data, displaying voice files allows for more flexible control over the output of voice data, thereby improving the flexibility of voice output. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1This is a schematic diagram of the architecture of an information interaction system provided in an embodiment of this application;

[0019] Figure 2 This is a flowchart illustrating the front-end and back-end interaction of an information interaction method provided in an embodiment of this application.

[0020] Figure 3 This is a flowchart illustrating an information interaction method provided in an embodiment of this application;

[0021] Figure 4a This is a schematic diagram of an information interaction interface provided in an embodiment of this application;

[0022] Figure 4b This is a schematic diagram of another information interaction interface provided in an embodiment of this application;

[0023] Figure 4c This is a schematic diagram of another information interaction interface provided in an embodiment of this application;

[0024] Figure 4d This is a schematic diagram of a process for generating voice interaction operations provided in an embodiment of this application;

[0025] Figure 5a This is a schematic diagram of a process for obtaining a first voice file provided in an embodiment of this application;

[0026] Figure 5b This is a schematic diagram of another process for obtaining a first voice file provided in an embodiment of this application;

[0027] Figure 5c This is a schematic diagram of a process for requesting user authorization provided in an embodiment of this application;

[0028] Figure 5d This is a schematic diagram of a process for collecting voice data provided in an embodiment of this application;

[0029] Figure 5e This is a schematic diagram of another process for collecting voice data provided in an embodiment of this application;

[0030] Figure 6a This is a schematic diagram of a process for generating a preview operation provided in an embodiment of this application;

[0031] Figure 6b This is a schematic diagram of a process for previewing a first audio file provided in an embodiment of this application;

[0032] Figure 6c This is a schematic diagram of another process for previewing a first audio file provided in an embodiment of this application;

[0033] Figure 6d This is a schematic diagram of a preview operation for a first audio file provided in an embodiment of this application;

[0034] Figure 6e This is another schematic diagram of a preview operation for a first audio file provided in an embodiment of this application;

[0035] Figure 6f This is a flowchart illustrating a process for editing text information in a first voice file, as provided in an embodiment of this application.

[0036] Figure 7a This is a schematic diagram of a process for outputting a first voice file provided in an embodiment of this application;

[0037] Figure 7b This is a schematic diagram of another process for outputting the first voice file provided in an embodiment of this application;

[0038] Figure 7c This is a schematic diagram of another process for outputting the first voice file provided in an embodiment of this application;

[0039] Figure 8 This is a schematic diagram of a scanning process for a first voice file provided in an embodiment of this application;

[0040] Figure 9 This is a schematic diagram of a process for controlling the playback of voice data provided in an embodiment of this application;

[0041] Figure 10a This is a schematic diagram of an interface for outputting a reply voice file, provided in an embodiment of this application;

[0042] Figure 10b This is a schematic diagram of another interface for outputting and replying to voice files provided in an embodiment of this application;

[0043] Figure 10c This is a schematic diagram of another interface for outputting and replying to voice files provided in an embodiment of this application;

[0044] Figure 11a This is a schematic diagram of an interface for outputting a reply voice file and a reply text, provided in an embodiment of this application;

[0045] Figure 11b This is a schematic diagram of another interface for outputting reply voice files and reply text provided in an embodiment of this application;

[0046] Figure 12a This is a schematic diagram of a process for sharing and replying to voice files provided in an embodiment of this application;

[0047] Figure 12b This is a schematic diagram of another process for sharing and replying to voice files provided in an embodiment of this application;

[0048] Figure 13This is a schematic diagram of an interface for exporting session content provided in an embodiment of this application;

[0049] Figure 14 This is a flowchart illustrating another information interaction method provided in an embodiment of this application;

[0050] Figure 15 This is a technical flowchart of a dialogue interaction scheme provided in an embodiment of this application;

[0051] Figure 16 This is a schematic diagram of the structure of an information interaction device provided in an embodiment of this application;

[0052] Figure 17 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0053] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0054] I. Introduction to the key terms used in this application.

[0055] (1) Interaction objects:

[0056] Interactive objects, as the name suggests, refer to objects that can interact with users (such as the first object). Classified by object type, interactive objects here can include any type of object such as real objects (such as users) and virtual objects (such as model objects). In this application, interactive objects mainly refer to model objects. A model object is a software program that integrates the functions and technologies of a model (such as a large language model). It can understand and learn human language to engage in dialogue and interact based on contextual information during the dialogue, thus achieving a truly human-like chat communication. Therefore, an interactive object is a generative language interactive object with user dialogue as its core function. For example, an interactive object can be an intelligent chatbot; it has the following characteristics: (1) Generative: It generates responses or replies based on user input and predicts the most context-appropriate text based on a probability model; (2) Interactive: The model design focuses on dynamic interaction with users, and can adjust the generated content according to the conversation history; it also supports multi-turn dialogues and retains the understanding and memory of the context; (3) Task flexibility: It can complete various language-related tasks, such as: knowledge question answering, translation, text creation, code explanation, email writing, etc. In this application, information interaction between users and interactive objects is supported. During the information interaction process, the interactive objects can directly engage in dialogue with the users; or, the interactive objects can also output a display interface (such as an information interaction interface) and engage in dialogue by outputting interactive information (such as a first voice file and a reply voice file) in real time within the information interaction interface.

[0057] (2) Large language model.

[0058] Large Language Models (LLMs), often shortened to large language models or large models, are neural network models that learn the statistical patterns and semantic information of language by training on large amounts of text data, thereby enabling them to predict the next word or sentence. Examples of LLM models include, but are not limited to, MoE (Mixture of Experts) and MLLM (Multimodal Large Language Model). LLM models can be generated using different training methods and model structures. Training methods include pre-training and fine-tuning. Pre-training involves large-scale training of the model under unsupervised or self-supervised conditions to give it a general understanding and expressive ability for language data. Fine-tuning, on the other hand, optimizes the model for specific application scenarios or tasks through supervised learning based on the pre-trained model. LLM models have wide applications in Natural Language Processing (NLP), such as human-computer interaction, machine translation, speech recognition, and text generation. This application mainly relates to the human-computer interaction function of LLM models; for example, it supports information interaction between users and LLM models, such as information interaction including but not limited to: voice interaction, text interaction, voice + text interaction, voice + image + text interaction, etc.

[0059] (3) Information exchange:

[0060] Information interaction refers to an interactive method in which two or more objects exchange information. This application primarily concerns information interaction between a user (such as the first object) and model objects of a large language model. For example, a user sends a message to an interaction object, which can then reply / respond to that message. Furthermore, this interacted information can be displayed in the information interaction interface. The types of information exchanged between objects can include, but are not limited to, voice, text, images, video, and files. Moreover, during information interaction, objects can exchange various types of information. This application primarily concerns information interaction between a user (such as the first object) and model objects of a large language model, where information interaction includes, but is not limited to, any one or more of voice interaction, text interaction, image interaction, and video interaction.

[0061] (4) In response to:

[0062] In response to: the conditions or states on which the operation performed (such as the information interaction operation performed by the first object) depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0063] (5) First voice file, reply voice file, second voice file.

[0064] The first voice file is an audio file generated based on voice data input by a user (such as a first object). For example, if the first object generates 10 seconds of voice data, then this 10-second voice data can constitute an audio file. In this application, the first voice file is used to control the playback of the first object's voice data. Here, the control method of the first object's voice data refers to controlling the playback attributes of the first object's voice data. For example, the playback attributes here include, but are not limited to, any one or more of the following: playback progress, playback speed, playback timbre, and playback volume. That is to say, based on the first voice file, the playback attributes of the first object's voice data, such as playback progress, playback speed, playback timbre, and playback volume, can be controlled, allowing for flexible playback control of the first object's voice data during voice playback, thereby meeting personalized voice playback needs.

[0065] A response audio file is an audio file generated by a large language model responding to the voice data input by a first object. In other words, the response audio file is used to respond to the first object's first voice file. In this application, the response audio file is used to control the playback of the interactive object's response data. Similarly, the control method for the interactive object's response data refers to controlling the playback attributes of the response data. These playback attributes include, but are not limited to, any one or more of the following: playback progress, playback speed, playback timbre, and playback volume. In other words, based on the ability of the response audio file to control the playback attributes of the interactive object's response data, flexible playback control of the interactive object's response data can be achieved during voice playback, thereby meeting personalized voice playback needs.

[0066] The file formats of the aforementioned first voice file and response voice file can be any audio format such as MP3 (Moving Picture Experts Group Audio Layer III), WAV (Waveform Audio File Format), or FLAC (Free Lossless Audio Codec). Based on this, this application displays the user's input voice data as a file in the information interaction interface, and the response data after the interactive object responds to the voice data is also displayed as a file in the information interaction interface. This allows the information interaction process between the user and the large language model to simulate the effect of real user voice dialogue as closely as possible, improving the user experience of human-computer interaction.

[0067] The second voice file refers to the interactive information displayed after the first voice file in the information input area is output. This second voice file is displayed in the conversation area of ​​the information interaction interface. The second voice file includes: the first voice file and response data (e.g., response voice file, response text).

[0068] (6) Pattern code.

[0069] A pattern code (commonly known as a QR code) is an encoding method that uses a black and white matrix to represent information (such as interactive information). It can be quickly identified and decoded by specific devices (such as cameras and QR code scanners) and is widely applicable to scenarios such as mobile payments, information sharing, and identity verification. Specifically, QR codes typically have the following characteristics:

[0070] 1) Large information storage capacity: It can store numbers, letters, special symbols, and even binary data.

[0071] 2) Fast reading: The QR code is designed to be decoded quickly, and can be recognized even if it is partially damaged.

[0072] 3) Strong fault tolerance: QR codes have built-in error correction function (using Reed-Solomon encoding), which allows them to be decoded even if some parts of the graphic are damaged.

[0073] 4) Multi-directional recognition: QR codes can be scanned from any direction.

[0074] 5) Wide range of applications: used for payment, website redirection, identity verification, logistics tracking, etc.

[0075] In this application, the first voice file of the user (such as the first object) and the reply voice file of the interactive object can be converted into corresponding pattern codes and displayed in the information interaction interface of the first client (such as the web terminal); and, it supports the second client (such as the mobile terminal or the app terminal) to scan the pattern code of the first voice file (or the pattern code of the reply voice file) and then play the first voice file (or reply voice file) output by the second client.

[0076] II. Introduction to this application proposal.

[0077] This application provides an information interaction scheme applicable to human-computer interaction scenarios. It primarily offers a solution supporting voice interaction between users and a large language model. The scheme displays the user-recorded first voice file in an input box (i.e., the information input area), allowing the user to preview and manipulate the voice file before outputting it, thus improving data accuracy during voice interaction. Furthermore, the large language model responds to the first voice file and outputs a response voice file (along with response text and other information), thereby achieving a voice dialogue effect between the user and the large language model. Specifically, the principle of the information interaction scheme provided in this application is roughly as follows:

[0078] (1) Display information interaction interface, which is used to interact with interaction objects (such as model objects of large language models), such as voice interaction, text interaction, image interaction, etc.; and the information interaction interface includes an information input area (commonly known as an input box).

[0079] (2) In response to an information interaction operation in the information input area, the first voice file of the first object is displayed in the information input area. The first voice file is generated based on the voice data input by the first object, and is used to trigger a preview operation on the first object's voice data in the information input area. For example, the information interaction interface has an information interaction entry point (such as a voice button). After the first object clicks the voice button, its voice data can be collected through the voice button, and a first voice file is generated based on the first object's voice data and displayed in the input box of the information interaction interface. Furthermore, the first object can preview the first voice file in the input box. The purpose of this preview is for the first object to verify whether the first voice file meets the expected requirements. If it does, the first voice file is output; if it does not, it can be deleted and re-recorded.

[0080] (3) In response to the output operation of the first voice file displayed in the information input area, the first voice file and the reply voice file are output on the information interaction interface. The reply voice file is generated based on the reply data generated by the large language model in response to the voice data. For example, the output methods of the reply voice file include: playing the reply voice file, or playing the reply voice file and simultaneously displaying the corresponding text information, etc. That is, the model's reply data can also be displayed on the information interaction interface in the form of a voice file, thereby realizing the voice-to-voice dialogue interaction effect between the user and the model.

[0081] Based on this, this application provides a solution to support voice dialogue between users and large language models. On one hand, this solution supports displaying the voice file input by the user (e.g., the first object) in the input box (i.e., the information input area), and supports previewing the user's input voice data through the voice file in the input box. This allows users to preview whether the input data meets expectations, thereby improving the reliability of the user's input data. On the other hand, during the human-computer interaction between the user and the model, it supports outputting the model's response data in the form of a voice file in the information interaction interface. Compared to outputting the model's response data in text form, this application can realize voice interaction between the user and the model, simulating the effect of human-to-human dialogue as much as possible during the human-computer interaction process, thereby improving the realism and flexibility of information interaction.

[0082] It should be noted that the information exchange scheme provided in this application requires special explanation of the following two points:

[0083] 1. This application involves relevant data during information interaction (e.g., voice data, response data, first voice file, response voice file, etc.). When the above embodiments of this application are applied to specific products or technologies, permission or consent from the target audience is required. Furthermore, the collection, use, and processing of relevant data must comply with relevant regional laws, regulations, and standards, adhering to the principles of legality, legitimacy, and necessity, and must not involve acquiring data types prohibited or restricted by laws and regulations. In some optional embodiments, the relevant data involved in the embodiments of this application is obtained after separate authorization from the target audience. Additionally, when obtaining separate authorization from the target audience, the purpose of the relevant data is explained to the target audience.

[0084] 2. It is understood that in this application, the term "at least one" refers to one or more, and "multiple" means two or more; for example, at least one task list displayed in an information interaction interface refers to one, two, or more task lists. The terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor is there any limitation on the quantity or execution order.

[0085] The following is combined Figure 1 and Figure 2 This application provides a detailed description of the information interaction system provided.

[0086] I. Introduction to the macro system architecture.

[0087] Please see Figure 1 , Figure 1 This is a schematic diagram of the architecture of an information interaction system provided in an embodiment of this application. Figure 1 As shown, the information interaction system may include multiple terminal devices (such as first terminal device 101, second terminal device 102, etc.) and server 103. This application does not specify the number of terminal devices and servers. For example, first terminal device 101 may be a terminal device used by a first object, second terminal device 102 may be a terminal device used by a second object, and server 103 runs a large language model to realize information interaction with users (such as the first object). Furthermore, each terminal device displays an information interaction interface and supports information interaction (such as voice interaction, text interaction, voice + text interaction, etc.) with the model object corresponding to the large language model in the information interaction program. In this system, any terminal device (such as the first terminal device 101) can be directly or indirectly connected to the server 103 via a network. The network may include, but is not limited to, wired networks and wireless networks. Wired networks may include local area networks (LANs), metropolitan area networks (MANs), and wide area networks (WANs). Wireless networks may include Bluetooth, Wi-Fi (Wireless Fidelity, a standard wireless LAN), and other networks that enable wireless communication. Furthermore, the first terminal device 101 can exchange information with other terminal devices (such as the second terminal device 102) connected to the server 103 through the server 103.

[0088] in, Figure 1Any terminal device in the information interaction system shown can include, but is not limited to: smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, in-vehicle terminals, smart wearable devices (such as smart bracelets, smart pedometers), virtual reality devices (such as VR (Virtual Reality) devices, AR (Augmented Reality) devices), smart robots, etc. Terminal devices are often equipped with display devices, which can be monitors, displays, touchscreens, etc., and touchscreens can be touchscreens, touch panels, etc. Furthermore, Figure 1 The server 103 in the information interaction system shown can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0089] Next, taking any terminal device (such as the first terminal device 101) as an example, the information interaction process between various devices in the information interaction system will be described accordingly:

[0090] ① The first terminal device 101 (such as the terminal device used by the first object) can display an information interaction interface, which is used to interact with the interaction object (such as the model object of the large language model) through information interaction (such as voice interaction, text interaction, image interaction), and the information interaction interface includes an information input area.

[0091] ② The first terminal device 101 responds to an information interaction operation in the information interaction interface by displaying a first voice file in the information input area. The first voice file is generated based on the voice data input by the first object, and is used to trigger a preview operation on the first object's voice data in the information input area. For example, the information interaction interface has an information interaction entry point (such as a voice button). After the first object clicks the voice button, an information interaction operation is generated, and the first object's voice data is collected through the voice button. Based on the first object's voice data, a first voice file is generated and displayed in the information interaction interface. The display method of the first voice file in the information interaction interface includes: playing the first voice file, or playing the first voice file and simultaneously displaying corresponding text information, etc.

[0092] ③ The first terminal device 101 sends the first voice file to the server 103.

[0093] ④ Server 103 calls the large language model to process the speech data in the first speech file and generate response data. For example, the large language model uses Automatic Speech Recognition (ASR) technology to convert the speech data in the first speech file into text, obtaining the text information corresponding to the first speech file; then it calls the LLM model to respond to the text information to obtain response data; and uses Text-to-Speech (TTS) technology to convert the response data into speech, obtaining the response speech file corresponding to the response data.

[0094] ⑤ The server 103 will return the reply voice file to the first terminal device 101.

[0095] ⑥ The first terminal device 101 outputs a response voice file on the information interaction interface. This response voice file is generated based on a large language model to respond to the voice data, and is used to control the playback of the response data. For example, users can adjust the playback attributes of the response data through the response voice file, such as playback progress, playback volume, playback speed, and playback timbre.

[0096] It should be noted that the above information interaction process is for illustrative purposes only and does not limit the specific execution process of the first terminal device and the server. Optionally, the first terminal device can execute the above automatic speech recognition process and text-to-speech recognition process, etc.; or, the complete process of the above information interaction can be executed by any terminal device or server alone.

[0097] In one possible implementation, the information interaction system provided in this application embodiment can be deployed in a blockchain network. For example, the first terminal device 101, the second terminal device 102, and the server 103 in the above system can all be regarded as node devices in the blockchain network, and these node devices together constitute the blockchain network. Therefore, the information interaction process executed in this application embodiment can be executed on the blockchain, which can ensure the fairness and impartiality of the information interaction process, make the process traceable, and ensure data security during the information interaction process, thereby improving the security and reliability of the entire information interaction process.

[0098] II. Introduce the complete interaction process of information exchange methods.

[0099] Furthermore, each terminal device in the aforementioned information interaction system runs a client application, which is used to run an application program (APP) with information interaction functions. An application program refers to a computer program designed to perform one or more specific tasks. Applications can be categorized according to different dimensions (such as the application's operating method and functions) to obtain the types of the same application under different dimensions, including: ① According to the application's operating method, applications may include, but are not limited to: clients installed on the terminal, mini-programs that can be used without downloading and installation, web applications opened through a browser, etc. ② According to the application's functional type, applications may include, but are not limited to: IM (Instant Messaging) applications, content interaction applications, etc. The client (also known as the front-end) displays an information interaction interface, which is used for users to interact with the model objects of the large language model (e.g., users interacting with the LLM model via voice). Additionally, the server runs a server-side component (also known as the back-end), which runs the large language model and provides backend technical support to the client through the large language model (such as data response, text-to-speech, speech-to-text, etc.).

[0100] Please see Figure 2 , Figure 2 This is a flowchart illustrating the front-end and back-end interaction process of an information interaction method provided in an embodiment of this application. The following is in conjunction with… Figure 2 The complete interaction process between the client (frontend), server (backend), and large language model is illustrated with an example. The interaction process includes the following steps:

[0101] 1. The user (e.g., the first object) clicks the record button in the information interaction interface displayed on the front end.

[0102] 2. The front end requests microphone authorization from the user.

[0103] 3. The user agrees to the authorization.

[0104] 4. A front-end continuous monitoring microphone is used to collect the voice data of the first object.

[0105] 5. The user stops recording or deletes and re-records.

[0106] 6. The user submits a recording (e.g., submitting the recording as the first audio file).

[0107] 7. The front end reads the first audio file as binary data.

[0108] 8. The front end uploads the first audio file to the back end server.

[0109] 9. The backend performs automatic speech recognition (ASR) processing on the first audio file to convert the first audio file into text information.

[0110] 10. The backend requests a large model (such as a large language model: LLM model), in which the request includes the text information corresponding to the first speech file.

[0111] 11. The large language model processes the text information of the first speech file to obtain response data.

[0112] 12. The large model returns text results (such as reply data) to the backend.

[0113] 13. The backend uses TTS (Text-to-Speech) to convert the reply data into reply audio files.

[0114] 14. The backend returns the audio result (responding with a voice file) to the frontend.

[0115] 15. The front-end displays the reply voice file in the information interaction interface; among them, users can scan the QR code of the reply voice file to preview it.

[0116] The information interaction system provided in this application offers a solution that supports voice dialogue between users and a large language model. On one hand, this solution supports displaying the voice file input by the user (e.g., the first object) in an input box (i.e., the information input area), and allows previewing of the user's input voice data through the voice file in the input box. This helps users preview whether the input data meets expectations, thereby improving the reliability of the user's input data. On the other hand, during the human-computer interaction between the user and the model, the system supports outputting the model's response data in the form of a voice file in the information interaction interface. Compared to outputting the model's response data in text form, this application enables voice interaction between the user and the model, simulating the effect of human-to-human dialogue as much as possible during the human-computer interaction process, thereby improving the realism and flexibility of information interaction.

[0117] It is understood that the information interaction system described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0118] The embodiments of the information interaction method provided in this application will be described in detail below from a product perspective.

[0119] Please see Figure 3 , Figure 3This is a flowchart illustrating an information interaction method provided in an embodiment of this application. This information interaction method can be implemented using a computer device (such as...). Figure 1 Execute on any of the terminal devices shown. Figure 3 As shown, the information interaction method includes the following steps S301-S303:

[0120] S301. Display an information interaction interface, which includes an information input area.

[0121] The information interaction interface is used for users to interact with the model objects corresponding to the large language model. For example, the information interaction methods here include, but are not limited to, different types of interaction such as voice interaction, text interaction, and image interaction. Of course, in the process of information interaction between the user and the interaction object, both unimodal and multimodal data interaction are supported. Unimodal data interaction refers to interaction using only one type of data, such as the user sending text and the model replying with text. Multimodal data interaction refers to the simultaneous interaction of multiple different types of data, such as supporting users to send text, images, and voice data, and the model can also reply with text, images, and voice data. In this embodiment, the main focus is on the voice interaction process between the user and the model objects of the large language model.

[0122] Specifically, an information interaction interface, as the name suggests, refers to an interface used to display interactive information. The information interaction interface involved in this application is a front-end interface displayed by the model object corresponding to the large language model. A front-end interface refers to an interface that allows direct interaction with the user; for example, a front-end interface can be a UI (User Interface), a Web (webpage), etc. The information interaction interface supports any user to interact with the model object corresponding to the large language model, such as through voice interaction, text interaction, image interaction, etc. Furthermore, this information interaction interface supports displaying real-time interaction information as well as historical interaction information. Please see [link to relevant documentation]. Figure 4a , Figure 4a This is a schematic diagram of an information interaction interface provided in an embodiment of this application. For example... Figure 4a As shown, the information interaction interface is divided into a conversation area 401 (commonly known as a dialog box) and an information input area 402 (commonly known as an input box). The information input area 402 is used to support user input of information (such as voice, text, images, videos, files, and other data of any type). The conversation area 401 is used to display the interaction information generated by the user and the interaction object. For example, the conversation area can display both the interaction information sent by the user and the response information generated by the interaction object in response to the above interaction information.

[0123] In one possible implementation, the information interaction interface includes a task display area and a session area. The task display area shows at least one task list, which in turn shows at least one historical session task that has been created. The session area mentioned here is used for information interaction with the model object corresponding to the large language model during the execution of session tasks. Please see [link to relevant documentation]. Figure 4b , Figure 4b This is a schematic diagram of another information interaction interface provided in an embodiment of this application. For example... Figure 4b As shown, the information interaction interface is divided into: a task display area 4001, a conversation area 4002, and an information input area 4003. The task display area 4001 displays at least one task list (e.g., test point 1, test point 2, and test point 3 are each represented as a task list); wherein, the conversations included in a task list (such as test point 1, test point 2, or test point 3) belong to the same knowledge domain. Optionally, the conversation process between the first object and the interaction object can include the following two conversation modes: (1) In response to the conversation creation operation, the newly created conversation task is displayed, and information interaction under the newly created conversation task is performed with the model object corresponding to the large language model in the conversation area of ​​the newly created conversation task; that is, this application supports creating a conversation under any task list and then performing information interaction with the interaction object. (2) In response to the selection operation for the target historical conversation task, the selected target historical conversation task is displayed, and information interaction under the target historical conversation task is performed with the model object corresponding to the large language model in the conversation area of ​​the target historical conversation task. Please refer to Figure 4c , Figure 4c This is a schematic diagram of another information interaction interface provided in an embodiment of this application. For example... Figure 4c As shown, you can select Session 2 in Test Point 1 (as shown in 404) as the target historical session from the task display area of ​​the information interaction interface, and display the historical session content of the target historical session in the session area (as shown in 405), thereby triggering information interaction with the interactive object under the target historical session (i.e., Session 2 in Test Point 1). In summary, this application supports creating a new session or selecting an existing historical session in the information interaction interface to trigger information interaction with the interactive object.

[0124] S302, in response to an information interaction operation in the information interaction interface, a first voice file is displayed in the information input area; the first voice file is generated based on the voice data input by the first object.

[0125] The information interaction operation refers to a voice input operation performed by the first object on the information interaction interface. For example, the information interaction interface has an operation area. When a preset operation is detected in the operation area, a voice input operation is generated. The preset operation includes, but is not limited to, any one of the following: gesture operation, single-click operation, double-click operation, and long-press operation. Alternatively, the first object can trigger a voice interaction entry on the information interaction interface to perform a voice input operation. In other words, the information interaction operation here is used to trigger the first object to perform voice input on the information interaction interface to generate a first voice file. Please refer to [link to relevant documentation]. Figure 4d , Figure 4d This is a schematic diagram illustrating a process for generating voice interaction operations, provided in an embodiment of this application. For example... Figure 4d As shown in (1), assuming that a gesture operation (such as drawing a "V") performed by a user (such as the first object) is detected in the information input area of ​​the information interface, it can be considered that the first object has performed an information interaction operation on the information interaction interface; optionally, as Figure 4d As shown in (2), the information input area of ​​the information interaction interface is provided with a voice interaction entry 400. When the first object clicks the voice interaction entry 400, it can be considered that the first object has performed an information interaction operation on the information interaction interface.

[0126] In one possible implementation, the information interaction interface includes an information interaction entry point. If this entry point is triggered (e.g., a user clicks on it), an information interaction operation can occur. Specifically, based on the type of information to be interacted with, the information interaction entry points include, but are not limited to: voice interaction entry points and information input controls; wherein, the voice interaction entry point is used to trigger the collection of voice data, and the information input controls are used to trigger the acquisition of data of types such as text, images, and videos. Figure 4a As shown, the information interaction interface includes a voice interaction entry 4022. If a user clicks (such as single-click, double-click, or long-press) on the voice interaction entry 4022, a first information interaction operation is considered to have occurred. This first information interaction operation is used to trigger the collection of the user's voice data. Additionally, the information interaction interface also includes an information input control 4021. If a user clicks (such as single-click, double-click, or long-press) on the information input control 4021, a second information interaction operation is considered to have occurred. This second information interaction operation is used to trigger the acquisition of data of types such as images, text, videos, and files. In this embodiment, the main focus is on the voice interaction between a user (such as the first object) and a model object of a large language model. Therefore, subsequent embodiments will mainly illustrate the voice interaction process between the user and the model object.

[0127] Specifically, the first audio file is generated based on the audio data input by the first object. For example, if the first object generates a 10-second audio data segment, then this 10-second audio data segment constitutes an audio file. Furthermore, the first audio file is used to control the playback of the first object's audio data. For example, the first audio file is used to control the playback attributes of the first object's audio data, such as playback progress, playback speed, playback timbre, and playback volume. The first audio file can be obtained in two main ways: Method 1, allowing the user to select the first audio file from pre-recorded audio files; Method 2, generating the first audio file through real-time recording. Examples of how to obtain the first audio file are given below.

[0128] Method 1: Supports users in selecting a first audio file from pre-recorded audio files. In one possible implementation, in response to a first object's triggering operation on the information input control in the information interaction interface, a file selection list is output, displaying at least one pre-recorded audio file; in response to a selection operation on a target audio file, the target audio file is displayed as the first audio file in the information input area. Please refer to [link to relevant documentation]. Figure 5a , Figure 5a This is a schematic diagram illustrating a process for obtaining a first audio file, provided in an embodiment of this application. Figure 5a As shown, the information input area of ​​the information interaction interface is equipped with an information input control 500. In response to a trigger operation by a first object on the information input control 500, a file selection list is output. For example, the file selection list includes at least one pre-recorded audio file (such as file 1, file 2, file 3, and file 4). The first object can select any audio file from the file selection list as the first audio file (such as file 1), thereby displaying the first audio file (such as file 4) in the information input area of ​​the information interaction interface. Figure 5a (See file 500 in the image). This implementation allows users to easily select the first audio file from pre-recorded audio files without real-time audio recording, thus improving the efficiency of audio file acquisition.

[0129] Method 2: Real-time recording to generate the first audio file. In one possible implementation, the information interaction interface has a voice interaction entry point, and the interface is divided into an information input area and a conversation area. In response to a first object's trigger operation on the voice interaction entry point, the first object's voice data is collected, and a first audio file is generated based on the collected voice data; the first audio file is displayed in the information input area of ​​the information interaction interface; when an output operation on the first audio file displayed in the information input area is detected, the first audio file is sent from the information input area to the conversation area for display. Please refer to [link to relevant documentation]. Figure 5b , Figure 5b This is a schematic diagram of another process for obtaining a first audio file provided in an embodiment of this application. For example... Figure 5b As shown, the information interaction interface includes a voice interaction entry 501. When a user (e.g., the first object) clicks this voice interaction entry 501, it can trigger the collection of the first object's voice data. Optionally, during the collection of the first object's voice data, a prompt window can be displayed in the information interaction interface. This prompt window displays a voice collection animation 502, which prompts the user that voice data is currently being collected. Next, a first voice file can be generated based on the collected voice data; and the first voice file is displayed in the information input area of ​​the information interaction interface (e.g., the first voice file displayed in the information input area is shown in the image). Figure 5b (As shown in 503). During the above voice interaction process, users can record audio in real time in the information input area, and the audio data collected during the recording process is displayed in the conversation area in the form of a first voice file. This method of real-time recording and obtaining voice files can meet the real-time requirements of users in the information interaction process.

[0130] Optionally, upon detecting an information interaction operation by the first object with the voice interaction entry point, a microphone authorization request can be requested from the first object to trigger recording. Please see [link to relevant documentation]. Figure 5c , Figure 5c This is a schematic diagram of a process for requesting user authorization provided in an embodiment of this application. Figure 5c As shown, the information interaction interface includes a voice interaction entry point. When the first user clicks on the voice interaction entry point, a prompt pop-up window is displayed. This pop-up window prompts the first user to grant microphone permissions. The pop-up window includes an "agree" control (520) and a "deny" control (510). If the first user clicks the "agree" control (520), it indicates that the first user agrees to the authorization, and the collection of the first user's voice data can be triggered. Conversely, if the first user clicks the "deny" control (510), it indicates that the first user does not agree to the authorization, and the collection of the first user's voice data cannot be triggered.

[0131] In one possible implementation, Figure 5bDuring the voice data acquisition process shown, the display form of the voice interaction entry point can be changed. The specific process is as follows: The aforementioned voice interaction entry point can be a voice acquisition control, which is displayed in the information interaction interface in a first display form; in response to a trigger operation by a first object on the voice acquisition control, the display form of the voice acquisition control is adjusted from the first display form to the second display form; when voice data is detected to be generated by the first object, the display form of the voice acquisition control is adjusted from the second display form to the third display form; in response to a voice data acquisition completion event, the display form of the voice acquisition control is adjusted from the third display form to the fourth display form; the acquisition completion event includes: the voice data acquisition completion operation is performed, or the voice data acquisition duration reaches a preset duration threshold. The first, second, third, and fourth display forms are all different; and the display form includes any one or more of color, shape, and size. Please refer to [link to relevant documentation]. Figure 5d , Figure 5d This is a schematic diagram illustrating a process for collecting voice data according to an embodiment of this application. Figure 5d As shown, the voice data acquisition process mainly includes the following steps S1-S6:

[0132] S1. The voice acquisition control is in an inactive state. In the inactive state, the voice acquisition control is displayed through the first display form (as shown in 5001).

[0133] S2. Display prompt text when inactive. When the user uses the mouse and moves it to the display location of the voice acquisition control, a prompt text (such as voice input) can be output. This prompt text is used to inform the user that voice data can be acquired by triggering the voice acquisition control.

[0134] S3. In the activated state, the voice acquisition control is displayed in the second display mode. When the user clicks the voice acquisition control, the voice acquisition control is activated. If no user voice data is detected at this time, the voice acquisition control is displayed in the second display mode. For example, the voice acquisition control in the second display mode is shown in 5002. The bar in the control icon has a uniform length.

[0135] S4. In the activated state, the voice acquisition control is displayed in the third display form; when the voice data generated by the user is detected, the voice acquisition control is changed from the second display form to the third display form. For example, the voice acquisition control in the third display form is shown in 5003, and the bar in the control icon will fluctuate.

[0136] S5. During the input of voice data, if the user moves the mouse to the display position of the voice acquisition control again, another prompt text (such as "Stop Input") can be output. This prompt text is used to remind the user that the voice acquisition control can be triggered to stop the acquisition of voice data.

[0137] S6. After stopping voice data acquisition, i.e., in response to a voice data acquisition completion event (e.g., clicking the voice acquisition control again during voice data acquisition generates a acquisition completion event, or when the voice data acquisition duration reaches a preset duration threshold (e.g., 1 minute), the display mode of the voice acquisition control is adjusted from the third display mode to the fourth display mode. For example, the fourth display mode is used to indicate that the voice acquisition control is deleted from the information input area, or the voice acquisition control can be grayed out. Subsequently, a first voice file can be generated based on the acquired voice data. The first voice file is displayed in the information input area of ​​the information interaction interface.

[0138] In one possible implementation, Figure 5b The voice data acquisition process shown supports displaying animation effects and outputting text information in real time. Please refer to [link / reference]. Figure 5e , Figure 5e This is a schematic diagram of another process for collecting voice data provided in an embodiment of this application. For example... Figure 5e As shown, the information interaction interface includes a voice interaction entry point. In response to a user's (e.g., the first object) triggering the voice interaction entry point 5100, a prompt window is output. This prompt window displays a first emoticon image 5200, which prompts the user to input voice. Further, when voice data input from the first object is detected, the first emoticon image 5200 can be changed to a second emoticon image 5300 in the prompt window. The second emoticon image differs from the first emoticon image in its display format (e.g., expression, style, action). Furthermore, text content corresponding to the user's input voice data (e.g., what characteristics a potato possesses) can be displayed around the second emoticon image 5300. This implementation supports displaying different animation effects through emoticon images during recording, thereby enabling interaction with the user and improving the user experience. Moreover, the text information corresponding to the voice data is output in real-time through emoticon images, allowing the user to understand the input voice content in real-time during the voice input process.

[0139] Based on this, during the above-mentioned voice data collection process, it is supported to display different stages of the voice collection process through different display forms of the voice interaction entry (such as voice collection control). For example, the display form when collecting voice data is different from the display form when not collecting voice data. This allows users to intuitively perceive the different stages of the voice collection process and improve the user experience.

[0140] Furthermore, after the first audio file is displayed in the information input area, it is supported to preview and operate on the first audio file; the preview process of the audio file is described in detail below.

[0141] In one possible implementation, after displaying the first audio file in the information input area of ​​the information interaction interface, in response to a preview operation on the first audio file displayed in the information input area, a preview of the audio data corresponding to the first audio file is performed. The audio preview refers to previewing the audio data generated by the first object, i.e., playing the audio data. The preview operation includes any of the following: a preview operation is generated when a trigger operation (such as a click) on the first audio file is detected; or a preview operation is generated when a preset audio command on the first audio file is detected; or, a preview operation is generated when the first audio file is dragged to a preview area (this preview area can be any area in the information interaction interface different from the information input area, such as the middle area, top area, bottom area, etc.; or, this preview area can also be an independent window floating in the information interaction interface, and the size of the preview area is not specifically limited). Please refer to [link to relevant documentation]. Figure 6a , Figure 6a This is a schematic diagram illustrating a process for generating a preview operation, provided in an embodiment of this application. For example... Figure 6a As shown in (1), when a user (such as the first object) clicks on the first audio file in the information input area, a preview operation on the first audio file is generated, and the audio file can then be previewed in the information input area; optionally, as shown in (1), Figure 6a As shown in (2), the information interaction interface has a preview area 600. When the user (such as the first object) drags the first voice file in the information input area to the preview area 600, a preview operation of the first voice file can be generated, and the first voice file can be previewed in the preview area.

[0142] During the voice preview process, the playback attributes of the voice data corresponding to the first voice file can be adjusted. These attributes include playback progress, playback speed, playback volume, and playback timbre. Please refer to [link / reference]. Figure 6b , Figure 6bThis is a schematic diagram illustrating a process for previewing a first audio file, provided in an embodiment of this application. Figure 6b As shown, a first audio file 6011 is displayed in the information input area of ​​the information interaction interface S601, and audio preview of the first audio file 6011 displayed in the information input area is supported. During the preview of the first audio file 6011, a first object can adjust the playback attributes (such as playback progress) of the audio data played by the first audio file 6011; for example, the first audio file displayed in the information input area includes a progress bar, and the first object can adjust the playback progress of the audio data played by the first audio file by dragging the progress bar left or right. In response to the adjustment operation of the progress bar of the first audio file (e.g., dragging the progress bar from left to right or from right to left), the first audio file 6021 with the adjusted playback progress is displayed in the information input area of ​​the information interaction interface S602. In this implementation, the playback progress of the first audio file can be conveniently adjusted by dragging the progress bar, which can meet the flexibility requirements during audio playback.

[0143] Similarly, the first object can also customize the voice preview of other playback attributes of the voice data (such as playback volume, playback timbre, etc.) in the information input area, and then use the voice preview process to detect whether the first voice file meets the first object's expected conditions. Optionally, please refer to Figure 6c , Figure 6c This is a schematic diagram illustrating another process for previewing a first audio file provided in an embodiment of this application. For example... Figure 6c As shown, a first audio file is displayed in the information input area of ​​the information interaction interface, and the interface supports audio preview of the first audio file displayed in the information input area. For example, it supports a first object adjusting the playback volume of the audio data played by the first audio file. Specifically, the first audio file includes a volume control 6031, and in response to a trigger operation on the volume control 6031, a volume bar 6032 is output. The first object can adjust the playback volume of the audio data played by the first audio file by dragging the volume bar 6032 (for example, dragging the volume bar upwards represents increasing the playback volume; dragging the volume bar downwards represents decreasing the playback volume). Based on this, it supports convenient adjustment of the playback audio of the first audio file by dragging the volume bar, which can also meet the flexibility requirements during audio playback.

[0144] Furthermore, during the preview of the first object's voice data, if the first voice file meets the expected conditions of the first object, then in response to the output operation on the first voice file, the first voice file is sent from the information input area to the session area for display. Please refer to [link to relevant documentation]. Figure 6d , Figure 6dThis is a schematic diagram illustrating a preview operation for a first audio file provided in an embodiment of this application. For example... Figure 6d As shown, an output control 6100 is provided in the information input area. If the first audio file displayed in the information input area meets the expected needs of the first object, the first object can click the output control 6100 to generate an output operation for the first audio file in the information input area. Optionally, please refer to... Figure 6e , Figure 6e This is a schematic diagram illustrating another preview operation for a first audio file provided in an embodiment of this application. For example... Figure 6e As shown, the information input area also has a delete control 6200. If the first voice file displayed in the information input area does not meet the expected needs of the first object, the first object can click the delete control 6200 to delete the first voice file in the information input area, and then record a new voice file.

[0145] In the above Figures 6a-6e The preview process shown supports previewing user-recorded voice files. During the preview, users can customize and adjust the playback attributes of the voice data. Only voice files that meet the user's expectations or needs can be sent to the conversation area for display; voice files that do not meet expectations can be deleted and re-recorded. This method of supporting file previewing within the input box (i.e., the information input area) provides users with a certain degree of error tolerance and improves the user experience.

[0146] In one possible implementation, during the voice preview process, text information can be displayed simultaneously with the playback of voice data, supporting synchronous adjustment of the voice file by editing the text information. The specific process is as follows: In response to a playback operation on a first voice file, the voice data of a first object is played, and the text information of the first voice file is displayed; during the playback of the voice data of the first object, in response to an editing operation on the text information, the voice data of the first object being played is adjusted; wherein, the voice data of the first object being played is adjusted to the data content corresponding to the edited text information. Please refer to [link to relevant documentation]. Figure 6f , Figure 6f This is a schematic diagram illustrating a process for editing text information in a first audio file, as provided in an embodiment of this application. Figure 6fAs shown, the information input area in the information interaction interface S610 can display the first audio file 6001. During the preview of the first audio file 6001, the corresponding text information 6002 (e.g., what characteristics a potato has) can be displayed simultaneously. Optionally, in response to the editing operation of the text information 6002, a text editing box 6003 is output in the information interaction interface S620, supporting the first object to perform content editing processing on the text information in the text editing box 6003 (such as deleting content, adding content, modifying content, etc.) to obtain the edited text information. Further, when the text information is detected to be edited, the first audio file (i.e., the audio data of the first object being played) can be adjusted simultaneously in the information interaction interface. That is, the audio data of the first object being played is adjusted to the data content corresponding to the edited text information. For example, the adjusted first audio file 6004 is displayed in the information interaction interface S630. As can be seen above, due to the addition of text information, the duration of the adjusted first audio file 6004 is changed from 1:24 to 2:24 compared to the original first audio file 6001. Based on this, during file preview, the text information can be displayed synchronously, and the audio file can be modified synchronously by editing the text. This avoids the need for users to re-record when they discover audio errors, and allows for synchronous modification of the audio file through text editing, thus improving the user experience.

[0147] In summary, in step S302, this embodiment of the application supports displaying the first voice file recorded by the user in the input box (i.e., the information input area) and supports audio preview of the first voice file in the input box, thereby meeting the user's flexibility needs in the information input process.

[0148] S303, In response to the output operation of the first voice file displayed in the information input area, the second voice file is displayed in the information interaction interface.

[0149] Specifically, in response to the output operation of the first voice file displayed in the information input area, the first voice file displayed in the information input area is sent to the conversation area for display; further, in response to the first voice file being output in the conversation area, the response data of the interactive object can be output in the conversation area of ​​the information interaction interface, and the response data includes at least one of the following: response voice file, response text, response image, and response video.

[0150] The following section will first provide a detailed explanation of how the first audio file is output in the conversation area.

[0151] In one possible implementation, the first voice file of the first object is output in the session area of ​​the information interaction interface, specifically including any of the following methods: (1) Outputting the first voice file of the first object in the session area of ​​the information interaction interface; see [link to relevant documentation]. Figure 7a , Figure 7a This is a schematic diagram of a process for outputting a first audio file provided in an embodiment of this application, such as... Figure 7a As shown, when the user clicks the send control in the information input area, the first voice file 7001 will be displayed in the conversation area of ​​the information interaction interface.

[0152] (2) Output the first audio file of the first object in the conversation area of ​​the information interaction interface, and display the text information obtained after the audio data of the first object has been converted into text; wherein, the display method between the text information and the first audio file includes, but is not limited to, any of the following methods: Method 1: Display all the text content of the text information in the information interaction interface while playing the first audio file; Method 2: Simultaneously display a portion of the text content corresponding to the played content in the first audio file while playing the first audio file; Method 3: Differentiately display a portion of the text content corresponding to the played content in the first audio file while playing the first audio file; wherein, differentiated display includes, any one of: highlighting, font differentiation, color differentiation, and animation display. Please refer to Figure 7b , Figure 7b This is a schematic diagram of another process for outputting the first audio file provided in an embodiment of this application, such as... Figure 7b As shown in Figure 7002, the first audio file can be displayed in the conversation area of ​​the information interaction interface, and the full text content of the text information obtained after the first audio file is converted into text can be displayed (e.g., what characteristics does a potato have).

[0153] (3) Output the first voice file of the first object in the conversation area of ​​the information interaction interface, and display the text information obtained by converting the voice data of the first object into text, as well as display the pattern code associated with the first voice file; wherein, the display position of the pattern code may include any of the following: an associated position around the first voice file (e.g., the pattern code is displayed in association after the display position of the first voice file), any position in the information interaction interface (e.g., the middle position of the interface), or a preset position in the information interaction interface (e.g., the upper right corner window position of the interface). Please refer to Figure 7c , Figure 7c This is a schematic diagram of another process for outputting the first audio file provided in an embodiment of this application, such as... Figure 7cAs shown in 7003, the first voice file can be displayed in the conversation area of ​​the information interaction interface, and the text information obtained after the first voice file is converted into text can be displayed (e.g., what characteristics a potato has), as well as the pattern code (such as a QR code) associated with the first voice file can be displayed.

[0154] Optionally, the QR code of the first audio file can be scanned and played by other clients. The specific process is as follows: An information interaction interface is displayed on the first client, and the first audio file is associated with a pattern control; in response to a trigger operation on the pattern control associated with the first audio file, the pattern code associated with the first audio file is displayed. The display method of the pattern code includes: magnified display, flashing display (e.g., flashing the border of the pattern code), and animated display (e.g., updating the pattern code every 1 minute); in response to a scanning operation of the pattern code by the second client, the first audio file is output to the second client. Please refer to [link to relevant documentation]. Figure 8 , Figure 8 This is a schematic diagram illustrating a scanning process for a first audio file provided in an embodiment of this application. Figure 8 As shown, a first client (e.g., a web client) displays an information interaction interface, and the conversation area of ​​the information interaction interface displays a first audio file and a pattern control 8001 associated with the first audio file. In response to a trigger operation on the pattern control 8001, a pattern code 8002 associated with the first audio file is displayed, wherein the pattern code 8002 of the first audio file is displayed in the information interaction interface S801 shown in the first client. Further, in response to a second client (e.g., a mobile phone) scanning the pattern code 8002 displayed in the information interaction interface S801, a file playback page S802 is displayed in the second client; wherein the second client displays the first audio file and supports playback of the first audio file in the second client. In this implementation, for users who find it inconvenient to listen to audio on a computer, the corresponding audio file can be displayed on the mobile phone by scanning a QR code, thereby supporting playback of the audio file on the mobile phone and improving the convenience of the audio playback process.

[0155] In one possible implementation, the first audio file displayed on the information interaction interface is used to control the playback of the audio data of the first object; the following is an example of how to control the playback process of the audio data. The specific process is as follows: In response to a playback operation on the first audio file, the audio data of the first object is played; during the playback of the audio data of the first object, in response to a playback adjustment operation on the first audio file, the playback attributes of the audio data of the first object are adjusted; wherein, the playback attributes include any one or more of the following: playback progress, playback speed, playback timbre, and playback volume. Please refer to... Figure 9 , Figure 9This is a schematic diagram illustrating a process for controlling the playback of voice data provided in an embodiment of this application. For example, see... Figure 9 As shown in the information interaction interface S901, this interface displays a first audio file. The first audio file is used to control the playback of audio data from a first object, and includes a progress bar, volume controls, icon controls, etc. The user (such as the first object) can drag the progress bar 9011 of the first audio file to trigger a playback adjustment operation, thereby adjusting the playback progress of the audio data. See also... Figure 9 As shown in the information interaction interface S902, the information interaction interface S902 displays a first audio file. The first object can click the volume control 9021 to display a volume bar. When the user adjusts the volume bar, it triggers a playback adjustment operation for an audio file, thereby adjusting the playback volume of the audio data.

[0156] In the above process, after the first voice file of the first object is output in the conversation area of ​​the information interaction interface, the first object can customize the playback attributes of the first object's voice data based on the first voice file, such as playback progress, playback speed, playback timbre, playback volume, etc., so as to meet the user's flexible needs for voice playback during the conversation.

[0157] Furthermore, in response to the output of the first audio file in the conversation area, reply data can be output in the conversation area of ​​the information interaction interface. This reply data includes any one or more of the following: reply audio file, reply text, reply image, and reply video. Specifically, the reply audio file is generated based on the large language model's response to the audio data. The reply data refers to the data generated by the model object after processing the first object's audio data during the information interaction between the interactive object (such as the model object corresponding to the large language model) and the user (such as the first object). The output method of the reply data is described in detail below.

[0158] In one possible implementation, the response data of the interactive object is output in the session area of ​​the information interaction interface, specifically including any one of the following three methods:

[0159] (1) Output the reply voice file in the conversation area of ​​the information interaction interface; please refer to Figure 10a , Figure 10a This is a schematic diagram of an interface for outputting a response voice file, provided in an embodiment of this application. For example... Figure 10a As shown, the reply voice file 1001 is displayed in the conversation area of ​​the information interaction interface.

[0160] (2) Output the reply audio file in the conversation area of ​​the information interaction interface and display the corresponding reply text. The display methods between the reply audio file and the reply text include, but are not limited to, any of the following: Method 1: Display the entire text content of the reply in the information interaction interface while the reply audio file is playing; Method 2: Simultaneously display a portion of the text content corresponding to the played content in the reply audio file while the reply audio file is playing; Method 3: Differentiatedly display a portion of the text content corresponding to the played content in the reply audio file while the reply audio file is playing; Differentiated display includes any one of the following: highlighting, font differentiation, color differentiation, and animation display. Please refer to [link to relevant documentation]. Figure 10b , Figure 10b This is a schematic diagram of another interface for outputting and replying to voice files provided in an embodiment of this application. For example... Figure 10b As shown in Figure 1002, the reply voice file is displayed in the conversation area of ​​the information interaction interface, and the reply text obtained after the corresponding text conversion is also displayed.

[0161] (3) Output the response voice file of the interactive object in the conversation area of ​​the information interaction interface, and display the response text obtained after the response data has been converted to text, as well as the pattern code associated with the response voice file. The display position of the pattern code can include any of the following: an associated position around the response voice file (e.g., the pattern code is displayed after the display position of the response voice file), any position in the information interaction interface (e.g., the center of the interface), or a preset position in the information interaction interface (e.g., the upper right corner of the window in the interface); and the display method of the pattern code associated with the response voice file includes any of the following: magnified display, flashing display (e.g., flashing the border of the pattern code), or animated display (e.g., updating the pattern code every 1 minute). Please refer to [link to relevant documentation]. Figure 10c , Figure 10c This is a schematic diagram of another interface for outputting and replying to voice files provided in an embodiment of this application. For example... Figure 10c As shown in Figure 1003, the reply voice file can be displayed in the conversation area of ​​the information interaction interface, along with the reply text obtained after the reply voice file has been converted to text, and the pattern code (such as a QR code) associated with the reply voice file.

[0162] Based on this, this application provides diverse output methods for the voice response files of the interactive object, such as voice + text output, or voice + text + QR code output, etc., thereby enriching the information interaction methods between users and models and improving the user experience in the information interaction process.

[0163] Optionally, when outputting an audio file and reply text, the display method between the reply audio file and reply text includes any of the following: (1) During the playback of the reply audio file, display the entire text content of the reply text in the information interaction interface; in this case, it is supported to directly display the entire text content of the reply audio file in the conversation area. (2) During the playback of the reply audio file, synchronously display a portion of the text content corresponding to the content being played in the reply audio file; please refer to Figure 11a , Figure 11a This is a schematic diagram of an interface for outputting reply voice files and reply text, provided in an embodiment of this application. For example... Figure 11a As shown, the conversation area of ​​the information interaction interface displays a reply voice file, and during the playback of the reply voice file, the text content 1101 corresponding to the played voice is displayed. (3) During the playback of the reply voice file, the text content corresponding to the played content in the reply voice file is displayed separately; wherein, the above-mentioned separate display includes any one or more of the following: highlighting, font differentiation, color differentiation, and animation display. Please refer to Figure 11b , Figure 11b This is a schematic diagram of another interface for outputting reply voice files and reply text provided in an embodiment of this application. For example... Figure 11b As shown, the conversation area of ​​the information interaction interface displays the reply voice file. During the playback of the reply voice file, it supports the differentiated display of the text content corresponding to the played voice. For example, in text content 1102, the text being played at this time is "(The leaves are pinnately compound leaves with an odd number of unequal digits, the flowers are white or blue-purple, and the berries are spherical)", then the above text content can be bolded and underlined. In this implementation, it supports displaying the model's reply content through a combination of voice and text, and provides multiple display methods between the two, which enriches the information display methods during the information interaction process.

[0164] Based on this, the conversation area supports the simultaneous output of the interactive object's audio file and the reply text. Users can see the text information generated by the model and also directly play the audio file to hear the reply content, realizing a multimodal interaction method combining audio and video. This approach can better meet the needs of different users in different scenarios. For example, when vision is limited or when focusing on other tasks, users can obtain information by listening to the audio file; while when they need to view or refer to the reply content in detail, they can refer to the text information, thereby improving the accuracy and effectiveness of information transmission and enhancing the interactive experience between users and the large model.

[0165] In one possible implementation, the voice file can be shared with other users. Specifically, the sharing of the reply voice file can be done in any of the following ways: (1) In response to a sharing operation on the reply voice file, a sharing link is generated and shared to a second object; see [link to relevant documentation]. Figure 12a , Figure 12a This is a schematic diagram of a process for sharing and replying to voice files provided in an embodiment of this application, such as... Figure 12a As shown, the conversation area of ​​the information interaction interface displays the reply voice file, and the information interaction interface has a sharing control 1201 for the reply voice file. In response to the sharing operation of the reply voice file (e.g., the first object clicks the sharing control 1201), the social object list of the first object is output. The social object list includes at least one user (e.g., user 1, user 2, user 3) who has a social relationship with the first object. When the first object selects user 2 from the social object list, a sharing link for the reply voice file can be generated and sent to user 2 (i.e., the second object). (2) In response to the sharing operation of the reply voice file, the pattern code of the reply voice file is shared to the second object. The pattern code has an expiration date, and the second object can play the reply voice file by scanning the pattern code within the expiration date. Please refer to Figure 12b , Figure 12b This is a schematic diagram of another process for sharing and replying to voice files provided in an embodiment of this application, such as... Figure 12b As shown, the conversation area of ​​this information interaction interface displays a reply voice file and an associated pattern code (such as a QR code). In response to a sharing operation of the pattern code, the pattern code of the reply voice file is shared with a second party. This pattern code has an expiration period (e.g., 10 minutes), allowing the second party to scan the pattern code within 10 minutes to obtain and play the reply voice file. Similarly, the first voice file displayed in the conversation area of ​​the information interaction interface also supports sharing, which will not be elaborated further here. With this implementation, voice files can be shared with other users via QR codes or sharing links, improving the convenience of information sharing.

[0166] In one possible implementation, the session content between the first object and the interaction object can be exported for model analysis. The specific process is as follows: In response to the session export operation on the task list, a candidate task list to be exported is obtained, which includes at least one created historical session task from the task display area; the candidate task list is output, where each historical session task in the candidate task list is used for model analysis of the large language model; wherein, the model analysis includes any one or more of the following: performance analysis, efficiency analysis, indicator analysis, and accuracy analysis.

[0167] Please see Figure 13 , Figure 13 This is a schematic diagram of an interface for exporting session content provided in an embodiment of this application. For example... Figure 13 As shown, the information interaction interface includes a conversation export control 1301. When this control is triggered, a conversation task selection interface is displayed. This interface shows at least one created historical conversation task, where each historical conversation task refers to a conversation task between the user (e.g., the first object) and the interaction object where dialogue has already taken place. The conversation task selection interface allows users to sort historical conversation tasks sequentially according to different test points (one test point corresponds to one knowledge domain, such as Chinese, mathematics, law, architecture, etc.), and supports users selecting a list of candidate tasks to export (e.g., selecting all or a portion of historical conversation tasks). Furthermore, in response to a conversation export operation on the candidate task list, the list is exported. Optionally, the exported historical conversation tasks can be used for model analysis of the large language model, such as analyzing the accuracy of the large language model's responses during dialogue interaction, analyzing the large language model's response efficiency during dialogue interaction, etc., to facilitate performance analysis and evaluation of the large language model, thereby improving its performance. This implementation supports exporting the conversation content between the user and the interactive object, and evaluates and analyzes the performance of the large language model (such as the accuracy and efficiency of the response content) based on the exported historical conversation tasks; thereby optimizing the information interaction capability of the large language model and improving its performance.

[0168] This application provides an interactive scheme that supports voice dialogue between a user and a large language model. On one hand, this scheme supports displaying a first voice file input by a first object in the information input area (input box) of the information interaction interface, and supports voice preview of the user's input voice data through the voice file in the input box. This helps the user check whether the input data meets expectations, thus satisfying the user's requirement for the reliability of the input data. On the other hand, the first voice file or reply voice file output in the information interaction interface supports customized adjustment of playback attributes (e.g., playback progress, playback speed, playback volume, playback timbre, etc.), thus satisfying the user's flexibility requirements during voice playback. Furthermore, in response to the output operation of the first voice file displayed in the information input area, the reply voice file of the interactive object can be output in the information interaction interface. It can be seen that in the human-computer interaction process between the user and the model, the scheme supports outputting the model's reply in the form of a voice file in the information interaction interface. Compared with outputting the model's reply in text form, this application can realize voice interaction between the user and the model, and can simulate the dialogue interaction effect between people as much as possible during the human-computer interaction process, thereby improving the realism and flexibility of information interaction.

[0169] The embodiments of the information interaction method provided in this application will be described in detail below from a technical perspective.

[0170] Please see Figure 14 , Figure 14 This is a flowchart illustrating an information interaction method provided in an embodiment of this application. This information interaction method can be implemented using a computer device (such as...). Figure 1 The server shown will execute this. Figure 14 As shown, the information interaction method includes the following steps S1401-S1403:

[0171] S1401. Receive an information interaction request sent by the client. The information interaction request is used to request a response to the first voice file. The first voice file is generated in response to the information interaction operation of the first object.

[0172] Specifically, the client here refers to a client running on a terminal device (such as a computer, mobile phone, etc.). This client displays an information interaction interface used to interact with model objects of the large language model, such as voice interaction, text interaction, voice + text interaction, or any other interaction method. Furthermore, when the client detects an information interaction operation by a user (such as the first object), it can obtain the first object's first voice file, and the client can generate an information interaction request based on the first voice file and send the request to the server. It should be noted that the detailed process of generating or outputting the first voice file in the client can be found in [reference needed]. Figure 3 The detailed description of step S301 in the embodiments will not be repeated here.

[0173] In one possible implementation, after receiving an information interaction request from the client, the server can obtain the identity information of the first object and perform verification processing on the first object based on the identity information. This verification processing includes one or more of the following: permission verification (e.g., verifying whether the first object has permission to initiate information interaction with the interaction object), legality verification (e.g., verifying whether the content of the first object's first voice file is legal), and security verification (e.g., verifying whether the first object's identity information is secure). Further, if the first object's identity verification is successful, the server responds to the information interaction request and triggers the execution of subsequent step S1402; optionally, if the first object's identity verification fails, the information interaction request is deleted. This implementation supports verification of the first object's identity information, thereby improving information security during the information interaction process.

[0174] S1402. Call the large language model to process the speech data in the first speech file, generate response data, and generate a response speech file based on the response data.

[0175] In one possible implementation, the server can respond to an information interaction request and parse the request to obtain a first speech file of the first object. Further, the server invokes a large language model to process the speech data in the first speech file, generating response data. The specific process is as follows:

[0176] (1) Automatic speech recognition technology (ASR) is used to convert the speech data in the first speech file into text to obtain the text information corresponding to the first speech file. Since large language models usually have NLP (natural language processing) capabilities, that is, the ability to process text, ASR technology can convert the user's language data into text information so that the model can understand and process it.

[0177] (2) Call the large language model to process the text information and obtain the response data; where the large language model here can be a model with information response function (such as understanding information and responding), such as: LLM model, MLLM model (Multimodal LLM, multimodal large language model), etc.

[0178] (3) Use text-to-speech technology (such as TTS technology) to process the reply data into speech and obtain the reply speech file corresponding to the reply data.

[0179] In the above processing, the user's voice data is first converted into text information using ASR technology, and a large language model is called to generate response data from the text information. Then, the response data undergoes TTS (Text-to-Speech) processing to generate the response audio file. This process aims to reproduce the voice-to-voice dialogue interaction between the user and the interacting object as accurately as possible.

[0180] S1403. Return the reply voice file to the client.

[0181] Specifically, the server can directly return the reply voice file to the client, allowing the client to output the reply voice file in the information interaction interface; alternatively, the server can encrypt the reply voice file and return the encrypted reply voice file to the client, and the client can output the reply voice file in the information interaction interface only after the file is decrypted, thereby improving information security during the information interaction process.

[0182] In summary, based on Figure 3 The product implementation process shown in the front end, and Figure 14The backend technical implementation process shown comprehensively realizes the complete process of information interaction (such as voice dialogue interaction) between the user and the large language model provided in this application. The following is a summary and explanation of the backend process implementation of the core technical points involved in the above information interaction process, specifically including the following key technical points:

[0183] (1) Integration of the input box (i.e., the information input area) with the voice acquisition control:

[0184] ① Design an input box at a designated location on the information interaction interface for users (such as the first object) to input text, voice, images, videos, or trigger other information interaction operations.

[0185] ② Integrate a voice capture control closely next to the input box. This control has a clear visual identifier (such as a microphone icon) so that users can easily identify it and click the control to trigger the start of recording.

[0186] ③ When the user clicks the trigger control, the Web Audio API (a web-based audio access interface) is called to detect and request the user to grant microphone access permissions.

[0187] ④ If the user agrees to grant the permission, the system will obtain the hardware access required for recording; if the user refuses, the system will prompt that the user does not have sufficient permissions and the recording function is unavailable.

[0188] ⑤ After successfully obtaining permissions, the system starts the recording function, and collects and records the user's voice data input in real time through the voice acquisition control.

[0189] (2) Recording function implementation:

[0190] ① When recording begins, the getUserMedia method of the Web Audio API is called via JavaScript to request access to the user's microphone, and the display of the voice capture control is changed to a fluctuating icon to give the user an intuitive visual sense that recording is in progress.

[0191] ② After obtaining microphone access permissions, create a Media Stream Audio Source Node (an audio interface), connect it to a Media Stream Destination Node, and start recording audio.

[0192] ③ The information interaction interface displays a progress bar that changes over time, making it easy for users to understand the recording duration.

[0193] (3) Complete the recording:

[0194] ① When the user moves the mouse over the voice capture control that is recording, a message "Stop voice input" will be displayed, and the background of the control will be grayed out to indicate that the user can end the recording by clicking the voice capture control again.

[0195] ②When the user clicks the stop voice capture control or the preset recording time limit is reached, the recording stops and the audio file (such as the first voice file) is saved.

[0196] ③ After successful recording, the audio is displayed using the HTML audio (Hypertext Markup Language) tag. It provides playback, pause, volume adjustment, and progress dragging functions to help users manage their newly recorded audio, and offers a deletion function to allow users to delete and re-record audio that does not meet their expectations or satisfaction.

[0197] (4) Audio file upload:

[0198] ① After the user (e.g., the first object) confirms that the recorded first voice file is correct, they can click the send button in the input box to transfer the first voice file to the server. Specifically, a Blob object (an immutable, raw data-like file object) can be used to read the first voice file as binary data, and the file data of the first voice file can be uploaded to the server via a POST request (a common request method in the HTTP protocol) through the Fetch API (a data retrieval interface).

[0199] ② During the upload of the first audio file, an animation can be displayed in the information interaction interface, and secondary input by the user is prohibited during this time. Optionally, the upload progress of the audio file can be displayed during the file upload process. If the upload fails, the user is informed of the reason for the failure and the restriction is lifted, allowing the user to record a new audio file and upload it again. Furthermore, if the file upload is successful, the backend returns a link to the successfully uploaded file for display in the conversation area of ​​the information interaction interface (e.g., the first audio file is displayed as...). Figure 5b (See file link shown in section 503). The first voice file sent by the user will be displayed on the right side of the conversation area window, representing the user's input in the dialogue; while the reply voice file from the interacting object will be displayed on the left side of the conversation area window (e.g., ...). Figure 10a (As shown).

[0200] (5) Voice playback and QR code generation:

[0201] In this application, for the first voice file that has been successfully sent, a QR code icon (such as...) can be provided in the information interaction interface. Figure 12bThe pattern control 1200 shown serves as a preview button for the QR code. When the mouse hovers over the QR code icon, the backend can draw a QR code image (pattern code) based on the link to the audio file (such as the first audio file or the reply audio file). Subsequently, it will support scanning the corresponding QR code with a mobile phone to trigger the playback of the audio file on the phone. In this way, for users who find it inconvenient to listen to audio on a computer, scanning the code with their mobile phone and playing the audio on their mobile device is more convenient.

[0202] (6) Calling the large language model:

[0203] ① The backend uses Automatic Speech Recognition (ASR) technology to preprocess the first voice file of the first object, and converts the first voice file uploaded by the user into text information after recognition; further, it calls a large language model to reply to the above text information and generate reply data.

[0204] ② Based on the response data (text format) returned by the large language model, TTS text-to-speech technology is used to convert the response data into audio files and return them to the front end.

[0205] (7) Display of the response content of the interactive object:

[0206] Specifically, the response content of the interactive object is displayed in two areas, upper and lower (e.g., Figure 10b As shown in 1002); the upper part displays the link to the audio file obtained by TTS text-to-speech technology from the content returned by the large language model; the lower part displays the original text content (i.e., the response data) returned by the large language model. This application allows users to customize the playback attributes of the audio file by playing, pausing, modifying the progress, adjusting the volume, etc., and also allows users to play the audio file on their mobile phones by scanning a QR code.

[0207] In summary, through the implementation of the above key technologies, this application provides a dialogue interaction scheme between a user and a large language model. Please refer to... Figure 15 , Figure 15 This is a technical flowchart of a dialogue interaction scheme provided in an embodiment of this application. Figure 15 As shown, the technical flowchart of this dialogue interaction scheme includes the following steps S1501-S1510:

[0208] S1501, Begin.

[0209] S1502. Real-time voice recording using the Web Audio API.

[0210] S1503. Process and save the first audio file.

[0211] S1504, Play and preview the first audio file.

[0212] S1505: The user confirms whether the first audio file is correct. If yes, proceed to S1506; otherwise, proceed to S1510.

[0213] S1506. Upload the first audio file to the server for processing.

[0214] S1507, Large Language Model Analysis and Text Processing.

[0215] S1508 and TTS technology convert reply data into reply voice files.

[0216] S1509. Output the reply voice file and generate a QR code.

[0217] S1510, Re-record the audio file.

[0218] This application provides a dialogue interaction scheme between a user and a large language model. This scheme integrates advanced speech processing technologies (such as ASR and TTS) with powerful large model capabilities to establish a novel voice interaction mode on the PC. Specifically, when a user enables the voice interaction function on the PC, they can directly input voice commands or questions through the microphone. The system transmits the voice information to the large language model in real time. The large language model recognizes, understands, and analyzes the speech, and generates corresponding voice responses. Thus, the entire information interaction process eliminates the need for manual text input by the user, achieving true voice-to-voice interaction. This interactive experience is more natural and efficient, allowing users to communicate with the large language model as if conversing with a human, greatly improving the speed and convenience of information exchange, thereby enhancing the user's experience of interacting with the model.

[0219] The following describes the relevant devices of the information interaction scheme provided in the embodiments of this application.

[0220] It should be noted that, in the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program with a predetermined function, which works together with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0221] Please see Figure 16 , Figure 16This is a schematic diagram of the structure of an information interaction device provided in an embodiment of this application. The information interaction device 1600 can be used to perform the functions described in this application. Figure 3 The corresponding steps in the information interaction method provided in the embodiment. Specifically, the information interaction device 1600 may include:

[0222] Display unit 1601 is used to display an information interaction interface, which is used to interact with an interactive object and includes an information input area.

[0223] Processing unit 1602 is configured to, in response to an information interaction operation in the information interaction interface, display a first voice file in the information input area; the first voice file is generated based on the voice data input by the first object;

[0224] The processing unit 1602 is also configured to, in response to an output operation of a first voice file displayed in the information input area, display a second voice file in the information interaction interface, the second voice file consisting of the first voice file and response data for responding to the voice data.

[0225] In one possible implementation, the information interaction operation is a voice input operation performed by the first object on the information interaction interface; or, the first object performs a voice input operation by triggering a voice interaction entry on the information interaction interface.

[0226] In one possible implementation, the information interaction interface includes a voice interaction entry point; the processing unit 1602, in response to an information interaction operation in the information interaction interface, displays a first voice file in the information input area to perform the following operations:

[0227] In response to a trigger operation on the voice interaction entry point on the information interaction interface, voice data of the first object is collected, and the first voice file corresponding to the voice data is displayed in the information input area; or,

[0228] In response to a trigger operation on the information input control set in the information interaction interface, a file selection list is output, which displays at least one audio file; in response to a selection operation on a target audio file among the at least one audio file, the target audio file is displayed as the first audio file in the information input area.

[0229] In one possible implementation, the information interaction interface further includes a session area, which is used to display the interaction information between the first object and the interaction object; the processing unit 1602 is also used to perform the following operations:

[0230] In response to a preview operation of the first audio file displayed in the information input area, a voice preview is performed on the audio data corresponding to the first audio file;

[0231] During the voice preview process, in response to the output operation of the first voice file, the first voice file is sent from the information input area to the conversation area for display; or, in response to the deletion operation of the first voice file, the first voice file is deleted from the information input area.

[0232] During the voice preview process, the playback attributes of the voice data can be adjusted, including at least one of the following: playback progress, playback speed, playback volume, and playback timbre.

[0233] In one possible implementation, the voice interaction entry point is a voice acquisition control, which is displayed on the information interaction interface in a first display form; the processing unit 1602 is also used to perform the following operations:

[0234] In response to the first object's trigger operation on the voice acquisition control, the display mode of the voice acquisition control is adjusted from the first display mode to the second display mode;

[0235] When voice data is detected from the first object, the display mode of the voice acquisition control is adjusted from the second display mode to the third display mode;

[0236] In response to the voice data acquisition completion event, the display mode of the voice acquisition control is adjusted from the third display mode to the fourth display mode; the acquisition completion event includes: the voice data acquisition completion operation is performed, or the voice data acquisition duration reaches the preset duration threshold;

[0237] Among them, the first display form, the second display form, the third display form, and the fourth display form are all different from each other; and the display form includes any one or more of the following: color, shape, and size.

[0238] In one possible implementation, after responding to the output operation of the first voice file displayed for the information input area, the processing unit 1602 is further configured to perform the following operations:

[0239] In response to a playback operation on the first audio file, the audio data corresponding to the first audio file is played in the session area;

[0240] During the playback of audio data, in response to the playback adjustment operation of the first audio file, the playback attributes of the audio data are adjusted;

[0241] The playback attributes include any one or more of the following: playback progress, playback speed, playback timbre, and playback volume.

[0242] In one possible implementation, the processing unit 1602 is further configured to perform the following operations:

[0243] In response to the playback operation of the first audio file, the audio data corresponding to the first audio file is played, and the text information of the first audio file is displayed;

[0244] During the playback of the first object's audio data, in response to the editing operation of the text information, the audio data corresponding to the first audio file being played is adjusted.

[0245] The adjustment refers to: adjusting the voice data of the first object to be played to the voice content corresponding to the edited text information.

[0246] In one possible implementation, the processing unit 1602 outputs the first voice file of the first object in the session area of ​​the information interaction interface to perform the following operations:

[0247] Output the first voice file of the first object in the conversation area of ​​the information interaction interface; or,

[0248] Output the first audio file of the first object in the conversation area of ​​the information interaction interface, and display the text information obtained by converting the audio data of the first object into text; or,

[0249] Output the first voice file of the first object in the conversation area of ​​the information interaction interface, and display the text information obtained after the voice data of the first object is converted into text, as well as the pattern code associated with the first voice file.

[0250] In one possible implementation, the information interaction interface is displayed on the first client, and the first voice file is associated with a pattern control; the processing unit 1602 is also used to perform the following operations:

[0251] In response to a trigger operation on the pattern control associated with the first audio file, the pattern code associated with the first audio file is displayed;

[0252] In response to the second client's scanning operation of the pattern code, the first voice file is output to the second client; wherein the first client and the second client are different types of clients.

[0253] In one possible implementation, the information interaction interface further includes a session area; the processing unit 1602 displays a second voice file on the information interaction interface for performing the following operations:

[0254] Output the first audio file in the session area; and,

[0255] In response to the output of the first audio file, the response data generated by the interactive object is output in the conversation area; wherein the response data includes at least one of the following: response audio file, response text corresponding to the response audio file, response image, and response video.

[0256] In one possible implementation, the response data includes a response voice file and a response text; the display method between the response voice file and the response text includes any of the following:

[0257] During the playback of the reply audio file, the full text content of the reply is displayed in the information interaction interface;

[0258] While the reply audio file is playing, the corresponding text content of the content played in the reply audio file is displayed simultaneously.

[0259] During the playback of the reply audio file, the text content corresponding to the played content in the reply audio file is displayed differently; the different display includes any one of the following: highlighting, font differentiation, color differentiation, and animation display.

[0260] In one possible implementation, the processing unit 1602 is further configured to perform the following operations:

[0261] In response to a sharing operation on the reply audio file, a share link is generated and shared to a second object; or...

[0262] In response to the sharing operation of the reply voice file, the pattern code of the reply voice file is shared to the second object; wherein, the pattern code has an expiration period, and the second object can play the reply voice file by scanning the pattern code within the expiration period.

[0263] In one possible implementation, the interaction object is the model object corresponding to the large language model; the information interaction interface includes a task display area and a conversation area, the task display area displays at least one task list, and the task list displays at least one historical conversation task that has been created; the conversation area is used for information interaction with the model object corresponding to the large language model during the execution of a conversation task; the processing unit 1602 is also used to perform the following operations:

[0264] In response to a session creation operation, the newly created session task is displayed, and information interaction under the new session task is performed with the model object corresponding to the large language model in the session area of ​​the new session task; or...

[0265] In response to the selection operation for the target historical session task, the selected target historical session task is displayed, and information interaction under the target historical session task is performed with the model object corresponding to the large language model in the session area of ​​the target historical session task.

[0266] In one possible implementation, the processing unit 1602 is further configured to perform the following operations:

[0267] In response to a session export operation on the task list, a candidate task list to be exported is obtained, which includes at least one created historical session task in the task display area.

[0268] Output a list of candidate tasks; each historical conversation task in the candidate task list is used to perform model analysis on the large language model; the model analysis includes any one or more of the following: performance analysis, efficiency analysis, indicator analysis, and accuracy analysis.

[0269] In one possible implementation, the processing unit 1602 is further configured to perform the following operations:

[0270] Obtain the first audio file of the first object;

[0271] The large language model is invoked to process the speech data in the first speech file and generate response data.

[0272] Generate a response audio file based on the response data.

[0273] In one possible implementation, the processing unit 1602 invokes a large language model to process the speech data in the first speech file, generating response data for performing the following operations:

[0274] Automatic speech recognition technology is used to convert the speech data in the first speech file into text, thereby obtaining the text information corresponding to the first speech file.

[0275] The large language model is invoked to process the text information and obtain the response data;

[0276] Text-to-speech technology is used to convert the response data into speech, resulting in the corresponding audio file.

[0277] This application provides an interactive scheme that supports voice dialogue between users and a large language model. On one hand, this scheme supports displaying a first voice file input by a first object in the information input area (input box) of the information interaction interface, and supports voice preview of the user's input voice data via the voice file in the input box. This allows users to check whether the input data meets expectations, thus satisfying their need for reliable input data. On the other hand, both the first voice file and the reply voice file output in the information interaction interface support customized adjustment of playback attributes (e.g., playback progress, playback speed, playback volume, playback timbre, etc.), thereby satisfying users' flexibility needs during voice playback. Furthermore, in response to the output operation of the first voice file displayed in the information input area, the reply voice file of the interactive object can be output in the information interaction interface. It can be seen that during the human-computer interaction between the user and the model, the scheme supports outputting the model's reply via a voice file in the information interaction interface. Compared to outputting the model's reply via text, this application can achieve voice interaction between the user and the model, simulating the dialogue interaction effect between people as much as possible during the human-computer interaction process, thereby improving the realism and flexibility of information interaction.

[0278] Please see Figure 17 , Figure 17 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device 1700 is used to execute the steps performed by the computer device in the aforementioned method embodiments. The computer device 1700 may include one or more devices (such as servers, nodes, terminal devices, etc.) or internal components (such as chips, software modules, or hardware modules). The computer device may include at least one processor 1701 and a communication interface 1702. Further optionally, the computer device may also include at least one memory 1703 and a bus 1704. Additionally, the processor 1701, communication interface 1702, and memory 1703 are connected via the bus 1704.

[0279] (1) The processor 1701 is a module that performs arithmetic and / or logical operations. Specifically, it may be one or a combination of processing modules such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor unit (MPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a complex programmable logic device (CPLD), a coprocessor (to assist the central processing unit in completing corresponding processing and applications), and a micro controller unit (MCU).

[0280] (2) Communication interface 1702 can be used to provide information input or output to at least one processor 1701. And / or, communication interface 1702 can be used to receive data sent externally and / or send data externally, and can be a wired link interface including such as an Ethernet cable, or a wireless link interface (Wi-Fi, Bluetooth, general wireless transmission, vehicle short-range communication technology, and other short-range wireless communication technologies, etc.). Communication interface 1702 can serve as a network interface.

[0281] (3) The memory 1703 is used to provide storage space, in which data such as the operating system and computer programs (including program instructions) can be stored. The memory 1703 can be one or a combination of random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), etc.

[0282] In specific implementation, the processor 1701 executes the following steps by running the computer program stored in the memory 1703:

[0283] The information interaction interface is displayed and used to interact with the interaction object. The information interaction interface includes an information input area.

[0284] In response to an information interaction operation in the information interaction interface, a first voice file is displayed in the information input area; the first voice file is generated based on the voice data input by the first object;

[0285] In response to the output operation of the first voice file displayed in the information input area, a second voice file is displayed in the information interaction interface. The second voice file consists of the first voice file and response data that responds to the voice data.

[0286] In one possible implementation, the information interaction operation is a voice input operation performed by the first object on the information interaction interface; or, the first object performs a voice input operation by triggering a voice interaction entry on the information interaction interface.

[0287] In one possible implementation, the information interaction interface includes a voice interaction entry point; in response to an information interaction operation in the information interaction interface, the processor 1701 displays a first voice file in the information input area for performing the following operations:

[0288] In response to a trigger operation on the voice interaction entry point on the information interaction interface, voice data of the first object is collected, and the first voice file corresponding to the voice data is displayed in the information input area; or,

[0289] In response to a trigger operation on the information input control set in the information interaction interface, a file selection list is output, which displays at least one audio file; in response to a selection operation on a target audio file among the at least one audio file, the target audio file is displayed as the first audio file in the information input area.

[0290] In one possible implementation, the information interaction interface further includes a session area, which is used to display the interaction information generated by the information interaction between the first object and the interaction object; the processing unit 1602 is also used to perform the following operations:

[0291] In response to a preview operation of the first audio file displayed in the information input area, a voice preview of the audio data of the first object is performed;

[0292] During the preview of the voice data of the first object, in response to the output operation of the first voice file, the first voice file is sent from the information input area to the conversation area for display; or, in response to the deletion operation of the first voice file, the first voice file is deleted from the information input area.

[0293] During the voice preview process, the playback attributes of the voice data can be adjusted, including at least one of the following: playback progress, playback speed, playback volume, and playback timbre.

[0294] In one possible implementation, the voice interaction entry point is a voice acquisition control, which is displayed on the information interaction interface in a first display form; the processor 1701 is also used to perform the following operations:

[0295] In response to the first object's trigger operation on the voice acquisition control, the display mode of the voice acquisition control is adjusted from the first display mode to the second display mode;

[0296] When voice data is detected from the first object, the display mode of the voice acquisition control is adjusted from the second display mode to the third display mode;

[0297] In response to the voice data acquisition completion event, the display mode of the voice acquisition control is adjusted from the third display mode to the fourth display mode; the acquisition completion event includes: the voice data acquisition completion operation is performed, or the voice data acquisition duration reaches the preset duration threshold;

[0298] Among them, the first display form, the second display form, the third display form, and the fourth display form are all different from each other; and the display form includes any one or more of the following: color, shape, and size.

[0299] In one possible implementation, after responding to the output operation of the first voice file displayed for the information input area, the processing unit 1602 is further configured to perform the following operations:

[0300] In response to a playback operation on the first audio file, the audio data corresponding to the first audio file is played in the session area;

[0301] During the playback of audio data, in response to the playback adjustment operation of the first audio file, the playback attributes of the audio data of the first object are adjusted;

[0302] The playback attributes include any one or more of the following: playback progress, playback speed, playback timbre, and playback volume.

[0303] In one possible implementation, processor 1701 is also used to perform the following operations:

[0304] In response to the playback operation of the first audio file, the audio data corresponding to the first audio file is played, and the text information of the first audio file is displayed;

[0305] During the playback of the first object's audio data, in response to the editing operation of the text information, the audio data corresponding to the first audio file being played is adjusted.

[0306] The adjustment refers to: adjusting the voice data of the first object to be played to the voice content corresponding to the edited text information.

[0307] In one possible implementation, processor 1701 outputs the first voice file of the first object in the session area of ​​the information interaction interface for performing the following operations:

[0308] Output the first voice file of the first object in the conversation area of ​​the information interaction interface; or,

[0309] Output the first audio file of the first object in the conversation area of ​​the information interaction interface, and display the text information obtained by converting the audio data of the first object into text; or,

[0310] Output the first voice file of the first object in the conversation area of ​​the information interaction interface, and display the text information obtained after the voice data of the first object is converted into text, as well as the pattern code associated with the first voice file.

[0311] In one possible implementation, the information interaction interface is displayed on a first client, and the first voice file is associated with a pattern control; the processor 1701 is also used to perform the following operations:

[0312] In response to a trigger operation on the pattern control associated with the first audio file, the pattern code associated with the first audio file is displayed;

[0313] In response to the second client's scanning operation of the pattern code, the first voice file is output to the second client; wherein the first client and the second client are different types of clients.

[0314] In one possible implementation, the information interaction interface further includes a session area; the processor 1701 displays a second voice file on the information interaction interface for performing the following operations:

[0315] Output the first audio file in the session area; and,

[0316] In response to the output of the first audio file, the response data generated by the interactive object is output in the conversation area; wherein the response data includes at least one of the following: response audio file, response text corresponding to the response audio file, response image, and response video.

[0317] In one possible implementation, the response data includes a response voice file and a response text; the display method between the response voice file and the response text includes any of the following:

[0318] During the playback of the reply audio file, the full text content of the reply is displayed in the information interaction interface;

[0319] While the reply audio file is playing, the corresponding text content of the content played in the reply audio file is displayed simultaneously.

[0320] During the playback of the reply audio file, the text content corresponding to the played content in the reply audio file is displayed differently; the different display includes any one of the following: highlighting, font differentiation, color differentiation, and animation display.

[0321] In one possible implementation, processor 1701 is also used to perform the following operations:

[0322] In response to a sharing operation on the reply audio file, a share link is generated and shared to a second object; or...

[0323] In response to the sharing operation of the reply voice file, the pattern code of the reply voice file is shared to the second object; wherein, the pattern code has an expiration period, and the second object can play the reply voice file by scanning the pattern code within the expiration period.

[0324] In one possible implementation, the interaction object is the model object corresponding to the large language model; the information interaction interface includes a task display area and a conversation area, the task display area displays at least one task list, and the task list displays at least one historical conversation task that has been created; the conversation area is used for information interaction with the model object corresponding to the large language model during the execution of a conversation task; the processor 1701 is also used to perform the following operations:

[0325] In response to a session creation operation, the newly created session task is displayed, and information interaction under the new session task is performed with the model object corresponding to the large language model in the session area of ​​the new session task; or...

[0326] In response to the selection operation for the target historical session task, the selected target historical session task is displayed, and information interaction under the target historical session task is performed with the model object corresponding to the large language model in the session area of ​​the target historical session task.

[0327] In one possible implementation, processor 1701 is also used to perform the following operations:

[0328] In response to a session export operation on the task list, a candidate task list to be exported is obtained, which includes at least one created historical session task in the task display area.

[0329] Output a list of candidate tasks; each historical conversation task in the candidate task list is used to perform model analysis on the large language model; the model analysis includes any one or more of the following: performance analysis, efficiency analysis, indicator analysis, and accuracy analysis.

[0330] In one possible implementation, processor 1701 is also used to perform the following operations:

[0331] Obtain the first audio file of the first object;

[0332] The large language model is invoked to process the speech data in the first speech file and generate response data.

[0333] Generate a response audio file based on the response data.

[0334] In one possible implementation, processor 1701 calls a large language model to process the speech data in the first speech file, generating response data for performing the following operations:

[0335] Automatic speech recognition technology is used to convert the speech data in the first speech file into text, thereby obtaining the text information corresponding to the first speech file.

[0336] The large language model is invoked to process the text information and obtain the response data;

[0337] Text-to-speech technology is used to convert the response data into speech, resulting in the corresponding audio file.

[0338] In this embodiment, after the computer device executes the above steps through the processor, the effect achieved can be referred to the description of the foregoing embodiments, and will not be repeated here.

[0339] According to one aspect of this application, embodiments of this application also provide a computer-readable storage medium storing a computer program. When a processor executes the computer program, it can perform the methods described in the preceding embodiments; therefore, further details will not be repeated here. For technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application. As an example, the computer program can be deployed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network.

[0340] According to one aspect of this application, embodiments of this application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, enabling the computer device to perform the methods described in the foregoing embodiments; therefore, further details will not be repeated here. For technical details not disclosed in the embodiments of the computer program product involved in this application, please refer to the description of the method embodiments of this application.

[0341] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product, which includes one or more computer programs. When the computer program is loaded and executed on a computer device, it generates, in whole or in part, the processes or functions described in the embodiments of this application; the computer device can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program can be stored in or transmitted through a computer-readable storage medium; the computer program can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible to the computer device or a data processing device such as a server or data center that integrates one or more available media; wherein, the available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0342] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. An information exchange method, characterized in that, include: The information interaction interface is displayed and is used to interact with the interaction object. The information interaction interface includes an information input area; In response to an information interaction operation in the information interaction interface, a first audio file is displayed in the information input area; The first audio file is generated based on the audio data input by the first object; In response to an output operation of the first audio file displayed in the information input area, a second audio file is displayed in the information interaction interface; The second voice file consists of the first voice file and response data that responds to the voice data.

2. The method as described in claim 1, characterized in that, The information interaction operation is a voice input operation performed by the first object on the information interaction interface; or, the first object realizes the voice input operation by triggering the voice interaction entry on the information interaction interface.

3. The method as described in claim 1 or 2, characterized in that, The step of displaying a first audio file in the information input area in response to an information interaction operation on the information interaction interface includes: In response to a trigger operation on the voice interaction entry point on the information interaction interface, voice data of the first object is collected, and the first voice file corresponding to the voice data is displayed in the information input area; or, In response to a trigger operation of the information input control set in the information interaction interface, a file selection list is output, the file selection list displaying at least one audio file; in response to a selection operation of a target audio file among the at least one audio file, the target audio file is displayed as the first audio file in the information input area.

4. The method according to any one of claims 1 to 3, characterized in that, The information interaction interface also includes a session area, which is used to display interaction information between the first object and the interaction object; the method further includes: In response to a preview operation of the first audio file displayed in the information input area, a voice preview is performed on the audio data corresponding to the first audio file; During the voice preview process, in response to the output operation of the first voice file, the first voice file is sent from the information input area to the conversation area for display; or, in response to the deletion operation of the first voice file, the first voice file is deleted from the information input area. During the voice preview process, the playback attributes of the voice data can be adjusted, including at least one of the following: playback progress, playback speed, playback volume, and playback timbre.

5. The method as described in claim 3, characterized in that, The voice interaction entry point is a voice acquisition control, which is displayed on the information interaction interface in a first display format; the method further includes: In response to a trigger operation by a first object on the voice acquisition control, the display mode of the voice acquisition control is adjusted from the first display mode to the second display mode; When voice data is detected generated by the first object, the display mode of the voice acquisition control is adjusted from the second display mode to the third display mode; In response to the voice data acquisition completion event, the display mode of the voice acquisition control is adjusted from the third display mode to the fourth display mode; the acquisition completion event includes: the voice data acquisition completion operation is performed, or the acquisition duration of the voice data reaches a preset duration threshold; The first display form, the second display form, the third display form, and the fourth display form are all different from each other; and the display form includes any one or more of the following: color, shape, and size.

6. The method according to any one of claims 1-5, characterized in that, Following the output operation in response to the first audio file displayed in the information input area, the method further includes: The first voice file is output in the conversation area of ​​the information interaction interface; In response to a playback operation on the first audio file, the audio data corresponding to the first audio file is played in the session area; During the playback of the audio data, in response to the playback adjustment operation of the first audio file, the playback attributes of the audio data are adjusted; The playback attributes include any one or more of the following: playback progress, playback speed, playback timbre, and playback volume.

7. The method as described in claim 6, characterized in that, The method further includes: In response to the playback operation of the first audio file, the audio data corresponding to the first audio file is played, and the text information of the first audio file is displayed; During the playback of the audio data, in response to the editing operation of the text information, the audio data corresponding to the first audio file being played is adjusted; The adjustment refers to: adjusting the audio data to be played to the audio content corresponding to the text information after the editing operation.

8. The method as described in claim 6, characterized in that, The step of outputting the first voice file in the conversation area of ​​the information interaction interface includes: Output the first audio file in the conversation area of ​​the information interaction interface; or, The first audio file is output in the conversation area of ​​the information interaction interface, and the text information obtained by converting the audio data of the first audio file into text is displayed; or, The first audio file is output in the conversation area of ​​the information interaction interface, and the text information obtained by converting the audio data of the first audio file into text is displayed, as well as the pattern code associated with the first audio file is displayed.

9. The method as described in claim 8, characterized in that, The information interaction interface is displayed on the first client, and the first audio file is associated with a pattern control; the method further includes: In response to a trigger operation on the pattern control associated with the first audio file, the pattern code associated with the first audio file is displayed; In response to the recognition operation of the pattern code by the second client, the first voice file is output in the second client; The first client and the second client are different types of clients.

10. The method as described in claim 1, characterized in that, The information interaction interface further includes a conversation area; the step of displaying the second audio file on the information interaction interface includes: Output the first audio file in the session area; and, In response to the output of the first audio file, the response data generated by the interactive object is output in the conversation area; wherein, the response data includes at least one of the following: a response audio file, a response text corresponding to the response audio file, a response image, and a response video.

11. The method as described in claim 10, characterized in that, The response data includes a response audio file and a response text; the display method between the response audio file and the response text includes any of the following: During the playback of the reply audio file, the full text content of the reply text is displayed on the information interaction interface; During the playback of the response audio file, the corresponding text content of the content being played in the response audio file is displayed simultaneously. During the playback of the response audio file, the text content corresponding to the content being played in the response audio file is displayed in a differentiated manner; wherein, the differentiated display includes any one of the following: highlighting, font differentiation, color differentiation, and animation display.

12. The method according to any one of claims 8-11, characterized in that, The method further includes: In response to a sharing operation on the reply audio file, a sharing link is generated, and the sharing link is shared to a second object; or... In response to the sharing operation of the reply voice file, the pattern code of the reply voice file is shared to the second object; wherein, the pattern code has an expiration period, and the second object can play the reply voice file by scanning the pattern code within the expiration period.

13. The method as described in claim 1, characterized in that, The interactive object is the model object corresponding to the large language model; the information interaction interface includes a task display area and a conversation area, the task display area displays at least one task list, and the task list displays at least one historical conversation task that has been created. The conversation area is used for information interaction with the model object corresponding to the large language model during the execution of a conversation task; the method further includes: In response to a session creation operation, the newly created session task is displayed, and information interaction under the newly created session task is performed with the model object corresponding to the large language model in the session area of ​​the newly created session task; or... In response to the selection operation for the target historical session task, the selected target historical session task is displayed, and information interaction under the target historical session task is performed with the model object corresponding to the large language model in the session area of ​​the target historical session task.

14. The method as described in claim 13, characterized in that, The method further includes: In response to a session export operation on the task list, a candidate task list to be exported is obtained, the candidate task list including at least one created historical session task in the task display area; Output the candidate task list; each historical conversation task in the candidate task list is used to perform model analysis on the large language model; The model analysis includes any one or more of the following: performance analysis, efficiency analysis, indicator analysis, and accuracy analysis.

15. The method as described in claim 1, characterized in that, The method further includes: Obtain the first audio file; The large language model is invoked to process the speech data in the first speech file and generate response data. A response audio file is generated based on the response data.

16. The method as described in claim 15, characterized in that, The step of calling the large language model to process the speech data in the first speech file and generating response data includes: Automatic speech recognition technology is used to convert the speech data in the first speech file into text to obtain the text information corresponding to the first speech file. The large language model is invoked to process the text information and obtain response data. Text-to-speech technology is used to convert the response data into speech, resulting in a corresponding audio file.

17. An information interaction device, characterized in that, include: The display unit is used to display an information interaction interface, which is used to interact with the interaction object. The information interaction interface includes an information input area; The processing unit is configured to display a first voice file in the information input area in response to an information interaction operation in the information interaction interface; The first audio file is generated based on the audio data input by the first object; The processing unit is further configured to, in response to an output operation of the first voice file displayed in the information input area, display a second voice file in the information interaction interface; the second voice file consists of the first voice file and response data that responds to the voice data.

18. A computer device, characterized in that, include: Memory and processor; A memory, wherein one or more computer programs are stored; A processor for loading one or more computer programs to implement the information interaction method as described in any one of claims 1-16.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and executed as described in any one of claims 1-16.

20. A computer program product, characterized in that, The computer program product includes a computer program adapted to be loaded by a processor and execute the information interaction method as described in any one of claims 1-16.