Voice playing system, method, device and apparatus

By processing the global time information of voice files on the client side, the coupling problem of the server-side speech recognition service is solved, enabling the continuous playback and synchronous display of voice files, and adapting to various user needs.

CN114582348BActive Publication Date: 2026-05-19ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA GROUP HOLDING LTD
Filing Date
2020-11-18
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In existing technologies, server-side speech recognition services have a high degree of coupling in their application of playing back the entire conference audio from multiple audio files and simultaneously displaying the recognized speech text, making it difficult to flexibly meet the variable needs of different users.

Method used

On the client side, it determines whether the audio files belong to the same activity, obtains local time information through atomic speech recognition service, and calculates global time information by combining the time information of multiple audio files, so as to realize the continuous playback and synchronous display of audio files.

Benefits of technology

It reduces the coupling between the server and the application, provides a flexible user experience, supports the merging and synchronous display of multiple voice data segments, and adapts to the variable needs of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114582348B_ABST
    Figure CN114582348B_ABST
Patent Text Reader

Abstract

The application discloses a voice playing system, which identifies whether a plurality of voice files belong to the same activity at a client, performs voice recognition on each voice file through an atomic voice recognition service at a server, obtains local time information of a word element in each voice recognition text relative to a starting point of the voice file, determines global time information of the word element in each voice recognition text relative to an activity starting point at the client, automatically opens a plurality of voice files in a voice file playing list in sequence in a voice playing controller, plays the multi-voice data of the entire activity in a coherent manner, and displays the voice recognition text corresponding to the voice playing progress of the entire activity, wherein the time information corresponding to the displayed voice recognition text is the global time information. By adopting the processing mode, the coupling of the voice recognition service of the server to the application can be effectively reduced, the user can play back the voice of the entire activity without sensing, and a good use experience of synchronously displaying the voice recognition text can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, specifically to speech playback systems, related methods and apparatus, and electronic devices. Background Technology

[0002] In various intelligent voice transcription scenarios such as meetings, interrogations, and interviews, there may be situations where voice recording is paused due to tea breaks or breaks in between. This results in multiple voice files being generated after the entire meeting. Later, it is necessary to play back the entire meeting's audio continuously based on these multiple audio files and simultaneously display the speech-recognized text.

[0003] Currently, a typical method for playing back multiple audio files from a conference in a coherent manner and simultaneously displaying the speech-recognized text involves the server continuously appending fragmented audio data generated during pauses in recording to a single audio file during speech recognition processing. This ensures that the multiple speech recognitions ultimately produce only one audio file, recognizing the complete speech of the entire conference. Then, the complete audio file and the complete speech recognition result are sent to the client, where the complete audio file is played and the corresponding text is displayed simultaneously, allowing users to compare the currently playing audio with the corresponding text.

[0004] However, in the process of realizing this invention, the inventors discovered that the above technical solution has at least the following problems: 1) The speech recognition server needs to perceive whether multiple speech recognition data segments need to be merged at the application level, and handle the logic of merging multiple speech recognition data segments. This results in high coupling between the server and the application of playing back the entire conference speech coherently from multiple audio files based on the entire conference and simultaneously displaying the speech recognition text. Therefore, it is difficult to provide atomic and general speech recognition services to multiple applications; 2) The server cannot flexibly respond to the variable needs of different users. For example, some application systems need to provide an overview of multiple audio segments and corresponding text for the entire conference, and also need to display segmented sub-topics. In summary, how to reduce the coupling of the server-side speech recognition service to the application of playing back the entire conference speech coherently from multiple audio files based on the entire conference and simultaneously displaying the speech recognition text has become an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] This application provides a voice playback system to address the problem in existing technologies where server-side speech recognition services have high coupling in their ability to coherently play back the entire conference audio from multiple audio files and simultaneously display the speech-recognized text. This application also provides a voice playback method and apparatus, as well as an electronic device.

[0006] This application provides a voice playback system, including:

[0007] The client is configured to: identify multiple audio files included in the target activity; receive speech recognition text corresponding to the audio files and local time information of word elements in the text relative to the start point of their respective audio files, sent by the server; determine global time information of the word elements relative to the start point of the target activity based on the time information of the multiple audio files and the local time information; and sequentially open the multiple audio files in the audio file playlist in the audio playback controller to continuously play multiple segments of audio data of the target activity corresponding to the multiple audio files; display speech recognition text corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition text includes global time information.

[0008] The server is used to perform speech recognition processing on the audio file and send the speech recognition text and local time information of the word elements to the client.

[0009] This application also provides a voice playback method, including:

[0010] Identify the multiple audio files included in the target activity;

[0011] Receive the speech recognition text corresponding to the speech file and the local time information of the word elements in the text relative to the start point of the speech file sent by the server;

[0012] Based on the time information of the multiple audio files and the local time information, the global time information of the word element relative to the starting point of the target activity is determined;

[0013] In the voice playback controller, the multiple voice files in the voice file playlist are opened sequentially to play multiple segments of voice data corresponding to the target activity in a continuous manner, and the voice recognition text corresponding to the voice playback progress of the target activity is displayed. The time information corresponding to the displayed voice recognition text includes global time information.

[0014] Optional, also includes:

[0015] Identify the target audio file;

[0016] The target audio file is opened in the audio playback controller, and the target speech recognition text corresponding to the audio playback progress of the target audio file is displayed. The time information corresponding to the displayed target speech recognition text includes local time information.

[0017] Optionally, the activity may include multiple activity themes;

[0018] The method further includes:

[0019] Determine the topic information of the audio file;

[0020] The determined target audio file includes:

[0021] Determine the target topic information;

[0022] Use the audio file that corresponds to the target topic information as the target audio file.

[0023] Optional, also includes:

[0024] The server sends information about the target activity, including the multiple audio files, topic information, and global time information, to the server. The server stores the global time information, the target activity information including the multiple audio files, and the topic information. This allows the server to respond to audio playback requests from other clients for a specific topic, sending the target audio file corresponding to the target topic, the target speech recognition text corresponding to the target audio file, and the local time information to other clients. This allows the clients to play the audio data for the target topic, display the target speech recognition text corresponding to the audio playback progress of the target audio file, and the time information corresponding to the displayed target speech recognition text includes local time information.

[0025] Optional, also includes:

[0026] The server sends information about the target activity, including the multiple audio files and global time information, to the server. The server stores the global time information and the information about the target activity, including the multiple audio files, so that the server can respond to audio playback requests for the target activity sent by other clients. The server sends the multiple audio files, the multiple speech recognition texts, the local time information, and the global time information to other clients so that other clients can play multiple audio data segments of the target activity continuously and display speech recognition texts corresponding to the audio playback progress of the target activity. The time information corresponding to the displayed speech recognition texts includes the global time information.

[0027] Optional, also includes:

[0028] Edit the speech-recognition text while playing multiple segments of speech data of the target activity in a continuous manner;

[0029] Determine the updated global time information and updated local time information of word elements in the edited speech recognition text.

[0030] Optionally, determining the updated global time information and updated local time information of word elements in the edited speech recognition text includes:

[0031] Determine the updated global time information;

[0032] Based on the updated global time information, the updated local time information is determined.

[0033] Optional, also includes:

[0034] The updated global time information and the updated local time information are sent to the server, so that the server updates the global time information and the local time information.

[0035] Optionally, the editing of the speech recognition text includes at least one of the following methods: modifying word elements, adding word elements, or deleting word elements.

[0036] This application also provides a voice playback method, including:

[0037] Receive speech recognition requests for multiple audio files in a target activity;

[0038] Perform speech recognition processing on the multiple audio files;

[0039] The system sends local time information of word elements in multiple speech recognition texts relative to the start point of their respective speech files to the client, so that the client can determine the global time information of the word elements relative to the start point of the target activity based on the time information of the multiple speech files and the local time information; and sequentially opens the multiple speech files in the speech file playlist in the speech playback controller to continuously play multiple segments of speech data of the target activity corresponding to the multiple speech files; and displays the speech recognition text corresponding to the speech playback progress of the target activity, and the time information corresponding to the displayed speech recognition text includes the global time information.

[0040] Optional, also includes:

[0041] The system stores local time information of word elements in the multiple speech files and multiple speech-recognized texts, as well as information and global time information of the target activity sent by the client, including the multiple speech files.

[0042] Receive voice playback requests for the target activity from other clients;

[0043] The system sends the target activity, including the multiple audio files, the multiple speech recognition texts, and the global time information, to other clients so that other clients can play multiple audio data segments of the target activity continuously, display speech recognition texts corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition texts includes global time information.

[0044] Optionally, the activity may include multiple activity themes;

[0045] Also includes:

[0046] Store the topic information of the voice file sent by the client;

[0047] The topic information is sent to other clients so that they can play the audio data of the target topic, display the target speech recognition text corresponding to the audio playback progress of the audio file of the target topic, and the time information corresponding to the target speech recognition text includes local time information.

[0048] Optional, also includes:

[0049] Based on the word element change information, updated local time information, and global time information sent by the client, the speech recognition text, word element local time information, and global time information are updated.

[0050] This application also provides a voice playback device, including:

[0051] The activity audio file determination unit is used to determine multiple audio files included in the target activity;

[0052] The data receiving unit is used to receive the speech recognition text corresponding to the speech file and the local time information of the word elements in the text relative to the start point of the speech file sent by the server.

[0053] A global time determination unit is used to determine the global time information of the word element relative to the target activity start point based on the time information of the multiple speech files and the local time information;

[0054] The synchronous display unit is used to sequentially open the multiple audio files in the audio file playlist in the audio playback controller to continuously play multiple segments of audio data of the target activity corresponding to the multiple audio files, and display the speech recognition text corresponding to the audio playback progress of the target activity. The time information corresponding to the displayed speech recognition text includes global time information.

[0055] This application also provides an electronic device, including:

[0056] Processor and memory;

[0057] The memory stores a program for implementing the voice playback method. After the device is powered on and the program for the method is run by the processor, the following steps are performed: determining multiple voice files included in the target activity; receiving speech recognition text corresponding to the voice files and local time information of word elements in the text relative to the start point of their respective voice files, sent by the server; determining global time information of the word elements relative to the start point of the target activity based on the time information of the multiple voice files and the local time information; sequentially opening the multiple voice files in the voice file playlist in the voice playback controller to continuously play multiple segments of voice data of the target activity corresponding to the multiple voice files, and displaying speech recognition text corresponding to the voice playback progress of the target activity, wherein the time information corresponding to the displayed speech recognition text includes global time information.

[0058] This application also provides a voice playback device, including:

[0059] The request receiving unit is used to receive speech recognition requests for multiple audio files in the target activity.

[0060] The speech recognition unit is used to perform speech recognition processing on the plurality of speech files;

[0061] The data sending unit is configured to send local time information of word elements in multiple speech recognition texts relative to the start point of their respective speech files to the client, so that the client can determine the global time information of the word elements relative to the start point of the target activity based on the time information of the multiple speech files and the local time information; and to sequentially open the multiple speech files in the speech file playlist in the speech playback controller to continuously play multiple segments of speech data of the target activity corresponding to the multiple speech files; and to display the speech recognition text corresponding to the speech playback progress of the target activity, wherein the time information corresponding to the displayed speech recognition text includes global time information.

[0062] This application also provides an electronic device, including:

[0063] Processor and memory;

[0064] The memory stores a program for implementing the voice playback method. After the device is powered on and the program for the method is run by the processor, it performs the following steps: receiving voice recognition requests for multiple voice files in a target activity; performing voice recognition processing on the multiple voice files; sending local time information of word elements in multiple voice recognition texts relative to the start point of their respective voice files to the client, so that the client can determine the global time information of the word elements relative to the start point of the target activity based on the time information of the multiple voice files and the local time information; and sequentially opening the multiple voice files in the voice file playlist in the voice playback controller to continuously play multiple segments of voice data of the target activity corresponding to the multiple voice files; displaying voice recognition text corresponding to the voice playback progress of the target activity, wherein the time information corresponding to the displayed voice recognition text includes global time information.

[0065] This application also provides a method for playing lecture audio, including:

[0066] Identify the multiple audio files included in the teaching process;

[0067] The system receives the teaching content text corresponding to the teaching audio file and the local time information of the word elements in the text relative to the start point of the corresponding audio file, sent by the server.

[0068] Based on the time information of the multiple teaching audio files and the local time information, the global time information of the word element relative to the start point of the teaching process is determined;

[0069] In the voice playback controller, the multiple teaching voice files in the voice file playlist are opened sequentially to play multiple segments of voice data in the teaching process corresponding to the multiple teaching voice files in a continuous manner, and the teaching content text corresponding to the voice playback progress in the teaching process is displayed. The time information corresponding to the teaching content text includes global time information.

[0070] Optionally, the teaching process includes multiple teaching topics, with different teaching audio files corresponding to different teaching topics;

[0071] The method further includes:

[0072] Define the target teaching topic;

[0073] Open the target lecture audio file corresponding to the target lecture topic in the audio playback controller, and display the target lecture content text corresponding to the audio playback progress of the target lecture audio file. The time information corresponding to the target lecture content text includes local time information.

[0074] This application also provides a method for playing live audio, including:

[0075] Identify the multiple audio files included in the live stream;

[0076] Receive the live content text corresponding to the live audio file and the local time information of the word elements in the text relative to the start point of the audio file sent by the server;

[0077] Based on the time information of the multiple live audio files and the local time information, the global time information of the word element relative to the start point of the live broadcast process is determined;

[0078] In the voice playback controller, the multiple live audio files in the audio file playlist are opened sequentially to play multiple segments of audio data during the live broadcast corresponding to the multiple live audio files, and the live content text corresponding to the audio playback progress during the live broadcast is displayed. The time information corresponding to the live content text includes global time information.

[0079] Optionally, the live streaming process includes multiple live streaming topics, with different live streaming audio files corresponding to different live streaming topics;

[0080] The method further includes:

[0081] Determine the target live stream topic;

[0082] Open the target live audio file corresponding to the target live topic in the audio playback controller, and display the target live content text corresponding to the audio playback progress of the target live audio file. The time information corresponding to the target live content text includes local time information.

[0083] This application also provides a method for playing conference audio, including:

[0084] Identify the multiple audio files included in the target meeting;

[0085] Receive the speech recognition text corresponding to the speech file and the local time information of the word elements in the text relative to the start point of the speech file sent by the server;

[0086] Based on the time information of the multiple audio files and the local time information, the global time information of the word element relative to the target meeting start point is determined;

[0087] In the voice playback controller, the multiple voice files in the voice file playlist are opened sequentially to play multiple segments of voice data from the target conference corresponding to the multiple voice files in a continuous manner, and the voice recognition text corresponding to the voice playback progress of the target conference is displayed. The time information corresponding to the displayed voice recognition text includes global time information.

[0088] Optionally, the target meeting includes multiple sub-topics, with different audio files corresponding to different sub-topics;

[0089] The method further includes:

[0090] Identify target sub-topics;

[0091] Open the target audio file corresponding to the target sub-topic in the audio playback controller, and display the target speech recognition text corresponding to the audio playback progress of the target audio file. The time information corresponding to the target speech recognition text includes local time information.

[0092] This application also provides a method for playing court hearing audio, including:

[0093] Identify the multiple audio files included in the court proceedings;

[0094] The server receives the text of the court hearing content corresponding to the court hearing audio file and the local time information of the word elements in the text relative to the start point of the audio file.

[0095] Based on the time information of the multiple court hearing audio files and the local time information, the global time information of the word element relative to the start point of the court hearing process is determined;

[0096] The voice playback controller sequentially opens multiple court hearing audio files in the audio file playlist to continuously play multiple segments of audio data of the court hearing process corresponding to the multiple court hearing audio files, and displays the court hearing content text corresponding to the audio playback progress in the court hearing process. The time information corresponding to the displayed court hearing content text includes global time information.

[0097] Optionally, the court hearing process includes multiple stages and themes, with different court hearing audio files corresponding to different stages and themes;

[0098] The method further includes:

[0099] Define the theme for the target phase;

[0100] Open the target court hearing audio file corresponding to the target stage theme in the audio playback controller, and display the target court hearing content text corresponding to the audio playback progress of the target court hearing audio file. The time information corresponding to the target court hearing content text includes local time information.

[0101] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the various methods described above.

[0102] This application also provides a computer program product including instructions that, when run on a computer, cause the computer to perform the various methods described above.

[0103] Compared with the prior art, this application has the following advantages:

[0104] The voice playback system provided in this application embodiment identifies whether multiple voice files belong to the same activity and whether these voice files need to be played continuously on the front-end application side, and synchronously displays the voice recognition text corresponding to the playback progress of the entire activity. On the server side, the system performs voice recognition on each voice file through atomic voice recognition service to obtain the local time information of each word element in the voice recognition text relative to the start point of the corresponding voice file. The voice recognition text and local time information are sent to the front-end application side, which then determines the global time information of the word elements in the recognition text of each voice file relative to the start point of the activity. The system automatically opens multiple voice files in the voice file playlist in the voice playback controller to play multiple segments of voice data of the entire activity continuously and displays the voice recognition text corresponding to the playback progress of the entire activity. The time information corresponding to the displayed voice recognition text is the global time information. This achieves the processing of merging and playing multiple segments of voice data of the same activity and synchronously highlighting the recognition text corresponding to the global playback progress. This approach avoids physically merging multiple audio files from the same activity through the server-side speech recognition module, without altering the atomic speech recognition logic provided by the server. Therefore, in application scenarios where multiple audio files from the entire activity are played back sequentially and the speech-recognized text is displayed simultaneously, the coupling of the server-side speech recognition service to the application is effectively reduced, achieving a seamless user experience where the entire activity's audio is played back and the speech-recognized text is displayed simultaneously. Furthermore, because this processing method does not physically merge multiple audio files from the same activity into a single audio file, but rather stores multiple audio files independently, along with their respective speech-recognition texts, it provides an effective data foundation for flexibly addressing various variable scenarios tailored to user needs within this application. Attached Figure Description

[0105] Figure 1 This application provides a schematic diagram of the structure of an embodiment of a voice playback system;

[0106] Figure 2 This application provides a schematic diagram of a scenario for an embodiment of a voice playback system;

[0107] Figure 3 This application provides a device interaction diagram of an embodiment of a voice playback system;

[0108] Figure 4 This application provides a segmented schematic diagram illustrating an embodiment of a voice playback system. Detailed Implementation

[0109] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0110] This application provides a speech recognition data processing system, method, and apparatus, as well as an electronic device. The various solutions are described in detail in the following embodiments.

[0111] First Embodiment

[0112] Please refer to Figure 1 This is a schematic diagram illustrating the structure of an embodiment of the voice playback system of this application. In this embodiment, the system may include: server 1 and client 2.

[0113] The server 1 can be a server deployed on a cloud server or a server dedicated to implementing speech recognition processing, which can be deployed in a data center. The server can be a cluster server or a single server.

[0114] The client 2 includes, but is not limited to, mobile communication devices, namely, mobile phones or smartphones, as well as personal computers, tablets, iPads and other terminal devices.

[0115] Please refer to Figure 2 This is a schematic diagram illustrating the scenario of the voice playback system described in this application. The server and client can connect via a network, such as the client connecting via Wi-Fi, etc. Figure 2 As shown, users can use a browser (such as Internet Explorer) installed on the client to continuously play multiple audio segments of the target activity stored in multiple audio files within a webpage. Although these audio segments are stored in different audio files, the user is unaware of this and does not perceive any interruption in the playback of multiple audio segments; instead, they perceive the entire complete audio of the activity being played directly. While the audio is playing through the browser, the client can use the browser's embedded webpage text editor (such as a rich text editor) to synchronously highlight (e.g., show) the text corresponding to the currently playing audio content of the entire activity, based on the speech-recognized text from each audio file provided by the server. This allows for a better association between the transcribed text content and the audio playback time, helping users focus on the currently playing content and check for any issues with the recognized text. When users find problems with the recognized text, they can edit the text online using the webpage text editor.

[0116] The activity can be a meeting, training course, live broadcast, court hearing, etc. The target activity may include multiple audio files, each storing audio data containing identifiable spoken content. These audio files are sequential in time, and the data from all audio files is concatenated to form the complete audio data of the entire activity. For example, in an educational training scenario, recording a teacher's lecture may result in multiple audio files during a single lecture due to various reasons. When students review the lecture content, they may want to play the complete lecture audio without interruption and simultaneously display the text of the currently playing content. Similarly, in a live-streaming e-commerce scenario, multiple audio files may be generated during a single live broadcast due to the host taking breaks. When consumers review the live broadcast content, they may want to play the complete live broadcast audio without interruption and simultaneously display the text of the currently playing content.

[0117] Please refer to Figure 3 This is a device interaction diagram of an embodiment of the voice playback system of this application. In this embodiment, the client is used to determine multiple voice files included in the target activity; receive speech recognition text corresponding to the voice files and local time information of word elements in the text relative to the start point of their respective voice files sent by the server; determine global time information of the word elements relative to the start point of the target activity based on the time information of the multiple voice files and the local time information; and sequentially open the multiple voice files in the voice file playlist in the voice playback controller to continuously play multiple segments of voice data of the target activity corresponding to the multiple voice files; display speech recognition text corresponding to the voice playback progress of the target activity, and the time information corresponding to the displayed speech recognition text includes global time information; the server is used to perform speech recognition processing on the voice files and send the speech recognition text and the local time information of the word elements to the client.

[0118] The client can determine the multiple audio files included in the target activity in the following way: for the target activity, directly collect multiple audio files, or collect multiple audio files in advance, and then specify multiple audio files corresponding to the target activity from the multiple audio files collected in advance.

[0119] After the client determines that the target activity includes multiple audio files, it can upload the multiple audio files to the server and request the server to perform speech recognition processing on these audio files; accordingly, the server performs speech recognition processing on each audio file separately to form speech recognition text corresponding to each audio file.

[0120] Speech recognition is the technology of converting speech into text. The input data for speech recognition algorithms (such as speech recognition models) can be speech data, and the algorithm outputs a recognition result, which is usually a string with timestamp information. In practice, various existing speech recognition algorithms can be used; since these algorithms are relatively mature existing technologies, they will not be elaborated upon here.

[0121] The result of speech recognition processing includes word element data; the speech recognition text of a speech file includes multiple word elements. In the prior art, a word element may include: word content information and time information. The time information may include: start time and end time. Since this time is the time information of the word element relative to the start point of its respective speech file, this application refers to it as local time information.

[0122] The system provided in this embodiment aims to unify the word element time information of multiple audio data segments from the same activity onto a single timeline, facilitating the simultaneous playback of complete audio and display of corresponding text. Therefore, the client can determine the time information of each word element relative to the start point of the entire activity; in this embodiment, this time information is referred to as global time information. The global time information can be the actual audio time of the word element, such as 15:30:08, or the duration relative to the start point of the activity, such as 25 minutes and 10 seconds. Table 1 shows the word element data of the speech recognition text in this embodiment.

[0123]

[0124]

[0125] Table 1. Word element data of speech recognition text

[0126] As shown in Table 1, one difference between the system provided in this application embodiment and the prior art is that the system provided in this application embodiment needs to determine both the local time information of each word element and the global time information. The local time information may include the start and end times of the word's speech segment within its corresponding speech file; the global time information may include the start and end times of the word's speech segment within the entire complete speech of the activity.

[0127] In this embodiment, the client determines the global time information of the word element based on the time information of multiple audio files and the local time information. Multiple audio files for an activity have a temporal sequence, and the data from all audio files are concatenated to form the complete audio data for the entire activity. The time information of the audio files can be a specific start time, such as 15:30:00, or a temporal sequence, such as the second audio file.

[0128] For example, conference A includes n audio files that are sequential in time, such as audio file 1 starting at 13:50, audio file 2 starting at 14:10, and so on. When recognizing text by combining each sentence from multiple audio files, it is necessary to consider not only the relative time of word elements with their corresponding audio files, but also the start time of the audio file. Based on these two pieces of information, the global time information corresponding to the word elements is calculated. For example... Figure 2 As shown, since the global time information is the actual timestamp, the speech recognition text and audio playback are aligned with the timeline according to the actual timestamp during playback, ensuring that the multiple segments of recognized text can be correctly associated with the visually "merged" complete conference audio playback.

[0129] In this embodiment, after the client determines the global time information, it can also send the global time information to the server, so that the server stores the global time information and forms the data shown in Table 1. This data can provide a speech recognition data foundation for ensuring the continuity of the shorthand interactive experience such as playback text association, text audio positioning, and text modification and editing.

[0130] like Figure 4 As shown, the system provided in this application differs from the prior art in that it does not merge multiple audio files of the same activity into a single complete audio file, nor does it merge the speech recognition text of each audio file into a single complete speech recognition text. Instead, it still stores multiple audio files and multiple speech recognition texts including local time information. This allows for flexible responses to the variable needs of different users. For example, some application systems need to provide a foundation of audio data for both a multi-segment overview of the audio data of the entire activity and the display of segmented sub-topics.

[0131] In this embodiment, the client can play multiple audio segments of the target activity through a web browser and simultaneously display the corresponding speech-recognized text. This establishes a correlation between the speech-recognized text content and the audio playback time, helping the user focus on the currently playing content and check for any issues with the corresponding recognized text. To this end, the server sends a speech-recognized text viewing webpage to the client. This webpage may include a web text editor, which can display the speech-recognized text corresponding to the playback progress. When the user wants to browse or edit the speech-recognized text, they can play multiple audio files of the activity consecutively through the client and receive the speech-recognized text editing page sent by the server. The user can listen to the audio while simultaneously viewing the entire speech-recognized text of the activity through the text editor on the page.

[0132] In this embodiment, the speech recognition text viewing webpage automatically opens multiple audio files of the same activity in the audio file playlist sequentially according to the start time information of multiple audio files of the same activity through the voice playback controller. That is, after playing one audio file, it automatically switches to the next audio file in the list. In this way, the user can watch the complete meeting audio played continuously. When playing the audio, the speech recognition text corresponding to the audio playback progress is displayed according to the global time information. The time information corresponding to the displayed speech recognition text includes the global time information.

[0133] In practice, multiple audio files can be managed as a playlist using audio application programming interfaces provided by web scripting languages, such as the JS Audio API in JavaScript. The same playback controller can be used on the user interface to control the positioning and playback of multiple audio files for the target activity.

[0134] In practice, the client can process the data using the following steps:

[0135] 1) Preload all audio files for the target activity in the playlist sequentially via the audio application programming interface, and obtain the playback duration of each audio file, such as... Figure 4 Two audio files generated during the same meeting due to an interruption; the total duration of the audio files in the list is calculated and displayed as the total duration on the playback controller's progress bar, such as... Figure 2 An interface for seamless, continuous playback of two audio files within the same meeting;

[0136] 2) Automatic switching of audio files is supported through the audio application programming interface's audio file playback end event (onended). When the current audio file finishes playing, if there is another audio file in the list, it will automatically switch to the next audio file to continue playing.

[0137] 3) Calculate the current playback time of the playlist based on the current playback time of the currently playing audio file and the total duration of its preceding audio files. This time will be displayed on the playback controller's progress bar and can also be used to calculate the current playback progress of the playlist. Figure 2 The text displayed corresponds to the progress of the full audio playback of the meeting;

[0138] 4) Supports locating specific audio playback positions based on playlist time. The playlist playback interval is divided according to the duration of a single playlist audio, and the location interval is used to determine the target audio file and the target audio position.

[0139] 5) Accept the start timestamp of the input audio. As the playback time changes, calculate the actual timestamp of the audio file based on the current time and the start timestamp of the audio file being played, so as to correspond to the specific position of the text content during playback.

[0140] Through the above steps 1 to 5, multiple audio segments and recognized texts generated from the same meeting can be presented to the user in their entirety, merging the playback experience and ensuring that the user is unaware of audio segmentation when playing back history, providing a seamless operating experience.

[0141] In this embodiment, the client receives the speech recognition text and the local time information sent by the server. Correspondingly, the client can also send information about the target activity, including the multiple speech files and global time information, to the server. This allows the server to store the global time information and the information about the target activity, including the multiple speech files, so that the server can respond to speech playback requests for the target activity sent by other clients. The server can send the multiple speech files, the multiple speech recognition text, the local time information, and the global time information to other clients, so that other clients can play multiple segments of speech data of the target activity consecutively and display the speech recognition text corresponding to the speech playback progress of the target activity. The time information corresponding to the displayed speech recognition text includes the global time information.

[0142] For example, the user of the client is an event administrator, and the users of the other clients are event followers. The event administrator can edit the speech recognition text of the entire event through the system and upload relevant information to the server. This information may include the target event, including information such as the multiple voice files, updated word elements, local time information, and global time information. Event followers can download relevant information from the server, play back the audio of the entire event based on this information, and watch the synchronously displayed speech recognition text.

[0143] In one example, the client can also edit the speech-recognized text while playing multiple segments of speech data of a target activity sequentially; it can determine the updated global and local time information of word elements in the edited speech-recognized text. This approach allows for further editing of the recognized text and updates to the speech-recognized text, local time information, and global time information on the server side.

[0144] The editing of speech recognition text includes, but is not limited to, at least one of the following methods: modifying word elements, adding word elements, and deleting word elements.

[0145] In specific implementation, determining the updated global time information and updated local time information of word elements in the edited speech recognition text may include the following sub-steps: determining the updated global time information; and determining the updated local time information based on the updated global time information. For example, based on the updated global time information of the word element and the position of the corresponding speech segment in the entire active speech, it is determined which speech file the word element belongs to, and then the updated local time information is determined based on the position of the speech segment corresponding to the word element in that speech file.

[0146] In this embodiment, the client receives the speech-recognized text and the local time information sent by the server; correspondingly, the client can also send updated global time information and updated local time information to the server, so that the server updates the global time information and local time information.

[0147] In one example, the client can also be used to identify a target audio file; open the target audio file in the audio playback controller; and display the target speech-recognition text corresponding to the audio playback progress of the target audio file. The time information corresponding to the displayed target speech-recognition text includes local time information. The target audio file belongs to the target activity, can be specified by the user, and is an audio segment of the activity that the user is interested in. Since the target audio file is played separately, the corresponding time information is local time information. For example, the word "Alibaba" in the target audio file starts at 15 seconds in its respective target audio file, but starts at 28 minutes and 15 seconds in the entire activity. This processing method can satisfy both the user's need for a multi-segment overview of the audio of the entire activity and the user's need for displaying segmented audio.

[0148] In one example, the activity includes multiple activity themes; the client is also used to determine the theme information of the audio file; the client can determine the target audio file in the following way: determine the target theme information; and use the audio file corresponding to the target theme information as the target audio file.

[0149] The theme of the event refers to the theme of multiple stages throughout the event, which may include sub-topics of the conference, knowledge points of training courses, and different products sold during the live broadcast.

[0150] For example, in an educational training scenario, a teacher's lecture audio is recorded, and the audio explanations of different knowledge points during the complete lecture are recorded in separate audio files. In this case, the activity theme is the lecture theme, and a complete lecture can include multiple lecture themes, which can be the names of knowledge points. Students can review the complete audio content of the lecture through the system, or they can specify to play a target lecture theme of interest. When a user specifies to play the complete lecture audio content, the client sequentially opens multiple lecture audio files of different knowledge points in the audio file playlist in the audio playback controller to play multiple segments of lecture audio data corresponding to the entire lecture process, and displays the lecture content text corresponding to the audio playback progress of the entire lecture process. The time information corresponding to the lecture content text includes global time information. When a user specifies a target lecture theme, the client plays the lecture audio of the target lecture theme accordingly and displays the lecture content text corresponding to the audio playback progress of the target lecture audio file. The time information corresponding to the lecture content text is the local time information of the word element in the audio file.

[0151] For example, in a live-streaming e-commerce scenario, the host's sales voice is recorded, and the descriptions of different products during the entire live-streaming sales process are stored in separate audio files. In this case, the event theme is the live-streaming theme, and a complete live-streaming process can include multiple live-streaming themes, where the theme can be the product name. Consumers can review the complete audio content of the live-streaming session through the system, or they can specify to play the audio content of a product they are interested in. When a user specifies to play the complete live-streaming audio content, the client sequentially opens multiple different product sales audio files from the audio file playlist in the audio playback controller to play the sales audio data of multiple products throughout the entire live-streaming process, and displays the product sales text corresponding to the audio playback progress of the entire live-streaming process. The time information corresponding to the product sales text includes global time information. When a user specifies a target product, the client plays the sales audio of the target product accordingly and displays the live-streaming content text corresponding to the audio playback progress of the target product's audio file. The time information corresponding to this live-streaming content text is the local time information of the word element in that audio file.

[0152] In specific implementation, the client can also be used to send information about the target activity, including the multiple audio files, topic information, and global time information, to the server. This allows the server to store the global time information, the information about the target activity, including the multiple audio files, and the topic information. This enables the server to respond to audio playback requests for the target topic sent by other clients, and to send the target audio file corresponding to the target topic, the target speech recognition text corresponding to the target audio file, and the local time information to other clients. This allows the clients to play the audio data of the target topic, display the target speech recognition text corresponding to the audio playback progress of the target audio file, and the time information corresponding to the displayed target speech recognition text includes local time information.

[0153] As can be seen from the above embodiments, the voice playback system provided in this application identifies whether multiple voice files belong to the same activity and whether these voice files need to be played continuously on the front-end application side, and synchronously displays the voice recognition text corresponding to the playback progress of the entire activity. On the server side, the system performs voice recognition on each voice file through atomic voice recognition service to obtain the local time information of each word element in the voice recognition text relative to the start point of the corresponding voice file. The voice recognition text and local time information are sent to the front-end application side, and the front-end application side determines the global time information of each word element in the recognition text of each voice file relative to the start point of the activity. The system automatically opens multiple voice files in the voice file playlist in the voice playback controller to play multiple segments of voice data of the entire activity continuously and displays the voice recognition text corresponding to the playback progress of the entire activity. The time information corresponding to the displayed voice recognition text is the global time information. This achieves the processing of merging and playing multiple segments of voice data of the same activity and synchronously highlighting the recognition text corresponding to the global playback progress. This approach avoids physically merging multiple audio files from the same activity through the server-side speech recognition module, without altering the atomic speech recognition logic provided by the server. Therefore, in application scenarios where multiple audio files from the entire activity are played back sequentially and the speech-recognized text is displayed simultaneously, the coupling of the server-side speech recognition service to the application is effectively reduced, achieving a seamless user experience where the entire activity's audio is played back and the speech-recognized text is displayed simultaneously. Furthermore, because this processing method does not physically merge multiple audio files from the same activity into a single audio file, but rather stores multiple audio files independently, along with their respective speech-recognition texts, it provides an effective data foundation for flexibly addressing various variable scenarios tailored to user needs within this application.

[0154] Second Embodiment

[0155] Corresponding to the aforementioned voice playback system, this application also provides a voice playback method. The execution subject of the method includes, but is not limited to, a client, and can also be any device capable of implementing the method. Since the method embodiments are basically similar to the system embodiments, the description is relatively simple; relevant details can be found in the descriptions of the system embodiments. The method embodiments described below are merely illustrative. Parts identical to those in the first embodiment will not be repeated; please refer to the corresponding parts in Embodiment 1.

[0156] In this embodiment, the method may include the following steps:

[0157] Step 1: Identify the multiple audio files included in the target activity;

[0158] Step 2: Receive the speech recognition text corresponding to the speech file and the local time information of the word elements in the text relative to the start point of the speech file sent by the server;

[0159] Step 3: Determine the global time information of the word element relative to the target activity start point based on the time information of the multiple audio files and the local time information;

[0160] Step 4: In the voice playback controller, sequentially open the multiple voice files in the voice file playlist to continuously play multiple segments of voice data corresponding to the target activity, and display the voice recognition text corresponding to the voice playback progress of the target activity. The time information corresponding to the displayed voice recognition text includes global time information.

[0161] In one example, the method may further include the following steps: sending information about the target activity, including the multiple audio files, and global time information to the server, so that the server stores the global time information and the information about the target activity, including the multiple audio files, so that the server can respond to audio playback requests for the target activity sent by other clients, and send the multiple audio files, the multiple speech-recognized text, the local time information, and the global time information to other clients, so that other clients can play multiple audio data segments of the target activity sequentially, display speech-recognized text corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech-recognized text includes the global time information. This processing method allows other clients to reuse the global time information, play multiple audio data segments of the target activity sequentially, and synchronously highlight the speech-recognized text corresponding to the audio playback progress of the target activity; therefore, it can effectively improve audio playback speed, thereby improving the user experience.

[0162] In one example, the method may further include the following steps: editing the speech recognition text while playing multiple segments of speech data of the target activity sequentially; and determining the updated global time information and updated local time information of word elements in the edited speech recognition text. This processing method helps users focus on the currently playing content, allowing them to check for problems with the corresponding recognized text. When users find problems with the recognized text, they can edit the text online using a web-based text editor, thus effectively improving the accuracy of the speech recognition text.

[0163] The editing of speech recognition text includes at least one of the following methods: modifying word elements, adding word elements, and deleting word elements.

[0164] In specific implementation, determining the updated global time information and the updated local time information of word elements in the edited speech recognition text may include the following sub-steps: determining the updated global time information; and determining the updated local time information based on the updated global time information.

[0165] In specific implementation, the method may further include the following steps: sending updated global time information and updated local time information to the server, so that the server updates the global time information and local time information. This processing method allows storing the word element information edited by the client user on the server, which can effectively improve the accuracy of other clients displaying speech-recognized text.

[0166] In one example, the method may further include the following steps: determining a target audio file; opening the target audio file in a voice playback controller and displaying the target speech-recognized text corresponding to the audio playback progress of the target audio file, wherein the time information corresponding to the displayed target speech-recognized text includes local time information. This processing method satisfies both the user's need for an overview of multiple audio segments and their corresponding text for the entire activity, and the user's need to display segmented audio segments and their corresponding text.

[0167] In one example, the activity includes multiple activity themes; the method may further include the following steps: determining the theme information of the audio file; determining the target audio file includes: determining target theme information; and using the audio file corresponding to the target theme information as the target audio file. This processing method allows for the display of audio and corresponding text for a specific theme within an activity of interest to the user.

[0168] In one example, the method may further include the following steps: sending information about the target activity, including the multiple audio files, topic information, and global time information, to the server, so that the server stores the global time information, the information about the target activity including the multiple audio files, and the topic information, so that the server can respond to audio playback requests for the target topic sent by other clients; sending the target audio file corresponding to the target topic, the target speech recognition text corresponding to the target audio file, and the local time information to other clients, so that the clients can play the audio data of the target topic, display the target speech recognition text corresponding to the audio playback progress of the target audio file, and the time information corresponding to the displayed target speech recognition text includes local time information. This processing method allows other clients to reuse the global time information and topic information, thus satisfying both the needs of other users to have an overview of multiple audio segments and corresponding texts of the entire activity, and the needs of other users to display audio segments and corresponding texts of topics of interest.

[0169] For example, the user of the client is the event administrator, and the users of the other clients are the event followers. The event administrator can edit the speech recognition text of the entire event through the system and upload relevant information (including information about the target event including the multiple voice files, global time information, etc.) to the server. The event followers can download relevant information from the server, play back the audio of the entire event based on this information, and watch the synchronously displayed speech recognition text.

[0170] As can be seen from the above embodiments, the voice playback method provided in this application identifies whether multiple voice files belong to the same activity and whether these voice files need to be played continuously on the front-end application side, and synchronously displays the voice recognition text corresponding to the playback progress of the entire activity. On the server side, the voice recognition service performs voice recognition on each voice file to obtain the local time information of each word element in the voice recognition text relative to the start point of the corresponding voice file. The voice recognition text and local time information are sent to the front-end application side, and the front-end application side determines the global time information of each word element in the recognition text of each voice file relative to the start point of the activity. The voice playback controller automatically opens multiple voice files in the voice file playlist in sequence to play multiple segments of voice data of the entire activity continuously and displays the voice recognition text corresponding to the playback progress of the entire activity. The time information corresponding to the displayed voice recognition text is the global time information. This realizes the processing of merging and playing multiple segments of voice data of the same activity and synchronously highlighting the recognition text corresponding to the global playback progress. This approach avoids physically merging multiple audio files from the same activity through the server-side speech recognition module, without altering the atomic speech recognition logic provided by the server. Therefore, in application scenarios where multiple audio files from the entire activity are played back sequentially and the speech-recognized text is displayed simultaneously, the coupling of the server-side speech recognition service to the application is effectively reduced, achieving a seamless user experience where the entire activity's audio is played back and the speech-recognized text is displayed simultaneously. Furthermore, because this processing method does not physically merge multiple audio files from the same activity into a single audio file, but rather stores multiple audio files independently, along with their respective speech-recognition texts, it provides an effective data foundation for flexibly addressing various variable scenarios tailored to user needs within this application.

[0171] Third Embodiment

[0172] In the above embodiments, a voice playback method is provided. Correspondingly, this application also provides a voice playback device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0173] This application also provides a voice playback device, including:

[0174] The activity audio file determination unit is used to determine multiple audio files included in the target activity;

[0175] The data receiving unit is used to receive the speech recognition text corresponding to the speech file and the local time information of the word elements in the text relative to the start point of the speech file sent by the server.

[0176] A global time determination unit is used to determine the global time information of the word element relative to the target activity start point based on the time information of the multiple speech files and the local time information;

[0177] The synchronous display unit is used to sequentially open the multiple audio files in the audio file playlist in the audio playback controller to continuously play multiple segments of audio data of the target activity corresponding to the multiple audio files, and display the speech recognition text corresponding to the audio playback progress of the target activity. The time information corresponding to the displayed speech recognition text includes global time information.

[0178] Fourth embodiment

[0179] In the above embodiments, a voice playback method is provided. Correspondingly, this application also provides an electronic device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0180] This embodiment provides an electronic device, comprising: a processor and a memory; the memory for storing a program implementing a voice playback method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: determining multiple voice files included in the target activity; receiving speech recognition text corresponding to the voice files and local time information of word elements in the text relative to the start point of their respective voice files, sent by a server; determining global time information of the word elements relative to the start point of the target activity based on the time information of the multiple voice files and the local time information; sequentially opening the multiple voice files in the voice file playlist in the voice playback controller to continuously play multiple segments of voice data corresponding to the target activity, and displaying speech recognition text corresponding to the voice playback progress of the target activity, wherein the time information corresponding to the displayed speech recognition text includes global time information.

[0181] Fifth embodiment

[0182] Corresponding to the aforementioned voice playback system, this application also provides a voice playback method. The execution entity of the method includes, but is not limited to, a server, and can also be any device capable of implementing the method. Since the method embodiments are basically similar to the system embodiments, the description is relatively simple; relevant details can be found in the descriptions of the system embodiments. The method embodiments described below are merely illustrative. Parts identical to those in the first embodiment will not be repeated; please refer to the corresponding parts in Embodiment 1.

[0183] In this embodiment, the method may include the following steps:

[0184] Step 1: Receive speech recognition requests for multiple audio files in the target activity.

[0185] The request may include the audio file, or it may include an identifier of the audio file. If the audio file is pre-stored on the server, the request may include the identifier of the audio file; if the audio file is stored on the client, the request may include the audio file.

[0186] Step 2: Perform speech recognition processing on the multiple speech files.

[0187] The method can perform speech recognition processing on each speech file separately using a speech recognition model to obtain the speech recognition text of each speech file. The recognition result includes local time information of word elements.

[0188] Step 3: Send local time information of word elements in multiple speech recognition texts relative to the start point of their respective speech files to the client, so that the client can determine the global time information of the word elements relative to the start point of the target activity based on the time information of multiple speech files and the local time information; and sequentially open the multiple speech files in the speech file playlist in the speech playback controller to play multiple segments of speech data of the target activity corresponding to the multiple speech files in a continuous manner; display the speech recognition text corresponding to the speech playback progress of the target activity, and the time information corresponding to the displayed speech recognition text includes the global time information.

[0189] In one example, the method may further include the following steps: storing local time information of word elements in the plurality of audio files and the plurality of speech-recognized texts; storing information about the target activity sent by the client, including the plurality of audio files, and global time information; and receiving audio playback requests for the target activity sent by other clients; sending the plurality of audio files, the plurality of speech-recognized texts, and the global time information included in the target activity to other clients, so that other clients can play multiple segments of audio data of the target activity consecutively, display speech-recognized text corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech-recognized text includes global time information. This processing method allows other clients to reuse global time information, play multiple segments of audio data of the target activity consecutively, and synchronously highlight speech-recognized text corresponding to the audio playback progress of the target activity; therefore, it can effectively improve audio playback speed, thereby improving user experience.

[0190] In one example, the activity includes multiple activity themes; the method may further include the following steps: storing the theme information of the audio file sent by the client; and sending the theme information to other clients so that other clients can play the audio data of the target theme, display the target speech-recognition text corresponding to the audio playback progress of the audio file of the target theme, and the time information corresponding to the target speech-recognition text includes local time information. This processing method allows other clients to reuse global time information and theme information, thus satisfying both the needs of other users to gain an overview of multiple audio segments and corresponding texts of the entire activity, and the needs of other users to display audio segments and corresponding texts of the themes they are interested in.

[0191] For example, the user of the client is the event administrator, and the users of the other clients are the event followers. The event administrator can edit the speech recognition text of the entire event through the system and upload relevant information (including information about the target event including the multiple voice files, global time information, etc.) to the server. The event followers can download relevant information from the server, play back the audio of the entire event based on this information, and watch the synchronously displayed speech recognition text.

[0192] In one example, the method may further include the following steps: updating the speech recognition text, the local time information of the word elements, and the global time information based on the word element change information sent by the client, the updated local time information, and the global time information. This processing method effectively improves the accuracy of speech recognition text by storing the word element information edited by the client user.

[0193] As can be seen from the above embodiments, the voice playback method provided in this application identifies whether multiple voice files belong to the same activity and whether these voice files need to be played continuously on the front-end application side, and synchronously displays the voice recognition text corresponding to the playback progress of the entire activity. On the server side, the voice recognition service performs voice recognition on each voice file to obtain the local time information of each word element in the voice recognition text relative to the start point of the corresponding voice file. The voice recognition text and local time information are sent to the front-end application side, and the front-end application side determines the global time information of each word element in the recognition text of each voice file relative to the start point of the activity. The voice playback controller automatically opens multiple voice files in the voice file playlist in sequence to play multiple segments of voice data of the entire activity continuously and displays the voice recognition text corresponding to the playback progress of the entire activity. The time information corresponding to the displayed voice recognition text is the global time information. This realizes the processing of merging and playing multiple segments of voice data of the same activity and synchronously highlighting the recognition text corresponding to the global playback progress. This approach avoids physically merging multiple audio files from the same activity through the server-side speech recognition module, without altering the atomic speech recognition logic provided by the server. Therefore, in application scenarios where multiple audio files from the entire activity are played back sequentially and the speech-recognized text is displayed simultaneously, the coupling of the server-side speech recognition service to the application is effectively reduced, achieving a seamless user experience where the entire activity's audio is played back and the speech-recognized text is displayed simultaneously. Furthermore, because this processing method does not physically merge multiple audio files from the same activity into a single audio file, but rather stores multiple audio files independently, along with their respective speech-recognition texts, it provides an effective data foundation for flexibly addressing various variable scenarios tailored to user needs within this application.

[0194] Sixth Embodiment

[0195] In the above embodiments, a voice playback method is provided. Correspondingly, this application also provides a voice playback device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0196] This application also provides a voice playback device, including:

[0197] The request receiving unit is used to receive speech recognition requests for multiple audio files in the target activity.

[0198] The speech recognition unit is used to perform speech recognition processing on the plurality of speech files;

[0199] The data sending unit is configured to send local time information of word elements in multiple speech recognition texts relative to the start point of their respective speech files to the client, so that the client can determine the global time information of the word elements relative to the start point of the target activity based on the time information of the multiple speech files and the local time information; and to sequentially open the multiple speech files in the speech file playlist in the speech playback controller to continuously play multiple segments of speech data of the target activity corresponding to the multiple speech files; and to display the speech recognition text corresponding to the speech playback progress of the target activity, wherein the time information corresponding to the displayed speech recognition text includes global time information.

[0200] Seventh Embodiment

[0201] In the above embodiments, a voice playback method is provided. Correspondingly, this application also provides an electronic device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0202] This embodiment provides an electronic device, comprising: a processor and a memory; the memory for storing a program implementing a voice playback method, wherein after the device is powered on and the program of the method is run by the processor, the following steps are performed: receiving voice recognition requests for multiple voice files in a target activity; performing voice recognition processing on the multiple voice files; sending local time information of word elements in multiple voice recognition texts relative to the start point of their respective voice files to a client, so that the client can determine the global time information of the word elements relative to the start point of the target activity based on the time information of the multiple voice files and the local time information; and sequentially opening the multiple voice files in the voice file playlist in a voice playback controller to continuously play multiple segments of voice data of the target activity corresponding to the multiple voice files; displaying voice recognition text corresponding to the voice playback progress of the target activity, wherein the time information corresponding to the displayed voice recognition text includes global time information.

[0203] Eighth embodiment

[0204] Corresponding to the aforementioned voice playback system, this application also provides a method for playing lecture voice. The execution subject of the method includes, but is not limited to, a client, and can also be any device capable of implementing the method. Since the method embodiments are basically similar to the system embodiments, the description is relatively simple; relevant details can be found in the descriptions of the system embodiments. The method embodiments described below are merely illustrative. Parts identical to those in the first embodiment will not be repeated; please refer to the corresponding parts in Embodiment 1.

[0205] In this embodiment, the method may include the following steps:

[0206] Step 1: Identify the multiple audio files included in the lecture;

[0207] Step 2: Receive the teaching content text corresponding to the teaching audio file sent by the server, and the local time information of the word elements in the text relative to the start point of the corresponding audio file;

[0208] Step 3: Based on the time information of the multiple teaching audio files and the local time information, determine the global time information of the word element relative to the start point of the teaching process;

[0209] Step 4: In the voice playback controller, sequentially open the multiple teaching voice files in the voice file playlist to continuously play multiple segments of voice data in the teaching process corresponding to the multiple teaching voice files, and display the teaching content text corresponding to the voice playback progress in the teaching process. The time information corresponding to the teaching content text includes global time information.

[0210] The teaching process may include multiple teaching topics, and different teaching audio files may correspond to different teaching topics. In one example, the method may further include the following steps: determining the target teaching topic; opening the target teaching audio file corresponding to the target teaching topic in the audio playback controller, and displaying the target teaching content text corresponding to the audio playback progress of the target teaching audio file, wherein the time information corresponding to the target teaching content text includes local time information.

[0211] Ninth Embodiment

[0212] Corresponding to the aforementioned voice playback system, this application also provides a live voice playback method. The execution subject of the method includes, but is not limited to, a client, and can also be any device capable of implementing the method. Since the method embodiments are basically similar to the system embodiments, the description is relatively simple; relevant details can be found in the descriptions of the system embodiments. The method embodiments described below are merely illustrative. The parts of this embodiment that are the same as those in the first embodiment will not be repeated; please refer to the corresponding parts in Embodiment 1.

[0213] In this embodiment, the method may include the following steps:

[0214] Step 1: Identify the multiple audio files included in the live stream;

[0215] Step 2: Receive the live content text corresponding to the live audio file and the local time information of the word elements in the text relative to the start point of the corresponding audio file sent by the server;

[0216] Step 3: Based on the time information of the multiple live audio files and the local time information, determine the global time information of the word element relative to the start point of the live broadcast process;

[0217] Step 4: In the voice playback controller, sequentially open the multiple live audio files in the audio file playlist to continuously play multiple segments of audio data during the live broadcast corresponding to the multiple live audio files, and display the live content text corresponding to the audio playback progress during the live broadcast. The time information corresponding to the live content text includes global time information.

[0218] The live streaming process may include multiple live streaming themes, and different live streaming audio files may correspond to different live streaming themes. In one example, the method may further include the following steps: determining the target live streaming theme; opening the target live streaming audio file corresponding to the target live streaming theme in the audio playback controller, and displaying the target live streaming content text corresponding to the audio playback progress of the target live streaming audio file, wherein the time information corresponding to the target live streaming content text includes local time information.

[0219] Tenth Embodiment

[0220] Corresponding to the aforementioned audio playback system, this application also provides a conference audio playback method. The execution subject of the method includes, but is not limited to, a client, and can also be any device capable of implementing the method. Since the method embodiments are basically similar to the system embodiments, the description is relatively simple; relevant details can be found in the descriptions of the system embodiments. The method embodiments described below are merely illustrative. The parts of this embodiment that are the same as those in the first embodiment will not be repeated; please refer to the corresponding parts in Embodiment 1.

[0221] In this embodiment, the method may include the following steps:

[0222] Step 1: Identify the multiple audio files included in the target meeting;

[0223] Step 2: Receive the speech recognition text corresponding to the speech file and the local time information of the word elements in the text relative to the start point of the speech file sent by the server;

[0224] Step 3: Based on the time information of the multiple audio files and the local time information, determine the global time information of the word element relative to the target meeting start point;

[0225] Step 4: In the voice playback controller, sequentially open the multiple voice files in the voice file playlist to continuously play multiple segments of voice data from the target conference corresponding to the multiple voice files, and display the voice recognition text corresponding to the voice playback progress of the target conference. The time information corresponding to the displayed voice recognition text includes global time information.

[0226] The target meeting may include multiple sub-topics, and different audio files may correspond to different sub-topics. In one example, the method may further include the following steps: determining the target sub-topic; opening the target audio file corresponding to the target sub-topic in the audio playback controller, and displaying the target speech recognition text corresponding to the audio playback progress of the target audio file, wherein the time information corresponding to the target speech recognition text includes local time information.

[0227] Eleventh Embodiment

[0228] Corresponding to the aforementioned audio playback system, this application also provides a method for playing court hearing audio. The execution entity of the method includes, but is not limited to, a client, and can also be any device capable of implementing the method. Since the method embodiments are basically similar to the system embodiments, the description is relatively simple; relevant details can be found in the descriptions of the system embodiments. The method embodiments described below are merely illustrative. Parts identical to those in the first embodiment will not be repeated; please refer to the corresponding parts in Embodiment 1.

[0229] In this embodiment, the method may include the following steps:

[0230] Step 1: Identify the multiple audio files included in the court proceedings;

[0231] Step 2: Receive the court hearing content text corresponding to the court hearing audio file sent by the server, and the local time information of the word elements in the text relative to the start point of the corresponding audio file;

[0232] Step 3: Based on the time information of the multiple court hearing audio files and the local time information, determine the global time information of the word element relative to the start point of the court hearing process;

[0233] Step 4: In the voice playback controller, sequentially open the multiple court hearing audio files in the audio file playlist to continuously play multiple segments of audio data of the court hearing process corresponding to the multiple court hearing audio files, and display the court hearing content text corresponding to the audio playback progress in the court hearing process. The time information corresponding to the displayed court hearing content text includes global time information.

[0234] The court hearing process may include multiple stages and themes, and different court hearing audio files may correspond to different stage themes. In one example, the method may further include the following steps: determining the target stage theme; opening the target court hearing audio file corresponding to the target stage theme in the audio playback controller, and displaying the target court hearing content text corresponding to the audio playback progress of the target court hearing audio file, wherein the time information corresponding to the target court hearing content text includes local time information.

[0235] Twelfth Embodiment

[0236] Corresponding to the various methods described above, this application also provides a computer program. Since this program embodiment is fundamentally similar to the method embodiment, it is described simply; relevant details can be found in the descriptions within the method embodiment. The program embodiment described below is merely illustrative.

[0237] The computer program provided in this application embodiment, when run on a computer, enables the computer to execute the various methods provided in the above embodiments.

[0238] The program includes, but is not limited to, applications deployed on servers or terminal devices, mobile applications (APPs) deployed on mobile devices, mini-programs within APPs, and other forms.

[0239] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.

[0240] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0241] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0242] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0243] 2. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A voice playback system, characterized in that, include: The client is used to determine multiple audio files included in the target activity, wherein the multiple audio files have a temporal order; receive speech recognition text corresponding to the audio files and local time information of word elements in the text relative to the start point of their respective audio files sent by the server; and determine global time information of the word elements relative to the start point of the target activity based on the time information of the multiple audio files and the local time information. Furthermore, the system sequentially opens multiple audio files in the audio file playlist within the audio playback controller to continuously play multiple segments of audio data corresponding to the target activity; displays speech recognition text corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition text includes global time information; and sends information about the target activity including the multiple audio files and global time information to the server, so that the server stores the global time information and the information about the target activity including the multiple audio files, so that the server can respond to audio playback requests for the target activity sent by other clients, and send the multiple audio files, the multiple speech recognition text, the local time information, and the global time information to other clients, so that other clients can continuously play multiple segments of audio data of the target activity, display speech recognition text corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition text includes global time information; the users of the clients are activity administrators, and the users of the other clients are activity followers; The server is used to perform speech recognition processing on the speech files, send the speech recognition text and local time information of the word elements to the client; store the local time information of the word elements in the multiple speech files and multiple speech recognition texts, as well as the information and global time information of the target activity sent by the client, including the multiple speech files; and receive speech playback requests for the target activity sent by other clients. The system sends the target activity, including the multiple audio files, the multiple speech recognition texts, and the global time information, to other clients so that other clients can play multiple audio data segments of the target activity continuously, display speech recognition texts corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition texts includes global time information.

2. A voice playback method, characterized in that, include: The target activity is identified as including multiple audio files, which are arranged in a chronological order. Receive the speech recognition text corresponding to the speech file and the local time information of the word elements in the text relative to the start point of the speech file sent by the server; Based on the time information of the multiple audio files and the local time information, the global time information of the word element relative to the starting point of the target activity is determined; In the voice playback controller, multiple voice files in the voice file playlist are opened sequentially to play multiple segments of voice data corresponding to the target activity. The voice recognition text corresponding to the playback progress of the target activity is displayed, and the time information corresponding to the displayed voice recognition text includes global time information. The server sends information about the target activity including the multiple voice files and global time information, allowing the server to store the global time information and the information about the target activity including the multiple voice files. This enables the server to respond to voice playback requests for the target activity from other clients, sending the multiple voice files, the multiple voice recognition text, the local time information, and the global time information to other clients. This allows other clients to play multiple segments of voice data of the target activity continuously, display the voice recognition text corresponding to the playback progress of the target activity, and the time information corresponding to the displayed voice recognition text includes global time information. The users of the clients are activity administrators, and the users of the other clients are activity followers.

3. The method according to claim 2, characterized in that, Also includes: Identify the target audio file; The target audio file is opened in the audio playback controller, and the target speech recognition text corresponding to the audio playback progress of the target audio file is displayed. The time information corresponding to the displayed target speech recognition text includes local time information.

4. The method according to claim 3, characterized in that, The target activities include multiple activity themes; The method further includes: Determine the topic information of the audio file; The determined target audio file includes: Determine the target topic information; Use the audio file that corresponds to the target topic information as the target audio file.

5. The method according to claim 4, characterized in that, Also includes: The server sends information about the target activity, including the multiple audio files, topic information, and global time information, to the server. The server stores the global time information, the target activity information including the multiple audio files, and the topic information. This allows the server to respond to audio playback requests from other clients for a specific topic, sending the target audio file corresponding to the target topic, the target speech recognition text corresponding to the target audio file, and the local time information to other clients. This allows the clients to play the audio data for the target topic, display the target speech recognition text corresponding to the audio playback progress of the target audio file, and the time information corresponding to the displayed target speech recognition text includes local time information.

6. The method according to claim 2, characterized in that, Also includes: Edit the speech-recognition text while playing multiple segments of speech data of the target activity in a continuous manner; Determine the updated global time information and updated local time information of word elements in the edited speech recognition text.

7. The method according to claim 6, characterized in that, The determination of the updated global time information and updated local time information of word elements in the edited speech recognition text includes: Determine the updated global time information; Based on the updated global time information, the updated local time information is determined.

8. The method according to claim 6, characterized in that, Also includes: The updated global time information and the updated local time information are sent to the server, so that the server updates the global time information and the local time information.

9. The method according to claim 6, characterized in that, The editing of speech recognition text includes at least one of the following methods: modifying word elements, adding word elements, and deleting word elements.

10. A method for playing voice messages, characterized in that, include: Receive speech recognition requests for multiple audio files in a target activity, wherein the multiple audio files are in chronological order; Perform speech recognition processing on the multiple audio files; Send local time information of word elements in multiple speech recognition texts relative to the start point of their respective speech files to the client, so that the client can determine the global time information of the word elements relative to the start point of the target activity based on the time information of multiple speech files and the local time information; Furthermore, the system sequentially opens multiple audio files in the audio file playlist within the audio playback controller to continuously play multiple segments of audio data corresponding to the target activity; displays speech recognition text corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition text includes global time information; stores local time information of word elements in the multiple audio files and multiple speech recognition texts, as well as information about the target activity including the multiple audio files and global time information sent by the client; and receives audio playback requests for the target activity sent by other clients. The system sends the target activity, including the multiple audio files, the multiple speech recognition texts, and the global time information, to other clients so that other clients can play multiple audio data segments of the target activity continuously, display speech recognition texts corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition texts includes global time information.

11. The method according to claim 10, characterized in that, The target activities include multiple activity themes; Also includes: Store the topic information of the voice file sent by the client; The topic information is sent to other clients so that they can play the audio data of the target topic, display the target speech recognition text corresponding to the audio playback progress of the audio file of the target topic, and the time information corresponding to the target speech recognition text includes local time information.

12. The method according to claim 11, characterized in that, Also includes: Based on the word element change information, updated local time information, and global time information sent by the client, the speech recognition text, word element local time information, and global time information are updated.

13. A voice playback device, characterized in that, include: An activity audio file determination unit is used to determine multiple audio files included in a target activity, wherein the multiple audio files have a temporal order. The data receiving unit is used to receive the speech recognition text corresponding to the speech file and the local time information of the word elements in the text relative to the start point of the speech file sent by the server. A global time determination unit is used to determine the global time information of the word element relative to the target activity start point based on the time information of the multiple speech files and the local time information; The synchronous display unit is used to sequentially open multiple audio files in the audio file playlist in the audio playback controller to continuously play multiple segments of audio data corresponding to the target activity, and display speech recognition text corresponding to the audio playback progress of the target activity. The time information corresponding to the displayed speech recognition text includes global time information. The unit also sends information about the target activity including the multiple audio files and global time information to the server, enabling the server to store the global time information and the information about the target activity including the multiple audio files. This allows the server to respond to audio playback requests for the target activity from other clients, sending the multiple audio files, the multiple speech recognition text, the local time information, and the global time information to other clients. This allows other clients to continuously play multiple segments of audio data of the target activity, display speech recognition text corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition text includes global time information. The users of the clients are activity administrators, and the users of the other clients are activity followers.

14. An electronic device, characterized in that, include: Processor and memory; The memory stores a program for implementing the voice playback method. After the device is powered on and the program for the method is run by the processor, the following steps are performed: determining multiple voice files included in the target activity, wherein the multiple voice files have a temporal order; receiving speech recognition text corresponding to the voice files sent by the server and local time information of word elements in the text relative to the start point of their respective voice files; determining global time information of the word elements relative to the start point of the target activity based on the time information of the multiple voice files and the local time information. In the voice playback controller, multiple voice files in the voice file playlist are opened sequentially to play multiple segments of voice data corresponding to the target activity. The voice recognition text corresponding to the playback progress of the target activity is displayed, and the time information corresponding to the displayed voice recognition text includes global time information. The server sends information about the target activity including the multiple voice files and global time information, allowing the server to store the global time information and the information about the target activity including the multiple voice files. This enables the server to respond to voice playback requests for the target activity from other clients, sending the multiple voice files, the multiple voice recognition text, the local time information, and the global time information to other clients. This allows other clients to play multiple segments of voice data of the target activity continuously, display the voice recognition text corresponding to the playback progress of the target activity, and the time information corresponding to the displayed voice recognition text includes global time information. The users of the clients are activity administrators, and the users of the other clients are activity followers.

15. A voice playback device, characterized in that, include: The request receiving unit is used to receive speech recognition requests for multiple speech files in a target activity, wherein the multiple speech files are in chronological order. The speech recognition unit is used to perform speech recognition processing on the plurality of speech files; The data sending unit is used to send local time information of word elements in multiple speech recognition texts relative to the start point of their respective speech files to the client, so that the client can determine the global time information of the word elements relative to the target activity start point based on the time information of multiple speech files and the local time information; Furthermore, the system sequentially opens multiple audio files in the audio file playlist within the audio playback controller to continuously play multiple segments of audio data corresponding to the target activity; displays speech recognition text corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition text includes global time information; stores local time information of word elements in the multiple audio files and multiple speech recognition texts, as well as information about the target activity including the multiple audio files and global time information sent by the client; and receives audio playback requests for the target activity sent by other clients. The system sends the target activity, including the multiple audio files, the multiple speech recognition texts, and the global time information, to other clients so that other clients can play multiple audio data segments of the target activity continuously, display speech recognition texts corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition texts includes global time information.

16. An electronic device, characterized in that, include: Processor and memory; The memory stores a program for implementing the voice playback method. After the device is powered on and the program for the method is run by the processor, the following steps are performed: receiving voice recognition requests for multiple voice files in a target activity, wherein the multiple voice files are in chronological order; performing voice recognition processing on the multiple voice files; and sending local time information of word elements in multiple voice recognition texts relative to the start point of their respective voice files to the client, so that the client can determine the global time information of the word elements relative to the start point of the target activity based on the time information of the multiple voice files and the local time information. Furthermore, the system sequentially opens multiple audio files in the audio file playlist in the audio playback controller to continuously play multiple segments of audio data corresponding to the target activity; displays speech recognition text corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition text includes global time information, local time information of word elements in the multiple audio files and multiple speech recognition text, and stores information about the target activity including the multiple audio files and global time information sent by the client; and receives audio playback requests for the target activity sent by other clients. The system sends the target activity, including the multiple audio files, the multiple speech recognition texts, and the global time information, to other clients so that other clients can play multiple audio data segments of the target activity continuously, display speech recognition texts corresponding to the audio playback progress of the target activity, and the time information corresponding to the displayed speech recognition texts includes global time information.

17. A method for playing lecture audio, characterized in that, include: The teaching process includes multiple audio files, which are arranged in a chronological order. The system receives the teaching content text corresponding to the teaching audio file and the local time information of the word elements in the text relative to the start point of the corresponding audio file, sent by the server. Based on the time information of the multiple teaching audio files and the local time information, the global time information of the word element relative to the start point of the teaching process is determined; In the voice playback controller, the multiple teaching voice files in the voice file playlist are opened sequentially to play multiple segments of voice data in the teaching process corresponding to the multiple teaching voice files in a continuous manner, and the teaching content text corresponding to the voice playback progress in the teaching process is displayed. The time information corresponding to the teaching content text includes global time information. The system sends information about a target activity, including multiple audio files and global time information, to the server. The server stores the global time information and the information about the target activity including the multiple audio files. This allows the server to respond to audio playback requests from other clients for the target activity by sending the multiple audio files, multiple speech-recognized texts, the local time information, and the global time information to other clients. This allows other clients to continuously play multiple audio segments of the target activity and display speech-recognized texts corresponding to the audio playback progress of the target activity. The time information corresponding to the displayed speech-recognized texts includes the global time information. The users of the clients are activity administrators, and the users of the other clients are activity followers.

18. The method according to claim 17, characterized in that, The teaching process includes multiple teaching topics, with different teaching audio files corresponding to different teaching topics; The method further includes: Define the target teaching topic; Open the target lecture audio file corresponding to the target lecture topic in the audio playback controller, and display the target lecture content text corresponding to the audio playback progress of the target lecture audio file. The time information corresponding to the target lecture content text includes local time information.

19. A method for playing live audio, characterized in that, include: The live broadcast process includes multiple live audio files, which are arranged in a chronological order. Receive the live content text corresponding to the live audio file and the local time information of the word elements in the text relative to the start point of the audio file sent by the server; Based on the time information of the multiple live audio files and the local time information, the global time information of the word element relative to the start point of the live broadcast process is determined; In the voice playback controller, the multiple live audio files in the audio file playlist are opened sequentially to play multiple segments of audio data in the live broadcast process corresponding to the multiple live audio files in a continuous manner, and the live content text corresponding to the audio playback progress in the live broadcast process is displayed. The time information corresponding to the live content text includes global time information. The system sends information about the target activity, including the multiple audio files and global time information, to the server. The server stores the global time information and the information about the target activity including the multiple audio files. This allows the server to respond to audio playback requests from other clients for the target activity, sending the multiple audio files, multiple speech-recognized texts, the local time information, and the global time information to other clients. This allows other clients to continuously play multiple segments of audio data related to the target activity, displaying speech-recognized texts corresponding to the audio playback progress of the target activity. The time information corresponding to the displayed speech-recognized texts includes the global time information. The users of the clients are activity administrators, and the users of the other clients are activity followers.

20. The method according to claim 19, characterized in that, The live streaming process includes multiple live streaming themes, and different live streaming audio files correspond to different live streaming themes; The method further includes: Determine the target live stream topic; Open the target live audio file corresponding to the target live topic in the audio playback controller, and display the target live content text corresponding to the audio playback progress of the target live audio file. The time information corresponding to the target live content text includes local time information.

21. A method for playing conference audio, characterized in that, include: The target meeting includes multiple audio files, which are arranged in chronological order. Receive the speech recognition text corresponding to the speech file and the local time information of the word elements in the text relative to the start point of the speech file sent by the server; Based on the time information of the multiple audio files and the local time information, the global time information of the word element relative to the target meeting start point is determined; In the voice playback controller, the multiple voice files in the voice file playlist are opened sequentially to play multiple segments of voice data of the target conference corresponding to the multiple voice files in a continuous manner, and the voice recognition text corresponding to the voice playback progress of the target conference is displayed. The time information corresponding to the displayed voice recognition text includes global time information. The system sends information about the target activity, including the multiple audio files and global time information, to the server. The server stores the global time information and the information about the target activity including the multiple audio files. This allows the server to respond to audio playback requests from other clients for the target activity, sending the multiple audio files, the multiple speech-recognized texts, the local time information, and the global time information to other clients. This allows other clients to continuously play multiple audio segments of the target activity and display speech-recognized texts corresponding to the audio playback progress of the target activity. The time information corresponding to the displayed speech-recognized texts includes the global time information. The users of the clients are activity administrators, and the users of the other clients are activity followers.

22. The method according to claim 21, characterized in that, The target meeting includes multiple sub-topics, with different audio files corresponding to different sub-topics; The method further includes: Identify target sub-topics; Open the target audio file corresponding to the target sub-topic in the audio playback controller, and display the target speech recognition text corresponding to the audio playback progress of the target audio file. The time information corresponding to the target speech recognition text includes local time information.

23. A method for playing court hearing audio, characterized in that, include: The court hearing process includes multiple audio files, which are arranged in chronological order. The server receives the text of the court hearing content corresponding to the court hearing audio file and the local time information of the word elements in the text relative to the start point of the audio file. Based on the time information of the multiple court hearing audio files and the local time information, the global time information of the word element relative to the start point of the court hearing process is determined; In the voice playback controller, the multiple court hearing voice files in the voice file playlist are opened sequentially to play multiple segments of voice data of the court hearing process corresponding to the multiple court hearing voice files in a continuous manner, and the court hearing content text corresponding to the voice playback progress in the court hearing process is displayed. The time information corresponding to the displayed court hearing content text includes global time information. The system sends information about the target activity, including the multiple audio files and global time information, to the server. The server stores the global time information and the information about the target activity including the multiple audio files. This allows the server to respond to audio playback requests from other clients for the target activity, sending the multiple audio files, the multiple speech-recognized texts, the local time information, and the global time information to other clients. This allows other clients to continuously play multiple audio segments of the target activity and display speech-recognized texts corresponding to the audio playback progress of the target activity. The time information corresponding to the displayed speech-recognized texts includes the global time information. The users of the clients are activity administrators, and the users of the other clients are activity followers.

24. The method according to claim 23, characterized in that, The court proceedings included multiple themes, with different audio files corresponding to different themes. The method further includes: Define the theme for the target phase; Open the target court hearing audio file corresponding to the target stage theme in the audio playback controller, and display the target court hearing content text corresponding to the audio playback progress of the target court hearing audio file. The time information corresponding to the target court hearing content text includes local time information.

25. A computer program product, characterized in that, When it is run on a computer, it causes the computer to perform the method according to any one of claims 2 to 12, or claims 17 to 24.