A method, apparatus and device for displaying subtitles

By establishing a long connection between the server and the device, the problems of low resource utilization and poor text display efficiency are solved, enabling efficient and accurate display of real-time captions.

CN116246192BActive Publication Date: 2026-05-05MASHANG CONSUMER FINANCE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MASHANG CONSUMER FINANCE CO LTD
Filing Date
2021-12-08
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

When there is a high demand for voice conversion, the existing technology frequently establishes multiple connections between the device and the server, resulting in low resource utilization, long text display time, and poor accuracy, which affects the synchronization and display effect of real-time subtitles.

Method used

By establishing a long-lived connection between the server and the device and continuously sending multiple data packets, the number of connection creations is reduced, and text conversion and correction of the voice stream are performed, thereby improving resource utilization and text display efficiency.

Benefits of technology

It reduces the resource consumption of connection creation, saves time, improves the accuracy and efficiency of text display, and enables real-time synchronous display of subtitles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246192B_ABST
    Figure CN116246192B_ABST
Patent Text Reader

Abstract

The embodiment of the specification discloses a subtitle display method, device and equipment, the method comprises the following steps: a long connection is established with a first device, and the target voice stream to be converted collected by the first device is received through the long connection; text conversion processing is performed on the target voice stream, first text data corresponding to the target voice stream is obtained, and correction processing is performed on the first text data, second text data corresponding to the target voice stream is obtained; based on the first text data and the second text data, it is determined that the subtitle is target text display data corresponding to the target voice stream. Through the above-mentioned subtitle display method, a long connection can be established with the first device, the resource utilization rate is improved, and the text data display efficiency and accuracy of the target voice stream are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of computer technology, and in particular to a method, apparatus, and device for displaying subtitles. Background Technology

[0002] With the rapid development of computer technology, the demand for real-time subtitle display is increasing. In scenarios such as watching online live courses and participating in video conferences, in order to enable users to more intuitively obtain the content explained by the presenter, it is necessary to convert the presenter's speech into subtitles for display.

[0003] Typically, an audio stream can be divided into multiple audio segments, and a connection can be established between each audio segment and the server. The audio segments are then sent to the server for text conversion processing to obtain text data corresponding to the audio stream. The subtitles are then displayed in real time based on the text data.

[0004] However, when the demand for speech conversion is large, the above method requires establishing a large number of connections between the devices (including speech acquisition devices, subtitle display devices, etc.) and the server. Due to the frequent establishment of a large number of connections between the devices and the server, and the need to divide the speech stream into multiple speech segments for separate recognition, the time required to obtain complete subtitles is very long, making it difficult to synchronize speech and subtitles, resulting in low text display accuracy. Therefore, a technical solution is needed to improve resource utilization and the efficiency and accuracy of text display for speech data. Summary of the Invention

[0005] The purpose of the embodiments in this specification is to provide a technical solution to improve resource utilization and the efficiency and accuracy of text display for voice data.

[0006] To achieve the above technical solution, the embodiments in this specification are implemented as follows:

[0007] This specification provides an embodiment of a method for displaying subtitles, the method comprising:

[0008] Establish a long connection with the first device, and receive the target speech stream to be converted collected by the first device through the long connection;

[0009] The target speech stream is subjected to text conversion processing to obtain first text data corresponding to the target speech stream, and the first text data is subjected to correction processing to obtain second text data corresponding to the target speech stream.

[0010] Based on the first text data and the second text data, the subtitle is determined to be the target text display data corresponding to the target speech stream.

[0011] This specification provides an embodiment of a subtitle display device, the device comprising:

[0012] The connection establishment module is configured to establish a long connection with the first device and receive the target speech stream to be converted collected by the first device through the long connection;

[0013] The data conversion module is configured to perform text conversion processing on the target speech stream to obtain first text data corresponding to the target speech stream, and to perform correction processing on the first text data to obtain second text data corresponding to the target speech stream.

[0014] The data determination module is configured to determine, based on the first text data and the second text data, that the subtitle is the target text display data corresponding to the target speech stream.

[0015] This specification provides an embodiment of a subtitle display device, which includes:

[0016] Processor; and

[0017] A memory configured to store computer-executable instructions, which, when executed, cause the processor to: establish a long connection with a first device and receive, through the long connection, a target speech stream to be converted collected by the first device;

[0018] The target speech stream is subjected to text conversion processing to obtain first text data corresponding to the target speech stream, and the first text data is subjected to correction processing to obtain second text data corresponding to the target speech stream.

[0019] Based on the first text data and the second text data, the subtitle is determined to be the target text display data corresponding to the target speech stream.

[0020] This specification also provides a storage medium for storing computer-executable instructions, which, when executed, implement the following process:

[0021] Establish a long connection with the first device, and receive the target speech stream to be converted collected by the first device through the long connection;

[0022] The target speech stream is subjected to text conversion processing to obtain first text data corresponding to the target speech stream, and the first text data is subjected to correction processing to obtain second text data corresponding to the target speech stream.

[0023] Based on the first text data and the second text data, the subtitle is determined to be the target text display data corresponding to the target speech stream.

[0024] The technical solution provided in this specification establishes a long connection between the server and the first device. A long connection means that multiple data packets can be continuously sent between the server and the first device during the connection period without establishing a connection for each data packet. Therefore, transmitting the target audio stream through a long connection can reduce the resource consumption of frequently creating connections, improve resource utilization, and save the time spent creating connections. This also improves the efficiency of obtaining target text display data for the target audio stream. In addition, the target text display data corresponding to the target audio stream, i.e., the subtitles to be displayed, can be accurately determined through the first text data and the second text data, thus improving the accuracy of determining the text display data for the target audio stream. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is an embodiment of a method for displaying subtitles as described in this specification;

[0027] Figure 2 This is a schematic diagram of a subtitle display system architecture for this specification;

[0028] Figure 3 This is yet another embodiment of a method for displaying subtitles as described in this specification;

[0029] Figure 4 This is yet another embodiment of a method for displaying subtitles as described in this specification;

[0030] Figure 5 This is a schematic diagram illustrating the display process of one type of subtitle in this instruction manual;

[0031] Figure 6 This is yet another embodiment of a method for displaying subtitles as described in this specification;

[0032] Figure 7 This is yet another embodiment of a method for displaying subtitles as described in this specification;

[0033] Figure 8 This is a schematic diagram illustrating the processing of a target speech stream as described in this specification;

[0034] Figure 9 This is yet another embodiment of a method for displaying subtitles as described in this specification;

[0035] Figure 10 This is a schematic diagram illustrating yet another method of displaying subtitles in this instruction manual;

[0036] Figure 11 This is an embodiment of a subtitle display device described in this specification;

[0037] Figure 12 This is an embodiment of a subtitle display device described in this specification. Detailed Implementation

[0038] This specification provides a method, apparatus, and device for displaying subtitles.

[0039] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0040] The inventive concept of this application is as follows: With the rapid development of computer technology, the demand for real-time subtitle display is increasing. In scenarios such as watching online live courses and participating in video conferences, in order to enable users to more intuitively obtain the content explained by the presenter, it is necessary to convert the presenter's speech into subtitles for display. For example, real-time text conversion of audio data can be achieved through HTTP(S) API calls. Specifically, an audio segment can be divided into multiple audio streams, and a connection can be created for each audio stream to obtain and display the text data for each stream. When the demand for audio conversion is large, multiple connections need to be established with the server to send the audio segments to be converted to the server for text conversion processing, obtaining the text data corresponding to the audio stream and displaying it. Because multiple connections need to be established frequently, resource utilization is low, resulting in data waste. At the same time, establishing multiple connections and transmitting data separately will result in poor accuracy of the text data for the audio stream, and the connection time will also affect the efficiency of text data acquisition, which will cause a delay in the display of subtitles. Furthermore, since HTTP(S) is a stateless protocol, the accuracy of obtaining text display data for audio streams through the above method is poor, affecting the real-time display effect and resulting in a poor user experience. Therefore, a technical solution is needed to improve resource utilization and the efficiency and accuracy of text display for voice data. This technical solution obtains the target voice stream to be converted by establishing a long connection between the server and the first device. Then, it performs text conversion processing on the target voice stream to obtain first text data. One or more first text data are corrected to obtain second text data. Finally, based on the first and second text data, the subtitles are determined to be the target text display data corresponding to the target voice stream. In this way, by establishing a long connection between the server and the first device for voice stream transmission, the resource consumption of frequently creating connections is reduced, improving resource utilization. At the same time, it also saves the time spent creating connections and improves the efficiency of obtaining target text display data. In addition, determining the target text display data through the first and second text data can also improve the accuracy of determining the text display data for the target voice stream.

[0041] like Figure 1 As shown in the embodiments of this specification, a method for displaying subtitles is provided. The execution subject of this method can be a server, which can be a single independent server or a server cluster composed of multiple different servers. The server can be a server that provides services for converting voice data into text data, etc., and the specific configuration can be determined according to the actual situation. This method can be applied to the processing of voice streams.

[0042] like Figure 2As shown in the embodiments of this specification, the system architecture corresponding to the subtitle display method may include a server 201 and one or more first devices 202. The server 201 communicates with each first device 202. The first device 202 can be a device capable of collecting voice data, such as a mobile terminal device with an audio acquisition device, such as a mobile phone or tablet computer; a terminal device with an audio acquisition device, such as a laptop computer; or a wearable device with an audio acquisition device, such as a smartwatch or bracelet. The first device 202 can send data processing requests such as long connection establishment requests and voice stream conversion requests to the server 201. After the server 201 detects the relevant information of the first device 202 through a pre-set processing mechanism and determines that the first device 202 can establish a connection with the server 201, it can establish a long connection with the first device 202 and perform corresponding data processing operations.

[0043] This method may specifically include the following steps:

[0044] In step S102, a long connection is established with the first device, and the target speech stream to be converted collected by the first device is received through the long connection.

[0045] The first device can be, as described above, the first device 202. This first device can be a mobile terminal device, wearable device, or other device capable of collecting voice data, and can be specifically configured according to actual conditions. A long connection refers to the ability to continuously send multiple data packets during the connection period without establishing a connection for each data packet. The target voice stream to be converted can be any voice stream. For example, the target voice stream can be the voice content collected in real-time by the first device when a presenter uses a microphone or other amplification equipment during a video conference.

[0046] In practice, with the rapid development of computer technology, the demand for real-time subtitle display is increasing to provide users with a better viewing experience. For example, in scenarios such as watching online live courses or participating in video conferences, it is necessary to convert the speaker's speech into subtitles for display so that users can more intuitively understand the content being presented. To improve resource utilization and the efficiency and accuracy of text display for audio data, this specification provides an achievable processing method, which may include the following:

[0047] The first device can send a long connection establishment request to the server. The server can respond to the request by authenticating the first device and establishing a long connection after successful authentication. After establishing the long connection, the server can also return the long connection establishment result to the first device (such as a message indicating successful or failed connection establishment).

[0048] After receiving a message from the server indicating that the long connection has been successfully established, the first device can send the collected target audio stream to be converted to the server through the long connection. In other words, the server can receive the target audio stream collected by the first device through the long connection.

[0049] Because a long-lived connection is established between the server and the first device, the first device can continuously send the target audio stream to the server through the long-lived connection while the connection is maintained, thus avoiding the resource waste caused by frequent connection establishment.

[0050] In addition, the first device can collect one or more target voice streams to be converted. During the long connection period between the server and the first device, the server can receive one or more target voice streams sent by the first device through the long connection, thereby improving resource utilization.

[0051] In step S104, the target speech stream is processed by text conversion to obtain the first text data corresponding to the target speech stream, and the first text data is processed by correction to obtain the second text data corresponding to the target speech stream.

[0052] In practice, the target speech stream can be converted into text based on a preset speech recognition algorithm (such as Automatic Speech Recognition (ASR)) to obtain the first text data corresponding to the target speech stream.

[0053] Furthermore, before performing text conversion on the target speech stream, the server can perform type analysis to determine the corresponding language type. After determining the language type, it can obtain the ASR algorithm corresponding to that speech type and perform text conversion on the target speech stream based on the obtained ASR algorithm.

[0054] The target speech stream can be in various languages, such as Mandarin Chinese, dialects (specifically, Shanghainese, Sichuanese, etc.), English, and Japanese.

[0055] For example, the server can obtain a speech segment of a preset length (e.g., 10 seconds) containing human voice data from the target speech stream, then determine the language type corresponding to the speech segment using a preset type determination algorithm, and obtain the ASR algorithm corresponding to the speech type, so as to perform text conversion processing on the target speech stream based on the obtained ASR algorithm.

[0056] Since data processing failures may occur during the text conversion process of the target speech stream, such as missing a speech data point and failing to convert it to text, or repeatedly converting a speech data point to text, the first text data can be corrected to obtain the second text data corresponding to the target speech stream.

[0057] For example, the first text data corresponding to the human voice data contained in the target speech stream can be obtained, and then the first text data can be corrected using a pre-trained text correction model to obtain the corresponding second text data. For instance, if the first text data can be "data", "of", "processing", or "management", inputting it into the pre-trained text correction model will result in the second text data being "data processing". The text correction model can be trained based on historical first and second text data, using a model constructed from a pre-defined machine learning algorithm.

[0058] The method for determining the first text data and the second text data described above is an optional and implementable method. In actual application scenarios, there can be a variety of different methods, which may vary depending on the specific application scenario. This specification does not specifically limit the methods used in this embodiment.

[0059] In step S106, based on the first text data and the second text data, the subtitle is determined to be the target text display data corresponding to the target audio stream.

[0060] In implementation, the target text display data corresponding to the target speech stream can be obtained by using a pre-trained display data determination model, first text data, and second text data. The display data determination model can be used to determine the target text display data corresponding to the target speech stream, and the display data determination model can be obtained by training a model constructed by a preset machine learning algorithm based on historical first text data and historical second text data.

[0061] Alternatively, the target text display data corresponding to the target speech stream can be determined by using preset text display data determination rules, first text data, and second text data. For example, a preset keyword extraction algorithm can be used to extract the first keyword from the first text data and the second keyword from the second text data. If the second keyword completely contains the first keyword, then the second text data can be determined as the target text display data corresponding to the target speech stream. Specifically, assuming the first text data can be "data", "information", "of", "place", or "place", and the second text data is "data processing", the first keyword corresponding to the first text data can be "data", "information", and "place", and the keyword corresponding to the second text data can be "data" and "processing". Since the second keyword completely contains the first keyword, "data processing" can be determined as the target text display data corresponding to the target speech stream, and the target text display data can be displayed as subtitles.

[0062] The method for determining the target text display data described above is an optional and feasible method. In actual application scenarios, there can be many different methods, which may vary depending on the specific application scenario. This specification does not specifically limit the methods used in this embodiment.

[0063] This specification provides a method for displaying subtitles. A long-lived connection is established with a first device, and the target audio stream to be converted is received from the first device via this long-lived connection. The target audio stream undergoes text conversion processing to obtain first text data corresponding to the target audio stream. The first text data is then corrected to obtain second text data corresponding to the target audio stream. Based on the first and second text data, the subtitles are determined to be the target text display data corresponding to the target audio stream. This method utilizes a long-lived connection between the server and the first device. A long-lived connection allows for the continuous transmission of multiple data packets between the server and the first device without requiring a separate connection for each data packet. Therefore, transmitting the target audio stream via a long-lived connection reduces the resource consumption of frequently creating connections, improving resource utilization. It also saves time associated with connection creation, increasing the efficiency of acquiring the target text display data for the target audio stream. Furthermore, the first and second text data accurately determine the target text display data corresponding to the target audio stream, improving the accuracy of the determination of the text display data for the target audio stream.

[0064] In one or more embodiments of this specification, the server can return determined target text display data to the first device for display, thereby achieving text display of the target audio stream, and correspondingly, as shown below. Figure 3 As shown, the following step S202 can also be performed.

[0065] In step S202, the target text display data is returned to the first device via a long connection, and the first device is triggered to display the target text display data corresponding to the target audio stream as subtitles.

[0066] The first device can be a device capable of collecting audio data and displaying text data. For example, the first device can be a terminal device such as a personal computer, or a mobile terminal device such as a mobile phone or tablet computer.

[0067] In practice, during the long connection period between the server and the first device, the server can return the determined target text display data to the first device through the long connection, and trigger the first device to display the target text display data corresponding to the target audio stream as subtitles.

[0068] Since the first device establishes a long-lived connection with the server, the resource waste caused by frequent connection creation can be avoided. At the same time, since the time spent creating the connection is saved, the first device can obtain and display the target text display data corresponding to the target audio stream in a timely manner, realizing real-time text display of the target audio stream and improving the user experience.

[0069] In one or more embodiments of this specification, the first device may be a voice acquisition device, and the server may return the determined target text display data to other devices for display, accordingly, such as Figure 4 As shown, the following step S302 can also be performed.

[0070] In step S302, a long connection is established with the second device, and the target text display data is sent to the second device through the long connection established with the second device, triggering the second device to display the target text display data corresponding to the target voice stream.

[0071] The second device can be a device capable of displaying text data. For example, the second device can be a terminal device such as a personal computer, a mobile terminal device such as a mobile phone or tablet computer, or an electronic device such as a projector or display screen. The specific device can be set according to the actual situation.

[0072] In implementation, such as Figure 5 As shown, the first device can be a device such as a mobile phone or a personal computer. The first device can send the target voice stream to the server through a long connection. After the server determines the target text display data corresponding to the target voice stream, it can establish a long connection with the second device and send the target text display data to the second device through the long connection, triggering the second device to display the target text display data corresponding to the target voice stream.

[0073] In this way, even when the first and second devices are not in the same area, or when the first device does not have text display capabilities, the target text display data corresponding to the target audio stream collected by the first device can be sent to the second device. Furthermore, the server can also send the target audio stream to the second device, triggering the second device to display the target text display data corresponding to the target audio stream while playing the target audio stream.

[0074] In addition, there can be multiple second devices. That is, after the server determines the target text display data, it can send the target text display data to the corresponding second devices and trigger the second devices to display the target text display data corresponding to the target voice stream.

[0075] Additionally, the server can determine the text display type of the target text display data based on the device identifier of the second device, and whether it matches the text display requirements of the second device. If the text display type of the target text display data does not match the text display requirements of the second device, the server can perform a type conversion on the target text display data and send the converted target text display data to the corresponding second device. For example, if the text display type of the target text display data can be Chinese, and assuming that the device identifier of the second device determines that the text display requirement of the second device is English, then the text display type of the target text display data can be converted from Chinese to English, and the converted target text display data can be sent to the corresponding second device.

[0076] Because the server establishes long-lived connections with the first device and the second device, the server only needs to establish one connection with the first device and one connection with the second device to achieve data interaction with the first device and the second device respectively. This avoids the waste of resources caused by frequent connection establishment and the problem of not being able to display text in real time for the target voice stream.

[0077] In one or more embodiments of this specification, the processing of determining the first text data and the second text data in step S104 above can be varied. The following provides one optional processing method, such as... Figure 6 As shown, the specific process may include the following steps S1042 to S1044.

[0078] In step S1042, the target speech stream is processed by text conversion based on the first time interval to obtain the first text data corresponding to the target speech stream.

[0079] The first time interval can be a time interval set for a preset number of characters. For example, the first time interval can be set to 1 second for a single character and 2 seconds for two characters. It can also be other intervals besides the above-mentioned intervals, which can be set according to the actual situation.

[0080] In implementation, for example, the first time interval can be set to 50ms for a single character, and text conversion processing can be performed on the target speech stream every 50ms. The first text data corresponding to the target speech stream can be determined based on the conversion processing result. Specifically, assuming the target speech stream is "data processing method", then, based on the first time interval, the text conversion processing of the target speech stream can yield multiple text data such as "data", "of", and "place".

[0081] In step S1042 above, the text conversion processing of the target speech stream based on the first time interval to obtain the first text data corresponding to the target speech stream can be performed in various ways. The following provides an optional processing method, which may specifically include the processing of steps A2 to A6.

[0082] In step A2, the target parameters sent by the first device are received via a long connection.

[0083] The target parameters can include format parameters and sampling parameters of the target speech stream. The format parameters can be used to represent the format of the target speech stream, such as mp3, cda, wav, etc., and the sampling parameters can be the sampling rate, etc.

[0084] In step A4, the target speech stream is validated based on the target parameters, and the validation result is used to determine whether the target speech stream can be converted into text.

[0085] In practice, since the target audio stream can have various formats, and the server may only be able to perform text conversion processing on audio streams of certain specific formats, when the target audio stream sent by the first device is received, the target audio stream can be verified according to the target parameters to determine whether the server can perform text conversion on the target audio stream, that is, to determine whether the format of the target audio stream is a format that the server can process, or to determine whether the server can perform format conversion processing on the target audio stream to convert the format of the target audio stream into a format that the server can process.

[0086] In step A6, if it is determined that the target speech stream can be converted into text, then the target speech stream is processed into text based on the first time interval to obtain the first text data corresponding to the target speech stream.

[0087] The specific processing method of step A6 above can be varied. The following provides another optional processing method, which may include the processing of steps A62 to A66.

[0088] In step A62, the target format conversion algorithm is determined based on the target parameters.

[0089] The target format conversion algorithm can be used to convert the format of the target speech stream. For example, the target format conversion algorithm can be an algorithm that can record and convert digital audio and video and convert them into a stream. There can be a variety of target format conversion algorithms, which can vary depending on the actual application scenario. The embodiments in this specification do not specifically limit this.

[0090] In step A64, the target speech stream is converted based on the target format conversion algorithm to obtain the converted target speech stream.

[0091] In practice, the target audio stream can be converted using a target format conversion algorithm to obtain the converted target audio stream. For example, FFmpeg can be used to convert the target audio stream to PCM format. FFmpeg is an open-source computer program that can be used to record and convert digital audio and video and convert them into streams.

[0092] In step A66, the target speech stream after format conversion is processed into text based on the first time interval to obtain the first text data corresponding to the target speech stream.

[0093] In S1044, the first text data is corrected to obtain the second text data corresponding to the target speech stream.

[0094] There are various ways to determine the first text data and the second text data in step S104 above. The following is one optional processing method, such as... Figure 7 As shown, the specific process may include the following steps S1046 to S10414.

[0095] In step S1046, one or more target speech segments in the target speech stream are determined.

[0096] The target speech segment can be a segment of speech data that includes human voice data.

[0097] In implementation, one or more target speech segments in the target speech stream can be determined based on a preset segment acquisition algorithm. For example, a Voice Activity Detection (VAD) algorithm can be used to identify silence periods (i.e., speech segments that do not contain human voice data) in the target speech stream, and one or more target speech segments in the target speech stream can be identified based on the silence periods. Specifically, for example... Figure 9 As shown, the target speech stream may include 3 silence periods and 3 target speech segments.

[0098] In step S1048, the target speech segment is divided into one or more sub-speech segments.

[0099] In implementation, the target speech segment can be divided into one or more sub-speech segments based on a first time interval. For example, assuming the first time interval is 50ms, a sub-speech segment can be determined every 50ms. Figure 8 As shown, the target speech segment can be divided into multiple sub-speech segments.

[0100] In step S10410, text conversion processing is performed on each sub-speech segment to obtain the intermediate text data corresponding to the target speech segment.

[0101] The text conversion processing of each sub-speech segment in step S10410 above to obtain the intermediate text data corresponding to the target speech segment can be done in various ways. The following provides an optional processing method, which may include the processing of steps B2 to B6.

[0102] In step B2, the target parameters sent by the first device are received via a long connection.

[0103] The target parameters can include the format parameters and sampling parameters of the target speech stream.

[0104] In step B4, the target text conversion model is determined based on the target parameters.

[0105] Among them, the target text conversion model can be used to convert target speech streams into text data.

[0106] In implementation, since the text conversion requirements of voice streams of different business types are different, the target text conversion model corresponding to the target parameters can be determined according to the preset correspondence between parameters and text conversion models.

[0107] In step B6, based on the target text conversion model, text conversion processing is performed on each sub-speech segment to obtain the intermediate text data corresponding to the target speech segment.

[0108] In S10412, the intermediate text data corresponding to the target speech segment is corrected to obtain the corrected text data corresponding to the target speech segment.

[0109] The processing procedure in step S10412 above can be found in step S104, and will not be repeated here.

[0110] In S10414, based on the intermediate text data corresponding to the target speech segment, the first text data corresponding to the target speech stream is determined, and based on the correction text data corresponding to the target speech segment, the second text data corresponding to the target speech stream is determined.

[0111] By performing text conversion processing on the segmented sub-speech segments, the accuracy of target speech stream recognition can be improved.

[0112] In step S106 above, the process of determining the subtitle as the target text display data corresponding to the target speech stream based on the first text data and the second text data can take many forms. The following provides one optional processing method, such as... Figure 7 As shown, the process may specifically include the following steps S1062 to S1066.

[0113] In step S1062, if the corrected text data completely contains the intermediate text data, then the intermediate text data is determined as the text display data corresponding to the target speech segment.

[0114] In practice, for example, assuming the intermediate text data can be “a”, “b”, and “c”, and the corrected text data is “abbc”, and the corrected text data completely contains the intermediate text data, then “abc” composed of “a”, “b”, and “c” can be determined as the text display data corresponding to the target speech segment.

[0115] In step S1064, if the corrected text data does not completely contain the intermediate text data, then based on the pre-trained semantic analysis model, the intermediate text data corresponding to the target speech segment, and the corrected text data, the text display data corresponding to the target speech segment is determined.

[0116] The pre-trained semantic analysis model can be obtained by training a model constructed by a preset semantic analysis algorithm based on historical intermediate text data and historical correction text data. The preset semantic analysis algorithm can be any analysis algorithm that can determine the text display data corresponding to the target speech segment based on the intermediate text data and correction text data.

[0117] In implementation, historical intermediate text data and historical corrected text data within the preset model training period can be obtained, and the model constructed by the preset semantic analysis algorithm can be trained to obtain a pre-trained semantic analysis model. Then, the intermediate text data and corrected text data corresponding to the target speech segment are input into the pre-trained semantic analysis model to determine the text display data corresponding to the target speech segment.

[0118] Furthermore, the processing of determining the text display data corresponding to the target speech segment in step S1064 above can be varied. The following provides an optional processing method, which may specifically include the processing of steps C2 to C8.

[0119] In step C2, the intermediate text data and correction text data corresponding to the historical speech segments are obtained, as well as the intermediate text data and correction text data corresponding to the first historical speech segment.

[0120] The first historical speech segment may include the speech segment preceding the historical speech segment and / or the speech segment following the historical speech segment.

[0121] In step C4, based on the intermediate text data and correction text data corresponding to the historical speech segments, and the intermediate text data and correction text data corresponding to the first historical speech segment, the model constructed by the preset speech analysis algorithm is trained to obtain the pre-trained semantic analysis model.

[0122] In practice, to improve the accuracy of determining the text display data corresponding to the target speech segment, the preceding and / or following speech segments that have a semantic relationship with the target speech segment can be obtained. The model can be trained based on the obtained speech segments to improve the accuracy of the model. That is, the model constructed by the preset speech analysis algorithm is trained by using the intermediate text data and correction text data corresponding to the historical speech segments, and the intermediate text data and correction text data corresponding to the first historical speech segment, so that the accuracy of the trained speech analysis model is higher.

[0123] In step C6, the first speech segment in the target speech stream is obtained.

[0124] The first speech segment includes the preceding target speech segment and / or the following target speech segment.

[0125] In implementation, for example, Figure 8 As shown, the first speech segment corresponding to target speech segment 1 is target speech segment 2, the first speech segment corresponding to target speech segment 2 is target speech segment 1 and target speech segment 3, and the first speech segment corresponding to target speech segment 3 is target speech segment 2.

[0126] In step C8, based on the pre-trained semantic analysis model, the intermediate text data and correction text data corresponding to the target speech segment, and the intermediate text data and correction text data corresponding to the first speech segment, the subtitles are determined to be the text display data corresponding to the target speech stream.

[0127] In implementation, the intermediate text data and correction text data corresponding to the target speech segment, and the intermediate text data and correction text data corresponding to the first speech segment, are input into the semantic analysis model trained by steps C2 and C4 above to obtain the text display data corresponding to the target speech stream. Thus, by using the intermediate text data and correction text data corresponding to the first speech segment, which have a semantic relationship with the target speech segment, and the pre-trained semantic analysis model to determine the text display data corresponding to the target speech stream, the accuracy of determining the text display data can be improved.

[0128] In step S1066, based on the text display data corresponding to the target speech segment, the subtitle is determined to be the target text display data corresponding to the target speech stream.

[0129] In implementation, the corresponding punctuation can be determined based on the duration of the silence period. Then, based on the text display data corresponding to the target speech segment and the punctuation during the silence period, the target text display data corresponding to the target speech stream can be determined. For example... Figure 9 As shown, if the duration of silence period 1 and silence period 2 is less than the preset duration threshold, then the label corresponding to silence period 1 and silence period 2 can be determined to be a comma. If the duration of silence period 3 is not less than the preset duration threshold, then the label corresponding to silence period 3 can be determined to be a period. Assuming that the text display data corresponding to target speech segment 1 is "data processing method", the text display data corresponding to target speech segment 2 is "audio conversion method", and the text display data corresponding to target speech segment 3 is "text data display method", then the target text display data corresponding to the target speech stream can be determined to be "data processing method, audio conversion method, text data display method".

[0130] The method for determining the target text display data corresponding to the target speech stream described above is an optional and feasible method. In addition, there are multiple methods available. The following provides another optional processing approach, such as... Figure 9 As shown, the process may specifically include the following steps S1068 to S10610.

[0131] In step S1068, the matching degree between the first text data and the second text data is obtained.

[0132] In practice, the matching degree between the first text data and the second text data can be obtained by using a preset matching degree determination method (such as determining the matching degree through regular expressions or string truncation).

[0133] In step S10610, based on the matching degree, the first text data, and the second text data, the subtitle is determined to be the target text display data corresponding to the target speech stream.

[0134] In implementation, for example, if the matching degree is greater than a preset matching degree threshold, the first text data can be determined as the target text display data corresponding to the target speech. If the matching degree is not greater than the preset matching degree threshold, the target text display data corresponding to the target speech stream can be determined based on a pre-trained semantic analysis model, the first text data, and the second text data, the probability value of the first text data conforming to the semantics, and the probability value of the second text data conforming to the semantics. Specifically, according to the pre-trained semantic analysis model, the probability value of the first text data conforming to the semantics is determined to be 0.8, and the probability value of the second text data conforming to the semantics is determined to be 0.75. Based on the above probability values, the first text data can be determined as the target text display data corresponding to the target speech stream. The pre-trained semantic analysis model is obtained by training a model constructed by a preset semantic analysis algorithm based on historical first text data and historical second text data.

[0135] The following section provides a detailed explanation of the display of the above subtitles through specific application scenarios, such as video conferencing and online live classes. These scenarios may include the following:

[0136] Multiple users can participate in video conferences or live online classes. The device used by the user giving the presentation can be considered the first device, and the device used by the user watching can be considered the second device, i.e. Figure 10 As shown, there can be multiple first devices and multiple second devices. The server can be configured with a first module and a second module. The first module can be a module within the server used for data reception, parameter verification, format conversion, and other processing. The second module can be a module within the server used for text conversion and text correction of the target speech stream, such as an Automatic Speech Recognition (ASR) module. The long-lived connection established between the first module and the second module can be a Netty-based WebSocket connection. The first device can initiate a WebSocket request to the server. After receiving the request, the server's first module can establish a Netty-based WebSocket connection between the first module and the first device and return the connection establishment result.

[0137] After receiving a successful connection establishment message, the first device can send the target parameters to the first module via a long connection. The target parameters may include format parameters and sampling parameters of the target audio stream. The first module can verify the target audio stream based on the target parameters and determine whether text conversion of the target audio stream is possible based on the verification result. If text conversion is possible, the module returns the parameter verification result to the first device. If the parameter verification result is successful, the first device can send the target audio stream, which is acquired in real-time through a preset program, to the first module on the server via a long connection. The first module can then perform format conversion on the target audio stream based on the target format conversion algorithm determined by the target parameters, obtaining the format-converted target audio stream.

[0138] A WebSocket connection is established between the first and second modules based on Netty. After the connection is successful, the target audio stream after format conversion is sent to the second module through the first module.

[0139] In addition, there can be multiple second modules. Each second module can be used to perform text conversion processing on a target speech stream of a specific language type. The first module can perform type analysis on the target speech stream to determine the language type corresponding to the target speech stream, and send the target speech stream to the second module corresponding to the language type for text conversion processing.

[0140] The first module can also obtain the device identifier of the second device to determine the text display requirements of the second device for the target text display data based on the device identifier of the second device, and determine the corresponding second module based on the text display requirements and the voice type of the target voice stream, so that the second module can perform text conversion processing on the target voice stream according to the voice type of the target voice stream and the text display requirements.

[0141] The server can trigger the second module to perform text conversion processing on the target speech stream after format conversion based on the first time interval, to obtain the first text data corresponding to the target speech stream, and to perform correction processing on one or more of the first text data to obtain the second text data corresponding to the target speech stream.

[0142] After the first text data and the second text data are returned to the first module through the second module, the server can trigger the first module to determine the subtitle as the target text display data corresponding to the target audio stream based on the first text data and the second text data. The method for determining the target text display data can be found in the specific content of the above embodiment, and will not be repeated here.

[0143] The server can establish a Netty-based WebSocket connection with the second device. After a successful connection, it sends the target text display data and the target audio stream to the second device via a persistent connection. Since Netty is an asynchronous event-driven Java open-source network application framework used for rapid development of maintainable, high-performance protocol servers and clients, Netty encapsulates the JDK's built-in NIO API, greatly reducing complexity. Developers can perform business logic development without worrying about the underlying technology, resulting in low latency and low resource consumption. WebSocket is an HTML5 protocol that enables full-duplex communication over a single TCP connection, used to solve real-time communication between clients and servers. It first sends an HTTP request via HTTP(S) to establish a handshake. After a successful handshake, a TCP connection for exchanging data is created. The client and server can then use this TCP connection for real-time communication. Once the WebSocket client and server have successfully established a handshake (i.e., after the communication connection is established), the previous HTTP request is no longer needed. Furthermore, after the connection is established, the header of the data packets used for protocol control during data exchange between the server and client is relatively small. Therefore, by establishing Netty-based WebSocket connections between the first device and the first module, the first module and the second module, and the second module and the second device, on the one hand, the overhead of frequently creating HTTP connections can be reduced, thereby saving the time spent on connection creation. On the other hand, asynchronous message processing can be achieved based on WebSocket, that is, messages are placed in Netty's internal queue and processed, distributed, and consumed through internal mechanisms, so that requests are not blocked, thereby improving the concurrency processing capacity.

[0144] Furthermore, the server can be a server cluster consisting of multiple different servers. To protect the source code security of the second module, the second module and the first module can be deployed on different servers in the server cluster. The server where the first module is located can be a server used to establish long-term connections with the first device and / or the second device. The server where the second module is located can be a server that can establish long-term connections with the server where the first module is located. The server where the first module is located can send the target voice stream to the server where the second module is located through remote calls. In this way, it can be realized that only data receiving and data sending services are provided to the outside world, ensuring that the technical capabilities of the second module are not leaked.

[0145] In addition, when there are multiple first devices, the server can receive multiple target audio streams. After determining the target text display data corresponding to each target audio stream, the server can send the device identifier of the first device, the target audio stream, and the corresponding target text display data to the second device. This will trigger the second device to display the device identifier of the first device (or the user identifier corresponding to the device identifier of the first device) and the target text display data in real time when playing the target audio stream data, so that the user of the second device can clearly identify the user corresponding to the target audio stream.

[0146] This specification provides a method for displaying subtitles. A long-lived connection is established with a first device, and the target audio stream to be converted is received from the first device via this long-lived connection. The target audio stream is then processed into text based on a first time interval to obtain first text data corresponding to the target audio stream. One or more of the first text data are then corrected to obtain second text data corresponding to the target audio stream. Based on the first and second text data, the subtitles are determined to be the target text display data corresponding to the target audio stream. This method utilizes a long-lived connection between the server and the first device. A long-lived connection allows for the continuous transmission of multiple data packets between the server and the first device during the connection period, eliminating the need to establish a connection for each data packet. Therefore, transmitting the target audio stream via a long-lived connection reduces the resource consumption of frequently creating connections, improving resource utilization. It also saves time associated with connection creation, increasing the efficiency of acquiring the target text display data for the target audio stream. Furthermore, the first and second text data accurately determine the target text display data corresponding to the target audio stream, improving the accuracy of the determination of the text display data for the target audio stream.

[0147] The above are the methods for displaying subtitles provided in the embodiments of this specification. Based on the same idea, such as... Figure 11 As shown in the embodiments of this specification, a subtitle display device is also provided to improve resource utilization and the efficiency and accuracy of text display for voice data. For specific implementation of this device, please refer to the relevant content of the subtitle display method. To avoid redundancy, it will not be repeated here.

[0148] Corresponding to the subtitle display method provided in the above embodiments, based on the same technical concept, this specification also provides a subtitle display device, which is used to execute the above-described subtitle display method. Figure 12 This is a hardware structure diagram of a subtitle display device according to various embodiments of this specification. Figure 12The subtitle display device 120 shown includes, but is not limited to, components such as: a radio frequency unit 121, a network module 122, an audio output unit 123, an input unit 124, a sensor 125, a user input unit 126, an interface unit 127, a memory 128, a processor 129, and a power supply 1210. Those skilled in the art will understand that... Figure 12 The structure of the subtitle display device shown in the figure does not constitute a limitation on the subtitle display device. The subtitle display device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0149] The interface unit 127 is used to establish a long connection with the first device and receive the target voice stream to be converted collected by the first device through the long connection;

[0150] The processor 129 is configured to perform text conversion processing on the target speech stream to obtain first text data corresponding to the target speech stream, and to perform correction processing on the first text data to obtain second text data corresponding to the target speech stream.

[0151] The processor 129 is further configured to determine, based on the first text data and the second text data, that the subtitle is target text display data corresponding to the target speech stream.

[0152] In this embodiment of the specification, the interface unit 127 is further configured to return the target text display data to the first device through the long connection, and trigger the first device to display the target text display data corresponding to the target audio stream as subtitles.

[0153] Interface unit 127 is also used to establish a long connection with the second device and send the target text display data to the second device through the long connection established with the second device, thereby triggering the second device to display the target text display data corresponding to the target voice stream.

[0154] In the embodiments described in this specification, the processor 129 is further configured to:

[0155] The target speech stream is processed by text conversion based on a first time interval to obtain the first text data corresponding to the target speech stream.

[0156] In the embodiments described in this specification, the processor 129 is further configured to:

[0157] The target parameters sent by the first device are received through the long connection, and the target parameters include the format parameters and sampling parameters of the target voice stream;

[0158] Based on the target parameters, the target speech stream is verified, and based on the verification result, it is determined whether the target speech stream can be converted into text.

[0159] If it is determined that the target speech stream can be converted into text, then the target speech stream is processed into text based on the first time interval to obtain the first text data corresponding to the target speech stream.

[0160] In the embodiments described in this specification, the processor 129 is further configured to:

[0161] Based on the target parameters, a target format conversion algorithm is determined, which is used to convert the target speech stream into a new format.

[0162] Based on the target format conversion algorithm, the target speech stream is converted to obtain the format-converted target speech stream;

[0163] Based on a first time interval, the target speech stream after format conversion is processed into text to obtain the first text data corresponding to the target speech stream.

[0164] In the embodiments described in this specification, the processor 129 is further configured to:

[0165] Obtain the matching degree between the first text data and the second text data;

[0166] Based on the matching degree, the first text data, and the second text data, the subtitle is determined to be the target text display data corresponding to the target speech stream.

[0167] In the embodiments described in this specification, the processor 129 is further configured to:

[0168] Identify one or more target speech segments in the target speech stream, wherein the target speech segment is a segment of speech data containing human voice data;

[0169] The target speech segment is divided into one or more sub-speech segments, and text conversion processing is performed on each sub-speech segment to obtain the intermediate text data corresponding to the target speech segment;

[0170] The intermediate text data corresponding to the target speech segment is corrected to obtain the corrected text data corresponding to the target speech segment.

[0171] Based on the intermediate text data corresponding to the target speech segment, the first text data corresponding to the target speech stream is determined, and based on the corrected text data corresponding to the target speech segment, the second text data corresponding to the target speech stream is determined.

[0172] In the embodiments described in this specification, the interface unit 127 is further used for:

[0173] The target parameters sent by the first device are received through the long connection, and the target parameters include the format parameters and sampling parameters of the target voice stream;

[0174] A target text conversion model is determined based on the target parameters, and the target text conversion model is used to convert the target speech stream into text data;

[0175] Based on the target text conversion model, text conversion processing is performed on each of the sub-speech segments to obtain the intermediate text data corresponding to the target speech segment.

[0176] In the embodiments described in this specification, the processor 129 is further configured to:

[0177] If the corrected text data completely contains the intermediate text data, then the intermediate text data is determined as the text display data corresponding to the target speech segment;

[0178] If the corrected text data does not completely contain the intermediate text data, then based on the pre-trained semantic analysis model, the intermediate text data corresponding to the target speech segment, and the corrected text data, the text display data corresponding to the target speech segment is determined. The pre-trained semantic analysis model is obtained by training a model constructed by a preset semantic analysis algorithm based on historical intermediate text data and historical corrected text data.

[0179] Based on the text display data corresponding to the target speech segment, the subtitle is determined to be the target text display data corresponding to the target speech stream.

[0180] In the embodiments described in this specification, the processor 129 is further configured to:

[0181] Obtain a first speech segment from the target speech stream, wherein the first speech segment includes the preceding target speech segment and / or the following target speech segment;

[0182] Based on the pre-trained semantic analysis model, the intermediate text data and corrected text data corresponding to the target speech segment, and the intermediate text data and corrected text data corresponding to the first speech segment, the text display data corresponding to the target speech stream is determined.

[0183] The subtitle display device in the embodiments of this specification establishes a long connection with a first device and receives the target audio stream to be converted collected by the first device through the long connection. Based on a first time interval, it performs text conversion processing on the target audio stream to obtain first text data corresponding to the target audio stream. It then performs correction processing on one or more of the first text data to obtain second text data corresponding to the target audio stream. Based on the first and second text data, the subtitles are determined to be the target text display data corresponding to the target audio stream. Thus, by establishing a long connection between the server and the first device, which allows for the continuous transmission of multiple data packets without establishing a connection for each data packet, the resource consumption of frequently creating connections can be reduced, improving resource utilization. Simultaneously, it saves the time spent creating connections, improving the efficiency of acquiring the target text display data for the target audio stream. Furthermore, the first and second text data can accurately determine the target text display data corresponding to the target audio stream, improving the accuracy of determining the text display data for the target audio stream.

[0184] It should be noted that the subtitle display device 120 provided in the embodiments of this specification can realize all the processes implemented by the subtitle display device in the above-described subtitle display method embodiments. To avoid repetition, it will not be described again here.

[0185] It should be understood that, in the embodiments of this specification, the radio frequency unit 121 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink data from upstream devices and processes it with the processor 129; additionally, it transmits uplink data to upstream devices. Typically, the radio frequency unit 121 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc. Furthermore, the radio frequency unit 121 can also communicate with networks and other devices through a wireless communication system.

[0186] The subtitle display device provides users with wireless broadband internet access via network module 122, enabling them to send and receive emails, browse web pages, and access streaming media.

[0187] The audio output unit 123 can convert audio data received by the radio frequency unit 121 or the network module 122 or stored in the memory 129 into audio signals and output them as sound. Furthermore, the audio output unit 123 can also provide audio output related to specific functions performed by the mobile terminal 120 (e.g., call signal reception sound, message reception sound, etc.). The audio output unit 123 includes a speaker, a buzzer, and a receiver, etc.

[0188] Input unit 124 is used to receive audio or video signals. Input unit 124 may include a graphics processing unit (GPU) 1241 and a microphone 1242. The GPU 1241 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on display unit 126. The image frames processed by GPU 1241 can be stored in memory 129 (or other storage medium) or transmitted via radio frequency unit 121 or network module 122. Microphone 1242 can receive sound and process such sound into audio data. The processed audio data can be converted into a format that can be transmitted to a mobile communication base station via radio frequency unit 121 in telephone call mode.

[0189] Interface unit 127 serves as an interface for connecting external devices to the subtitle display device 120. For example, external devices may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 127 can be used to receive input from external devices (e.g., data, power, etc.) and transmit the received input to one or more components within the mobile terminal 120, or it can be used to transmit data between the mobile terminal 120 and the external device.

[0190] The memory 128 can be used to store software programs and various data. The memory 128 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 128 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0191] The processor 129 is the control center of the mobile terminal. It connects various parts of the mobile terminal via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 128, and by calling data stored in the memory 128, it performs various functions and processes data of the subtitle display device, thereby providing overall monitoring of the subtitle display device. The processor 129 may include one or more processing units; preferably, the processor 129 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 129.

[0192] The subtitle display device 120 may also include a power supply 1210 (such as a battery) that supplies power to various components. Preferably, the power supply 1210 can be logically connected to the processor 129 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system.

[0193] In addition, the subtitle display device 120 includes some functional modules not shown, which will not be described in detail here.

[0194] Preferably, this specification also provides a subtitle display device, including a processor 129, a memory 128, and a computer program stored in the memory 128 and executable on the processor 129. When the computer program is executed by the processor 129, it implements the various processes of the above-described subtitle display method embodiment and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0195] Furthermore, based on the above Figures 1 to 10 The method shown in this specification, along with one or more embodiments, also provides a storage medium for storing computer-executable instruction information. In one specific embodiment, the storage medium can be a USB flash drive, optical disc, hard disk, etc. When the computer-executable instruction information stored in the storage medium is executed by a processor, it can achieve the following process:

[0196] Establish a long connection with the first device, and receive the target speech stream to be converted collected by the first device through the long connection;

[0197] The target speech stream is subjected to text conversion processing to obtain first text data corresponding to the target speech stream, and the first text data is subjected to correction processing to obtain second text data corresponding to the target speech stream.

[0198] Based on the first text data and the second text data, the subtitle is determined to be the target text display data corresponding to the target speech stream.

[0199] In the embodiments of this specification, the method further includes:

[0200] The target text display data is returned to the first device via the long connection, and the first device is triggered to display the target text display data corresponding to the target audio stream as subtitles.

[0201] In the embodiments of this specification, the method further includes:

[0202] A long connection is established with the second device, and the target text display data is sent to the second device through the long connection established with the second device, triggering the second device to display the target text display data corresponding to the target voice stream.

[0203] In the embodiments of this specification, the step of performing text conversion processing on the target speech stream to obtain the first text data corresponding to the target speech stream includes:

[0204] Based on a first time interval, the target speech stream is subjected to text conversion processing to obtain first text data corresponding to the target speech stream, and one or more of the first text data are subjected to correction processing to obtain second text data corresponding to the target speech stream.

[0205] In this embodiment of the specification, the step of performing text conversion processing on the target speech stream based on a first time interval to obtain the first text data corresponding to the target speech stream includes:

[0206] The target parameters sent by the first device are received through the long connection, and the target parameters include the format parameters and sampling parameters of the target voice stream;

[0207] Based on the target parameters, the target speech stream is verified, and based on the verification result, it is determined whether the target speech stream can be converted into text.

[0208] If it is determined that the target speech stream can be converted into text, then the target speech stream is processed into text based on the first time interval to obtain the first text data corresponding to the target speech stream.

[0209] In this embodiment of the specification, the step of performing text conversion processing on the target speech stream based on the first time interval to obtain the first text data corresponding to the target speech stream includes:

[0210] Based on the target parameters, a target format conversion algorithm is determined, which is used to convert the target speech stream into a new format.

[0211] Based on the target format conversion algorithm, the target speech stream is converted to obtain the format-converted target speech stream;

[0212] Based on the first time interval, the target speech stream after format conversion is processed into text to obtain the first text data corresponding to the target speech stream.

[0213] In this embodiment of the specification, determining that the subtitle is the target text display data corresponding to the target speech stream based on the first text data and the second text data includes:

[0214] Obtain the matching degree between the first text data and the second text data;

[0215] Based on the matching degree, the first text data, and the second text data, the subtitle is determined to be the target text display data corresponding to the target speech stream.

[0216] In the embodiments of this specification, the step of performing text conversion processing on the target speech stream to obtain first text data corresponding to the target speech stream, and performing correction processing on the first text data to obtain second text data corresponding to the target speech stream, includes:

[0217] Identify one or more target speech segments in the target speech stream, wherein the target speech segment is a segment of speech data containing human voice data;

[0218] The target speech segment is divided into one or more sub-speech segments, and text conversion processing is performed on each sub-speech segment to obtain the intermediate text data corresponding to the target speech segment;

[0219] The intermediate text data corresponding to the target speech segment is corrected to obtain the corrected text data corresponding to the target speech segment.

[0220] Based on the intermediate text data corresponding to the target speech segment, the first text data corresponding to the target speech stream is determined, and based on the corrected text data corresponding to the target speech segment, the second text data corresponding to the target speech stream is determined.

[0221] In the embodiments of this specification, the step of performing text conversion processing on each of the sub-speech segments to obtain intermediate text data corresponding to the target speech segment includes:

[0222] The target parameters sent by the first device are received through the long connection, and the target parameters include the format parameters and sampling parameters of the target voice stream;

[0223] A target text conversion model is determined based on the target parameters, and the target text conversion model is used to convert the target speech stream into text data;

[0224] Based on the target text conversion model, text conversion processing is performed on each of the sub-speech segments to obtain the intermediate text data corresponding to the target speech segment.

[0225] In this embodiment of the specification, determining that the subtitle is the target text display data corresponding to the target speech stream based on the first text data and the second text data includes:

[0226] If the corrected text data completely contains the intermediate text data, then the intermediate text data is determined as the text display data corresponding to the target speech segment;

[0227] If the corrected text data does not completely contain the intermediate text data, then based on the pre-trained semantic analysis model, the intermediate text data corresponding to the target speech segment, and the corrected text data, the text display data corresponding to the target speech segment is determined. The pre-trained semantic analysis model is obtained by training a model constructed by a preset semantic analysis algorithm based on historical intermediate text data and historical corrected text data.

[0228] Based on the text display data corresponding to the target speech segment, the subtitle is determined to be the target text display data corresponding to the target speech stream.

[0229] In the embodiments of this specification, determining the text display data corresponding to the target speech segment based on the pre-trained semantic analysis model, the intermediate text data corresponding to the target speech segment, and the corrected text data includes:

[0230] Obtain a first speech segment from the target speech stream, wherein the first speech segment includes the preceding target speech segment and / or the following target speech segment;

[0231] Based on the pre-trained semantic analysis model, the intermediate text data and corrected text data corresponding to the target speech segment, and the intermediate text data and corrected text data corresponding to the first speech segment, the text display data corresponding to the target speech stream is determined.

[0232] This specification provides a storage medium that establishes a long connection with a first device and receives a target speech stream to be converted from the first device through the long connection. Based on a first time interval, the target speech stream is processed into text to obtain first text data corresponding to the target speech stream. One or more of the first text data are then corrected to obtain second text data corresponding to the target speech stream. Based on the first and second text data, subtitles are determined to be target text display data corresponding to the target speech stream. Thus, by establishing a long connection between the server and the first device, which allows for the continuous transmission of multiple data packets without requiring a connection for each packet, the resource consumption of frequently creating connections is reduced, improving resource utilization. Simultaneously, it saves the time spent creating connections, improving the efficiency of acquiring target text display data for the target speech stream. Furthermore, the first and second text data accurately determine the target text display data corresponding to the target speech stream, improving the accuracy of determining the text display data for the target speech.

[0233] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0234] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using a hardware physical module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0235] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0236] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0237] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.

[0238] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0239] Embodiments in this specification are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable parallel device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable parallel device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0240] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable fraud device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0241] These computer program instructions can also be loaded onto a computer or other programmable device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0242] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0243] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0244] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0245] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0246] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0247] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0248] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0249] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.

Claims

1. A method for displaying subtitles, the method comprising: Establish a long connection with the first device, and receive the target speech stream to be converted collected by the first device through the long connection; The target speech stream is subjected to text conversion processing to obtain first text data corresponding to the target speech stream, and the first text data is subjected to correction processing to obtain second text data corresponding to the target speech stream. Based on the first text data and the second text data, the subtitle is determined to be the target text display data corresponding to the target speech stream; The step of determining the subtitle as target text display data corresponding to the target speech stream based on the first text data and the second text data includes: If the corrected text data does not completely contain the intermediate text data, then based on the pre-trained semantic analysis model, the intermediate text data corresponding to the target speech segment, and the corrected text data, the text display data corresponding to the target speech segment is determined. The first text data is determined based on the intermediate text data, and the second text data is determined based on the corrected text data. The corrected text data is obtained by correcting the intermediate text data corresponding to the target speech segment in the target speech stream. The intermediate text data is obtained by performing text conversion processing on the sub-speech segments in the target speech segment. The target speech segment is a segment of speech data containing human voice data in the target speech stream. If the corrected text data completely contains the intermediate text data, then the intermediate text data is determined as the text display data corresponding to the target speech segment; Based on the text display data corresponding to the target speech segment, the subtitle is determined to be the target text display data corresponding to the target speech stream.

2. The method according to claim 1, further comprising: The target text display data is returned to the first device via the long connection, and the first device is triggered to display the target text display data corresponding to the target audio stream as subtitles.

3. The method according to claim 1, further comprising: A long connection is established with the second device, and the target text display data is sent to the second device through the long connection established with the second device, triggering the second device to display the target text display data corresponding to the target voice stream.

4. The method according to claim 1, wherein performing text conversion processing on the target speech stream to obtain the first text data corresponding to the target speech stream includes: The target speech stream is processed by text conversion based on a first time interval to obtain the first text data corresponding to the target speech stream.

5. The method according to claim 4, wherein performing text conversion processing on the target speech stream based on a first time interval to obtain first text data corresponding to the target speech stream includes: The target parameters sent by the first device are received through the long connection, and the target parameters include the format parameters and sampling parameters of the target voice stream; Based on the target parameters, the target speech stream is verified, and based on the verification result, it is determined whether the target speech stream can be converted into text. If it is determined that the target speech stream can be converted into text, then the target speech stream is processed into text based on the first time interval to obtain the first text data corresponding to the target speech stream.

6. The method according to claim 5, wherein performing text conversion processing on the target speech stream based on the first time interval to obtain the first text data corresponding to the target speech stream includes: Based on the target parameters, a target format conversion algorithm is determined, which is used to convert the target speech stream into a new format. Based on the target format conversion algorithm, the target speech stream is converted to obtain the format-converted target speech stream; Based on the first time interval, the target speech stream after format conversion is processed into text to obtain the first text data corresponding to the target speech stream.

7. The method according to claim 1, wherein the method for determining the intermediate text data includes: The target parameters sent by the first device are received through the long connection, and the target parameters include the format parameters and sampling parameters of the target voice stream; A target text conversion model is determined based on the target parameters, and the target text conversion model is used to convert the target speech stream into text data; Based on the target text conversion model, text conversion processing is performed on each of the sub-speech segments to obtain the intermediate text data corresponding to the target speech segment.

8. The method according to claim 1, wherein determining the text display data corresponding to the target speech segment based on a pre-trained semantic analysis model, intermediate text data corresponding to the target speech segment, and corrected text data comprises: Obtain a first speech segment from the target speech stream, wherein the first speech segment includes the preceding target speech segment and / or the following target speech segment; Based on the pre-trained semantic analysis model, the intermediate text data and corrected text data corresponding to the target speech segment, and the intermediate text data and corrected text data corresponding to the first speech segment, text display data corresponding to the target speech stream is determined. The pre-trained semantic analysis model is obtained by training a model constructed by a preset semantic analysis algorithm based on historical intermediate text data and historical corrected text data.

9. A subtitle display device, the device comprising: The connection establishment module is configured to establish a long connection with the first device and receive the target speech stream to be converted collected by the first device through the long connection; The data conversion module is configured to perform text conversion processing on the target speech stream to obtain first text data corresponding to the target speech stream, and to perform correction processing on the first text data to obtain second text data corresponding to the target speech stream. The data determination module is configured to determine, based on the first text data and the second text data, that the subtitle is the target text display data corresponding to the target speech stream; The data determination module is configured as follows: If the corrected text data does not completely contain the intermediate text data, then based on the pre-trained semantic analysis model, the intermediate text data corresponding to the target speech segment, and the corrected text data, the text display data corresponding to the target speech segment is determined. The first text data is determined based on the intermediate text data, and the second text data is determined based on the corrected text data. The corrected text data is obtained by correcting the intermediate text data corresponding to the target speech segment in the target speech stream. The intermediate text data is obtained by performing text conversion processing on the sub-speech segments in the target speech segment. The target speech segment is a segment of speech data containing human voice data in the target speech stream. If the corrected text data completely contains the intermediate text data, then the intermediate text data is determined as the text display data corresponding to the target speech segment; Based on the text display data corresponding to the target speech segment, the subtitle is determined to be the target text display data corresponding to the target speech stream.

10. A subtitle display device, the subtitle display device comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, cause the processor to: Establish a long connection with the first device, and receive the target speech stream to be converted collected by the first device through the long connection; The target speech stream is subjected to text conversion processing to obtain first text data corresponding to the target speech stream, and the first text data is subjected to correction processing to obtain second text data corresponding to the target speech stream. Based on the first text data and the second text data, the subtitle is determined to be the target text display data corresponding to the target speech stream; The step of determining the subtitle as target text display data corresponding to the target speech stream based on the first text data and the second text data includes: If the corrected text data does not completely contain the intermediate text data, then based on the pre-trained semantic analysis model, the intermediate text data corresponding to the target speech segment, and the corrected text data, the text display data corresponding to the target speech segment is determined. The first text data is determined based on the intermediate text data, and the second text data is determined based on the corrected text data. The corrected text data is obtained by correcting the intermediate text data corresponding to the target speech segment in the target speech stream. The intermediate text data is obtained by performing text conversion processing on the sub-speech segments in the target speech segment. The target speech segment is a segment of speech data containing human voice data in the target speech stream. If the corrected text data completely contains the intermediate text data, then the intermediate text data is determined as the text display data corresponding to the target speech segment; Based on the text display data corresponding to the target speech segment, the subtitle is determined to be the target text display data corresponding to the target speech stream.

11. A storage medium for storing computer-executable instructions, which, when executed by a processor, perform the following process: Establish a long connection with the first device, and receive the target speech stream to be converted collected by the first device through the long connection; The target speech stream is subjected to text conversion processing to obtain first text data corresponding to the target speech stream, and the first text data is subjected to correction processing to obtain second text data corresponding to the target speech stream. Based on the first text data and the second text data, the subtitle is determined to be the target text display data corresponding to the target speech stream; in, The step of determining the subtitle as target text display data corresponding to the target speech stream based on the first text data and the second text data includes: If the corrected text data does not completely contain the intermediate text data, then based on the pre-trained semantic analysis model, the intermediate text data corresponding to the target speech segment, and the corrected text data, the text display data corresponding to the target speech segment is determined. The first text data is determined based on the intermediate text data, and the second text data is determined based on the corrected text data. The corrected text data is obtained by correcting the intermediate text data corresponding to the target speech segment in the target speech stream. The intermediate text data is obtained by performing text conversion processing on the sub-speech segments in the target speech segment. The target speech segment is a segment of speech data containing human voice data in the target speech stream. If the corrected text data completely contains the intermediate text data, then the intermediate text data is determined as the text display data corresponding to the target speech segment; Based on the text display data corresponding to the target speech segment, the subtitle is determined to be the target text display data corresponding to the target speech stream.

Citation Information

Patent Citations

  • Chinese online audio and video caption generation method

    CN109257547A

  • Real-time playing method and device

    CN111479124A

  • Subtitle generation method and device, computer readable storage medium and electronic equipment

    CN113225612A