Radio and television voice translation retrieval system
By using a voice translation and retrieval system based on smartphones and cloud computing at broadcasting stations, the problem of high costs has been solved, achieving efficient voice-to-text retrieval, improving regulatory efficiency and accuracy, and making it suitable for real-time and historical program monitoring.
Patent Information
- Application Number
- CN202511198684.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-25
AI Technical Summary
Existing radio and television stations face difficulties in large-scale application of speech-to-text monitoring methods due to high costs, resulting in low efficiency in retrieving audio and video materials and an inability to effectively monitor the integrity and authenticity of program content.
This invention provides a broadcast television speech-to-text retrieval system that combines smartphones and cloud computing. It uses voice input software to convert speech to text and adds timestamps to the data system. It supports real-time and historical translation modes to improve retrieval efficiency.
It significantly reduced system setup costs, achieved efficient voice-to-text retrieval, improved the efficiency and accuracy of regulatory work, and enabled the rapid location and playback of designated programs, allowing for timely monitoring of program broadcasts.
Smart Images

Figure CN121012941A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to a broadcast television voice translation retrieval system. BACKGROUND
[0002] Preventing potential illegal interruption or program tampering is one of the core tasks of broadcast program supervision. Once malicious destruction occurs, the integrity and authenticity of the program content will be severely compromised, and it may further cause adverse effects in the society and erode the credibility of the country. Therefore, in addition to real-time monitoring, the broadcast industry must also equip itself with post-evidence collection techniques such as program audio recording to ensure comprehensive and effective supervision. In addition, historical audio-visual materials are not only valuable historical evidence and cultural carriers, but also embody the cultural identity of the people and the shaping of the national spirit. They contain important potential for industrial development and economic transformation value, not only serving the fields of social education and public service, but also providing valuable inspiration and guidance for subsequent artistic creation and technological innovation.
[0003] In the era of information flooding, massive videos are pouring in, especially new media emerging like mushrooms after rain, and the problem of low efficiency of audio-visual material retrieval is imminent and needs to be solved. Currently, audio-visual material retrieval still mainly relies on manual viewing and playing, even with the help of fast-forwarding, the operation is extremely inconvenient. Especially for broadcast programs, their special nature is that they cannot rely on associated information such as images for identification due to their pure audio properties, and mainly rely on program schedules for manual retrieval. With the rapid development of artificial intelligence, more efficient and accurate technical solutions are provided for program sound retrieval.
[0004] The current scheme for converting broadcast program voice content into text is to use high-performance servers to build a voice large model system, and through a voice recognition engine, the voice is converted into text in real time. And for each program and each monitoring node, if a separate system is deployed, the construction cost and energy consumption are particularly high. In order to reduce costs, an improved method is a centralized intelligent translation solution based on cloud native architecture, which replaces the traditional single-node independent deployment mode by building a distributed voice recognition micro-service cluster to realize unified access and intelligent routing processing of multi-channel program audio streams. The core technology uses a lightweight ASR engine optimized by model distillation, combined with streaming transmission and intelligent caching mechanism, which can reduce more than 50% of computing resource consumption while ensuring more than 98% of recognition accuracy. However, despite this, it is still costly and ordinary radio and television stations cannot afford it. Therefore, the monitoring method of real-time conversion of broadcast program voice content into text has not been widely applied. SUMMARY
[0005] The application provides a broadcast television voice translation retrieval system to solve the problems in the prior art that the application of a large model is expensive and has high cost, which makes it difficult for a broadcast television station to bear, and leads to the difficulty of wide and large-scale application of the monitoring method for translated text.
[0006] The broadcast television voice translation retrieval system comprises a recording system and a text retrieval system.
[0007] The recording system is used in two application scenarios of real-time program monitoring and historical program review monitoring through two working modes of real-time translation and historical translation.
[0008] In the real-time translation mode, when real-time audio and video data of a program are received, the recording system saves the audio and video data as an audio and video file, simultaneously translates the sound content of the program into text in real time, and saves the time when the text is captured as a timestamp and the translated text in a data system, which is used for time positioning of retrieval and playback of the text retrieval system.
[0009] In the historical translation mode, the recording system is used to play back an audio and video file and extract the sound of a program, and then translate the extracted sound of the program into text, and record the timestamp of the translated text, i.e., the historical time when the audio and video file is recorded, which is used for query of the text retrieval system.
[0010] The text retrieval system selects the voice translation text corresponding to a program from a text query list output by the recording system, and realizes playback of the audio and video content when the original playing voice of the program is played.
[0011] The broadcast television voice translation retrieval system has the following advantages:
[0012] The broadcast television voice translation retrieval system according to the application is based on the two key interfaces of upstream voice signal input and downstream text content information output, and is connected with a free voice input method software on a smart phone to realize the function of broadcast television voice text translation. In this way, the expensive artificial intelligence research and development work is changed into the screening and testing of various existing mobile phone voice input methods, and the deployment of a high-cost server cluster is changed into the specific use of a smart phone, which obviously improves the system building and fund use efficiency, and can quickly convert technical investment into business value.
[0013] The broadcast television voice translation retrieval system provided by the application provides a very low-cost voice translation text solution. Compared with video and audio content retrieval, the application has the advantages of directness and high efficiency by translating audio signals into text content in real time. By connecting with an artificial intelligence system, the voice content of a broadcast television program can be translated into text, and a timestamp can be added synchronously and stored in a data system. When a user performs retrieval work, in the precise retrieval mode, the user can quickly locate and obtain the video and audio content of the broadcast television program by using the text timestamp information attached to the program; in the fuzzy retrieval mode, the user can obtain the video and audio of the broadcast television program of interest or association; and in the auxiliary positioning retrieval mode, the user can locate and retrieve the video and audio of the broadcast television program according to the characteristic information of the program.
[0014] The broadcast television voice translation retrieval system provided by the application translates the voice of a broadcast television program into text and then performs text retrieval, which has the significant advantage of being much faster than the traditional retrieval mode of playing a broadcast television program. In the broadcast television supervision work, the system can help staff quickly locate and play a specified program, timely supervise and confirm the program broadcast, and collect evidence, thereby greatly improving the efficiency and accuracy of the supervision work. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 The structure diagram of the recording system in the real-time translation mode of the broadcast television voice translation retrieval system provided by the application is shown in the figure.
[0016] Figure 2 The structure diagram of the recording system in the historical translation mode is shown in the figure. DETAILED DESCRIPTION
[0017] Specific implementation one, combination Figure 1 and Figure 2The present embodiment illustrates a broadcast television voice translation retrieval system, which comprises a recording system and a text retrieval system. The recording system is used to capture the audio and video signals of a program synchronously and convert the program sound into text information in real time. In order to facilitate the subsequent retrieval work, the recording system will mark the timestamp information for each translated text while performing the recording task, which provides accurate positioning information for subsequent retrieval and accurate playback. In the retrieval system, the staff can search in the recorded text information with the help of keywords. This function seems to search in the vast audio and video files, but actually it is to find specific records in the text paragraphs generated by translation. Once the interested content is found, the system can quickly respond and accurately play back the audio and video content of the original program when the voice is played, by clicking the hyperlink corresponding to the conclusion information in the retrieval query output list, that is, selecting the voice translation text corresponding to the program from the recorded text query list.
[0018] The recording system has two working modes: real-time translation and historical translation, which can be applied to real-time program supervision and historical program review supervision. In real-time translation mode, when receiving real-time audio and video data of a program, the recording system saves the audio and video data as a recording file while synchronously converting the program sound content into text in real time. During this period, the recording system also captures the time when the text is translated out and saves it as a timestamp together with the translated text in the data system, providing time positioning for subsequent retrieval and playback. In historical translation mode, the recording system does not receive real-time audio and video data of a program, but performs the function of playing back historical audio and video files of the program, and extracts the program sound during playback and then converts it into text. Unlike real-time translation mode, the recording system in historical translation mode needs to record the text timestamp, which is not the time when the text is translated out, but the historical time when the program sound is recorded in the historical program file. That is, the historical translation mode is to play back the recording file once, and then add an index system containing voice text and timestamp to the recording file. With this index, the historical program files that have not been translated can also be queried by the text retrieval system.
[0019] As shown in Figure 1 , in the present embodiment, the recording system in real-time translation mode is composed of a broadcast television signal source, an interface circuit board, a smart phone, a communication network and a translation recording computer;
[0020] In real-time translation mode, the program signal emitted by the broadcast television signal source is received and converted into a program data stream by the interface circuit board, and in the process, it can also accurately extract the sound stream data from the program data stream. The interface circuit board adopts a double-output interface parallel strategy, on the one hand, it uses the USB interface to transmit the sound stream data to the smart phone, and on the other hand, it uses the communication network to send the program data stream to the translation recording computer; the translation recording computer is responsible for completing the recording function, ensuring that the video and audio content of the program is completely retained.
[0021] A voice translation transmission software is run in the smart phone, which is used to translate the sound stream data into text, and then send it to the translation recording computer through the communication network. The translation recording computer saves the text information at the same time, and for each sentence of the translated text, it takes the exact time when the text is received as a timestamp, and attaches a special "time label" to each sentence, and saves it with the corresponding text information. With these timestamps, reliable time positioning basis can be provided for subsequent retrieval and playback work, so that the program video and audio playback and the translated text can be accurately traced and presented in the time dimension. The computer in the communication network can remotely log in to the translation recording computer through the communication network, click on the hyperlink corresponding to the conclusion information in the retrieval query output list, and accurately play back the video and audio content when the program originally played the voice.
[0022] As Figure 2 shown, in the historical translation mode, the recording system is composed of an audio and video recording server, a player, an interface circuit board, a smart phone, a communication network, a translation computer and a query computer, Figure 2 is a structural diagram of the historical translation retrieval system.
[0023] In the historical translation mode, the historical video and audio files of the program are saved in the audio and video recording server, and the playback is performed by the player. The audio signal in the playback program is converted into program sound data stream through the interface circuit board, and then sent to the smart phone. The text output by the smart phone is sent to the translation computer through the communication network. For each sentence of the translated text, it takes the time when the voice in the historical video and audio file was recorded as a timestamp, and saves it with the corresponding text information. The main difference between this historical translation process and the real-time translation process is the generation mechanism of the timestamp. The timestamp of real-time translation is the time when the translated text is generated, while the timestamp of historical translation is the time when the historical audio and video file is recorded, that is, the time when the sentence is spoken.
[0024] Computers within a communication network can remotely log in and query other computers. These computers retrieve the translated text stored in the translation computer based on keywords entered by the user. When a user clicks a hyperlink corresponding to a conclusion in the search query output list, the original audio and video content of that segment is precisely edited from the recording server and transmitted via the communication network to the user's computer for playback.
[0025] In this embodiment, after downloading the voice input method software to the smartphone, a self-developed text forwarding software also needs to be run. This text forwarding software sends the text translated by the input method to a translation computer via a communication network. The text forwarding software has a text input box. The voice input method receives voice data from the interface circuit board, converts it into text, and writes it directly into this text input box. The text forwarding software can then immediately obtain the voice translation result and send this text to the corresponding computer (translation recording computer, translation computer) via the communication network.
[0026] The smartphone communicates with the corresponding computer using the TCP protocol. The smartphone acts as the server and supports multiple client software connections simultaneously. Each client can receive the same content information. This allows the text translated by a monitoring node to be used for retrieval functions, as well as for public opinion monitoring, comparative monitoring, and many other applications.
[0027] In this embodiment, the text forwarding software has only one function: to monitor text input messages. Whenever text is input, it is sent to the computer via the communication network. Regardless of whether the text box content is entered manually using a keyboard or any other input method, and regardless of the type of text, all new text input is sent out, thus ensuring compatibility with all mobile phone input methods.
[0028] Because most voice input methods now have AI-powered backtracking and correction capabilities, meaning they can correct previously entered text based on context, the text forwarding software must also support this function. This backtracking and correction is implemented using a backspace and rewriting method. Specifically, when the input method performs backtracking and correction, the text forwarding software sends backspace keys to the corresponding computer connected to the communication network. The number of backspace keys equals the number of characters corrected by the input method. After all backspace keys have been sent, all corrected text is resent to the corresponding computer. A key advantage of this approach is that when the corresponding computer displays all received text directly in a single editing box in real time, no additional processing is required; the smartphone input method software can directly correct the translated text received by the computer in real time.
[0029] In this embodiment, a translation and storage software runs on the translation computer, which can establish a real-time data communication channel with the text forwarding software on the smartphone, seamlessly receiving and storing all text data streams from the smartphone. The translation and storage software has a built-in dynamic text processing engine with two core functions: one is that when a backspace key is detected, it uses a last-in-first-out (STACK) principle to backtrack and delete the most recently received character, ensuring the continuity of text editing; the other is a smart punctuation recognition module that automatically triggers a text segmentation mechanism when it captures the three key punctuation marks: comma, period, or question mark. This parses the continuous text data stream into short sentence units with independent semantics and adds a timestamp to each newly generated sentence fragment, forming an index structure with a mapping relationship between text content and text time. As an important module of the entire text retrieval system, the translation and storage software realizes the functions of receiving, processing, and storing network text data in real time, and is particularly suitable for cross-platform application scenarios that require real-time text translation, semantic segmentation, and time tracing.
[0030] In this embodiment, a transcoding and storage mechanism is implemented in real-time translation mode;
[0031] After the system receives the audio and video data stream of a program from the network, it needs to perform transcoding and storage operations. Specifically, this involves converting the received audio and video data stream into a specific format (audio encoding format, audio compression encoding format, or MP3 audio format), then saving it as an audio / video file and storing it on the disk. The core function of transcoding is to reduce the bitrate of the data stream while meeting monitoring and management requirements. To achieve this, the system adjusts the size or resolution of the image and further compresses the audio. This reduces both the network transmission bitrate and the disk space occupied by the recorded files. Essentially, transcoding is a method of reducing program data traffic at the expense of program quality. Although this may affect the viewing experience to some extent, it results in more efficient transmission and storage of the program. Live programs saved in this way eventually become historical files, and subsequent retrieval functions are based on this historically saved information.
[0032] In this embodiment, a file saving and querying mechanism is implemented in the historical translation mode;
[0033] File save path rules are fundamental for retrieving program recording history. In history translation mode, audio and video files are pre-stored on the disk. Two common rules for planning file save paths are the root directory date method and the root directory channel method. The root directory date method uses the program's recording date as the root directory name, and the program's physical channel as the second-level subdirectory name. A significant advantage of this method is its ease of use in querying all programs within a specified date range. The root directory channel method uses the program's physical channel as the subdirectory name, and within that subdirectory, the program's date is used as the second-level subdirectory name. Its advantage is its ease of use in querying specific programs. Since there is a mapping relationship between physical channels and program names, in practical applications, the program name can also be used directly instead of the physical channel. In the second-level subdirectories, the program's audio and video files are typically saved using the program's recording time (hour, minute, second) as the filename.
[0034] This embodiment provides examples of saving program recording files, querying broadcast audio files, and querying television video files; the details are as follows:
[0035] Example of saving program recording files:
[0036] Suppose a radio program audio recording was made using the root directory date method, and the full path name of the saved file is "D:\April 21, 2025\Channel 04\123456.MP3". Here, "D:" represents the drive letter; "April 21, 2025" is the recording date; "Channel 04" indicates the physical channel numbered 4; and "123456.MP3" indicates that the program was recorded starting at 12:34:56 and saved as an MP3 file.
[0037] If the root directory channel method is used, the full path name of the saved video / audio recording of a TV program is "E:\CCTV-1\20250421\012345.WMV". Here, "E:" is the drive letter; "CCTV-1" is the physical channel mapped by the program name; "20250421" is the abbreviation of the recording date; and "012345.WMV" indicates that the program was recorded starting at 01:23:45 and saved as a WMV file.
[0038] This implementation includes file querying and editing. Regardless of the path rule used, the key information for querying audio / video files is always only two: the query time period and the channel number. To meet this query requirement, a media query and editing software is designed. This software can locate the specified channel number in the disk file system and filter out files containing valid data within the specified query time period. Subsequently, these files are sorted chronologically. Program data in the first file that is earlier than the start time of the query time period but does not belong to the query time period is removed. All program data in the other intermediate files is integrated. Finally, program data that is earlier than the end time of the query time period and belongs to the query time period is extracted from the last file. By simply merging all the processed program data into a new file, an audio / video clip file generated according to the specified query conditions can be generated.
[0039] In online query scenarios, the edited files generated based on the query information can typically be set up as virtual files, without needing to be physically stored on disk. Especially in streaming media applications, all file information matching the query timeframe can be stored in memory first. This information includes the full file path, the storage timeframe of the program data in each file, and the start and end positions of the valid data, without needing to save the actual program data. When playing the program content to the user, the system dynamically and gradually performs the reading operation of the corresponding file data according to the program playback progress. In other words, it does not read all the data at once before the program starts playing, but rather reads it in real time as needed during playback, thereby improving the efficiency and flexibility of data processing.
[0040] Example of searching for radio recording files:
[0041] Let's take querying a radio recording file as an example. Suppose a user wants to query radio recording files from a specific channel within a specific time period. The media query and editing software first searches the disk file system based on the channel number to find all files corresponding to that channel. Then, it filters and sorts these files according to the time period specified by the user. For files that meet the criteria, it performs data trimming and merging operations according to the file query and editing generation method described above, ultimately generating an audio recording clip file that meets the user's query requirements. During playback, the media playback software reads and plays the corresponding program data in real time based on the file information stored in memory.
[0042] In this embodiment, two media query and editing software are used: one is an audio file query service software used for broadcast audio files, and the other is a video file query service software used for television program audio and video recording files.
[0043] Specifically, design a software program for querying recorded audio files. This program runs on computers that store recorded program files and provides a TCP network interface (port 3333) for querying historical recordings and streaming playback. From any client computer on the network, users can use media playback software to enter the following URL in the interface input box to query and play the recorded program.
[0044] The URL `http: / / 192.168.0.33:3333 / A16 / 2007-1-1T4-5-6 / 2007-1-1T01-59-59` uses the HTTP protocol. In this URL, "http:" indicates the use of the HTTP playback protocol, "192.168.0.33:3333" is the IP address and network port number of the recording file query service software, "A16" indicates the program's channel number is 16, and "2017-1-1T4-5-6 / 2017-01-01T05-06-59" indicates the query period is from 4:05:06 AM to 5:06:59 AM on January 1, 2017. A flexible URL parsing strategy is employed, supporting zero-padding for timestamps (e.g., 04-05-06 is equivalent to 4-5-6) and mixed-case input for channel identifiers, ensuring improved user operation tolerance while maintaining semantic consistency.
[0045] Example of searching for TV recording files:
[0046] The process of searching for TV program audio and video recordings is similar to that of searching for radio recordings. It involves filtering, sorting, and editing files containing data for a specific time period from a particular channel.
[0047] Specifically, design a video recording file query service software that runs on the computer storing the program recording files. It provides an external function for querying and playing historical RTSP recordings via network port 5568. From any client computer on the network, users can use media playback software to enter the following URL address in the interface input box to query and play the recordings.
[0048] The URL rtsp: / / 192.168.0.33:5568 / d4 / 2019-6-19T13-24-0 / 2019-6-19T13-34-0 indicates that the video file query service software uses the RTSP streaming media playback protocol. "192.168.0.33" is the IP address of the software, "5568" is the network port number provided by the software, "d4" indicates that the program's channel number is 4, and "2019-6-19T13-24-0 / 2019-6-19T13-34-0" indicates that the query period is from 13:24:00 on June 19, 2019 to 13:34:00 on June 19, 2019.
[0049] This implementation also includes query linkage; after the user enters query keywords into the text retrieval system, the system will perform a search on the text data stored in the translation computer or translation recording computer based on three retrieval strategies: precise matching, semantic-assisted positioning, and fuzzy semantics. The search results are displayed to the user terminal in list form. Each result adopts a horizontal column layout and simultaneously presents the following core information: 1) The keyword matching statement is highlighted, and a hidden variable is used to store the channel number of the program in which the matching statement is located; 2) The dynamically generated time positioning interval. Because the speech translation function has a backtracking correction mechanism, the timestamp of the translated text is later than the actual time of the program sound, and it may also backtrack and correct more than one sentence. Therefore, the end time of the time interval is directly taken from the translation completion time of the current sentence, while the start time is backtracked until the previous sentence has a different timestamp; 3) For TV program scenarios, the system will also automatically associate and display the video frame thumbnail corresponding to the translation time.
[0050] When a user clicks the hyperlink corresponding to the conclusion information in the search query output list, the text retrieval system first extracts the value of the hidden variable to obtain the channel number of the program, then obtains the time positioning interval, and then synthesizes the URL for audio file query or video file query based on whether the queried program type belongs to broadcast or television, and automatically inputs it into the media playback software to play back the historical program obtained by the user.
[0051] In this embodiment, the precise retrieval is taken as an example of regulatory work. When it is necessary to listen to and confirm the broadcast call sign of a radio program in historical recording files, staff often need to manually input the possible time period of the broadcast call sign in the recording system based on the time information on the program broadcast schedule, and then play it for confirmation. However, this working method has significant drawbacks: the program broadcast schedule often has errors due to temporary insertions, program adjustments, etc. If the broadcast call sign cannot be located through the preset time, staff have to manually drag the playback progress bar and blindly search in the vast recordings. Since the recording files lack other effective auxiliary information besides time information, this blind search method is not only time-consuming and laborious, but also extremely inefficient, seriously affecting the overall effectiveness of regulatory work. If artificial intelligence technology is used to translate the audio content of the program into text and add corresponding timestamp information, the difficulty of retrieving and confirming the broadcast call sign of the radio program will be greatly reduced. Specifically, software is used to directly associate the timestamp information of possible broadcast call signs in the database with the edited playback time period of the recording file. In this way, when a user clicks on a hyperlink that appears to be a broadcast call sign, the system can automatically play the corresponding audio clip, and staff only need to listen to confirm it, which will greatly improve work efficiency and accuracy.
[0052] In this embodiment, the assisted location retrieval specifically includes:
[0053] The above precise retrieval mode can play a significant advantage in specific application scenarios. It is especially applicable to the situation where the program voice can be accurately translated into text, and news commentary programs are typical representatives. Such programs usually have no interference from background music, with excellent signal quality and no obvious noise, providing good conditions for voice translation. However, when faced with music programs or advertisement retrieval work with background sounds, the background music will seriously affect the accuracy of voice recognition, especially the lyrics are often difficult to be perfectly translated into text. To address this issue, the relevance of radio and television program schedules can be used to assist in location retrieval. Radio and television programs have the characteristic of regular rebroadcasts. For any program to be retrieved, among the programs broadcast before or after it, there must be content that can be accurately translated into text within a certain time range. For example, assume that it is necessary to monitor and confirm a certain advertisement, but the voice translation result of the advertisement is not accurate enough. At this time, among all the translated text information, a program text that is more accurately translated before or after the advertisement can be selected as a keyword for retrieval. After the retrieval is completed, several translated texts near the keyword, including several translated texts before and after the keyword, as well as the corresponding timestamps, are all presented. Subsequently, the staff manually judges whether the advertisement program to be retrieved exists within this time period. Through this method, the relevant position of the advertisement can be assisted in location, thereby compensating for the problem of insufficient recognition rate of voice recognition in complex audio environments to a certain extent and improving the accuracy and efficiency of retrieval.
[0054] In this embodiment, the fuzzy retrieval is specifically as follows:
[0055] In the retrieval work of radio and television programs, programs are also retrieved based on certain specific keywords. However, text noise is inevitably generated during the voice translation process, which will greatly reduce the effect of precise retrieval. In this case, fuzzy retrieval becomes an effective solution. It can locate information under non-exact matching conditions. With the triple technologies of character fault tolerance, semantic expansion, and index optimization, an efficient query path is built in the uncertainty of data to help accurately find the programs of interest. This fuzzy location for non-exact matching requirements can be achieved through the following technical framework:
[0056] Character fault tolerance processing can adopt the edit distance algorithm to tolerate spelling errors in the translated text. For example, when the user inputs "News Link Broadcast", the system can match "News Broadcast". In practical applications, the threshold can be set to a difference of 1-2 characters to adapt to different degrees of spelling mistakes. The pinyin fault tolerance algorithm can also be adopted to support pinyin initial letter matching. For example, when the user inputs "xwlb", it can match "News Broadcast". At the same time, it can also handle the conversion problem of polyphonic characters. Polyphonic characters like "industry (háng / xíng)" can also be accurately recognized, thereby enhancing the flexibility and accuracy of retrieval.
[0057] The semantic expansion strategy constructs a thesaurus, combines it with a broadcast terminology database (e.g., expanding "CCTV channel" to "CCTV"), and leverages a domain knowledge graph to expand query terms. This allows the system to understand the underlying semantics of user queries and return more comprehensive search results. Furthermore, by utilizing context awareness and analyzing program metadata, such as the category tag "sports-football," the system can associate semantically similar words. For example, when a user searches for "World Cup," the system can also associate it with related terms like "UEFA Champions League," thereby expanding the search scope and improving the comprehensiveness of the search.
[0058] The index optimization mechanism utilizes a hierarchical inverted index to build a timestamp-linked index structure for program audio text based on word segmentation results (e.g., "evening / news / summary"). This hierarchical inverted index structure supports skip list-based quick location, significantly improving retrieval speed and allowing users to find the desired program segments more quickly. Combined with dynamic weight adjustments, high-frequency words in the program (e.g., program names) are de-weighted to improve the retrieval performance of long-tail content. In this way, some relatively less popular but potentially relevant program content can also be better retrieved, preventing search results from being overly dominated by high-frequency words.
[0059] These technologies effectively address text noise generated during speech-to-text translation, ensuring retrieval recall while also improving retrieval accuracy, providing strong technical support for radio and television program retrieval.
[0060] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0061] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A broadcast television speech translation and retrieval system, characterized by: The system includes a recording system and a text retrieval system; The recording system operates in two modes: real-time translation and historical translation, and is used in two application scenarios: real-time program monitoring and historical program review and monitoring. In real-time translation mode, when real-time audio and video data of a program is received, the recording system saves the audio and video data as an audio and video recording file, and simultaneously translates the program's audio content into text in real time; and stores the moment when the captured text is translated as a timestamp along with the translated text in the data system, which is used by the text retrieval system to provide time positioning for retrieval and playback. In historical translation mode, the recording system is used to play back audio and video files and extract program audio, and then translate the extracted program audio into text; at the same time, it records the timestamp of the translated text, that is, the historical time when the audio and video files were recorded is used as the timestamp. Used for text retrieval system queries; The text retrieval system selects the corresponding speech-to-text text of the program from the text query list output by the recording system, thereby enabling the playback of the original audio and video content of the program.
2. The broadcast television speech translation and retrieval system according to claim 1, characterized in that: In real-time translation mode, the recording system includes a broadcast television signal source, an interface circuit board, a smartphone, a communication network, and a translation and recording computer; The program signal emitted by the broadcast television signal source is received by the interface circuit board, which extracts the audio stream data. The interface circuit board transmits the audio stream data to a smartphone via a USB interface and sends the audio stream data to a transcribing and recording computer via a communication network. The transcribing and recording computer is used to complete the audio recording and video recording.
3. The broadcast television speech translation and retrieval system according to claim 2, characterized in that: The smartphone runs voice translation and transmission software to translate audio stream data into text, which is then sent to a translation and recording computer via a communication network. The translation and recording computer saves the text information and uses the time of receiving the text information as a timestamp to provide a time location basis for the text retrieval system to play back the audio.
4. The broadcast television speech translation and retrieval system according to claim 3, characterized in that: A computer set up in the communication network can remotely log in to the translation and recording computer through the communication network, click on the hyperlink corresponding to the conclusion information in the search query output list, and accurately play back the audio and video content of the original playback of the corresponding program.
5. The broadcast television speech translation and retrieval system according to claim 1, characterized in that: In the historical translation mode, the recording system includes an audio / video recording server, a player, an interface circuit board, a smartphone, a communication network, a translation computer, and a query computer; The program's historical audio and video files are stored on the recording server and played back by the player; The audio signal from the playback program is converted into a program sound data stream through an interface circuit board and then transmitted to a smartphone. The smartphone translates and outputs the text, which is then sent to a translation computer via a communication network. The translation computer adds a timestamp to the translated text. The timestamp is the moment when the historical audio and video file was recorded, and the timestamp is attached to the corresponding text information.
6. The broadcast television speech translation and retrieval system according to claim 5, characterized in that: Computers in a communication network remotely log in to the query computer via the communication network. The query computer searches the translated text stored in the translation computer based on the keywords entered by the user. When the user clicks the hyperlink corresponding to the conclusion information in the search query output list, the audio and video content of the original program when the audio was played is accurately edited from the audio and video recording server and transmitted to the computer being operated by the user via the communication network for playback.
7. The broadcast television speech translation and retrieval system according to claim 6, characterized in that: Translation and storage software runs on the translation computer, and the translation and storage software receives and stores the text data stream transmitted by the smartphone; The translation and saving software has a built-in dynamic text processing engine, which is used to backtrack and delete the most recently received character when a backspace key is detected; It is also equipped with an intelligent punctuation recognition module. When the three key punctuation marks, comma, period, or question mark are captured, the text segmentation mechanism is automatically triggered to parse the continuous text stream into short sentence units with independent semantics, and a timestamp is added to each newly generated short sentence fragment to form an index structure with a text content and text time mapping relationship.
8. The broadcast television speech translation and retrieval system according to claim 1, characterized in that: The smartphone downloads voice input method software and runs text forwarding software; the text forwarding software has a text input box, the voice input method converts the received voice data into text and writes it into the text input box to obtain the text result of voice translation, and sends the text result to the corresponding computer through the communication network.
9. The broadcast television speech translation and retrieval system according to claim 8, characterized in that: The text forwarding software is equipped with a backtracking and correction function; When the voice input method software performs backspace correction, the text forwarding software sends backspace keys to the corresponding computer through the communication network. The number of backspace keys is equal to the number of characters that the voice input method corrected. After all the backspace keys have been sent, all the corrected text is resent to the corresponding computer.
10. The broadcast television speech translation and retrieval system according to claim 1, characterized in that: In real-time translation mode: When the recording system receives the audio and video data stream of the program, it performs transcoding and storage operations; it converts the received program audio and video data into an audio encoding format, then saves it as an audio and video recording file and stores it on the disk.