Intelligent display voice-to-text method and system, and medium

By converting audio and video files into text, splitting sentences according to punctuation marks, looking up merged text paragraphs in a dictionary, rendering and highlighting sensitive words, the system solves the problem of users quickly locating illegal content in audio and video, thus improving work efficiency.

CN116320614BActive Publication Date: 2026-07-21SHENZHEN CRAFTSMAN NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN CRAFTSMAN NETWORK TECH CO LTD
Filing Date
2022-12-02
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, users need to play the audio and video frame by frame and second by second to determine whether there are problems with the wording, which is inefficient and lacks the function of highlighting and searching for sensitive words and prohibited words, making it difficult for users to quickly locate illegal content.

Method used

By converting audio and video files into text, splitting sentences according to pre-defined punctuation marks, searching dictionaries to obtain historical data or merging text paragraphs, rendering and highlighting sensitive and banned words, and combining with the search function to quickly locate illegal content.

Benefits of technology

It synchronizes audio and video playback progress with rendering statements, allowing users to quickly identify the location of problems, reducing the workload of compliance processing and improving work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116320614B_ABST
    Figure CN116320614B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent display voice-to-text method and system and a medium, and the method comprises the following steps: when an audio and video file is captured, the audio and video file is converted into text; each sentence of text is segmented according to a pre-agreed punctuation mark to obtain segmented sentence text information; a pre-established dictionary table is searched; if historical data corresponding to the segmented sentence text information is found in the dictionary table, a corresponding previously combined text paragraph is obtained from the dictionary table; the text paragraph is rendered; single characters in the rendered text paragraph content are processed according to sensitive words, forbidden words or search words; and the processed single characters are synchronized with the text paragraph on a page for display. Through the application, a user can quickly find out where a violation is located, and can also quickly jump to a corresponding progress for manual auditing to confirm whether a problem exists, thereby greatly reducing the workload of the user for compliance processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to an intelligent speech-to-text method, system, and medium. Background Technology

[0002] Currently, most services on the market provide video monitoring of media stream data. If users need to determine whether there are problems with the dialogue in the audio and video, they need to play it frame by frame and second by second to confirm. This is undoubtedly a waste of manpower and very inefficient in the long audio and video.

[0003] Currently, established vendors such as iFlytek and Tencent Cloud can be used to retrieve audio-related text from media streams. However, there are no features such as custom sentence combinations, sensitive word and banned word highlighting, or search functionality. Users still face considerable difficulties in finding illegal words and locating the corresponding playback points. Summary of the Invention

[0004] The main objective of this invention is to propose an intelligent speech-to-text method, system, and medium, which aims to synchronize the audio and video playback progress with the rendered sentences. Simultaneously, it highlights words from a specific dictionary of sensitive and prohibited words, providing users with a more intuitive experience. Furthermore, it allows users to quickly identify violations through a search function and jump to the corresponding progress for manual review to confirm the existence of the problem, greatly reducing the workload of compliance processing and improving work efficiency.

[0005] To achieve the above objectives, the present invention provides an intelligent speech-to-text conversion method, the method comprising the following steps: When an audio or video file is captured, the audio or video file is converted into text. Each sentence is segmented according to pre-agreed punctuation marks to obtain the segmented sentence text information; Search a pre-built dictionary table; If historical data corresponding to the segmented sentence text information is found in the dictionary table, the corresponding previously merged text segment is retrieved from the dictionary table. Render the text paragraph; Process single characters in the rendered text paragraphs based on sensitive words, banned words, or search terms; The processed single character is displayed synchronously on the page along with the text paragraph.

[0006] A further technical solution of the present invention is that, after the step of searching the pre-established dictionary table, the method further includes: If no historical data corresponding to the segmented sentence text information is found in the dictionary table, then the text increase number is obtained; Based on the increase in the number of sentence text increments, a corresponding strategy is used to merge the segmented sentence text information into paragraphs.

[0007] A further technical solution of the present invention is that the step of merging the segmented sentence text information into paragraphs using a corresponding strategy based on the increase in sentence text size includes: Compare the increase in the number of sentences with a preset threshold. If the increase in the number of statement texts is greater than or equal to the preset threshold, then a preset number of statement texts are processed every preset time interval to merge paragraphs. If the increase in the amount of the statement text is greater than or equal to the preset threshold, then the paragraphs are directly merged.

[0008] A further technical solution of the present invention is that the preset threshold is 100 sentences and the preset duration is 200ms.

[0009] A further technical solution of the present invention is that the step of processing single characters in the rendered text paragraph content based on sensitive words, banned words, or search terms includes: Based on changes in sensitive words, banned words, or search terms, assign a corresponding type set to each character in the paragraph and store it.

[0010] To achieve the above objectives, the present invention also proposes an intelligent speech-to-text display system, the system comprising a memory, a processor, and an intelligent speech-to-text display program stored on the processor, wherein the intelligent speech-to-text display program, when run by the processor, performs the following steps: When an audio or video file is captured, the audio or video file is converted into text. Each sentence is segmented according to pre-agreed punctuation marks to obtain the segmented sentence text information; Search a pre-built dictionary table; If historical data corresponding to the segmented sentence text information is found in the dictionary table, the corresponding previously merged text segment is retrieved from the dictionary table. Render the text paragraph; Process single characters in the rendered text paragraphs based on sensitive words, banned words, or search terms; The processed single character is displayed synchronously on the page along with the text paragraph.

[0011] A further technical solution of the present invention is that, when the intelligent speech-to-text program is run by the processor, it further performs the following steps: If no historical data corresponding to the segmented sentence text information is found in the dictionary table, then the text increase number is obtained; Based on the increase in the number of sentence text increments, a corresponding strategy is used to merge the segmented sentence text information into paragraphs.

[0012] A further technical solution of the present invention is that, when the intelligent speech-to-text program is run by the processor, it further performs the following steps: Compare the increase in the number of sentences with a preset threshold. If the increase in the number of statement texts is greater than or equal to the preset threshold, then a preset number of statement texts are processed every preset time interval to merge paragraphs. If the increase in the amount of the statement text is greater than or equal to the preset threshold, then the paragraphs are directly merged.

[0013] A further technical solution of the present invention is that the preset threshold is 100 sentences and the preset duration is 200ms.

[0014] To achieve the above objectives, the present invention also proposes a computer-readable storage medium storing an intelligent speech-to-text program, wherein the intelligent speech-to-text program, when run by a processor, performs the steps of the method described above.

[0015] The beneficial effects of this invention's intelligent speech-to-text method, system, and medium are as follows: Upon capturing an audio / video file, the invention converts the audio / video file into text; it segments each sentence of text according to pre-defined punctuation marks to obtain segmented sentence text information; it searches a pre-established dictionary table; if historical data corresponding to the segmented sentence text information is found in the dictionary table, the corresponding previously merged text segment is retrieved from the dictionary table; the text segment is rendered; single characters in the rendered text segment are processed based on sensitive words, banned words, or search terms; and the processed single characters are displayed synchronously on the page along with the text segment. When users perform compliance processing of media streams, they can start with the text rendered from speech-to-text, quickly identify violations through intelligent segmentation, highlighted sensitive and banned words, and a search function, and quickly jump to the corresponding progress for manual review to confirm the existence of the problem, greatly reducing the workload of users in compliance processing. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a preferred embodiment of the intelligent speech-to-text method of the present invention; Figure 2 This is a screenshot showing the display effect of the text converted from the media stream and returned sentence by sentence by the backend; Figure 3 This is a screenshot of the page rendering. Figure 4 This is a schematic diagram of the overall process of the intelligent speech-to-text method of the present invention.

[0018] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0020] Please refer to Figures 1 to 4 This invention proposes an intelligent speech-to-text display method. A preferred embodiment of this intelligent speech-to-text display method includes the following steps: Step S10: When an audio or video file is captured, the audio or video file is converted into text.

[0021] It should be noted that in this embodiment, the backend returns the text converted from the media stream sentence by sentence, as shown in the following figure. Figure 2 As shown.

[0022] Step S20: Segment each sentence text according to the pre-agreed punctuation marks to obtain the segmented sentence text information.

[0023] Step S30: Search the pre-established dictionary table.

[0024] Step S40: If historical data corresponding to the segmented sentence text information is found in the dictionary table, the corresponding previously merged text segment is obtained from the dictionary table.

[0025] It should be noted that this embodiment segments each sentence of text according to agreed-upon punctuation marks, which involves processing large amounts of data. Because browser rendering performance has bottlenecks, excessive computation when returning a certain amount of data can cause the page to crash. This embodiment, however, uses a space-for-time trade-off: the data for each paragraph's rendering, as well as the HTML content, are stored in a dictionary table using a specific key. If historical data is encountered, it is retrieved and rendered directly.

[0026] In this embodiment, after step S30, which involves searching a pre-established dictionary table, the method further includes: Step S301: If no historical data corresponding to the segmented sentence text information is found in the dictionary table, then the text increase number is obtained.

[0027] Step S302: Based on the increase in the number of sentence text increments, the segmented sentence text information is merged into paragraphs using the corresponding strategy.

[0028] Step S302, the step of merging the segmented sentence text information into paragraphs using a corresponding strategy based on the increase in sentence text size, specifically includes: The increase in the amount of the statement text is compared with a preset threshold.

[0029] If the increase in the number of statement texts is greater than or equal to the preset threshold, then a preset number of statement texts are processed at preset intervals to merge paragraphs.

[0030] If the increase in the amount of the statement text is greater than or equal to the preset threshold, then the paragraphs are directly merged.

[0031] As one implementation scheme, in this embodiment, the preset threshold is 100 sentences and the preset duration is 200ms.

[0032] In this embodiment, if no historical data corresponding to the segmented sentence text information is found in the dictionary table, two cases are handled based on the increase in the number of sentences (historical data must be used when the increase decreases): 1. If the increase is too large (e.g., more than 100 sentences), a certain number of sentences (e.g., every 100 sentences) are processed every certain time interval (e.g., 200ms) to merge paragraphs; 2. If the increase is small (e.g., less than 100 sentences), the paragraphs are merged directly.

[0033] Step S50: Render the text paragraph.

[0034] Step S60: Process single characters in the rendered text paragraph content based on sensitive words, banned words, or search terms.

[0035] Specifically, the step of processing single characters in the rendered text paragraph content based on sensitive words, banned words, or search terms includes: Based on changes in sensitive words, banned words, or search terms, assign a corresponding type set to each character in the paragraph and store it.

[0036] This embodiment assigns a corresponding type set to each character in a paragraph based on changes in sensitive words, banned words, and search terms. For example, if the current sensitive word is "most" and the search term is also "most," then when the character "most" is encountered in the paragraph, a class name related to the sensitive word and search term is assigned to that character (for demonstration purposes). Of course, this is also stored, and if the vocabulary set corresponding to the paragraph has not changed during each rendering, historical data will be retrieved for rendering.

[0037] Step S70: The processed single character is displayed synchronously on the page along with the text paragraph.

[0038] The beneficial effects of the intelligent speech-to-text method of this invention are as follows: Upon capturing an audio / video file, the invention converts the audio / video file into text; it segments each sentence of text according to pre-agreed punctuation marks to obtain segmented sentence text information; it searches a pre-established dictionary table; if historical data corresponding to the segmented sentence text information is found in the dictionary table, the corresponding previously merged text segment is retrieved from the dictionary table; the text segment is rendered; single characters in the rendered text segment are processed according to sensitive words, banned words, or search terms; the processed single characters are displayed synchronously on the page along with the text segment. When users perform compliance processing of media streams, they can start with the text rendered from speech-to-text, quickly identify violations through intelligent segmentation, highlighted sensitive and banned words, and a search function, and quickly jump to the corresponding progress for manual review to confirm the existence of the problem, greatly reducing the workload of users in compliance processing.

[0039] To achieve the above objectives, the present invention also proposes an intelligent speech-to-text display system, the system comprising a memory, a processor, and an intelligent speech-to-text display program stored on the processor, wherein the intelligent speech-to-text display program, when run by the processor, performs the following steps: When an audio or video file is captured, the audio or video file is converted into text. Each sentence is segmented according to pre-agreed punctuation marks to obtain the segmented sentence text information; Search a pre-built dictionary table; If historical data corresponding to the segmented sentence text information is found in the dictionary table, the corresponding previously merged text segment is retrieved from the dictionary table. Render the text paragraph; Process single characters in the rendered text paragraphs based on sensitive words, banned words, or search terms; The processed single character is displayed synchronously on the page along with the text paragraph.

[0040] Furthermore, when the intelligent speech-to-text program is run by the processor, it also performs the following steps: If no historical data corresponding to the segmented sentence text information is found in the dictionary table, then the text increase number is obtained; Based on the increase in the number of sentence text increments, a corresponding strategy is used to merge the segmented sentence text information into paragraphs.

[0041] Furthermore, when the intelligent speech-to-text program is run by the processor, it also performs the following steps: Compare the increase in the number of sentences with a preset threshold. If the increase in the number of statement texts is greater than or equal to the preset threshold, then a preset number of statement texts are processed every preset time interval to merge paragraphs. If the increase in the amount of the statement text is greater than or equal to the preset threshold, then the paragraphs are directly merged.

[0042] Furthermore, the preset threshold is 100 sentences, and the preset duration is 200ms.

[0043] The beneficial effects of the intelligent speech-to-text system of this invention are as follows: Upon capturing an audio / video file, the system converts the audio / video file into text; it segments each sentence of text according to pre-defined punctuation marks to obtain segmented sentence text information; it searches a pre-established dictionary table; if historical data corresponding to the segmented sentence text information is found in the dictionary table, the corresponding previously merged text segment is retrieved from the dictionary table; the text segment is rendered; single characters in the rendered text segment are processed according to sensitive words, banned words, or search terms; and the processed single characters are displayed synchronously on the page along with the text segment. When users perform compliance processing of media streams, they can start with the text rendered from speech-to-text, quickly identify violations through intelligent segmentation, highlighted sensitive and banned words, and a search function, and quickly jump to the corresponding progress for manual review to confirm the existence of the problem, greatly reducing the workload of users in compliance processing.

[0044] To achieve the above objectives, the present invention also proposes a computer-readable storage medium storing an intelligent speech-to-text program. When the intelligent speech-to-text program is run by a processor, it executes the steps described in the above embodiments, which will not be repeated here.

[0045] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. All equivalent structural transformations made under the concept of the present invention using the contents of the present invention specification and drawings, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.

Claims

1. A method for intelligently displaying speech-to-text conversion, characterized in that, The method includes the following steps: When an audio or video file is captured, the audio or video file is converted into text. Each sentence is segmented according to pre-agreed punctuation marks to obtain the segmented sentence text information; Search a pre-built dictionary table; If historical data corresponding to the segmented sentence text information is found in the dictionary table, the corresponding previously merged text paragraph is retrieved from the dictionary table; the data rendered for each paragraph and the HTML content are stored in the dictionary table according to the corresponding key, and if historical data is encountered, the data is retrieved and rendered directly; Render the text paragraph; Process single characters in the rendered text paragraphs based on sensitive words, banned words, or search terms; The processed single character is displayed synchronously on the page along with the text paragraph; The step of searching the pre-established dictionary table is followed by: If no historical data corresponding to the segmented sentence text information is found in the dictionary table, then the text increase number is obtained; Based on the increase in the number of sentence text increments, the segmented sentence text information is merged into paragraphs using a corresponding strategy. The step of merging the segmented sentence text information into paragraphs using a corresponding strategy based on the increase in the number of sentence text increments includes: Compare the increase in the number of sentences with a preset threshold. If the increase in the number of statement texts is greater than or equal to the preset threshold, then a preset number of statement texts are processed every preset time interval to merge paragraphs. If the increase in the amount of the statement text is greater than or equal to the preset threshold, then the paragraphs are directly merged.

2. The intelligent speech-to-text method according to claim 1, characterized in that, The preset threshold is 100 sentences, and the preset duration is 200ms.

3. The intelligent speech-to-text method according to claim 1, characterized in that, The steps for processing single characters in the rendered text paragraph content based on sensitive words, banned words, or search terms include: Based on changes in sensitive words, banned words, or search terms, assign a corresponding type set to each character in the paragraph and store it.

4. An intelligent speech-to-text display system, characterized in that, The system includes a memory, a processor, and an intelligent speech-to-text program stored on the processor. When the intelligent speech-to-text program is run by the processor, it performs the following steps: When an audio or video file is captured, the audio or video file is converted into text. Each sentence is segmented according to pre-agreed punctuation marks to obtain the segmented sentence text information; Search a pre-built dictionary table; If historical data corresponding to the segmented sentence text information is found in the dictionary table, the corresponding previously merged text paragraph is retrieved from the dictionary table; the data rendered for each paragraph and the HTML content are stored in the dictionary table according to the corresponding key, and if historical data is encountered, the data is retrieved and rendered directly; Render the text paragraph; Process single characters in the rendered text paragraphs based on sensitive words, banned words, or search terms; The processed single character is displayed synchronously on the page along with the text paragraph; When the intelligent speech-to-text program is run by the processor, it also performs the following steps: If no historical data corresponding to the segmented sentence text information is found in the dictionary table, then the text increase number is obtained; Based on the increase in the number of sentence text increments, the segmented sentence text information is merged into paragraphs using a corresponding strategy. When the intelligent speech-to-text program is run by the processor, it also performs the following steps: Compare the increase in the number of sentences with a preset threshold. If the increase in the number of statement texts is greater than or equal to the preset threshold, then a preset number of statement texts are processed every preset time interval to merge paragraphs. If the increase in the amount of the statement text is greater than or equal to the preset threshold, then the paragraphs are directly merged.

5. The intelligent speech-to-text system according to claim 4, characterized in that, The preset threshold is 100 sentences, and the preset duration is 200ms.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an intelligent speech-to-text program, which, when run by a processor, performs the steps of the method as described in any one of claims 1 to 3.