Telephone subtitle system and application method thereof

Real-time conversion and display of voice information through the telephone subtitle system, solving the call needs of hearing-impaired people, improving the accuracy and operational convenience of voice recognition, ensuring information security, and improving the communication efficiency and quality of life of hearing-impaired people.

CN120356471APending Publication Date: 2025-07-22AVIC HUADONG OPTOELECTRONICS (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510330724.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing communication technology cannot effectively meet the call needs of people with hearing impairments. Voice recognition technology has shortcomings in accuracy and reliability, and the existing subtitle calling system still needs to be improved in terms of operational convenience and privacy protection.

Method used

It provides a telephone subtitle system, which establishes voice calls with target communication devices through user communication devices, uses display devices and voice recognition result acquisition devices to realize real-time conversion and display of voice information, supports manual and automatic voice recognition mode, and combines SIM card phone number association technology to ensure information security and privacy protection.

Benefits of technology

Provide stable and reliable subtitle call services for people with hearing impairments, improve the accuracy and reliability of voice recognition, simplify operational processes, improve communication efficiency and quality of life, enhance information security, and promote barrier-free communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356471A_ABST
    Figure CN120356471A_ABST
Patent Text Reader

Abstract

The invention provides a telephone subtitle system and an application method thereof, and the system comprises user communication equipment which builds voice communication with target communication equipment through a mobile network and obtains voice information of an opposite side; the display equipment comprises a voice information receiving unit, a voice information sending unit, a voice recognition result receiving unit and a display unit, and is used for receiving the voice information, sending the voice information to the voice recognition result acquisition equipment, receiving a voice recognition result and displaying the voice recognition result in real time; the voice recognition result acquisition equipment is used for transmitting the voice information to an operator or voice recognition equipment and receiving a generated voice recognition result; and the voice recognition equipment is used for converting the voice information into text data. Both the user communication equipment and the target communication equipment are provided with SIM cards, and call connection is established through telephone numbers recorded in the SIM cards. The communication efficiency can be improved, and the security and privacy of information are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and more particularly, to a telephone captioning system and its application method. Background Art

[0002] In the development of modern communication technologies, voice calls, as an important means of communication, are widely used in various communication scenarios, including traditional voice communications based on telecommunications networks, Internet-based voice communications (such as VoIP), video communications, etc. Although these communication methods meet the needs of users to a certain extent, there are still some limitations in actual use. For example, traditional telephone communications mainly rely on voice transmission, lack additional functions, do not support instant messaging, and it is difficult for people with hearing impairments to obtain call content. Internet-based voice communications, although rich in functions, have high requirements for network conditions, large fluctuations in voice quality, and risks of privacy leakage. In addition, although existing speech recognition technologies can convert speech into text, in actual applications, their accuracy and stability still need to be improved, especially in complex environments or usage scenarios of special populations.

[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: existing communication technologies cannot effectively meet the call needs of people with hearing impairments, speech recognition technologies have deficiencies in accuracy and reliability, and existing captioning call systems still need to be improved in terms of operation convenience and privacy protection. Summary of the Invention

[0004] The present invention provides a telephone captioning system and its application method.

[0005] In a first aspect of the present invention, a telephone captioning system is provided, including:

[0006] A user communication device, configured to establish a voice call with a target communication device through a mobile network and obtain the voice information of the target communication device;

[0007] A display device, including a voice information receiving unit, a voice information sending unit, a speech recognition result receiving unit, and a display unit, wherein:

[0008] The voice information receiving unit receives the voice information transmitted by the user communication device through wireless communication;

[0009] The voice information sending unit sends the voice information to a speech recognition result acquisition device;

[0010] The speech recognition result receiving unit receives the speech recognition result returned by the speech recognition result acquisition device;

[0011] The display unit displays the speech recognition result in real time;

[0012] The speech recognition result acquisition device is used to transmit the speech information to the operator or the speech recognition device and receive the speech recognition result generated by the operator or the speech recognition device;

[0013] The speech recognition device is used to convert the speech information into text data;

[0014] Both the user communication device and the target communication device are configured with SIM cards and establish a call connection through the phone numbers recorded in the SIM cards.

[0015] Further, the display device and the user communication device are of a split structure or integrated into the housing of the user communication device.

[0016] Further, the speech recognition result acquisition device transmits the speech information to the operator or the speech recognition device according to the operation object selected by the user.

[0017] Further, the display unit of the display device includes operation object selection buttons for switching the operation object to the operator or the speech recognition device;

[0018] When the operator is selected, the operator generates a speech recognition result through manual input;

[0019] When the speech recognition device is selected, the speech recognition device generates a speech recognition result through automatic recognition.

[0020] Further, the display unit adopts a differentiated display method for the speech recognition result generated by the operator and the speech recognition result generated by the speech recognition device, including at least one of font, color or background.

[0021] Further, the speech recognition result acquisition device pauses obtaining the speech recognition result when it detects that the speech silent time exceeds a preset threshold.

[0022] Further, the display device is directly connected to the Internet through wireless communication and conducts data interaction with the speech recognition result acquisition device.

[0023] Further, the display unit also displays the speech recognition result of the speech information of the user communication device and displays it in a partitioned manner with the speech recognition result of the target communication device.

[0024] Further, when the speech recognition result acquisition device detects that the speech recognition result contains a preset keyword, it triggers the display unit to display a warning message.

[0025] In a second aspect of the present invention, there is provided a method for applying telephone subtitles, including:

[0026] Establish a voice call through a user communication device and a target communication device, and obtain the voice information of the target communication device;

[0027] Transmit the voice information to a display device through wireless communication;

[0028] Send the voice information to a voice recognition result collection device through the display device;

[0029] Transmit the voice information to an operator or a voice recognition device through the voice recognition result collection device, and receive the generated voice recognition result;

[0030] Return the voice recognition result to the display device;

[0031] The voice recognition result is displayed in real time through the display unit of the display device.

[0032] According to the above embodiments of the present invention, it has at least the following beneficial effects: The telephone subtitle system of the present invention can provide stable and reliable subtitle call services for hearing-impaired people. Through the SIM card phone number association technology, the user operation process is simplified, and the subtitle call function can be started with one key, and the subtitles are displayed in real time, enabling the user to instantly obtain the content of the other party's speech. The system adopts a dual-device collaborative working mode, efficiently converts the voice information in the call into text subtitles, ensures a high accuracy conversion from voice to text, and supports both manual and automatic voice recognition modes. The user can flexibly select according to the importance of the call, further improving the accuracy and reliability of voice recognition.

[0033] In addition, the present invention can also provide barrier-free communication support for hearing-impaired people in multiple scenarios such as home, work, and social interactions, significantly improving their communication efficiency and quality of life. The system conducts calls and subtitle transmission through the phone number in the SIM card, ensuring the security and privacy protection of information. Users do not need to worry about information leakage or unauthorized access. At the same time, the system can generate certain social benefits, promote barrier-free communication in society, and enhance the care and inclusion for special groups. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, wherein:

[0035] Figure 1 is a schematic structural diagram of a telephone subtitle system provided by an embodiment of the present invention;

[0036] Figure 2 The flowchart of the application method of telephone subtitles provided by an embodiment of the present invention;

[0037] Figure 3 Schematically shows the structural diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0038] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present invention, rather than limiting the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to be able to fully convey the scope of the present invention to those skilled in the art.

[0039] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, a device, an equipment, a method or a computer program product. Therefore, the present invention can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0040] It should be noted that the number of any element in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0041] The following reference Figure 1 , Figure 1 The structural diagram of the telephone subtitle system provided by an embodiment of the present invention. As Figure 1 shown, a telephone subtitle system includes:

[0042] A user communication device, configured to establish a voice call with a target communication device through a mobile network and obtain the voice information of the target communication device;

[0043] A display device, including a voice information receiving unit, a voice information sending unit, a voice recognition result receiving unit and a display unit, wherein:

[0044] The voice information receiving unit receives the voice information transmitted by the user communication device through wireless communication;

[0045] The voice information sending unit sends the voice information to a voice recognition result collection device;

[0046] The voice recognition result receiving unit receives the voice recognition result returned by the voice recognition result collection device;

[0047] The display unit displays the voice recognition result in real time;

[0048] A voice recognition result collection device, which is used to transmit the voice information to an operator or a voice recognition device and receive the voice recognition result generated by the operator or the voice recognition device;

[0049] The voice recognition device is used to convert the voice information into text data;

[0050] Both the user communication device and the target communication device are configured with SIM cards, and a call connection is established through the phone numbers recorded in the SIM cards.

[0051] It should be noted that the phone caption system of the present invention establishes a voice call through the user communication device and the target communication device, and obtains the voice information of the target communication device. The user communication device refers to a device that can make a voice call through a mobile network, such as a smart phone, a tablet computer, or a smart watch supporting a SIM card, etc. Its core function is to realize voice calls and transmit voice information. The display device includes a voice information receiving unit, a voice information sending unit, a voice recognition result receiving unit, and a display unit, which are used to receive voice information and send it to the voice recognition result collection device, and at the same time receive and display the voice recognition result in real time. Voice information refers to the voice signal generated during a call, which can be transmitted and recognized after being collected and processed. The voice recognition result collection device is a key component in the system, which is responsible for transmitting the voice information to an operator or a voice recognition device and receiving the voice recognition result generated by it. Here, the operator refers to a manual operator, and the voice recognition device is an automatic speech recognition system based on artificial intelligence technology.

[0052] Specifically, both the user communication device and the target communication device are configured with SIM cards, and a call connection is established through the phone numbers recorded in the SIM cards. The SIM card (Subscriber Identity Module) is a module used to store user identity information in a mobile communication device, which can realize the connection and authentication between the device and the mobile network. The voice information receiving unit receives the voice information transmitted by the user communication device through wireless communication. The wireless communication can adopt methods such as Bluetooth, Wi-Fi, or Zigbee, etc. The specific parameter settings need to be optimized according to the communication protocol and transmission distance requirements of the device. The voice information sending unit sends the voice information to the voice recognition result collection device, which can be a cloud server or a local server. Its function is to process the received voice information and allocate it to an operator or a voice recognition device. The voice recognition device uses advanced voice recognition algorithms to convert the voice information into text data. Its recognition accuracy and response speed are key indicators of the system performance.

[0053] Preferably, after receiving the voice information, the voice recognition result collection device can transmit the voice information to the operator or the voice recognition device according to the user's selection. The user can switch through the operation object selection button on the display device to select the manual input or automatic recognition mode. When the operator is selected, the operator generates the voice recognition result through manual input. Although this method is slower, it has a higher accuracy rate; when the voice recognition device is selected, the system generates the voice recognition result through automatic recognition. This method is fast, but there may be a certain misrecognition rate. In addition, the voice recognition result collection device can also pause obtaining the voice recognition result when it detects that the voice silence time exceeds the preset threshold. For example, the silence time is set to 5 seconds to avoid unnecessary voice recognition processing during the call gap, thereby saving system resources and improving efficiency.

[0054] In some embodiments, the display device and the user communication device are of a split structure or integrated into the housing of the user communication device.

[0055] It should be noted that in the telephone caption system of the present invention, the display device and the user communication device can adopt a split structure or be integrated into the housing of the user communication device. The split structure means that the display device and the user communication device are two independent physical entities that work together through wireless communication or other connection methods. This design allows for flexible selection of the device combination according to the user's needs and usage scenarios. For example, in some cases, the user may want to separate the display device from the communication device for more convenient viewing of the captions; in other cases, the integrated design can provide a more compact and portable usage experience. This flexibility provides users with diverse options to meet the usage requirements in different scenarios.

[0056] Specifically, the display device of the split structure can be connected to the user communication device through wireless communication technologies such as Bluetooth and Wi-Fi to achieve the transmission of voice information and the display of voice recognition results. In this structure, the display device can be an independent tablet computer, smart watch or other portable device, and its display unit can be optimized according to the screen size and resolution of the device. For example, for a larger tablet computer screen, a larger display area can be set to improve the readability of the captions; for a smaller smart watch screen, scrolling display or segmented display can be adopted. The integrated design embeds the display device into the housing of the user communication device. For example, a dedicated display area is reserved in the housing of a smart phone or tablet computer for real-time display of voice recognition results. This design not only saves space but also improves the integrity and aesthetics of the device.

[0057] Preferably, for a split structure, the display device can be used in cooperation with the user communication device through magnetic connection or a dedicated bracket for quick installation and disassembly when needed. For example, the user can fix the display device beside the communication device during a call for convenient subtitle viewing. For an integrated design, the display unit of the display device can adopt a liquid crystal display (LCD) or an organic light-emitting diode (OLED) screen with adjustable brightness and contrast to adapt to different lighting environments. In addition, the display device can also set the font size, color, and background according to the user's preferences to further enhance the user experience. For example, the user can select a high-contrast display mode to clearly view subtitles even in strong light.

[0058] In some embodiments, the voice recognition result acquisition device transmits the voice information to the operator or the voice recognition device according to the operation object selected by the user.

[0059] It should be noted that in the telephone subtitle system of the present invention, the voice recognition result acquisition device can transmit voice information to the operator or the voice recognition device according to the operation object selected by the user. The operation object refers to the voice recognition method selected by the user during a call, that is, manual recognition or automatic recognition. The operator refers to professional trained artificial voice recognition personnel who can convert voice information into text by dictation; while the voice recognition device is a system based on artificial intelligence technology that can automatically convert voice signals into text. The role of the voice recognition result acquisition device is to distribute voice information to the corresponding recognition entity according to the user's selection and receive the generated voice recognition result. This design provides users with flexible voice recognition options to meet the needs in different scenarios.

[0060] Specifically, after receiving the voice information transmitted by the display device, the voice recognition result acquisition device will decide whether to send the voice information to the operator or the voice recognition device according to the selection made by the user through the operation object selection button on the display device. For example, the user can select the manual recognition or automatic recognition mode by clicking the button on the display device. When manual recognition is selected, the voice recognition result acquisition device will transmit the voice information to the operator, and the operator will input the voice content through the keyboard and generate a text result; when automatic recognition is selected, the voice information will be sent to the voice recognition device, which will generate text through voice recognition algorithms. The voice recognition result acquisition device needs to have high-efficiency data transmission and processing capabilities to ensure that the voice information can be transmitted in a timely manner and the recognition result can be returned. For example, the voice recognition result acquisition device can adopt a high-speed network connection, such as 5G or Wi-Fi 6, to ensure low latency and high reliability of data transmission.

[0061] Preferably, the speech recognition result acquisition device can recommend or automatically switch the speech recognition mode according to the user's preference or the importance of the call. For example, when the call content is more important or complex, the system can preferably recommend the manual recognition mode to ensure higher recognition accuracy; while in ordinary call scenarios, it can be automatically switched to the automatic recognition mode to improve efficiency. In addition, the speech recognition result acquisition device can also dynamically adjust the speech recognition mode according to the real-time performance indicators (such as misrecognition rate, response time, etc.) of the speech recognition device feedback by the user. For example, when the misrecognition rate of the feedback speech recognition device exceeds the preset threshold, the system can automatically prompt the user to switch to the manual recognition mode. In addition, the speech recognition result acquisition device can support the access of multiple speech recognition devices, such as different cloud service providers or local servers, to provide more choices for users.

[0062] In some embodiments, the display unit of the display device includes an operation object selection button for switching the operation object to the operator or the speech recognition device;

[0063] When the operator is selected, the operator generates a speech recognition result through manual input;

[0064] When the speech recognition device is selected, the speech recognition device generates a speech recognition result through automatic recognition.

[0065] It should be noted that the display unit of the display device is equipped with an operation object selection button in the present invention for switching the operation object to the operator or the speech recognition device. The operation object selection button is a key interface element for user-system interaction, allowing the user to flexibly switch between manual recognition and automatic recognition. When the user selects the operator, the operator converts the voice information into text by manual dictation; while when the user selects the speech recognition device, the speech recognition device completes this process through an automatic speech recognition algorithm. This design enables the user to choose the most suitable speech recognition method according to the call scenario and requirements, thereby improving the accuracy and efficiency of the call.

[0066] Specifically, the operation object selection button can be designed as an intuitive graphical user interface element, such as a toggle button with manual and automatic options. The user can select the desired speech recognition mode through touch, click, or other interaction methods. On the display unit of the display device, the position and size of the operation object selection button can be optimized according to the screen size of the device and the user's operation habits. For example, the button can be placed at the bottom or side of the screen for easy switching by the user during a call. In addition, to improve the user experience, the button can be designed to have obvious visual feedback, such as color change or animation effects, to prompt the user of the currently selected mode. At the same time, the system can set a default mode, such as defaulting to the speech recognition device mode when the device is started, to adapt to the usage habits of most users.

[0067] Preferably, the display unit of the display device can further optimize the function of the operation object selection button. For example, the button can support long-press operations. When the user long-presses the button, they can enter the settings menu to further adjust the speech recognition parameters, such as the sensitivity of the speech recognition device or the input speed of the operator. In addition, the system can automatically remember the user's preference settings according to the user's usage habits. For example, if the user frequently selects the manual recognition mode during multiple calls, the system can automatically recommend this mode in subsequent calls. As an alternative, the display device can also support switching the operation object by voice command. The user can complete the mode switch by voice command (such as switching to manual recognition), further enhancing the convenience of operation.

[0068] In some embodiments, the display unit uses a differentiated display method for the speech recognition results generated by the operator and the speech recognition results generated by the speech recognition device, including at least one of font, color, or background.

[0069] It should be noted that the display unit uses a differentiated display method for the speech recognition results generated by the operator and the speech recognition results generated by the speech recognition device. The purpose of this design is to enable the user to intuitively distinguish between the speech recognition results from two different sources. The differentiated display method can be achieved through visual elements such as font, color, or background, thereby helping the user quickly identify whether the currently displayed text is generated by the operator or the speech recognition device. For example, the text generated by the operator can be displayed in blue font, while the text generated by the speech recognition device can be displayed in black font, or distinguished by different background colors. This differentiated display method not only improves the usability of the system but also enhances the user's trust in the speech recognition results.

[0070] Specifically, differential display can be achieved through various visual elements. For example, in terms of fonts, different font sizes, styles (such as bold or italic), or font types (such as Song typeface, Kai typeface, etc.) can be set to distinguish between two speech recognition results. In terms of colors, a specific color, such as green or blue, can be set for the text generated by the operator, while another color, such as gray or black, can be set for the text generated by the speech recognition device. In terms of backgrounds, a light background can be set for the text generated by the operator, while a dark background can be set for the text generated by the speech recognition device, or different background patterns can be used for differentiation. In addition, the display unit can further enhance the differentiation effect through other visual effects such as blinking, underlining, or borders. The specific parameter settings can be adjusted according to user preferences and the display capabilities of the display device. For example, the font size can be optimized according to the screen size and the user's vision requirements, and the color selection can consider color contrast to ensure the readability of the text.

[0071] Preferably, the display unit can provide a user-customizable differential display setting function, allowing users to adjust parameters such as fonts, colors, and backgrounds according to their preferences. For example, users can select their favorite color combinations or font styles through the settings menu of the display device, and the system can provide multiple preset display modes for users to choose from, such as high-contrast mode or night mode. In addition, the display unit can automatically adjust the display effect according to the accuracy of the speech recognition results. For example, when the text generated by the speech recognition device is detected by the system as possibly being misrecognized, its font color can be automatically changed to red to alert the user. As an alternative, the display unit can support voice and text prompt functions. When it detects a change in the source of the speech recognition result, it can prompt the user through voice and text whether the currently displayed text is generated by the operator or the speech recognition device, further enhancing the user experience.

[0072] In some embodiments, when the speech recognition result acquisition device detects that the speech silence time exceeds a preset threshold, it pauses obtaining speech recognition results.

[0073] It should be noted that when the speech recognition result acquisition device detects that the speech silence time exceeds a preset threshold, it can pause obtaining speech recognition results. The speech silence time refers to the time period during a call when the speech signal is continuously interrupted or the volume is lower than a certain threshold. The preset threshold is a system-set time parameter used to determine when to pause the speech recognition process to avoid unnecessary speech recognition operations during call gaps or when there is no speech input, thereby saving system resources and improving efficiency. For example, when the user briefly leaves or pauses speaking during a call, the system can automatically pause speech recognition until the speech signal resumes.

[0074] Specifically, the detection of the voice silent time can be achieved through the audio processing module in the voice recognition result acquisition device. This module can monitor the intensity and duration of the voice signal in real time. When the volume of the voice signal is lower than the set threshold (such as -40 dB) and the duration exceeds the preset threshold (such as 5 seconds), the system determines it as the silent state. The preset threshold can be adjusted according to the actual usage scenario and user requirements. For example, in a noisy environment, the silent time threshold can be appropriately extended to avoid misjudgment. In addition, the voice recognition result acquisition device can also combine the Voice Activity Detection (VAD) algorithm to more accurately judge the start and end of the voice signal, so as to more accurately control the pause and resume of voice recognition.

[0075] Preferably, the voice recognition result acquisition device can further optimize the silent detection mechanism. For example, the system can dynamically adjust the silent time threshold according to the historical data of the call. If the user frequently has short silences during the call, the system can automatically shorten the silent time threshold to respond more quickly to the user's voice input; conversely, if the silent time is long and frequent, the system can appropriately extend the threshold to avoid frequent pauses and resumptions of voice recognition. In addition, after pausing the voice recognition, the system can also inform the user that the voice recognition has been paused through a short beep or text prompt. When the voice signal resumes, the user is prompted again that the voice recognition has been restarted, enhancing the user's interaction experience.

[0076] In some embodiments, the display device is directly connected to the Internet through wireless communication and performs data interaction with the voice recognition result acquisition device.

[0077] It should be noted that the display device can be directly connected to the Internet through wireless communication and perform data interaction with the voice recognition result acquisition device. This design allows the display device to complete the data transmission task independently of the user communication device, thereby improving the flexibility and response speed of the system. Wireless communication refers to the technology of transmitting data through radio waves. Common wireless communication methods include Wi-Fi, Bluetooth, 4G / 5G, etc. The Internet connection provides the display device with the ability to exchange data with external servers or other devices, ensuring that the voice recognition results can be transmitted and displayed quickly and accurately.

[0078] Specifically, the display device can be connected to the Internet through a built-in wireless communication module. For example, when using a Wi-Fi module, the display device can be connected to the wireless network at home or in public places; if a 4G / 5G module is used, the mobile network can be utilized for data transmission. In terms of parameter settings, the display device needs to support multiple wireless communication standards to adapt to different network environments. For example, the Wi-Fi module should support the common 2.4GHz and 5GHz frequency bands, while the 4G / 5G module needs to support the mainstream network frequency bands and systems. In addition, the display device also needs to have the function of automatically connecting and disconnecting from the network to optimize power consumption and the use of network resources. For example, when detecting the interruption of the connection between the user's communication device and the display device, the display device can automatically switch to an independent Internet connection mode to continue to complete the data interaction task.

[0079] Preferably, the display device can support the automatic switching of multiple wireless communication methods. For example, when the Wi-Fi signal is weak, the system can automatically switch to the 4G / 5G network to ensure the continuity of data transmission. In addition, the display device can also optimize the data transmission efficiency through intelligent network management algorithms. For example, according to the current network bandwidth and latency, dynamically adjust the priority and rate of data transmission. As an alternative, the display device can be equipped with a wired network interface, such as an Ethernet interface, so as to access the Internet through a wired connection when needed, further improving the stability and security of data transmission.

[0080] In some embodiments, the display unit also displays the speech recognition result of the voice information of the user's communication device and displays it in a separate area from the speech recognition result of the target communication device.

[0081] It should be noted that the display unit can not only display the speech recognition result of the voice information of the target communication device, but also display the speech recognition result of the voice information of the user's communication device and display the two in separate areas. This design enables the user to view the speech recognition results of both parties at the same time, so as to more comprehensively understand the conversation content. Separate area display means dividing different display areas on the screen of the display device to respectively display the speech recognition texts of the user and the other party, so that the user can clearly distinguish and compare the conversation contents of both parties.

[0082] Specifically, the split - area display can be achieved by setting two independent display areas on the screen of the display device. For example, the upper half of the screen can be used to display the speech recognition results of the user's communication device, while the lower half is used to display the speech recognition results of the target communication device. These two areas can be distinguished by different background colors, border styles, or font sizes. In terms of parameter settings, the size and position of the display area can be adjusted according to the screen size of the display device and user preferences. For example, for a larger screen, the display area can be set in a left - right layout for the user to view in landscape mode; for a smaller screen, a top - bottom layout can be adopted. In addition, the display device can also provide a scroll bar or paging function to handle longer speech recognition texts.

[0083] Preferably, the display device can provide multiple layout options for split - area display, and users can choose the most suitable layout according to their usage habits and screen size. For example, in addition to the top - bottom layout and left - right layout, a grid layout or a fold - out layout can also be provided to adapt to different usage scenarios. In addition, the display device can support real - time synchronous updates of speech recognition results. When the user or the other party speaks, the corresponding speech recognition results will be immediately displayed in the corresponding area. As an alternative, the display device can also provide a voice - broadcast function for speech recognition results, and users can obtain the speech content of the other party through voice prompts to further enhance the interaction experience.

[0084] In some embodiments, when the speech recognition result acquisition device detects that the speech recognition result contains a preset keyword, it triggers the display unit to display a warning message.

[0085] It should be noted that when the speech recognition result acquisition device detects that the speech recognition result contains a preset keyword, it can trigger the display unit to display a warning message. Here, the preset keyword refers to specific words or phrases predefined by the system, and these words may be related to fraud, harassment, or other content that requires user attention. For example, when keywords such as "transfer", "remittance", "winning a prize" that may be involved in fraud appear in the speech recognition result, the system will automatically identify and trigger a warning. This design aims to help users promptly identify potential risks during a call, thereby improving the security and reliability of the call.

[0086] Specifically, the speech recognition result collection device will monitor and analyze the received speech recognition results in real time. The preset keywords can be managed through the keyword library built into the system, and this keyword library can be customized and updated according to different application scenarios and user needs. For example, the system can set up a general fraud keyword library, and at the same time allow users to add custom keywords according to their personal needs. In terms of parameter settings, the system can set a threshold. For example, when the frequency of the preset keyword appears more than a certain number of times, a warning will be triggered. In addition, the display method of the warning information can include pop-up prompts, flashing warning signs, or special background colors, etc., to ensure that users can quickly notice potential risks.

[0087] Preferably, the speech recognition result collection device can further optimize the keyword detection mechanism. For example, the system can combine context semantic analysis to determine whether the keyword really involves risks. For example, when the word "transfer" appears in a normal financial business scenario, the system may not trigger a warning, but if it appears in a suspicious fraud context, it will immediately remind the user. In addition, the system can also support multi-language keyword detection to adapt to call requirements in different language environments. As an alternative, in addition to displaying warning information, the system can also use voice prompts or vibration to remind the user, further enhancing the intuitiveness of the interaction.

[0088] The above-mentioned various embodiments of the present invention have the following beneficial effects: The telephone caption system of the present invention can provide a convenient caption call function for hearing-impaired people. Through the voice call connection between the user communication device and the target communication device, combined with the display device receiving and sending voice information, the real-time conversion and display of voice information are realized. Users can choose manual or automatic speech recognition methods to generate captions as needed, and the system supports differential display to facilitate users to distinguish speech recognition results from different sources. In addition, when the system detects that the voice silent time exceeds the preset threshold, it can pause obtaining speech recognition results to avoid resource waste, and at the same time the display device can directly connect to the Internet for data interaction, improving the flexibility and response speed of the system.

[0089] Furthermore, the present invention can realize the regional display of the voice information of the user communication device and the target communication device, which is convenient for users to view the speech recognition results of both parties. When the speech recognition result contains a preset keyword, the system can trigger the display unit to display a warning message, effectively preventing users from being defrauded by telecommunications. Through these functions, the system can not only improve the call experience of hearing-impaired people, but also enhance the information security during the call, providing more reliable and efficient communication support for users.

[0090] As Figure 2 shown, an application method 200 of a telephone caption in some embodiments, the method 200 includes:

[0091] Step 201, establish a voice call through the user communication device and the target communication device, and obtain the voice information of the target communication device;

[0092] Step 202, transmit the voice information to the display device through wireless communication;

[0093] Step 203, send the voice information to the voice recognition result collection device through the display device;

[0094] Step 204, transmit the voice information to the operator or the voice recognition device through the voice recognition result collection device, and receive the generated voice recognition result;

[0095] Step 205, return the voice recognition result to the display device;

[0096] Step 206, display the voice recognition result in real time through the display unit of the display device.

[0097] It can be understood that the steps described in the application method 200 of the telephone caption correspond to the respective modules in the telephone caption system described in the reference Figure 1 Therefore, the modules, features, and beneficial effects described above for the telephone caption system are equally applicable to the application method 200 of the telephone caption and the operations included therein, and will not be elaborated herein.

[0098] Next, refer to Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present invention. The electronic devices in some embodiments of the present invention may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The terminal device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present invention.

[0099] As Figure 3As shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0100] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 3 Each block shown in may represent one device or, as required, multiple devices.

[0101] Furthermore, the storage medium of the embodiments of the present application stores program instructions capable of implementing all the above methods. Among them, the program instructions may be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, a tablet, etc.

[0102] The above description is only some preferred embodiments of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present invention.

Claims

1. A telephone captioning system, characterized in that, Comprising: A user communication device, configured to establish a voice call with a target communication device via a mobile network and obtain voice information of the target communication device; A display device, including a voice information receiving unit, a voice information sending unit, a voice recognition result receiving unit, and a display unit, wherein: The voice information receiving unit receives the voice information transmitted by the user communication device via wireless communication; The voice information sending unit sends the voice information to a voice recognition result collection device; The voice recognition result receiving unit receives the voice recognition result returned by the voice recognition result collection device; The display unit displays the voice recognition result in real time; A voice recognition result collection device, configured to transmit the voice information to an operator or a voice recognition device, and receive the voice recognition result generated by the operator or the voice recognition device; The voice recognition device is configured to convert the voice information into text data; Both the user communication device and the target communication device are configured with SIM cards, and a call connection is established through the phone numbers recorded in the SIM cards.

2. The telephone captioning system according to claim 1, wherein The display device and the user communication device are of a split structure or integrated within the housing of the user communication device.

3. The telephone captioning system according to claim 1, characterized in that, The voice recognition result collection device transmits the voice information to the operator or the voice recognition device according to the operation object selected by the user.

4. The telephone captioning system according to claim 3, wherein The display unit of the display device includes operation object selection buttons for switching the operation object to the operator or the voice recognition device; When the operator is selected, the operator generates a voice recognition result through manual input; When the voice recognition device is selected, the voice recognition device generates a voice recognition result through automatic recognition.

5. The telephone captioning system according to claim 1, characterized in that, The display unit adopts a differentiated display method for the voice recognition result generated by the operator and the voice recognition result generated by the voice recognition device, including at least one of font, color, or background.

6. The telephone captioning system according to claim 1, wherein The voice recognition result collection device suspends obtaining the voice recognition result when detecting that the voice silence time exceeds a preset threshold.

7. The telephone captioning system according to claim 1, wherein The display device is directly connected to the Internet via wireless communication and performs data interaction with the voice recognition result collection device.

8. The telephone captioning system according to claim 1, wherein The display unit also displays the voice recognition result of the voice information of the user communication device and displays it in a partitioned manner from the voice recognition result of the target communication device.

9. The telephone captioning system according to claim 1, wherein The voice recognition result collection device triggers the display unit to display a warning message when detecting that the voice recognition result contains a preset keyword.

10. A method for applying a telephone caption system, characterized in that, Including the following steps: Establish a voice call with a target communication device through a user communication device and obtain voice information of the target communication device; Transmit the voice information to the display device via wireless communication; Send the voice information to the voice recognition result collection device through the display device; Transmit the voice information to an operator or a voice recognition device through the voice recognition result collection device and receive the generated voice recognition result; Return the voice recognition result to the display device; The display unit of the display device displays the voice recognition result in real time.