Instant Messaging Speech Recognition via Environment Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In scenarios where it is inconvenient to play speech input, such as noisy environments or lack of playback devices, recipients face challenges in directly acquiring speech content during instant communication.
Innovation Solution
A speech recognition method that receives speech information, assesses the environment, and converts it to text if playback is not necessary, using either a cloud-based or built-in recognition module, allowing users to acquire speech content as text for easier handling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech information is directly played for the receiver, then the receiver can hear the speech content, but the receiver cannot access the speech content in scenarios without playback devices or in noisy environments
Solution Approach 1:
The patent introduces text information as an intermediary form between speech information and the receiver. When speech information is transmitted, it is converted to text information that can be displayed on the receiver's screen, allowing the receiver to access the content without requiring audio playback capabilities or being able to hear the speech directly.
Solution Approach 2:
The patent changes the parameter of information representation from acoustic (speech) to visual (text). By converting speech information into text information, the system adapts to different reception scenarios where audio playback may not be feasible, such as when the receiver has no playback device or is in a noisy environment.
2Adaptability or versatility
If speech information is converted to text information, then the receiver can access speech content in various scenarios, but the system complexity increases due to environment assessment and conversion processes
Solution Approach 1:
The patent performs environment assessment before transmitting speech information. The system evaluates factors such as noise level, receiver's playback capabilities, and communication scenario in advance, and pre-determines whether to convert the speech to text based on this preliminary assessment, avoiding unnecessary conversion processes.
Solution Approach 2:
The patent implements a dynamic decision-making mechanism that adjusts the information transmission method based on real-time environment assessment. The system can switch between direct speech playback and text conversion modes depending on the current conditions, making the complexity manageable through adaptive rather than fixed processing.
3Productivity
If speech information is transmitted as-is, then the transmission process is simple, but the receiver cannot process the speech content in noisy environments or without playback devices
Solution Approach 1:
The patent uses text information as an intermediary that bridges the gap between speech transmission and receiver accessibility. The text representation serves as a universal medium that can be processed regardless of the receiver's audio capabilities or environmental conditions, significantly improving content accessibility.
Solution Approach 2:
The patent replaces the mechanical/audio-based speech playback system with a text-based information processing system. Instead of relying on the receiver's ability to hear and process audio signals, the system converts speech to text that can be read and processed through visual interfaces, working around the limitations of audio-based communication.
Data Source
AI summary
The present disclosure discloses a speech recognition method and a terminal, which belong to the field of communications. The method comprises: receiving speech information inputted by a user; acquiring the current environment information, and judging whether the speech information needs to be played according to the current environment information; and recognizing the speech information as text information, when it is judged that the speech information needs not to be played. The terminal comprises an acquisition module, a judgment module and a recognition module. The present disclosure provides the speech receiver with a speech recognition function, when the speech information of the instant messaging is received by the terminal, it can help the receiver to normally acquire the content to be expressed by the speech sender under an inconvenient situation.


