Speech recognition device for digital media

By designing a multi-channel structure with filtering components and blocking components in the speech recognition device, and controlling the switching of the side channels according to the sound wave velocity, the problem of interference from interfering sound waves on speech recognition is solved, achieving high-precision speech signal processing and subsequent digital media support.

CN223450558UActive Publication Date: 2025-10-17SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202521974477.X
Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
2035-09-15

AI Technical Summary

Technical Problem

Existing speech recognition systems are subject to a large amount of interference from external sound waves, which increases recognition errors and makes it difficult to meet the high accuracy requirements of digital media processing.

Method used

Design a voice recognition device that includes a filter and has multiple channel structures and blocking components inside. The device controls the switching of the side channels by controlling the sound wave flow rate to eliminate high-speed interference sound waves and ensure effective voice sound wave transmission.

Benefits of technology

It effectively reduces the impact of interfering sound waves on recognition, improves the purity and recognition accuracy of voice signals, and supports the efficient implementation of subsequent digital media processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN223450558U_ABST
    Figure CN223450558U_ABST
Patent Text Reader

Abstract

The utility model discloses a voice recognition device for digital media, and relates to the technical field of voice recognition devices. The voice recognition module is mounted at the top end of the base, and a voice input port is formed in one end of the voice recognition module; the filtering piece is detachably installed on the peripheral side of the voice recognition module through a support and located on one side of the voice input port, and the filtering piece is provided with a receiving face and a discharging face which are oppositely arranged; intelligent switch control is achieved through the blocking piece according to the sound wave flow speed (sound pressure), the side channel is opened only when the sound wave flow speed is higher than a set threshold value (high-speed interference sound waves), and the sound waves are discharged out of the filtering piece; effective voice sound waves (low flow speed) can be normally transmitted along a main channel formed by the first straight channel, the arc-shaped channel and the second straight channel, high-speed interference is accurately eliminated, and the influence of interference signals on follow-up recognition is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The utility model relates to voice recognition device technical field, specifically be a voice recognition device for digital media. BACKGROUND

[0002] At the moment of the rapid development of digital media technology, voice recognition technology as the core link of man-machine interaction, digital signal processing has been widely used in intelligent terminal, media content production, remote control and other scenes. Its core demand is to accurately convert external voice signal into digital signal, to provide reliable data support for subsequent digital media content editing, storage and transmission, therefore the purity and integrity of voice signal directly determines the accuracy of voice recognition and the efficiency of subsequent digital media processing.

[0003] However, the existing voice recognition system in practical application faces many technical bottlenecks: one, there are a large number of high-speed interference sound waves (such as sudden noise, explosive sound, high-frequency mechanical noise, etc.) in the external environment, these interference sound waves and effective voice sound wave mixed propagation, directly into the voice recognition module, will seriously interfere with the extraction and analysis of effective signal by recognition algorithm, leading to increased recognition error, even appear recognition failure, it is difficult to meet the high requirements of signal accuracy of digital media processing. UTILITARY MODEL

[0004] In view of the deficiencies existing in the prior art, the utility model provides a voice recognition device for digital media.

[0005] In order to realize the above purpose, the technical scheme of the utility model is as follows:

[0006] A voice recognition device for digital media, comprising:

[0007] Base;

[0008] Voice recognition module, installed at the top of the base, and one end is provided with a voice input port;

[0009] Filter piece, through the support, can be detachably installed on the side of the voice recognition module and located on one side of the voice input port, the filter piece has oppositely arranged receiving surface and discharge surface;

[0010] Wherein, the filter piece is internally provided with a plurality of sequentially communicated first straight channel, arc channel and second straight channel, the first straight channel extends inward from the receiving surface, the second straight channel extends inward from the discharge surface and communicates with the convex side of the arc channel;

[0011] The side of the filter piece is also provided with a side channel communicated with the arc channel, for making the sound wave with flow rate higher than the set threshold value discharge through the side channel.

[0012] Preferably, a blocking piece is arranged in the connection between the arc-shaped channel and the side channel.

[0013] Preferably, the blocking piece is configured to rotate to open or close the entrance of the side channel under the action of sound pressure.

[0014] Preferably, the radius of curvature of the arc-shaped channel is 3-5 cm.

[0015] Preferably, the side channel is staggered with the first straight channel and the second straight channel.

[0016] Preferably, the diameter of the first straight channel is larger than that of the arc-shaped channel, and the diameter of the arc-shaped channel is smaller than that of the side channel.

[0017] Compared with the prior art, the present application has the following advantages:

[0018] 1. The blocking piece realizes intelligent switch control according to the sound wave flow rate (sound pressure), and only when the sound wave flow rate is higher than the set threshold (high-speed interference sound wave), the side channel is opened to discharge it outside the filter; effective voice sound waves (low flow rate) can be normally transmitted along the main channel composed of the first straight channel, the arc-shaped channel and the second straight channel, high-speed interference is accurately excluded, and the influence of interference signals on subsequent identification is greatly reduced.

[0019] 2. The radius of curvature of the arc-shaped channel is 3-5 cm, and the diameter is smaller than that of the first straight channel and the side channel, the motion track difference of sound waves of different frequencies and flow rates is used to make the effective voice sound waves (stable track) transmit along the convex side, and the high-speed interference sound waves (track offset) approach the side channel connection position, realizing "preliminary shunt and then exclusion", and further improving the accuracy of filtering. BRIEF DESCRIPTION OF DRAWINGS

[0020] The disclosure of the present application will be described with reference to the accompanying drawings. It should be understood that the drawings are only for illustrative purposes, and are not intended to limit the scope of protection of the present application. In the drawings, the same reference numerals are used to refer to the same parts. Among them:

[0021] Figure 1 It is a three-dimensional structure schematic view of the voice recognition device for digital media of the present application;

[0022] Figure 2 It is a three-dimensional structure schematic view of the filter of the voice recognition device for digital media of the present application;

[0023] Figure 3 It is a cross-sectional view of the filter of the voice recognition device for digital media of the present application;

[0024] Figure 4The utility model is used for the blocking piece setting structure schematic drawing of voice recognition device of digital media.

[0025] The figure mark explanation: 1, base; 2, voice recognition module; 3, filter piece; 31, first straight passage; 32, arc channel; 33, second straight passage; 34, side passage; 35, blocking piece. DETAILED DESCRIPTION

[0026] It is easy to understand that according to the technical scheme of the utility model, a plurality of structure modes and implementation modes that can be replaced with each other can be proposed by the general technical personnel in the art without changing the utility model essential spirit. Therefore, the following specific embodiment and drawing are only exemplary description to the technical scheme of the utility model, and should not be regarded as the whole of the utility model or regarded as the limitation or restriction of the technical scheme of the utility model.

[0027] EMBODIMENT

[0028] As Figures 1-4 shown, a voice recognition device for digital media comprises:

[0029] Base 1: base 1 is the basic support component of the whole device, provides stable mounting platform for other components of the device. Its material selects high-strength, corrosion-resistant material, ensures that the device remains stable during use, is not easy to produce displacement or sway due to external force, guarantees the relative position stability between various components, so as to not affect the normal progress of voice recognition.

[0030] Voice recognition module 2: voice recognition module 2 is installed at the top end of base 1, is the component for realizing the voice recognition function of the device. The module is provided with a voice input port at one end, for receiving the voice signal transmitted from the outside. It has advanced voice recognition algorithm, can quickly process and analyze the received voice signal, accurately recognize the voice content, and convert the recognition result into corresponding digital signal, to provide support for subsequent digital media processing work. At the same time, voice recognition module 2 also has good compatibility, can be connected with various digital media equipment and system and data interaction.

[0031] Filter piece 3: filter piece 3 is detachably installed on the side of voice recognition module 2 through a support, and is located on one side of the voice input port, mainly for filtering the sound wave entering the voice input port, removing the interference sound wave, and improving the quality of the voice signal. Filter piece 3 selects disc or multi-prism, and filter piece 3 has oppositely arranged receiving surface and discharge surface, the receiving surface is used for receiving external sound wave, and the discharge surface transmits the effective sound wave after filtering to the voice input port of voice recognition module 2.

[0032] Internal passage structure

[0033] The filter 3 is internally provided with a plurality of sequentially connected first straight channels 31, arc-shaped channels 32 and second straight channels 33, which work cooperatively to realize accurate screening of sound waves.

[0034] The first straight channel 31 extends inward from the receiving surface of the filter 3, and has a diameter larger than that of the arc-shaped channel 32. The main function of the channel is to preliminarily receive and guide external sound waves. Since the diameter is large, more sound waves can be accommodated, and at the same time, it also makes good preparation for the subsequent transmission of sound waves in the arc-shaped channel 32, reducing the energy loss of sound waves in the transmission process.

[0035] The arc-shaped channel 32 is connected with the first straight channel 31, and has a curvature radius of 3-5 cm. The arc-shaped channel 32 utilizes the propagation characteristics of sound waves. When the sound waves are transmitted in the arc-shaped channel 32, sound waves of different frequencies and flow rates will produce different motion trajectories and energy changes. Compared with straight channels, the arc-shaped channel can better split and screen sound waves, laying a foundation for subsequent exclusion of interfering sound waves. Moreover, the diameter of the arc-shaped channel 32 is smaller than that of the first straight channel 31 and the side channel 34. This size design helps to control the transmission speed and direction of sound waves, further improving the filtering effect.

[0036] The second straight channel 33 extends inward from the discharge surface of the filter 3 and is connected with the convex side of the arc-shaped channel 32. After being screened in the arc-shaped channel 32, the effective sound waves will enter the second straight channel 33, and then be transmitted to the voice input port of the voice recognition module 2 through the discharge surface. The arrangement of the second straight channel 33 ensures that the effective sound waves can be accurately and stably transmitted to the core recognition component, reducing interference and loss in the transmission process.

[0037] The side channel 34 is provided on the side of the filter 3 and is connected with the arc-shaped channel 32. The side channel 34 is staggered with the first straight channel 31 and the second straight channel 33. The main function of the channel is to discharge sound waves with a flow rate higher than a set threshold through the side channel 34, so as to exclude the influence of high-speed interfering sound waves on effective voice signals. Since the side channel 34 is staggered with the first straight channel 31 and the second straight channel 33, mutual interference between the channels during sound wave transmission can be avoided, ensuring the independence and stability of the functions of the channels. At the same time, the diameter of the side channel 34 is larger than that of the arc-shaped channel 32. This size design is conducive to quickly discharging high-speed interfering sound waves and improving the filtering efficiency.

[0038] Further, the blocking piece 35 is rotatably arranged inside the connection between the arc-shaped channel 32 and the side channel 34. The cross-sectional size of the blocking piece 35 is smaller than that of the side channel 34. The blocking piece 35 is configured to rotate to open or close the entrance of the side channel 34 under the action of sound pressure, and is a key component for controlling the working state of the side channel 34.

[0039] When the incoming sound wave flow rate is lower than the set threshold, the sound pressure generated by the sound wave is not enough to push the blocking piece 35 to rotate, and the blocking piece 35 remains in a state of closing the entrance of the side channel 34, at this time the sound wave will be transmitted along the path of the first straight channel 31, the arc-shaped channel 32 and the second straight channel 33 to the voice input port; when the sound wave flow rate is higher than the set threshold, the sound pressure generated by the sound wave will push the blocking piece 35 to rotate, open the entrance of the side channel 34, and the high-speed sound wave will be discharged out of the filter 3 through the side channel 34, thereby realizing effective filtering of high-speed interference sound waves.

[0040] Sound wave receiving and preliminary guidance: when the external voice signal (including effective voice sound wave and interference sound wave) propagates to the device, it first contacts the filter 3 which is detachably installed on the side of the voice recognition module 2. The filter 3 receives external sound waves through the receiving surface, at this time the sound waves enter the first straight channel 31 opened inside the filter 3. Since the first straight channel 31 extends inward from the receiving surface and has a diameter larger than the subsequent arc-shaped channel 32, it can efficiently accommodate and preliminarily guide the sound waves to transmit inward, while minimizing the energy loss of the sound waves in the initial transmission stage, thereby laying a good foundation for subsequent filtering and screening.

[0041] Sound wave shunt screening of arc-shaped channel: the sound waves entering the first straight channel 31 then flow into the arc-shaped channel 32 (curvature radius 3-5 cm, diameter smaller than the first straight channel 31 and the side channel 34) connected thereto. By utilizing the propagation characteristics of sound waves, sound waves of different frequencies and flow rates will form different motion trajectories and energy changes in the arc-shaped channel 32: effective voice sound waves (flow rate usually lower than interference sound waves) can continue to transmit along the convex side of the arc-shaped channel 32 due to stable motion trajectory; while part of the interference sound waves (especially high-speed interference sound waves) gradually approach the connection position of the arc-shaped channel 32 and the side channel 34 due to the deviation of the motion trajectory, thereby preparing for subsequent targeted exclusion.

[0042] Intelligent switch control of blocking piece and high-speed interference discharge: the blocking piece 35 (cross-sectional size smaller than the side channel 34) inside the connection between the arc-shaped channel 32 and the side channel 34 realizes intelligent control according to the sound pressure corresponding to the sound wave flow rate:

[0043] When the sound wave flow rate is lower than the set threshold (i.e. effective voice sound wave), the sound pressure generated by the sound wave is not enough to push the blocking piece 35 to rotate, and the blocking piece 35 remains in a state of closing the entrance of the side channel 34, ensuring that the effective sound wave will not be shunted to the side channel, but will continue to transmit along the arc-shaped channel 32 to the second straight channel 33;

[0044] When the flow rate of the sound wave is higher than the set threshold (i.e. high-speed interference sound wave), the sound pressure generated thereby can push the blocking piece 35 to rotate and open the side passage 34 entrance. Since the side passage 34 has a larger diameter than the arc-shaped passage 32 and is staggered with the first straight passage 31 and the second straight passage 33 (to avoid mutual interference), the high-speed interference sound wave can be quickly discharged out of the filter 3 through the side passage 34, thereby achieving accurate filtering of the high-speed interference signal.

[0045] Effective sound wave transmission and voice recognition: After being screened by the arc-shaped passage 32 and interference excluded by the side passage 34, the remaining effective voice sound wave enters the second straight passage 33 connected to the convex side of the arc-shaped passage 32. The second straight passage 33 extends inward from the discharge surface of the filter 3 and can stably guide the effective sound wave to pass through the discharge surface and accurately transmit to the voice input port of the voice recognition module 2.

[0046] After receiving the effective voice signal, the voice recognition module 2 processes and analyzes the signal quickly through the built-in advanced voice recognition algorithm, accurately identifies the voice content, and converts the recognition result into a digital signal that meets the digital media processing requirements. At the same time, thanks to its good compatibility, the voice recognition module 2 can connect and interact with external digital media devices or systems, providing core support for subsequent digital media content editing, storage, transmission, etc., and ultimately completing the entire voice recognition and signal conversion process.

[0047] Efficient filtering of high-speed interference sound waves and improvement of voice signal purity: The blocking piece 35 realizes intelligent switch control according to the flow rate (sound pressure) of the sound wave, and only when the flow rate is higher than the set threshold (high-speed interference sound wave) does it open the side passage 34 to discharge it out of the filter 3; effective voice sound waves (low flow rate) can be normally transmitted along the main passage composed of the first straight passage 31, the arc-shaped passage 32, and the second straight passage 33, accurately excluding high-speed interference and greatly reducing the influence of interference signals on subsequent recognition.

[0048] Reducing the initial transmission energy loss of sound waves and ensuring signal integrity: The first straight passage 31 of the filter 3 is designed to "extend inward from the receiving surface and have a larger diameter than the subsequent arc-shaped passage 32", which can efficiently accommodate external sound waves and preliminarily guide transmission, thereby minimizing energy loss in the initial stage and laying a foundation for subsequent filtering, screening, and effective signal transmission.

[0049] Utilizing the characteristics of sound waves to achieve accurate shunting and improve the targeting of filtering: The arc-shaped passage 32 has a curvature radius of 3-5 cm and a diameter smaller than that of the first straight passage 31 and the side passage 34. By taking advantage of the differences in the motion trajectories of sound waves of different frequencies and flow rates, effective voice sound waves (stable trajectory) are transmitted along the convex side, and high-speed interference sound waves (trajectory deviation) are attracted to the connection position of the side passage 34, thereby achieving "shunting first and then excluding", and further improving the accuracy of filtering.

[0050] Reasonable layout of channel structure to avoid signal interference: the diameter of the side channel 34 is larger than that of the arc-shaped channel 32, and is staggered with the first straight channel 31 and the second straight channel 33, which not only ensures that the high-speed interference sound waves can be quickly discharged, but also effectively avoids mutual interference between the side channel 34 and the sound waves in the main channel (the first straight channel 31 and the second straight channel 33), thereby ensuring the stability of effective signal transmission.

[0051] Support for subsequent digital media processing and strong compatibility: after the voice recognition module 2 receives effective sound waves, it can quickly process and convert them into digital signals through advanced algorithms, and can connect with external digital media devices or systems to realize data interaction, thereby providing core support for subsequent work such as editing, storage, and transmission of digital media content, and adapting to various digital media processing scenarios.

[0052] Filtering element can be installed and removed for easy maintenance and replacement: the filtering element 3 is installed on the side of the voice recognition module 2 in a detachable manner. When the filtering element 3 is worn out or needs to be adapted to different use environments, it can be easily removed, maintained, or replaced, thereby reducing the overall maintenance cost and operation difficulty of the device.

[0053] One specific implementation parameter setting:

[0054] 1. Arc-shaped channel parameters: precise control of "centrifugal force threshold" to only divert strong blast airflow

[0055] The core is to use "radius of curvature" and "channel cross-sectional area" to divert only "strong blast airflow with flow rate > 3 m / s", and weak airflow and normal sound waves propagate along the straight channel:

[0056] Radius of curvature: designed as a gentle arc (about 1 / 4 of a circle) with R = 3-5 cm:

[0057] Too small curvature (R < 3 cm): too sharp a bend, too large centrifugal force, easy to mis-divert weak airflow;

[0058] Too large curvature (R > 5 cm): too slow bend, insufficient centrifugal force, strong blast airflow may "break through" the bend and still propagate along the straight channel;

[0059] Channel cross-sectional area: follows the "gradual scaling" principle of "straight channel → arc segment → lateral channel":

[0060] Inlet straight channel cross-sectional area (S1): slightly larger than the cross-sectional area of the microphone pickup head (e.g., microphone pickup head diameter 1 cm, S1 = 1.2 cm²), to ensure that the sound waves are not blocked;

[0061] Arc segment cross-sectional area (S2): S2 = 0.8 × S1, slightly shrink the airflow to increase the flow rate of the airflow in the arc segment (the faster the flow rate, the greater the centrifugal force, and the more complete the diversion);

[0062] Lateral passage cross-sectional area (S3): S3 = 1.5 x S1, a large enough outlet allows the strong airflow after splitting to be quickly discharged, avoiding the formation of "vortex" in the passage (vortex will produce additional low-frequency noise).

[0063] 2. Passage inner wall treatment: eliminate sound wave reflection, avoid high frequency distortion

[0064] Material selection: prefer "matte ABS plastic + inner wall flocking":

[0065] Matte material itself has low reflectivity, and flocking can further absorb high-frequency sound waves (especially above 10 kHz), avoiding reflection-induced comb filtering;

[0066] Avoid using metal or shiny plastic: metal reflectivity > 80%, shiny plastic reflectivity > 60%, prone to high-frequency noise;

[0067] Intersection treatment: the intersection of straight passages and arc segments, arc segments and lateral passages must be "R = 0.5 cm fillet transition":

[0068] Eliminate sharp corners to avoid low-frequency sound waves diffracting into the lateral passage (fillet allows sound waves to propagate more smoothly along the straight passage, reducing diffraction loss);

[0069] At the same time, prevent airflow from "hitting corners and creating turbulence" (turbulence will be converted into low-frequency noise, which will affect sound quality).

[0070] 3. Straight passage outlet design: ensure that sound waves "directly reach the diaphragm without phase difference"

[0071] Outlet position: the straight passage outlet needs to "directly face the center of the microphone pickup head", and the distance to the diaphragm is only 2-3 cm:

[0072] Too far (> 5 cm): sound wave propagation path becomes longer, which may cause a slight phase difference with the small amount of environmental sound leaked from the lateral passage (although it is difficult for the human ear to perceive, it will slightly affect the "cohesion" of the sound);

[0073] Too close (< 2 cm): the passage outlet may block the diaphragm, causing the sound wave incidence angle to deviate (affecting the microphone pickup polarity, such as the sensitivity of the heart-shaped microphone decreasing on the front side);

[0074] Outlet shape: designed as "slight horn mouth" (the outlet diameter is 10% larger than the inlet):

[0075] Expand the sound wave receiving range to avoid the straight passage outlet edge "blocking" part of the dispersed sound waves (especially mid-frequency sound waves, 200-3 kHz, affecting the clarity of human voice).

[0076] The technical scope of the utility model is not only limited to the content in the above description, and the person skilled in the art can make various deformations and modifications to the above embodiment without departing from the technical thought of the utility model, and these deformations and modifications should all belong to the protection scope of the utility model.

Claims

1. A speech recognition device for digital media, characterized in that: include: Base (1); A voice recognition module (2) is mounted on the top of the base (1), and a voice input port is provided at one end thereof; A filter element (3) is detachably mounted on the peripheral side of the voice recognition module (2) via a bracket and is located on one side of the voice input port, the filter element (3) having a receiving surface and a discharge surface that are arranged opposite to each other; The filter element (3) is provided with a plurality of first straight channels (31), curved channels (32) and second straight channels (33) that are sequentially connected. The first straight channels (31) extend inward from the receiving surface, and the second straight channels (33) extend inward from the discharge surface and are connected to the convex side of the curved channels (32). A side channel (34) communicating with the arc-shaped channel (32) is also provided on the peripheral side of the filter element (3) for discharging sound waves with a flow rate higher than a set threshold through the side channel (34).

2. The speech recognition device for digital media according to claim 1, wherein: A blocking member (35) is rotatably provided inside the connection between the arc-shaped channel (32) and the side channel (34), and the cross-sectional dimension of the blocking member (35) is smaller than the cross-sectional dimension of the side channel (34).

3. The speech recognition device for digital media according to claim 2, wherein: The blocking member (35) is configured to be capable of rotating under the action of sound pressure to open or close the entrance of the side channel (34).

4. The speech recognition device for digital media according to claim 3, wherein: The curvature radius of the arc-shaped channel (32) is 3-5 cm.

5. The speech recognition device for digital media according to claim 4, characterized in that: The side channel (34) is arranged in an alternating manner with the first straight channel (31) and the second straight channel (33).

6. The speech recognition device for digital media according to claim 5, characterized in that: The diameter of the first straight channel (31) is larger than the diameter of the arc-shaped channel (32), and the diameter of the arc-shaped channel (32) is smaller than the diameter of the side channel (34).