Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

20 results about "Voice frequency" patented technology

A voice frequency (VF) or voice band is one of the frequencies, within part of the audio range, that is being used for the transmission of speech. In telephony, the usable voice frequency band ranges from approximately 300 Hz to 3400 Hz. It is for this reason that the ultra low frequency band of the electromagnetic spectrum between 300 and 3000 Hz is also referred to as voice frequency, being the electromagnetic energy that represents acoustic energy at baseband. The bandwidth allocated for a single voice-frequency transmission channel is usually 4 kHz, including guard bands, allowing a sampling rate of 8 kHz to be used as the basis of the pulse code modulation system used for the digital PSTN. Per the Nyquist–Shannon sampling theorem, the sampling frequency (8 kHz) must be at least twice the highest component of the voice frequency via appropriate filtering prior to sampling at discrete times (4 kHz) for effective reconstruction of the voice signal.

Intelligent physique identification system based on lingual surface infrared thermal imaging and traditional Chinese medicine syndrome differentiation and application method

The invention provides an intelligent constitution identification system based on lingual surface infrared thermal imaging and traditional Chinese medicine syndrome differentiation and an application method. According to the system, an infrared thermal image partition acquisition module, a lingual surface image normalization acquisition module and a voice feature analysis module are fused, and an intelligent database integrating traditional Chinese medicine syndrome differentiation knowledge and cases is constructed. Through a deep learning algorithm, multi-modal features such as a human body infrared thermal structure, lingual texture color, voice frequency tone and the like are extracted, an improved cross-modal attention mechanism is applied to realize data deep fusion, a dialectical analysis model is established in combination with a traditional Chinese medicine viscera meridian theory, and dynamic tracking and health risk prediction are performed on physique change by using an LSTM network. The application method covers multi-modal data synchronous acquisition, joint feature extraction, fusion dialectical analysis, dynamic health assessment and other processes, can generate a personalized diagnosis and treatment scheme including medicine and food compatibility, meridian guidance and emotion regulation, and supports doctor-patient interaction optimization.
Owner:陆丁鹏

Radio analog telephone communication encryption method, device, equipment and medium

The invention discloses a radio analog telephone communication encryption method, device and equipment and a medium, and the method comprises the steps: generating a frequency spectrum segment number and a pseudo-random sorting sequence through a preset encryption algorithm; and carrying out controllable segmentation on the total spectrum of the analog voice according to a continuous sequence from low frequency to high frequency, carrying out sequence recombination on each sub-spectrum according to a pseudo-random sequence, finally combining and outputting encrypted voice, and recovering the original voice by utilizing the same pseudo-random sequence reverse recombination at a receiving end. Therefore, the time-frequency continuity and statistical characteristics of the voice spectrum in the traditional radio analog communication are fundamentally changed. According to the method, frequency spectrums are recombined as a whole according to a pseudo-random rule, so that an eavesdropper can only obtain disordered voice signals after the frequency spectrums are rearranged even if the eavesdropper intercepts the signals, and the difficulty of passive cracking means such as frequency spectrum analysis, filtering reduction and voice recognition is remarkably increased; the problem that in the prior art, radio analog telephone communication is insufficient in encryption and easy to crack is solved.
Owner:SHENZHEN TIANHAI COMM CO LTD

Voice false wake-up processing method and device, equipment and storage medium

The application provides a voice false wake-up processing method, device and equipment and a storage medium, and relates to the technical field of smart home. The method comprises the following steps: determining the response state of each smart device in a target space receiving a voice signal; determining a pre-wake-up signal of a smart device identified as a pre-wake-up state and a non-wake-up device identified as a non-wake-up state; determining that the voice signal in the target space is a valid wake-up voice based on the pre-wake-up signal and the non-wake-up device; and waking up a target smart device in the target space based on the valid wake-up voice. The application solves the defects of low voice recognition accuracy and high false wake-up voice frequency in the prior art, reduces the dependence on wake-up audio training in a specific space environment, and improves the voice recognition accuracy by determining the valid wake-up voice.
Owner:QINGDAO HAIER TECH +3

Acoustic biomimetic voiceprint recognition method and device, electronic equipment and storage medium

The application provides an auditory bionic voiceprint recognition method and device, electronic equipment and a storage medium. The auditory bionic voiceprint recognition method comprises the following steps: obtaining a to-be-recognized voice signal, and performing cutting processing on the to-be-recognized voice signal to obtain a plurality of to-be-recognized voice segments; inputting each to-be-recognized voice segment into a cochlea bionic filter to obtain a voice frequency feature; inputting the voice frequency feature into a pre-trained voiceprint recognition model to obtain a voiceprint recognition result; and the voiceprint recognition model is obtained based on a cochlea nucleus bionic network. The cochlea bionic filter and the voiceprint recognition model provided by the application have higher adaptability to noise and complex environments, thereby improving the reliability and efficiency of voiceprint recognition.
Owner:PEKING UNIV

Headphone speech listening

Microphone signals of a primary headphone are processed and either a first transparency mode of operation is activated or a second transparency mode of operation. In another aspect, a processor enters different configurations in response to estimated ambient acoustic noise being lower or higher than a threshold, wherein in a first configuration a transparency audio signal is adapted via target voice and wearer voice processing (TVWVP) of a microphone signal to boost detected speech frequencies in the transparency audio signal, and in a second configuration the TVWVP is controlled to, as the estimated ambient acoustic noise increases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal. Other aspects are also described and claimed.
Owner:APPLE INC

Mechanically adjustable acoustic chamber microphone

The application discloses a mechanical adjustable sound cavity microphone, which comprises a shell, a microphone assembly arranged on the inner side of the sound cavity hole at the bottom of the shell, a microphone, a sound cavity, a microphone position adjusting assembly, a supporting rod, a limiting buckle, a supporting cup, a rack, a driven gear, an adjusting gear, an adjusting wheel and a rack rail bottom plate. The position of the microphone in the sound cavity is adjusted through mechanical gear cooperation, the sound cavity can be adjusted and optimized according to the voice frequency characteristics of the user, the sound cavity frequency response characteristics of specific users are optimized, the user adjusts the microphone sound cavity structure while reading a piece of sample text according to the prompt, and the structure is solidified when the best frequency response characteristics are obtained, so that the recognition and interaction effect of the voice is greatly improved. The sound cavity can also be automatically adjusted through the adjustment of the driving device, the adjustment of the driving device is controlled through software and a central control system, the sound cavity can be intelligently and automatically adjusted, and the voice recognition of specific users can be adapted.
Owner:SHENZHEN VALLEY VENTURES

Audio compensation method in intercom system

The invention belongs to the field of building intercom, and particularly relates to an audio compensation method in an intercom system, which comprises the following steps of: S1, filling two empty mute frames; s2, setting a statistical flag bit, wherein the statistical flag bit is used for entering S4 to record the starting time when the voice frequency just establishes a call; s3, judging whether the statistical mark is set or not, if the statistical mark is set, entering S4, and if the statistical mark is not set, entering S5; s4, recording an initial millisecond number S, wherein the initial millisecond number S is used for judging whether statistics is cut off or not in S7; s5, the length L of the pcm audio data is obtained and used for S12, the actual playing amount G is calculated; and S6, recording an ending millisecond number E. According to the method, a statistical model is established and used for discovering and positioning audio trembling, and after a problem is positioned, sound is compensated through a mute data feeding method so as to eliminate the problem of sound trembling.
Owner:XIAMEN DNAKE INTELLIGENT TECH CO LTD

Signal processing method for enhancing human voice, electronic device, and storage medium

The application provides a signal processing method for enhancing human voice, an electronic device and a storage medium, and relates to the technical field of signal processing. In the scene of playing audio and video by the electronic device, the human voice gain parameter can be determined according to the environmental noise and the background voice frequency domain signal and the human voice frequency domain signal corresponding to the multi-channel audio stream, then the human voice frequency domain signal is gain adjusted according to the human voice gain parameter and the normalized spectral flux, and the enhanced human voice frequency domain signal is obtained. Since the gain of the human voice signal can be dynamically adjusted for the input multi-channel audio signal under different environmental noises, the effect of enhancing the human voice is achieved, and the phenomenon of large and small hearing is avoided, so the application scheme can achieve the effect that the video voice can be heard clearly without disturbing others when playing the video in a quiet environment, or the user can hear the video voice clearly when playing the video in a noisy environment, and the user experience can be improved.
Owner:HONOR DEVICE CO LTD

Headphone speech listening based on ambient noise

Microphone signals of a primary headphone are processed and either a first transparency mode of operation is activated or a second transparency mode of operation. In another aspect, a processor enters different configurations in response to estimated ambient acoustic noise being lower or higher than a threshold, wherein in a first configuration a transparency audio signal is adapted via target voice and wearer voice processing (TVWVP) of a microphone signal to boost detected speech frequencies in the transparency audio signal, and in a second configuration the TVWVP is controlled to, as the estimated ambient acoustic noise increases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal. Other aspects are also described and claimed.
Owner:APPLE INC

Dual-band high-directivity topological acoustic wave receiving antenna

The application belongs to the technical field of acoustic antennas, and particularly relates to a dual-band high-directivity topological acoustic receiving antenna. The left side surface, the upper side surface, the lower side surface and the rear end surface are mutually spliced to form a single-port acoustic waveguide, so that the right side of the single-port acoustic waveguide is an open port, the single-port acoustic waveguide is a parallelogram structure, and a topological phononic crystal and a sound-absorbing sponge are arranged in the single-port acoustic waveguide. Through the structure, high-directivity reception of two different voice frequency band acoustic waves can be realized, and an anti-interference high-security acoustic communication function is realized. Meanwhile, the acoustic receiving antenna has the characteristics of dual working frequency bands, small size, light weight and convenience in carrying, has high-directivity receiving effect on two independent wide frequency band acoustic signals, provides a feasible solution for voice frequency band acoustic signal directivity anti-interference and security transmission, and can be applied to the field of artificial intelligence robots.
Owner:NANJING NANDA ELECTRONIC INTELLIGENT SERVICE ROBOT RES INST CO LTD +1

Child intelligent development detection bracelet based on multi-mode mouth shape analysis

The invention discloses an intelligent child development detection bracelet based on multi-modal mouth shape analysis. The bracelet comprises a wearable bracelet main body, the multi-mode sensor module comprises a visual sensor used for capturing mouth shape dynamic images and an audio sensor used for collecting voice; the data processing unit is used for preprocessing sensor data and executing multi-mode mouth shape analysis; the intelligent development evaluation module is used for generating a development score and a risk index according to the analysis result; the user interface module is used for displaying information and providing feedback; according to the invention, a high-resolution visual sensor and a multi-microphone array are integrated on the bracelet, so that multi-modal data collaborative acquisition of dynamic images and synchronous voices of mouth shapes of children is realized, and fusion analysis is carried out by adopting a model based on a convolutional neural network and a recurrent neural network. The time sequence relevance among the lip motion trail, the mouth opening amplitude and the voice frequency spectrum feature can be extracted, and therefore the problem of one-sidedness of single-mode monitoring is solved.
Owner:CHANGZHOU TEXTILE GARMENT INST

Control method and control equipment for preventing mistaken door opening of motor train unit

The invention discloses a control method and control equipment for preventing mistaken door opening of a motor train unit, relates to the technical field of high-speed rail traffic safety detection and is used for preventing a driver from opening a train door mistakenly. In the method, the control equipment firstly performs rapid and low-power-consumption preliminary screening by utilizing rhythm characteristics of broadcast prompt tones to determine a pre-selected opening side, so that simultaneous physical detection on two sides is avoided. Then, aiming at the preselected opening side, through emitting a modulation infrared pulse and analyzing a mean value and a variance of a reflection signal, the existence of a static entity platform is physically confirmed from two dimensions of whether an object exists and whether the object is static, and under the condition that the static entity platform exists on the preselected opening side, the preselected opening side is opened; and the control equipment outputs an unlocking pulse signal to the door opening button protection device on the pre-selected opening side, so that a door opening mistaken accident caused by voice frequency misrecognition or driver fatigue is prevented, and the safety of passengers getting on and off the bus is improved.
Owner:CHENGDU HOUYOU TECHNOLOGY CO LTD

Sea wave height measuring device and measuring method using wireless data transmission

The application discloses a kind of wireless data transmission formula sea wave height measuring device and measuring method, mainly solve the problem of high cost of existing sea wave height measuring equipment.The device includes float component and measuring component, two are fixed as a whole, float component includes liquid level sensor circuit, electromagnetic valve, high-pressure carbon dioxide cabin and air bag air bag, liquid level sensor contacts seawater after electromagnetic valve is opened, carbon dioxide gas rushes into air bag, so that it floats on sea surface;Measuring component includes acceleration sensor, data processor, voice conversion chip, ultrashort wave transmitter, power amplifier circuit, antenna and power supply.Measuring component keeps the state of stable motion under float component, and calculates sea wave height by collecting sea wave acceleration data, and converts it into voice frequency modulation signal and emits to the air above sea surface, for pilot to listen to the voice broadcast of sea wave height.The device is small in size, light in weight, low in measurement cost, and can be used to ensure the safe take-off and landing of aircraft on water surface.
Owner:SHAANXI CHANGLING ELECTRONICS TECH

Air traffic control command end-to-end speech recognition method under high noise condition

The application relates to the field of air traffic control and speech recognition technology, and particularly relates to an air traffic control instruction end-to-end speech recognition method under high noise conditions, which comprises the following steps: pre-processing to-be-recognized speech to extract original speech frequency features; adopting an adaptive attention noise reduction module to perform noise reduction intensity control, frequency band gain control and pitch tracking processing on the original speech frequency features to obtain noise reduction speech; pre-processing the noise reduction speech to extract noise reduction speech frequency features; adopting a speech recognition module to perform shared encoder coding, connection time sequence classification decoder decoding and attention-based encoder-decoder decoding on the noise reduction speech frequency features to obtain output speech recognition transcription text; and the application can improve the speech recognition accuracy of air traffic control instructions under high noise conditions.
Owner:BEIHANG UNIV

Adaptive noise estimation

In some embodiments, a method includes: dividing an audio input into speech segments and non-speech segments using at least one processor; estimating a time-varying noise spectrum of the non-speech segment for each frame in each non-speech segment using at least one processor; estimating a speech spectrum of the speech segment for each frame in each speech segment using at least one processor; identifying one or more non-speech frequency components in the speech spectrum for each frame in each speech segment; comparing the one or more non-speech frequency components with one or more corresponding frequency components in a plurality of estimated noise spectra; and selecting an estimated noise spectrum from the plurality of estimated noise spectra based on the comparison results.
Owner:DOLBY LABORATORIES LICENSING CORP +1

System for processing text, image, and audio signals using artificial intelligence and method thereof

A system (100, 200) for processing at least concurrent audio signals (AS) and image signals (IS) to generate corresponding analytic data comprising sentiment measurements is disclosed. The system (100, 200) comprises a computing device (102) comprising at least an audio processing module (APM, 104) and an image processing module (IPM, 106) for processing the AS and IS. Each module (104, 106) uses an artificial intelligence (Al) algorithm. The IPM (106) is configured to process facial image information present in the IS to identify a plurality of critical facial image points indicative of facial expressions and generate temporal facial state data (TFSD), and to identify a plurality of critical body image points indicative of body languages and generate temporal body language state data (TBLSD). The APM (104) is configured to process speech present in the AS by parsing the speech to associate with a database of words (108) to generate text data (TD) and processing the speech to determine temporal speech frequency information (TSFI) to temporally associate the TSFI with the TD. The computing device (102) further includes an analysis module (110) that uses an AI algorithm to process the TFSD, the TBLSD, the TD, and the TSFI with an emotion model to generate interpretations of the AS and IS, thereby generating analysis data including emotion measurements.
Owner:KAI CONVERSATIONS LTD

Earphone

To provide an earphone adopting a sound transmission system via a cartilage with excellent articulation without deteriorating a voice frequency band by improving the transmission efficiency of vibration in a state that an ear hole is open so as to provide a sufficient sound pressure.SOLUTION: An earphone (1) is provided with an earphone body (2) including a vibrating unit having a vibrator (23) for generating vibration, a vibrating plate (24) for holding the vibrator and vibrating due to the vibration generated by the vibrator, and a vibration transmission section (25) provided to the vibrating plate, a case (22) for containing at least the vibrating unit, and a vibration-proof material (51) for supporting the vibrating unit and preventing the vibration generated by the vibrator from being transmitted to the case, and a vibrating body (31) having a tip section (33) extending in one axial direction (- Z-axis direction), an inclined surface (35) connected to the base end of the tip section and inclined with respect to the one axial direction, an annular base section (34) expanding in the directions (X-axis direction and Y-axis direction) orthogonal to the one axial direction, and a vibration receiving section (32) to which the vibration is transmitted from the vibration transmission section and which is in contact with the ear cartilage and transmits the vibration to the contact portion of the ear cartilage.SELECTED DRAWING: Figure 4
Owner:MEI CO LTD(JP)

SNN unsupervised multi-modal classification method based on neural coding specificity

The invention discloses an SNN unsupervised multi-modal classification method based on neural coding specificity, and belongs to the field of artificial intelligence and brain-like computing. The method comprises the following steps: generating an image pulse sequence and a voice pulse sequence of each sample, respectively inputting the image pulse sequence and the voice pulse sequence into a pulse neural network to obtain an image pulse response sequence and a voice pulse response sequence, and then extracting an image time coding feature, an image frequency coding feature, a voice time coding feature and a voice frequency coding feature. Then, based on the coding features, calculating an image specificity index and a voice specificity index of the sample, and according to the two indexes, carrying out weighted fusion on an image time coding feature, an image frequency coding feature, a voice time coding feature and a voice frequency coding feature to obtain a unified fusion feature vector of the sample; and finally, performing unsupervised classification on all samples by using an improved density peak clustering algorithm according to the unified fusion feature vector, the image specificity index and the voice specificity index.
Owner:HEBEI UNIV OF TECH

Voice interaction method and device, electronic equipment and computer program product

PendingCN122637777AEngineeringLoudspeaker
The application discloses a voice interaction method and device, electronic equipment and computer program product, and relates to the technical field of communication. The method comprises the following steps: detecting a voice frequency band of a user; determining a target frequency band of a loudspeaker based on the voice frequency band and a working frequency band of the loudspeaker, wherein the target frequency band is within the range of the working frequency band and does not overlap with the voice frequency band; controlling the loudspeaker to play a first voice signal according to the target frequency band, and in response to a second voice signal collected by a microphone, extracting a third voice signal corresponding to the voice frequency band from the second voice signal, and outputting the third voice signal as a user voice signal, wherein the microphone and the loudspeaker are arranged in the same sound field area. Through the application, the problem of low echo cancellation rate of AEC in a full-duplex voice interaction system is solved, and the technical effect of improving the echo cancellation rate in the full-duplex voice interaction process is achieved.
Owner:GOERTEK INC

Customer service auxiliary decision-making system based on knowledge base

The invention discloses a knowledge base-based customer service auxiliary decision-making system, and relates to the technical field of customer service auxiliary decision-making, and the system comprises the steps: carrying out the voice activity detection of incoming call audio, and calculating the signal-to-noise ratio and one-way time delay of a current session; based on a pre-calibrated selection time parameter set and a logarithmic relationship, obtaining selection time under different candidate suggested number; calculating a re-speaking time cost in combination with the signal-to-noise ratio, and determining a time delay penalty according to the one-way time delay; in a preset suggested number range, the selection time, the re-speaking time cost, the one-way time delay and the time delay penalty are synthesized into a unified time cost, and the hit income which is progressively decreased from margins along with the increase of the suggested number is introduced to determine the optimal suggested number; and shrinking the optimal number of suggestions under the session comfort time delay constraint to obtain a final number of suggestions, and retrieving and intercepting a suggestion set consistent with the final number of suggestions from the knowledge base document. According to the method, the suggestion scale can be adaptively controlled when the noise environment and the time delay condition dynamically change, and the customer service handling efficiency is improved.
Owner:SHENZHEN XIANGLIN EDUCATION TECH CO LTD