Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

429 results about "Timbre" patented technology

In music, timbre (/ˈtæmbər, ˈtɪm-/ TAM-bər, TIM-, French: [tɛ̃bʁ]), also known as tone color or tone quality (from psychoacoustics), is the perceived sound quality of a musical note, sound or tone. Timbre distinguishes different types of sound production, such as choir voices and musical instruments, such as string instruments, wind instruments, and percussion instruments. It also enables listeners to distinguish different instruments in the same category (e.g., an oboe and a clarinet, both woodwind instruments).

Rhythm migration method and device, electronic equipment and storage medium

The invention relates to the technical field of voice processing, and provides a rhythm migration method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a decoupled rhythm feature based on a source rhythm voice, and a decoupled tone feature based on the voice of a target speaker, the decoupled rhythm feature represents the rhythm of the source rhythm voice, and the decoupled tone feature represents the tone of the target speaker; the decoupled timbre features represent the timbre of the voice of the target speaker; generating a target voice vector sequence based on the text features of the target text, the voice features of the voice of the target speaker, the decoupled rhythm features and the decoupled timbre features; and synthesizing a target audio based on the target voice vector sequence. According to the method and the device, the decoupled rhythm features and the decoupled timbre features are acquired, and the target voice is generated based on the features, so that the problem of feature mixing is effectively relieved, the timbre purity of the target speaker in cross-person rhythm migration is ensured, the expressive force of rhythm migration is improved, and the synthesized audio is more natural and vivid.
Owner:IFLYTEK CO LTD

Multi-dimensional scoring auxiliary system and method in piano teaching

The invention relates to the technical field of piano teaching and computers, discloses a multi-dimensional scoring auxiliary system and method in piano teaching, and aims to solve the problems that existing piano teaching is single in evaluation dimension, one-sided in data acquisition and lagged in feedback. According to the method, a five-dimensional scoring model covering rhythm, strength, timbre, posture and fingering is constructed by synchronously collecting a key triggering time sequence, an audio frequency spectrum, limb posture and finger joint tracks, and quantitative scores are output through time alignment, feature extraction and neural network fusion calculation. The system comprises a key sensing module, an audio acquisition module, a posture capture module, a score calculation module and an AR visual feedback module, and supports teaching suggestion generation, historical trend analysis, difficulty self-adaption and multi-person comparison functions. Through multi-modal data closed-loop evaluation and intelligent intervention, the teaching accuracy and individuation level are remarkably improved, and scientific quantification and efficient improvement of the piano playing ability are achieved.
Owner:JILIN NORMAL UNIV

Intelligent identification and visual playing system for ancient music score

The invention relates to the technical field of ancient music score intelligent processing, and discloses an ancient music score intelligent identification and visual playing system, which comprises an image preprocessing module, a stroke enhancement module, a symbol segmentation module, a symbol identification module, a semantic analysis module, a visualization module, an acoustic synthesis module and the like. Through a specially designed image processing and symbol analysis method, the aging problem of the ancient music score can be accurately repaired, adhered symbols are separated, a complex structure is identified, music semantics are inferred in combination with a knowledge graph, and a multi-modal visualization effect and a high-sampling ancient charm tone are generated. The system supports feedback optimization and automatic workflow, significantly improves the efficiency and accuracy of ancient music score recognition, translation and presentation, and provides technical support for ancient music research and propagation.
Owner:ANHUI UNIV

Intelligent voice interaction system and method based on streaming multi-mode fusion and equipment control protocol

PendingCN121260156ASpeech recognitionSpeech synthesisSpeech comprehensionEngineering
The embodiment of the invention discloses an intelligent voice interaction system and method based on streaming multi-mode fusion and an equipment control protocol, the system comprises a voice input processing module, a voice understanding and generating module and a voice synthesis module, the voice input processing module is used for converting an audio signal into a first token sequence, and the first token sequence is used for converting the audio signal into a second token sequence; the voice understanding and generating module is used for determining a response token sequence according to the first token sequence on the basis of a multi-modal Transform architecture so as to realize voice understanding and generation; and the voice synthesis module is used for synthesizing the response token sequence into an output audio so as to carry out at least one of the following adjustments on the converted audio of the response token sequence: emotion parameter adjustment, tone adjustment and rhythm adjustment. By adopting the embodiment of the invention, low-delay and high-naturalness intelligent voice interaction can be realized, multi-modal fusion and equipment control are supported, and the user experience is remarkably improved.
Owner:SHENZHEN HUANZHI TECHNOLOGY CO LTD

Speech synthesis method and device, computer equipment and storage medium

The invention discloses a speech synthesis method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring multi-mode background sound condition input data; performing modal integrity detection on the multi-modal background sound condition input data to obtain a detection result; generating an environment background sound feature embedding vector according to a detection result; obtaining to-be-synthesized text data and speaker reference audio data, and performing feature extraction to obtain text semantic features and speaker timbre features; inputting the environment background sound feature embedded vector, the text semantic feature and the speaker timbre feature into an acoustic model to generate a Mel spectrum; and converting the Mel spectrum into a target voice waveform to obtain synthetic voice data. By implementing the method, scene requirements can be deeply matched, diversified scene types can be covered, accurate matching of background sounds and voice semantics is realized, and the technical scheme can be applied to the fields of finance and medical health.
Owner:PING AN TECH (SHENZHEN) CO LTD

Personalized tone migration and synthesis method based on virtual singer

The invention relates to a personalized timbre migration and synthesis method based on a virtual singer, and the method comprises the steps: obtaining a reference singing audio of a target virtual singer, and extracting a timbre identity benchmark feature used for representing the timbre identity stability through timbre coding; obtaining personalized timbre migration demand information for the virtual singer, wherein the demand information comprises a to-be-migrated timbre attribute and a corresponding target change amplitude or change direction; calculating a compatibility score according to the timbre identity reference feature and a migration demand, generating a timbre migration risk indication value, and comparing the risk indication value with a preset threshold value or a preset threshold value interval to determine that the migration is high-risk migration or low-risk migration or conservative-risk migration; according to the method, the tone identity benchmark features in the reference singing audio are extracted, and the tone identity invariant set with stable transpitch and sounding intensity is further constructed, so that the core tone identity of the virtual singer is accurately described.
Owner:CHANGSHA NORMAL UNIV

Video data processing method and device and electronic equipment

The invention provides a video data processing method and apparatus, and an electronic device. The method comprises the steps of obtaining initial video data; the initial video data comprises initial video stream data and initial audio stream data corresponding to the initial video stream data; determining a role feature of at least one audio generation role in the initial video data; determining scene features of the initial video data; based on the role features, the scene features and the content translation text corresponding to the initial audio stream data, determining input data; inputting the input data into a preset voice generation model, and generating target audio stream data through the voice generation model; and generating target video data based on the initial video stream data and the target audio stream data. According to the mode, in the process of generating the translated audio corresponding to the video, tone cloning and emotion restoration of the translated audio are realized through the role features of the speaker of the original video and the scene features embodied by the video, the translation effect of the video is improved, and the video translation cost is reduced.
Owner:WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD

Method for training speech synthesis model, speech synthesis method, and electronic device

A method for training a speech synthesis model includes obtaining training data; obtaining an initial speech synthesis model; training a semantic encoding network and a semantic decoding network in the speech synthesis model respectively based on a style sample speech, a timbre sample speech, an input sample text, and an output sample speech in training samples of the training data, to obtain a trained speech synthesis model.
Owner:BAIDU INT TECH (SHENZHEN) CO LTD

Gravity pressing structure applied to electronic guitar

The utility model relates to a gravity pressing structure applied to an electronic guitar, which comprises a fixed plate fixedly arranged on a panel of the electronic guitar. The PCB is fixedly arranged on the fixed plate; the string plucking structure comprises a plurality of string plucking bodies, a base is arranged on the bottom surface of each string plucking body, the bases are arranged on the PCB, the bases are provided with gravity sensors, the gravity sensors are electrically connected with the PCB, and the string plucking bodies are made of soft rubber materials. When a player slightly touches, presses or strikes the string plucking body with fingers, the gravity sensor captures the process, converts the magnitude of force into a corresponding level signal, and transmits the level signal to an audio processing system of the electronic guitar through the PCB. And the system adjusts pitch, volume or timbre parameters according to the parameters, and simulates sound similar to wood guitar soft string or string striking. The string plucking body is made of a soft rubber material, so that the string plucking body can be pressed, beaten, flexibly rubbed, toggled and other skills like a traditional string, and in cooperation with the gravity sensor, sound generated when the traditional string uses the skills can be simulated.
Owner:SHENZHEN WENTAI MICROELECTRONICS CO LTD

Optimization algorithm model and method based on elder care emotion accompanying

The invention discloses an adjusting and optimizing algorithm model based on elder care emotion accompanying. The model comprises a multi-mode emotion perception module for receiving voice, images and physiological signals and extracting features such as timbre, facial expression and pulse to generate emotion feature vectors; the emotion recognition module is used for predicting the real-time emotion state of the old people through fusion of CNN and LSTM and multi-modal features; the strategy adjusting and optimizing module adopts a Bayesian optimization or genetic algorithm to automatically adjust and optimize the voice intonation, the interaction rhythm and the dialogue strategy of the accompanying system; the feedback learning module iteratively updates the model based on the old people feedback data, and optimizes the interaction effect; and the personalized accompanying generation module generates accompanying strategies such as voice consolation and music recommendation according to the tuning result. The multi-mode emotion perception module is composed of a voice monitoring sub-module, an expression monitoring sub-module and a physiological signal monitoring sub-module. The strategy tuning module comprises a self-adaptive parameter search sub-module and a multi-target optimization sub-module; and the feedback learning module comprises a user feedback collection and reinforcement learning sub-module, so that the interaction experience and the emotion adaptability of the accompanying system are improved.
Owner:闫中举

Sound source localization and identification system for electric power intelligent service and operation method of sound source localization and identification system

The invention discloses a sound source positioning identification system for electric power intelligent service and an operation method, and the system comprises an input module which is used for receiving an interaction demand of a user, and uploading the interaction demand of the user to an identification module; the identification module receives a user interaction demand, identifies and positions a user sound position based on the user interaction demand, and uploads an identification and positioning result to the processing module; one end of the processing module is connected to the recognition module, the other end of the processing module is connected to the output module, the processing module analyzes and processes the user question according to the received recognition and positioning result, and the output module answers the user question based on the analysis and processing result. According to the method, after the noise doped in the utterance spoken by the user is removed, the timbre, the speech speed, the audio frequency and the like in the statement of the user are recorded, and meanwhile, the user is captured, so that the virtual digital human can more accurately recognize and locate the user through a sound source, the interaction error is reduced, and the naturalness during interaction is improved.
Owner:GUANGXI POWER GRID CORP

Cross-culture melody feature extraction and musical instrument matching system based on deep learning

The invention discloses a cross-culture melody feature extraction and musical instrument matching system based on deep learning, and belongs to the technical field of music information processing and artificial intelligence. The system comprises an audio acquisition module, an audio preprocessing module, a melody feature extraction module, a cultural context understanding module, an emotion semantic analysis module, a musical instrument timbre database, a musical instrument matching recommendation module, a distributor scheme generation module, a user interaction module and an audio synthesis module. Multi-dimensional melody features are extracted through a multi-scale convolutional neural network and a bidirectional long-short term memory network, cultural semantic understanding is realized through reasoning on a music knowledge graph by using a graph neural network, and musical instrument matching is performed by using a multi-objective optimization algorithm in combination with three dimensions of timbre integrating degree, cultural consistency and emotional expressive power. According to the system, automation of the whole process from humming melody to musical instrument configuration is achieved, the efficiency of the instrument is improved by 400%, the accuracy rate reaches 92%, and an intelligent tool is provided for cross-culture music creation and national music modernization reorganization.
Owner:SHENZHEN UNIV

Multi-language intelligent dubbing generation system based on sound cloning and emotion migration

The invention discloses a multi-language intelligent dubbing generation system based on sound cloning and emotion migration. The multi-language intelligent dubbing generation system comprises a sound cloning module, a cross-language synthesis module, an emotion migration module and a lip shape synchronization module. By constructing a few-sample speaker encoder, a cross-language rhythm migration module, a fine-grained emotion control module and a video lip shape synchronization module, end-to-end automatic generation from original dubbing audio to multi-language target dubbing is realized, and tone consistency, emotion authenticity and picture synchronism are kept.
Owner:JIANGSU HOPERUN SOFTWARE CO LTD

A voice conversion method, device, equipment and readable storage medium

The application provides a speech conversion method, device and equipment and a readable storage medium. The method comprises the following steps: obtaining speech information to be processed; based on a three-head encoder, encoding and modeling speech content, environmental noise and fundamental frequency information in the speech information to be processed respectively to obtain encoded and modeled speech information; changing the time sequence of the encoded and modeled speech information to adjust the speech speed of the encoded speech information; inputting the speech information with adjusted speech speed into a previously trained timbre conversion model corresponding to a target user to obtain target acoustic features, wherein the timbre of the target acoustic features is the same as the timbre of the target user. Thus, the timbre conversion of the speech information to be processed can be performed as required, and the conversion method is more efficient and accurate. Multi-dimensional encoding of the speech information can improve the robustness of the speech in a noisy environment. Speech speed control can make the speech more in line with user requirements.
Owner:MIGU CO LTD +1

Detachable guitar

ActiveCN224177095UReduce local stressimprove integrityGuitarsNoiseMechanical engineering
The utility model discloses a detachable guitar, and relates to the technical field of guitar structures, and the detachable guitar comprises a guitar body and two handrails. The first ends of the two handrails are connected with the guitar body, and the second ends of the two handrails are directly connected to form a closed frame structure, so that on one hand, vibration can be more smoothly conducted to the guitar body through the continuous frame structure to improve the tone consistency, and on the other hand, resonance at a specific frequency can be inhibited, noise is reduced, the tone is purer, and the tone quality is improved. The closed frame structure can effectively disperse external force, reduce local stress of the guitar body and reduce the deformation risk, on the other hand, the second ends of the two handrails are connected to the lower end of the guitar body through the threaded connecting piece after being stacked, compared with the mode that the two handrails are independently connected to the guitar body, the stacking design can reduce the number of open holes, and the number of the open holes is reduced. The completeness of the guitar body is protected.
Owner:CHANGSHA GANYIN TECHNOLOGY CO LTD

Method, device and software for applying an audio effect

The present invention provides a method for processing music audio data, comprising the steps of providing input audio data representing a first piece of music containing a mixture of predetermined musical timbres, decomposing the input audio data to generate at least a first audio track representing a first musical timbre selected from the predetermined musical timbres, and a second audio track representing a second musical timbre selected from the predetermined musical timbres, applying a predetermined first audio effect to the first audio track, applying no audio effect or a predetermined second audio effect, which is different from the first audio effect, to the second audio track, and obtaining recombined audio data by recombining the first audio track with the second audio track.
Owner:ALGORIDDIM GMBH

Story audio timbre processing method and related device

The invention discloses a story audio timbre processing method and a related device, and relates to the technical field of audio processing, and the method comprises the steps: extracting a story voice from a to-be-processed story audio before carrying out the timbre processing of the to-be-processed story audio through employing a reference audio, and then carrying out the timbre conversion processing of the story voice based on the reference audio, thereby achieving the timbre processing of the to-be-processed story audio. And finally, determining a final story audio based on the target story voice. In the whole processing process, voice recognition is not needed, the influence of voice recognition accuracy on the tone processing effect is avoided, the to-be-processed audio is subjected to story voice extraction and story voice processing, the influence of story background voice on the tone processing effect is avoided, and therefore the voice processing efficiency is improved. According to the scheme, the tone processing effect of the story audio can be improved.
Owner:HEFEI IFLYTEK TOYCLOUD TECH

Media data generation method and device, equipment, medium and product

The invention provides a media data generation method and device, equipment, a medium and a product, and the method comprises the steps: receiving a first text and first description information, the first text comprises the session content of at least one session participant, and the first description information comprises the session content of the at least one session participant; the first description information is at least used for reflecting first session state information of the at least one session participant; performing information processing on the first text and the first description information based on a first model to obtain first media data; wherein the first media data at least comprises first audio data, the first audio data is related to the session content in the first text, the tone information of the same session participant in the first audio data is the same, and the second session state information of the session participant is related to the first session state information; the dynamic adaptation of the acoustic characteristics of the first media data and the session state information is realized, so that the audio expression effect of the media data is improved.
Owner:BYTEDANCE TECHNOLOGY CO LTD +2

Graphical user interface for generative adversarial network music synthesizer

An information processing system that receives input sound and pitch information; extracts a timbre feature amount from the input sound; and generates information of a musical instrument sound with a pitch based on the timbre feature amount and the pitch information.
Owner:SONY GROUP CORP

A sound box

ActiveCN224503460USound waveBass (sound)
A kind of sound box, including cabinet and the protective cover being arranged in the front of cabinet, first acoustic cavity is equipped in the cabinet, high pitch loudspeaker assembly and bass horn assembly are equipped on the first acoustic cavity, second acoustic cavity is also equipped in the cabinet, the first acoustic cavity is communicated with second acoustic cavity by loudspeaker pipe passing through partition, the loudspeaker pipe is fixedly connected with cabinet, one end of the loudspeaker pipe is arranged in the back of high pitch loudspeaker assembly, the other end of the loudspeaker pipe is inserted into second acoustic cavity. By adding a second acoustic cavity in the cabinet and the loudspeaker pipe that is communicated with the first acoustic cavity and the second acoustic cavity, and one end of the loudspeaker pipe is arranged in the back of the high pitch loudspeaker assembly, the manufacturer selects loudspeaker pipe with different pipe diameter and different length according to actual situation, so that the inside of the loudspeaker pipe helps the reflection and resonance of sound waves, so that the utility model shows more rich tone, and then the sound performance comparable to real musical instrument.

Pluggable target speaker speech recognition method and system

ActiveCN121963713ARetain original universal recognition capabilitiesreduce error rateSpeech recognitionSemantic alignmentFeature extraction
The invention discloses a voice recognition method and system for a pluggable target speaker, is applied to the technical field of data processing, and designs a feature domain direct connection pluggable two-stage training architecture for the target voice recognition requirement of a multi-speaker scene. The method comprises the following steps: firstly, carrying out Mel spectrum preprocessing on multiple types of audios, and extracting a global timbre embedding vector through an adaptive voiceprint network; and then an adaptive convolutional coding extraction network is established, and basic target feature extraction is realized through AdaLayer modulation and L1 loss physical alignment training. The extraction module is cascaded with a pre-training ASR, parameters of the extraction module and the pre-training ASR are frozen, a corresponding loss function is matched according to an ASR framework, and semantic alignment is completed only by fine tuning a front-end module. Pluggable adaptation of an extraction module is realized through accurate dimension alignment, and a frozen ASR is taken as a static semantic discriminator, so that target speech semantic recognition and accurate extraction with low computing power overhead and high recognition rate are finally realized.
Owner:XIAMEN LIMAYAO NETWORK TECH CO LTD

Multi-dimensional timbre perception space model based on electroencephalogram features, modeling method, device and storage medium

ActiveCN118568466BAudiometeringPsychotechnic devicesAuditory stimuliFeature vector
The application discloses a kind of multidimensional timbre perception space model modeling method based on electroencephalogram characteristics, which comprises the following steps: collecting a variety of musical instrument single timbre samples and pretreating;To timbre sample, extract acoustic time-frequency domain feature, and extract psychological perception feature by behavior psychology experiment;Multiple musical instrument timbre samples are used as auditory stimulus to carry out electroencephalogram experiment, and corresponding event-related potential ERP signal is extracted as electroencephalogram feature;The dissimilarity of different timbre characteristics is represented by the Euclidean distance of different timbre points;The distance between different timbre points is calculated;According to the Euclidean distance matrix between sample timbre points, a low-dimensional space with mutually orthogonal dimensions is fitted, the dissimilarity of timbre feature vector is transformed into low-dimensional space, the timbre feature vector is directly mapped into low-dimensional space, a point set is formed, and the similarity and dissimilarity of each musical instrument timbre are directly displayed.The application represents the mapping relationship between acoustic characteristics, psychological perception characteristics and electroencephalogram characteristics.
Owner:TIANJIN UNIV

A multi-modal digital human generation method

The present application belongs to the field of images, the field of speech and the field of digital people, and particularly relates to a digital person generation method based on multi-modal. The method first acquires a sound video of different images under the same text, separates the audio and video and extracts facial features to construct a data set; then builds and trains a digital person image cloning model and a timbre cloning model, respectively realizing the mapping from audio to facial features, facial features to silent video and timbre cloning; finally, the two models are integrated to realize digital person question and answer communication with the help of a large language model. Compared with traditional single modal generation technology, the present application solves the problem of inconsistency between the appearance and timbre of virtual people and the inaccuracy of emotional expression through multi-modal data fusion, improves the realism and naturalness of digital people, enhances their expressiveness in virtual anchor, intelligent customer service and other scenes, and promotes the development of digital people technology.
Owner:JINAN FALAI TECH CO LTD

Voice conversion method, apparatus, medium, and device

The application discloses a speech conversion method, device, medium and equipment, the speech conversion method comprises the following steps: determining the task data input into different encoders in a speech conversion model according to a speech conversion type, and encoding the input task data through the encoders in the speech conversion model to obtain the encoding features output by each encoder; obtaining the public features corresponding to the encoding features output by a public classifier, and obtaining the adversarial features corresponding to the encoding features output by an adversarial classifier; performing decoding processing to obtain a target mel spectrum and a target high-frequency contour; and synthesizing a target audio based on the target mel spectrum and the target high-frequency contour through a vocoder. In one-time speech conversion, different representation styles are transmitted respectively, that is, according to the speech conversion type, speech conversion of timbre or timbre+pitch is realized respectively, so that the converted speech maintains the naturalness and expressiveness of the source speech.
Owner:SHENZHEN RAISOUND TECH

Emotional speech synthesis method, system, device and medium

The invention discloses an emotional speech synthesis method, system and device and a medium, and relates to the technical field of computer-aided dubbing, and the method comprises the steps: obtaining speech sample audios of various emotions and a target sound ray of a target object; performing timbre migration processing on the plurality of speaking sample audios based on the target sound ray to obtain a target multi-emotion long audio; acquiring a voice attribute control parameter input by a user; and performing voice synthesis based on the voice attribute control parameter and the target multi-emotion long audio to obtain a target emotion voice. The problem that an existing speech synthesis technology is remarkably insufficient in flexibility, fineness and adaptability of emotion expression is solved.
Owner:WUHAN JIANSHI TECH CO LTD

Speech synthesis method and device, electronic equipment and readable storage medium

The present disclosure relates to a sound synthesis method and device, electronic equipment and readable storage medium, wherein the present scheme realizes text-to-target timbre audio conversion through a pre-trained speech synthesis model, the speech synthesis model comprising a first feature extraction sub-model and a second feature extraction sub-model, wherein the first feature extraction sub-model outputs first acoustic features comprising bottleneck features according to inputted text to be processed; the second feature extraction sub-model outputs mel-frequency spectrum features corresponding to the text to be processed according to inputted first acoustic features; and target audio corresponding to the text to be processed is obtained according to the mel-frequency spectrum features corresponding to the text to be processed, the target audio having a target timbre. The present scheme decouples the speech synthesis model into two models through first acoustic features comprising bottleneck features, realizes relatively independent control of timbre and other features on speech synthesis, and meets the demand of users for personalized speech synthesis.
Owner:FACE CUTE CO LTD

A bone conduction speech conversion method based on spectral envelope mapping

The application discloses a bone conduction speech conversion method based on spectral envelope mapping, comprising the following steps: pre-processing and linear prediction analysis of the bone conduction speech signal, and calculating the LP filter coefficient; mapping the LSF coefficient of the air conduction speech signal corresponding to the bone conduction speech signal by using the trained neural network; controlling the minimum error of the original bone conduction speech signal and the synthesized air conduction speech signal according to the timbre weighting characteristic of the bone conduction speech characteristic; estimating the integer pitch of the bone conduction speech signal through the timbre weighting filter, obtaining the reference signal by passing the linear prediction residual signal through the timbre weighting filter, estimating the fractional pitch to obtain the adaptive codebook vector; obtaining the new reference signal by subtracting the adaptive codebook vector from the reference signal, searching for the optimal excitation in the fixed codebook; synthesizing the air conduction speech signal by combining the optimal excitation and the LP filter of the air conduction speech signal, and correcting the air conduction speech signal.
Owner:DALIAN UNIV OF TECH

Song timbre matching method and system based on audio spectrum analysis

The application discloses a song tone matching method and system based on audio spectrum analysis, and relates to the technical field of data processing.The method comprises the following steps: obtaining original singing audio and user audio, obtaining song sentences, obtaining first matching frequency spectrum segments corresponding to the song sentences, and obtaining second matching frequency spectrum segments corresponding to the song sentences; obtaining first frequency spectrum peak points, dividing the first matching frequency spectrum segments into first front frequency spectrum segments and first rear frequency spectrum segments, obtaining second frequency spectrum peak points, dividing the second matching frequency spectrum segments into second front frequency spectrum segments and second rear frequency spectrum segments; obtaining front frequency spectrum difference values, obtaining rear frequency spectrum difference values, and judging whether the front frequency spectrum difference values and the rear frequency spectrum difference values are smaller than a preset threshold value; if not, generating target front frequency spectrum segments and target rear frequency spectrum segments, constructing target sentence audio according to the target front frequency spectrum segments and the target rear frequency spectrum segments, and obtaining matching audio according to the target sentence audio.The application has the advantages of good song tone matching effect, rhythm deviation repair, original singing style reservation and user tone matching effect.
Owner:CHENGDU YINYUE CHUANGXIANG TECH CO LTD

Intelligent monitoring system and method for transmission equipment based on noise feature analysis

A kind of transmission equipment intelligent monitoring system and method based on noise feature analysis, through the method, expert can input the sound evaluation information heard from console, the sound evaluation value of expert corresponds to the state of specific reduction mechanism, the loudness and frequency characteristic of noise signal output by spectrum analyzer as input layer, the sound evaluation value of expert as output layer, mapping relationship between the two is established through data accumulation and neural network algorithm. Through the database of abnormal noise experts as a reference, warning is found by comparing similar abnormal noise data, ensure that noise meets the standard, through the noise evaluation of experts and neural network algorithm, the timbre can be evaluated, and then the specific type of fault can be judged. Moreover, the influence of on-site noise on sound signal collection is excluded, and the accuracy of machine detection is further improved.
Owner:HANGZHOU JIE DRIVE TECH

Information generation method capable of objectively improving timbre on basis of human body shape and sound data processing device capable of intuitively and easily changing impression of sound

A system S according to an exemplary embodiment of the present invention acquires head-related transfer functions (HRTFs) obtained by emitting sound toward the head of an individual P from multiple sound emission directions D in an anechoic chamber, averages the plurality of HRTFs, and calculates a target response curve (TRC) for the individual P, which is referred to as TPTRC, and which clarifies information related to sound timbre. The system S generates TPTRC(W) by multiplying the TPTRC by each of different multipliers W, presents sounds conforming to the respective TPTRC(W) to individual P, to prompt individual P to select from among the sounds a preferred sound, storing the TPTRC(W) selected by individual P as TPTRCadj. The system S generates a generic TPTRCadj by averaging TPTRCadj stored for each of a plurality of different individuals P1 to Pn. A sound data processing device according to the exemplary embodiment of the present invention superimposes the generic TPTRCadj on input sound data and outputs the processed sound data to a sound-emitting device such as earphones.
Owner:FINAL INC