Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

620 results about "Timbre" patented technology

In music, timbre (/ˈtæmbər, ˈtɪm-/ TAM-bər, TIM-, French: [tɛ̃bʁ]), also known as tone color or tone quality (from psychoacoustics), is the perceived sound quality of a musical note, sound or tone. Timbre distinguishes different types of sound production, such as choir voices and musical instruments, such as string instruments, wind instruments, and percussion instruments. It also enables listeners to distinguish different instruments in the same category (e.g., an oboe and a clarinet, both woodwind instruments).

Video translation method and system based on artificial intelligence

The invention discloses a video translation method and system based on artificial intelligence. The method relates to the technical field of video translation and comprises the following steps of original sound track extraction, target AI speaker adaptation, AI dubbing generation and mouth shape synchronization and video synthesis. According to the method, independent audio and video streams are obtained by adopting an audio and video separation technology, and multiple original sound tracks are extracted through a voice separation model; matching or generating an adaptive target AI speaker module in a preset tone library; converting the original language voice into a text, translating the text into a target language text, and synthesizing an AI dubbing audio track in combination with a target AI speaker module; and finally, the independent video stream and the multi-AI dubbing audio track are input into the mouth shape synchronization model to output a translated video, so that the timbre fitting degree, the voice quality and the voice consistency of the same speaker of AI dubbing are improved, and meanwhile, the resource utilization rate of video translation and the processing efficiency under batch tasks are improved. The problem that in the prior art, video translation is low in quality and efficiency is solved.
Owner:BEIJING DEEP LOGIC INTELLIGENT TECHNOLOGY CO LTD

Rhythm migration method and device, electronic equipment and storage medium

The invention relates to the technical field of voice processing, and provides a rhythm migration method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a decoupled rhythm feature based on a source rhythm voice, and a decoupled tone feature based on the voice of a target speaker, the decoupled rhythm feature represents the rhythm of the source rhythm voice, and the decoupled tone feature represents the tone of the target speaker; the decoupled timbre features represent the timbre of the voice of the target speaker; generating a target voice vector sequence based on the text features of the target text, the voice features of the voice of the target speaker, the decoupled rhythm features and the decoupled timbre features; and synthesizing a target audio based on the target voice vector sequence. According to the method and the device, the decoupled rhythm features and the decoupled timbre features are acquired, and the target voice is generated based on the features, so that the problem of feature mixing is effectively relieved, the timbre purity of the target speaker in cross-person rhythm migration is ensured, the expressive force of rhythm migration is improved, and the synthesized audio is more natural and vivid.
Owner:IFLYTEK CO LTD

Real-time sound duplicating method and system based on end-cloud fusion

The invention provides a real-time sound copying method and system based on end-cloud fusion. The method comprises the following steps: a cloud end carries out real-time tone copying and voice synthesis on a small amount of voice data of a user based on an AI large model; timbre samples are collected when a user registers voice audio data, and user timbre voice data of a preset text are synchronously generated by a large model and are used as fine tuning training data of an end-side voice synthesis model; user tone voice data of a preset text and voice audio data registered by a user are utilized to carry out migration fine tuning training on an end-side voice synthesis model to adapt to the personalized tone of the user, so that high-quality output of the end-side voice synthesis model is ensured, and personalized voice replication is realized; and issuing the trained end-side speech synthesis model to the user equipment, and independently completing speech replication in a network-free or weak network environment. According to the invention, through automatic generation of the user tone data and adaptive fine tuning of the model, the tone of the user is deployed to the end side after fine tuning, and high-quality and high-adaptability sound replication of end-cloud collaboration is realized.
Owner:PACHIRA TIMES (ZHUHAI HENGQIN) INFORMATION TECH CO LTD

Multi-dimensional scoring auxiliary system and method in piano teaching

The invention relates to the technical field of piano teaching and computers, discloses a multi-dimensional scoring auxiliary system and method in piano teaching, and aims to solve the problems that existing piano teaching is single in evaluation dimension, one-sided in data acquisition and lagged in feedback. According to the method, a five-dimensional scoring model covering rhythm, strength, timbre, posture and fingering is constructed by synchronously collecting a key triggering time sequence, an audio frequency spectrum, limb posture and finger joint tracks, and quantitative scores are output through time alignment, feature extraction and neural network fusion calculation. The system comprises a key sensing module, an audio acquisition module, a posture capture module, a score calculation module and an AR visual feedback module, and supports teaching suggestion generation, historical trend analysis, difficulty self-adaption and multi-person comparison functions. Through multi-modal data closed-loop evaluation and intelligent intervention, the teaching accuracy and individuation level are remarkably improved, and scientific quantification and efficient improvement of the piano playing ability are achieved.
Owner:JILIN NORMAL UNIV

Intelligent identification and visual playing system for ancient music score

The invention relates to the technical field of ancient music score intelligent processing, and discloses an ancient music score intelligent identification and visual playing system, which comprises an image preprocessing module, a stroke enhancement module, a symbol segmentation module, a symbol identification module, a semantic analysis module, a visualization module, an acoustic synthesis module and the like. Through a specially designed image processing and symbol analysis method, the aging problem of the ancient music score can be accurately repaired, adhered symbols are separated, a complex structure is identified, music semantics are inferred in combination with a knowledge graph, and a multi-modal visualization effect and a high-sampling ancient charm tone are generated. The system supports feedback optimization and automatic workflow, significantly improves the efficiency and accuracy of ancient music score recognition, translation and presentation, and provides technical support for ancient music research and propagation.
Owner:ANHUI UNIV

Video dubbing language conversion method and system and related equipment

The invention provides a video dubbing language conversion method, a video dubbing language conversion system and related equipment. The method comprises the following steps: acquiring audio track data from a video to be converted; carrying out human voice extraction on the audio track data and classifying according to roles to obtain a single speaker audio of each role; performing voice-to-text conversion on the single speaker audio of each role to obtain an original language copywriting of each role; performing sound cloning on the single speaker audio of each role to obtain a timbre model of each role; performing target language translation on the original language copywriting of each role to obtain a translated copywriting of each role; based on the translation copywriting of each role and the tone model of each role, performing text-to-voice conversion to obtain a translation audio of each role; and performing replacement of each role translation audio on the audio track data in the to-be-converted video to obtain a dubbing conversion video. According to the technical scheme, language video dubbing conversion combined with the tone of the speaker is achieved, the video is more diversified, and the user requirements can be better met.
Owner:SHENZHEN MAIFENG TECH CO LTD

Intelligent voice interaction system and method based on streaming multi-mode fusion and equipment control protocol

PendingCN121260156ASpeech recognitionSpeech synthesisSpeech comprehensionEngineering
The embodiment of the invention discloses an intelligent voice interaction system and method based on streaming multi-mode fusion and an equipment control protocol, the system comprises a voice input processing module, a voice understanding and generating module and a voice synthesis module, the voice input processing module is used for converting an audio signal into a first token sequence, and the first token sequence is used for converting the audio signal into a second token sequence; the voice understanding and generating module is used for determining a response token sequence according to the first token sequence on the basis of a multi-modal Transform architecture so as to realize voice understanding and generation; and the voice synthesis module is used for synthesizing the response token sequence into an output audio so as to carry out at least one of the following adjustments on the converted audio of the response token sequence: emotion parameter adjustment, tone adjustment and rhythm adjustment. By adopting the embodiment of the invention, low-delay and high-naturalness intelligent voice interaction can be realized, multi-modal fusion and equipment control are supported, and the user experience is remarkably improved.
Owner:SHENZHEN HUANZHI TECHNOLOGY CO LTD

End side voice model deployment method and device, equipment and storage medium

The invention relates to the technical field of end-side model deployment, in particular to an end-side voice model deployment method and device, equipment and a storage medium. Comprising the steps of recording reference audio at an end side; extracting a reference semantic token and an embedded vector of the reference audio; obtaining a training text set, integrating and inputting the training text code, the embedded vector and the reference semantic token corresponding to each training text in the training text set into a preset language model, and outputting a comprehensive token sequence corresponding to the training text; performing Mel spectrum conversion on each comprehensive token sequence to obtain a first spectrum representation corresponding to the comprehensive token sequence; generating an audio signal corresponding to the training text according to each first spectrum representation, and integrating to obtain a training data set; inputting the training data set into a to-be-trained model, and training to obtain a lightweight voice model which only retains timbre modeling parameters; and updating the lightweight voice model to the end side to complete end side deployment. According to the invention, the voice synthesis model deployment of the end-side equipment can be realized.
Owner:SHENZHEN RAISOUND TECH

Speech synthesis method and device, computer equipment and storage medium

The invention discloses a speech synthesis method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring multi-mode background sound condition input data; performing modal integrity detection on the multi-modal background sound condition input data to obtain a detection result; generating an environment background sound feature embedding vector according to a detection result; obtaining to-be-synthesized text data and speaker reference audio data, and performing feature extraction to obtain text semantic features and speaker timbre features; inputting the environment background sound feature embedded vector, the text semantic feature and the speaker timbre feature into an acoustic model to generate a Mel spectrum; and converting the Mel spectrum into a target voice waveform to obtain synthetic voice data. By implementing the method, scene requirements can be deeply matched, diversified scene types can be covered, accurate matching of background sounds and voice semantics is realized, and the technical scheme can be applied to the fields of finance and medical health.
Owner:PING AN TECH (SHENZHEN) CO LTD

Personalized tone migration and synthesis method based on virtual singer

The invention relates to a personalized timbre migration and synthesis method based on a virtual singer, and the method comprises the steps: obtaining a reference singing audio of a target virtual singer, and extracting a timbre identity benchmark feature used for representing the timbre identity stability through timbre coding; obtaining personalized timbre migration demand information for the virtual singer, wherein the demand information comprises a to-be-migrated timbre attribute and a corresponding target change amplitude or change direction; calculating a compatibility score according to the timbre identity reference feature and a migration demand, generating a timbre migration risk indication value, and comparing the risk indication value with a preset threshold value or a preset threshold value interval to determine that the migration is high-risk migration or low-risk migration or conservative-risk migration; according to the method, the tone identity benchmark features in the reference singing audio are extracted, and the tone identity invariant set with stable transpitch and sounding intensity is further constructed, so that the core tone identity of the virtual singer is accurately described.
Owner:CHANGSHA NORMAL UNIV

Video data processing method and device and electronic equipment

The invention provides a video data processing method and apparatus, and an electronic device. The method comprises the steps of obtaining initial video data; the initial video data comprises initial video stream data and initial audio stream data corresponding to the initial video stream data; determining a role feature of at least one audio generation role in the initial video data; determining scene features of the initial video data; based on the role features, the scene features and the content translation text corresponding to the initial audio stream data, determining input data; inputting the input data into a preset voice generation model, and generating target audio stream data through the voice generation model; and generating target video data based on the initial video stream data and the target audio stream data. According to the mode, in the process of generating the translated audio corresponding to the video, tone cloning and emotion restoration of the translated audio are realized through the role features of the speaker of the original video and the scene features embodied by the video, the translation effect of the video is improved, and the video translation cost is reduced.
Owner:WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD

Voice conversion method and related equipment

The embodiment of the invention discloses a voice conversion method and related equipment. The related equipment can comprise a voice conversion device, electronic equipment, a computer program product and a computer readable storage medium. After at least one to-be-converted voice and a target timbre identifier corresponding to the to-be-converted voice are acquired, audio content features and acoustic features are extracted from the to-be-converted voice, and the target timbre feature corresponding to the to-be-converted voice is determined based on the target timbre identifier; extracting audio conversion features from the audio content features, extracting rhythm features from the acoustic features, fusing the target timbre features, the audio conversion features and the rhythm features to obtain target audio features, and then generating target voice corresponding to the target timbre identifier based on the target audio features; according to the scheme, the voice conversion accuracy can be improved. The embodiment of the invention can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method for training speech synthesis model, speech synthesis method, and electronic device

A method for training a speech synthesis model includes obtaining training data; obtaining an initial speech synthesis model; training a semantic encoding network and a semantic decoding network in the speech synthesis model respectively based on a style sample speech, a timbre sample speech, an input sample text, and an output sample speech in training samples of the training data, to obtain a trained speech synthesis model.
Owner:BAIDU INT TECH (SHENZHEN) CO LTD

Audio communication method, audio conversion method, apparatus, electronic device, computer-readable storage medium, and computer program product

PCT designated stageWO2025237010A1Speech analysisComputer hardwareFeature coding
An audio communication method, an audio conversion method, a bitstream processing method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The audio communication method comprises: in response to a first communication request for an audio signal, acquiring, from a plurality of communication modes, a voice transformation mode for the audio signal (101); performing feature coding on the audio signal, so as to obtain a coded feature of the audio signal (102); acquiring a target timbre corresponding to the voice transformation mode, and determining a timbre feature of the target timbre (103); performing timbre conversion on the coded feature on the basis of the timbre feature, so as to obtain a target coded feature (104); and performing signal coding on the target coded feature, so as to obtain a target audio bitstream conforming to the target timbre, and transmitting the target audio bitstream to a decoding terminal (105).
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Gravity pressing structure applied to electronic guitar

The utility model relates to a gravity pressing structure applied to an electronic guitar, which comprises a fixed plate fixedly arranged on a panel of the electronic guitar. The PCB is fixedly arranged on the fixed plate; the string plucking structure comprises a plurality of string plucking bodies, a base is arranged on the bottom surface of each string plucking body, the bases are arranged on the PCB, the bases are provided with gravity sensors, the gravity sensors are electrically connected with the PCB, and the string plucking bodies are made of soft rubber materials. When a player slightly touches, presses or strikes the string plucking body with fingers, the gravity sensor captures the process, converts the magnitude of force into a corresponding level signal, and transmits the level signal to an audio processing system of the electronic guitar through the PCB. And the system adjusts pitch, volume or timbre parameters according to the parameters, and simulates sound similar to wood guitar soft string or string striking. The string plucking body is made of a soft rubber material, so that the string plucking body can be pressed, beaten, flexibly rubbed, toggled and other skills like a traditional string, and in cooperation with the gravity sensor, sound generated when the traditional string uses the skills can be simulated.
Owner:SHENZHEN WENTAI MICROELECTRONICS CO LTD

Optimization algorithm model and method based on elder care emotion accompanying

The invention discloses an adjusting and optimizing algorithm model based on elder care emotion accompanying. The model comprises a multi-mode emotion perception module for receiving voice, images and physiological signals and extracting features such as timbre, facial expression and pulse to generate emotion feature vectors; the emotion recognition module is used for predicting the real-time emotion state of the old people through fusion of CNN and LSTM and multi-modal features; the strategy adjusting and optimizing module adopts a Bayesian optimization or genetic algorithm to automatically adjust and optimize the voice intonation, the interaction rhythm and the dialogue strategy of the accompanying system; the feedback learning module iteratively updates the model based on the old people feedback data, and optimizes the interaction effect; and the personalized accompanying generation module generates accompanying strategies such as voice consolation and music recommendation according to the tuning result. The multi-mode emotion perception module is composed of a voice monitoring sub-module, an expression monitoring sub-module and a physiological signal monitoring sub-module. The strategy tuning module comprises a self-adaptive parameter search sub-module and a multi-target optimization sub-module; and the feedback learning module comprises a user feedback collection and reinforcement learning sub-module, so that the interaction experience and the emotion adaptability of the accompanying system are improved.
Owner:闫中举

System and method for creating timbres

A method of building a new voice having a new timbre using a timbre vector space includes receiving timbre data filtered using a temporal receptive field. The timbre data is mapped in the timbre vector space. The timbre data is related to a plurality of different voices. Each of the plurality of different voices has respective timbre data in the timbre vector space. The method builds the new timbre using the timbre data of the plurality of different voices using a machine learning system.
Owner:MODULATE INC

Sound source localization and identification system for electric power intelligent service and operation method of sound source localization and identification system

The invention discloses a sound source positioning identification system for electric power intelligent service and an operation method, and the system comprises an input module which is used for receiving an interaction demand of a user, and uploading the interaction demand of the user to an identification module; the identification module receives a user interaction demand, identifies and positions a user sound position based on the user interaction demand, and uploads an identification and positioning result to the processing module; one end of the processing module is connected to the recognition module, the other end of the processing module is connected to the output module, the processing module analyzes and processes the user question according to the received recognition and positioning result, and the output module answers the user question based on the analysis and processing result. According to the method, after the noise doped in the utterance spoken by the user is removed, the timbre, the speech speed, the audio frequency and the like in the statement of the user are recorded, and meanwhile, the user is captured, so that the virtual digital human can more accurately recognize and locate the user through a sound source, the interaction error is reduced, and the naturalness during interaction is improved.
Owner:GUANGXI POWER GRID CORP

Multi-timbre emotion adaptive speech synthesis method, device, equipment and medium

The invention discloses a multi-tone emotion adaptive speech synthesis method and device, equipment and a medium. The method comprises the following steps: acquiring a driving assistance prompt text; performing timbre coding on the driving auxiliary prompt text to obtain an auxiliary prompt timbre vector, and generating auxiliary prompt timbre data based on the auxiliary prompt timbre vector; performing emotional rhythm control on the auxiliary prompt tone data through an emotional rhythm mapping model and the target driving scene to obtain emotional expression tone data; according to the vehicle driving noise data, optimizing the emotion expression timbre data to obtain noise optimization timbre data; performing acoustic characteristic compensation on the noise optimization tone data, and performing spatial acoustic optimization on the noise optimization tone data to obtain acoustic optimization tone data; and performing context emotion expression adjustment on the acoustic optimization tone data to generate auxiliary prompt synthetic voice of the intelligent vehicle corresponding to the target driving scene. According to the embodiment of the invention, the user viscosity can be effectively improved.
Owner:SHANGHAI JIDOU TECH CO LTD

Cross-culture melody feature extraction and musical instrument matching system based on deep learning

The invention discloses a cross-culture melody feature extraction and musical instrument matching system based on deep learning, and belongs to the technical field of music information processing and artificial intelligence. The system comprises an audio acquisition module, an audio preprocessing module, a melody feature extraction module, a cultural context understanding module, an emotion semantic analysis module, a musical instrument timbre database, a musical instrument matching recommendation module, a distributor scheme generation module, a user interaction module and an audio synthesis module. Multi-dimensional melody features are extracted through a multi-scale convolutional neural network and a bidirectional long-short term memory network, cultural semantic understanding is realized through reasoning on a music knowledge graph by using a graph neural network, and musical instrument matching is performed by using a multi-objective optimization algorithm in combination with three dimensions of timbre integrating degree, cultural consistency and emotional expressive power. According to the system, automation of the whole process from humming melody to musical instrument configuration is achieved, the efficiency of the instrument is improved by 400%, the accuracy rate reaches 92%, and an intelligent tool is provided for cross-culture music creation and national music modernization reorganization.
Owner:SHENZHEN UNIV

Multi-language intelligent dubbing generation system based on sound cloning and emotion migration

The invention discloses a multi-language intelligent dubbing generation system based on sound cloning and emotion migration. The multi-language intelligent dubbing generation system comprises a sound cloning module, a cross-language synthesis module, an emotion migration module and a lip shape synchronization module. By constructing a few-sample speaker encoder, a cross-language rhythm migration module, a fine-grained emotion control module and a video lip shape synchronization module, end-to-end automatic generation from original dubbing audio to multi-language target dubbing is realized, and tone consistency, emotion authenticity and picture synchronism are kept.
Owner:JIANGSU HOPERUN SOFTWARE CO LTD

Editing method and device, equipment and storage medium

According to the embodiment of the invention, a method and device for editing, equipment and a storage medium are provided. The method comprises the following steps: in response to an operation of adding a target timbre on a timbre configuration interface, recording a reference audio for reading a reference text with the target timbre; in response to an operation of triggering audio generation on the editing interface, obtaining an input text for audio generation; in response to a selection operation on the target tone, generating a target audio based on the reference audio and the input text, the target audio including a voice for reading the input text with the target tone; and generating an audio editing result or a video editing result based on the target audio. In this manner, at least a portion of the target audio may be generated using timbres in the speech data. This may facilitate a user to use a desired tone in media content generation, thereby advantageously improving the efficiency of editing.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Real-time interaction method and terminal based on large language model

The invention discloses a real-time interaction method and terminal based on a large language model, and the method comprises the steps: obtaining user question information in a live broadcast interaction event in response to the live broadcast interaction event of a user, and extracting question keywords in the user question information; if the question keyword exists in a preset core keyword list, using a preset reply corresponding to the question keyword as reply content of the user question information, otherwise, using a large language model to generate the reply content of the user question information; and the reply content is converted into the voice reply consistent with the tone of the target sound by using the trained live broadcast voice model, so that the richness and flexibility of the reply content can be improved, and the interaction quality and efficiency are improved.
Owner:FUJIAN TQ ONLINE INTERACTIVE INC

Speech synthesis method and device, electronic equipment and storage medium

The invention provides a speech synthesis method and device, electronic equipment and a storage medium, and relates to the technical field of speech synthesis, and the method introduces a target attribute text in a speech synthesis process, can support speech synthesis with an audio attribute corresponding to the target attribute text, and improves the speech synthesis efficiency. Therefore, the expressive force and rhythm of the target synthetic speech can be controlled according to the user demand, so that the target synthetic speech better meets the user demand, and the user experience is improved. A speech synthesis model is obtained through training of a text sample with an attribute tag, so that the speech synthesis model has an audio attribute control capability during speech synthesis, and audio attributes such as tone, speech style, emotional expression, human settings, mood, rhythm and the like of target synthesis speech can be controlled at the same time; and audio attributes such as language switching, environment sound effect, dialect generation and the like can be supported, and controllability is improved while the generation quality of the speech synthesis model is ensured.
Owner:IFLYTEK CO LTD

Time-length-controllable end-to-end speech translation method and translation system

The invention discloses an end-to-end speech translation method and translation system capable of controlling duration, which effectively reduces the translation delay by introducing a speech end-to-end scheme. By constructing a brand new token and introducing a token alignment scheme, the length of a translation result is effectively controlled; by controlling the token duration and adding a length control variable, cross-language timbre and rhythm cloning is guided, so that relatively excellent cross-language timbre and rhythm cloning is achieved.
Owner:ZHILING WORKSHOP (HUZHOU) INTELLIGENT TECHNOLOGY CO LTD

Tone object separation method and device for audio data

The invention discloses a timbre object separation method and device for audio data. The method comprises the following steps: acquiring original video and audio data; if the subtitle data is received, cutting the original video and audio data according to the timestamp of the subtitle data to obtain a plurality of audio coarse slice segments; respectively extracting at least one audio fine slice segment from each audio coarse slice segment by adopting a sliding window with a preset size; extracting audio features from each audio fine slice segment and clustering the audio features, and determining an object tag corresponding to each audio fine slice segment; according to the invention, each object label is backfilled to the audio coarse slice segment, and the tone object corresponding to each audio coarse slice segment is determined, so that different tones are effectively separated in a plurality of main speaker scenes through a mode of determining object label backfilling in audio coarse cutting and fine cutting processes, the requirements in complex scenes are met, and the audio separation accuracy is improved.
Owner:SHANGHAI LINGGUANG ZHAXIAN TECHNOLOGY CO LTD

A voice conversion method, device, equipment and readable storage medium

The application provides a speech conversion method, device and equipment and a readable storage medium. The method comprises the following steps: obtaining speech information to be processed; based on a three-head encoder, encoding and modeling speech content, environmental noise and fundamental frequency information in the speech information to be processed respectively to obtain encoded and modeled speech information; changing the time sequence of the encoded and modeled speech information to adjust the speech speed of the encoded speech information; inputting the speech information with adjusted speech speed into a previously trained timbre conversion model corresponding to a target user to obtain target acoustic features, wherein the timbre of the target acoustic features is the same as the timbre of the target user. Thus, the timbre conversion of the speech information to be processed can be performed as required, and the conversion method is more efficient and accurate. Multi-dimensional encoding of the speech information can improve the robustness of the speech in a noisy environment. Speech speed control can make the speech more in line with user requirements.
Owner:MIGU CO LTD +1

Rhythm mode intelligent identification and video singing rhythm training system and method thereof

The invention discloses a rhythm mode intelligent identification and sightsinging rhythm training system and a method thereof, and relates to the technical field of music rhythm data processing, and the method comprises the steps: collecting multi-dimensional audio data played by a user, generating physical deviation data through time sequence comparison, judging a rhythm position and a timbre attribute in combination with sound intensity and frequency spectrum data, and carrying out the intelligent identification of the rhythm mode and the sightsinging rhythm training. Calculating deviation data by using the sensing weight coefficient; identifying a core rhythm problem based on the deviation data, and generating a personalized prescription exercise song in combination with the music style preference of the user; determining whether to add interference scene data or not according to a threshold value, generating interference music segments through an algorithm composition engine, and mixing and outputting the interference music segments; the interference intensity and the rhythm density are dynamically adjusted by comparing the change of deviation data before and after the change of the simulated exercise curve. According to the invention, personalized dynamic adjustment and closed-loop optimization of rhythm training are realized, and rhythm evaluation accuracy and training effect are effectively improved.
Owner:SHANDONG PETROCHEMICAL INST

Detachable guitar

ActiveCN224177095UReduce local stressimprove integrityGuitarsNoiseMechanical engineering
The utility model discloses a detachable guitar, and relates to the technical field of guitar structures, and the detachable guitar comprises a guitar body and two handrails. The first ends of the two handrails are connected with the guitar body, and the second ends of the two handrails are directly connected to form a closed frame structure, so that on one hand, vibration can be more smoothly conducted to the guitar body through the continuous frame structure to improve the tone consistency, and on the other hand, resonance at a specific frequency can be inhibited, noise is reduced, the tone is purer, and the tone quality is improved. The closed frame structure can effectively disperse external force, reduce local stress of the guitar body and reduce the deformation risk, on the other hand, the second ends of the two handrails are connected to the lower end of the guitar body through the threaded connecting piece after being stacked, compared with the mode that the two handrails are independently connected to the guitar body, the stacking design can reduce the number of open holes, and the number of the open holes is reduced. The completeness of the guitar body is protected.
Owner:CHANGSHA GANYIN TECHNOLOGY CO LTD

Music understanding model training method, audio processing method, equipment and medium

The invention discloses a music understanding model training method, an audio processing method, equipment and a medium, which are applied to the technical field of computers, and comprise the following steps: compressing a song audio through an initial music understanding model to obtain an audio compression representation; the audio compression representation comprises any one or more of timbre information, emotion information and semantic information; performing coding processing on the pitch sequence of the song audio through the initial music understanding model to obtain melody compression representation containing melody information; constructing a fusion compression representation based on the audio compression representation and the melody compression representation; and optimizing the initial music understanding model based on the fusion compression representation and the audio compression representation to obtain a music understanding model. According to the method, the pitch track of the melody is reserved through the pitch sequence, redundant information is directly stripped, semantic focusing of melody characteristics is achieved, melody compression representation is obtained based on the pitch sequence, the characteristic of strong association of the music before and after is reserved, and the method is more suitable for characteristic learning of the music melody.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD