Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22results about How to "Improve expressiveness" patented technology

Artificial intelligence-based speech synthesis method and device, computer equipment and medium

The application is suitable for the technical field of speech synthesis, and particularly relates to a speech synthesis method and device based on artificial intelligence, computer equipment and a medium. The application extracts a text feature vector of a target text through a feature extraction model, predicts the text feature vector through a stress predictor, outputs a stress prediction vector, adds the stress prediction vector to the text feature vector to obtain a text stress vector, predicts the text feature vector through a pause predictor, outputs a pause prediction vector, adds the pause prediction vector to the text feature vector to obtain a text pause vector, predicts the text stress vector and the text pause vector through a prosody predictor, outputs a text prosody vector, matches the text prosody vector with a phoneme sequence of the target text, obtains a phoneme sequence with prosody labels, performs speech conversion on the phoneme sequence with prosody labels, obtains synthesized speech, and through the prediction of stress, pause and prosody, the expressiveness, naturalness and accuracy of the synthesized speech are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Page interaction method and device, electronic equipment, storage medium and program product

This disclosure relates to page interaction methods, devices, electronic devices, storage media, and program products, including displaying media content corresponding to target text; when the media content is triggered, displaying a target page corresponding to the target text, the target page being used to display the content corresponding to the target text, and the target page being in video playback mode; and playing a target video corresponding to the target text on the target page in video playback mode. This disclosure can improve the efficiency and reach of information dissemination in text.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Vehicle sound field control method and device, computer equipment and storage medium

PendingCN121940688AImplement targeted offsetsStable and clear sound field experienceSignal processingTransducer circuitsSound energyEngineering
The invention relates to the technical field of vehicle control, and discloses a sound field control method and device of a vehicle, computer equipment and a storage medium, the method comprises the following steps: in response to a scene mode trigger signal, obtaining a first real-time opening condition of a vehicle door and / or a vehicle window, determining a target sound field projection area in the vehicle based on the first real-time opening condition, generating an audio processing parameter corresponding to a loudspeaker in the vehicle; the audio signal is processed according to the audio processing parameters, and the loudspeaker is driven to project sound energy corresponding to the processed audio signal to the target sound field projection area. The target sound field projection area outside the vehicle is determined by sensing the real-time opening state of the vehicle door / window and combining the preset scene mode, and the loudspeaker parameters are adjusted; therefore, the sound energy is guided to the position of the user, the problem that the sound field in the vehicle leaks or is poor in effect when the vehicle door and the vehicle window are opened is effectively solved, and the audio experience during outdoor use is improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Violin string tension adjusting control device

The invention belongs to the technical field of string adjustment, and particularly relates to a violin string tension adjustment control device which comprises an adjustment chamber, adsorption units are symmetrically arranged at one end of the adjustment chamber, positioning units are arranged on the sides, close to the adjustment chamber, of the adsorption units, an adjustment unit is arranged on one side in the adjustment chamber, and a feedback unit is arranged on one side of the adjustment unit. Through the lifting effect of the electric telescopic rod and the face-to-face or back-to-back movement between the idler wheels, lifting and sudden release of strings are achieved, then radial jumping of the strings is degraded and buffered through the multi-degree-of-freedom string roller, dominant impact between the lower electrode cap and the electrode ball is controlled through the prying roller short section, and therefore the electrode ball can be driven to move stably and stably. The high-frequency impact state between the strings and the long section of the prying roller is expressed, that is, impact feedback of the strings in a low-frequency state is explicitly and prominently expressed through an extended range, the change of tension is accurately sensed, excessive or insufficient adjustment is avoided, the overall tone balance is optimized, and the harmony degree and expressive force of the tone of the violin are improved.
Owner:MENGZHOU INTELLIGENT TECHNOLOGY (NANTONG) CO LTD

Guzheng soundboard bending and shaping die device

The utility model discloses a Chinese zither soundboard bending and shaping die device which comprises a bending die and a shaping structure. The bending die is composed of an upper die and a lower die, a cavity matched with the bent shape of the zither soundboard is formed between the upper die and the lower die, and the bending die is used for applying uniform bending force to the heated and softened zither soundboard to form the needed bent shape. The shaping structure comprises tensioning belts and a pressing device, the tensioning belts are arranged on the two sides of the bent zither soundboard, the pressing device presses the tensioning belts, and the bent shape of the soundboard is kept. According to the device, the heating, bending and shaping processes are accurately controlled, the machining precision and efficiency are remarkably improved, and the consistency of the sound quality effect of the soundboard is ensured. Besides, the device has the advantages of being high in adaptability, easy to maintain and upgrade, environment-friendly, energy-saving and the like, can be adjusted according to the shape, size and material of the zither soundboard, and meets the processing requirements of different types and specifications of soundboards.
Owner:NORTHWEST A & F UNIV

A method and apparatus for synthesizing audio recordings with different emotions

ActiveCN115762466BEmotionalImprove expressivenessDigital data information retrievalBiological models
This invention provides a method and apparatus for synthesizing audio with different emotions, including a training phase and an inference phase. The training phase includes the following steps: S11, collecting training corpora, including audio from different speakers and corresponding text, as well as emotion tags, and extracting the spectral features of the corresponding speech; S12, training an emotion speech feature extraction model based on the speech spectral features and corresponding emotion tags; S13, extracting the emotion feature vector of the training corpora and the text encoding vector of the corresponding training corpora; S14, combining the text encoding vector with the emotion feature vector of the speech, and training a speech synthesis model through the acoustic features of the corresponding speech; S15, training an emotion feature prediction model using the emotion feature vector of the speech and the text encoding vector as input; S16, training a vocoder through the acoustic features of the speech and the corresponding speech. This invention solves the problem of bland tone and unclear emotion in traditional speech synthesis.
Owner:四川启睿克科技有限公司 +1

Chinese character to speech method based on multi-modal large model offline construction of unit library

This invention provides a Chinese text-to-speech method based on an offline unit library constructed using a multimodal large model, relating to the field of text-to-speech technology. The method includes the following steps: receiving the target text to be synthesized; parsing the target text for scene tags, sentiment tags, and stress positions using a lightweight semantic parsing model; querying a pre-constructed offline unit library based on the parsing results, and retrieving corresponding candidate acoustic prosodic units for each Chinese character or word; constructing a state network with the candidate acoustic prosodic units as nodes; calculating the optimal path based on the inter-unit transition cost and semantic matching degree in the offline unit library using the Viterbi algorithm, and selecting the globally optimal acoustic prosodic unit sequence; and smoothly concatenating the candidate acoustic prosodic units in the optimal acoustic prosodic unit sequence to generate and output the target speech.
Owner:BEIJING FUNSHION ONLINE TECH LTD +1

Multi-subwoofer sound channel mapping method and device

The invention discloses a multi-subwoofer sound channel mapping method and device, and relates to the technical field of intelligent automobiles, and the method comprises the steps: obtaining a playing mapping relation; under the condition that the audio comprises a subwoofer sound channel, controlling the playing device to play the audio based on the playing mapping relation; wherein the playing device comprises a plurality of subwoofer sound channels arranged in different areas of a playing space where the playing device is located, and each subwoofer sound channel is correspondingly provided with at least one loudspeaker; the playing mapping relation comprises that the subwoofer sound channel of the audio is mapped to at least two subwoofer sound channels of the multiple subwoofer sound channels. According to the method, the problem that mega bass expressive force in different areas is not uniform when the audio is played can be solved.
Owner:BEIJING CO WHEELS TECH CO LTD

Multi-modal expression generation system and dynamic optimization method in virtual-real fusion scene

PendingCN121999100AEnsure strict consistencyEliminate microscopic delays in audio and videoTelevision system detailsCharacter and pattern recognitionEngineeringInteraction technology
The invention belongs to the technical field of computer graphics and human-computer interaction, and particularly discloses a multi-modal expression generation system and a dynamic optimization method in a virtual-real fusion scene, and the method comprises the steps: analyzing an audio rhythm to generate a phase reference signal, driving visual collection and calculation, and achieving the sound and picture synchronization of a microscopic time domain; calculating a dense optical flow field, and decoupling a facial muscle movement change vector through a vector dot product by using a muscle movement direction template; and mapping the audio acoustic feature into an audio driving energy value representing the sound production intensity. An energy conservation attenuation or compensation enhancement correction is performed on the expression vector by comparing the total amount of visual displacement with the audio drive energy to generate a corrected expression vector. And finally, superposing the corrected expression change vector to the state parameter of the previous frame, and generating a control instruction to drive the virtual avatar to render in real time. According to the method, the problems of sound and picture time sequence dislocation and physical constraint missing are solved, and the sense of reality and expressive force of virtual expressions are improved.
Owner:SHENZHEN XINGHUO MUTUAL ENTERTAINMENT DIGITAL TECH CO LTD

Digital human emotional speech generation method based on emotional semantic modeling

PendingCN122369428AImplementing fine-grained emotion modelingAccurately identify differential contributionsGraph neural networksSemantic feature
This invention relates to the field of artificial intelligence technology, specifically to a digital human emotional speech generation method based on emotion semantic modeling. The method includes multimodal data preprocessing, multimodal semantic feature extraction, emotion semantic unit partitioning, graph-based emotion interaction, speech expression parameter generation, emotional speech generation, and digital human emotion collaborative driving. By constructing emotion semantic units and introducing an emotion contribution evaluation mechanism, fine-grained emotion modeling of the speech semantic structure is achieved, accurately identifying the differential contributions of different semantic segments in a sentence to the overall emotional expression, thereby significantly improving the accuracy and controllability of emotional speech generation. Furthermore, by constructing a dual-graph structure of emotion interaction graph and emotion evolution graph, and combining graph neural networks to jointly model the interaction relationships and evolutionary processes between emotion semantic units, a structured expression of complex dynamic emotional changes is achieved, enhancing the generated speech's ability to exhibit emotional coherence and naturalness.
Owner:BEIJING XILIANLIAN TECHNOLOGY CO LTD

Coal mine image segmentation model, method and construction method based on VMamba and multi-expert hybrid network

ActiveCN121639706BImprove feature extractionFast and precise extractionAlgorithmFeature learning
The application discloses a coal mine image segmentation model and method based on a VMamba and multi-expert hybrid network and a construction method thereof. The coal mine image segmentation model is constructed. The image block encoding layer output of an encoder is taken as the input of the first encoding layer of a first VSS network, and the output of the first encoding layer of the first VSS network is taken as the input of the output end feature learning module. The input of the output end feature learning module is taken as the input of the first decoding layer of a decoder, and the output of the first decoding layer of the decoder is taken as the input of the second decoding layer of the decoder. The input of the second decoding layer of the decoder is taken as the input of a segmentation head, and the output of the segmentation head is taken as a segmented image. The application can significantly improve the segmentation precision and calculation efficiency of the coal image.
Owner:SHANGHAI XINLIJI SEMICON CO LTD

Training methods and devices, electronic equipment and storage media for lesion detection models

This application provides a training method, apparatus, electronic device, and storage medium for a lesion detection model, belonging to the field of image processing technology. It involves acquiring a first medical image and a second medical image, where the first image is unlabeled and the second image is labeled. A preset teacher model is used to detect lesions in the first medical image, and a preset student model is used to detect lesions in both the first and second medical images. Based on the lesion detection results, a loss calculation is performed to obtain target loss data. The first network parameters of the preset student model are adjusted based on the target loss data, and the second network parameters of the preset teacher model are adjusted based on the first network parameters to train the preset student model, thus obtaining a lesion detection model. This lesion detection model is then used to detect lesions in the target medical image, obtaining the location and type of the target lesion, thereby improving the accuracy of lesion detection.
Owner:SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Data organization and time series prediction method and apparatus based on spatiotemporal knowledge graph

This application provides a data organization and time series prediction method and apparatus based on spatiotemporal knowledge graphs, which can be applied to the fields of knowledge graph and time series prediction technology. The method includes: constructing a spatiotemporal knowledge graph based on multi-source heterogeneous data, wherein the node types of the spatiotemporal knowledge graph include at least entity nodes and event nodes; analyzing and representing the nodes, relationships, and spatiotemporal relationships of the spatiotemporal knowledge graph to dynamically construct a spatiotemporal graph network; and training a graph neural network model based on the spatiotemporal graph network for target time series prediction. This method not only improves the accuracy of the model in predicting target time series within future periods but also enhances the interpretability of the results.
Owner:AEROSPACE INFORMATION RES INST CAS

Digital human expression interaction generation method and device based on organizational dynamics law

The embodiment of the application discloses a digital human expression interaction generation method and device based on the law of organizational dynamics, comprising: extracting the rhythm features and semantic features of the to-be-processed voice content; combining the personalized expression base and the weighting coefficient through the linear skinning formula to generate linear expression features, wherein the personalized expression base is generated by combining the initial expression base with the individual feature vector, and the weighting coefficient is generated by encoding the rhythm features through a voice time sequence feature encoder; the semantic features and the linear expression features are extracted through a conditional prior encoder, and are jointly input into a preset diffusion model as a condition to generate nonlinear expression features and virtual skeleton control parameters; the nonlinear expression features and the virtual skeleton control parameters are input into a preset reinforcement learning model, and through left-right symmetry constraint, the final generated expression result is generated, and the final generated expression result is the digital human expression corresponding to the to-be-processed voice content.
Owner:SHENSTRONTIUM TECH (BEIJING) CO LTD +1

High-resolution image inpainting method, device, and storage medium

The application discloses a high-resolution image repairing method and device and a storage medium, relates to the technical field of computer vision, and comprises the following steps: performing structural feature extraction on a damaged image to obtain a corresponding edge graph and a line frame graph; determining a multi-channel input tensor based on the damaged image, the edge graph, the line frame graph and a binary mask graph; inputting the multi-channel input tensor into a high-resolution structure repairing network to generate a high-resolution edge graph and a high-resolution line frame graph; performing fusion processing on the high-resolution edge graph, the high-resolution line frame graph, a mask image and the binary mask graph to obtain a structural visual fusion feature graph; inputting the structural visual fusion feature graph into a structure-enhanced texture repairing network to generate a high-resolution enhanced feature graph, and generating a high-resolution repaired image according to the high-resolution enhanced feature graph, so that the structural reconstruction capability of high-resolution repairing is improved, the texture repairing quality is enhanced, and a repaired result with reasonable structure and consistent texture semantics is generated.
Owner:SHENZHEN SHIXI TECH CO LTD

Artificial intelligence-based voice conversion method and device, computer device and medium

The application is suitable for the field of financial technology, and particularly relates to a voice conversion method and device based on artificial intelligence, computer equipment and medium.The application extracts optimized semantic features and optimized prosodic features of the voice to be converted and reference speaker features of the reference voice through an encoder, uses a decoder to obtain target converted voice, extracts first semantic features, first prosodic features and first speaker features of the voice to be converted and second semantic features and second prosodic features of the augmented voice through the encoder, inputs the first semantic features, the first prosodic features and the first speaker features into the decoder to obtain reconstructed voice, and calculates model loss to train the encoder and the decoder, thereby improving the accuracy of the encoder and the decoder, improving the accuracy of voice conversion, providing customers with robot customer service with higher naturalness, expressiveness and richness in the field of financial technology, and improving service quality and customer experience.
Owner:PING AN TECH (SHENZHEN) CO LTD

Down feather identification method and system based on deep learning

The invention discloses a down feather recognition method and system based on deep learning, and the method comprises the steps: S1, obtaining an original down feather image, carrying out the preprocessing of the obtained original down feather image, and obtaining a processed down feather feature map; s2, the obtained feature map is input into an improved residual network WT-ResNet model, the WT-ResNet model at least comprises a WT-ResNet layer, and the WT-ResNet layer processes the input feature map and then outputs a final feature map containing deep semantic information; and S3, inputting the output final feature map into a classification activation layer, obtaining a probability value of down feather image classification through activation function processing, and judging whether the down feather image is fresh down feather or recycled down feather based on the probability value.
Owner:CHINA JILIANG UNIV +2

Digital human generation method and device based on multi-modal large model

ActiveCN120107427BGenerate delicate and naturalgenerate flexibleCharacter and pattern recognitionAnimationNoiseIntent recognition
The application provides a digital human generation method and device based on a multi-modal large model, relates to the technical field of digital human generation, breaks through the generation limitation of traditional key points or three-dimensional representation based on posture, and can generate a digital human image that is natural and delicate in target parts. The method comprises the following steps: acquiring multi-modal data input by a user and performing intention recognition and emotion analysis based on a multi-modal large model to determine audio sequence data corresponding to response text data; acquiring visual feature vector representation of reference image data; determining mask feature vector representation corresponding to the target part according to the position of the target part of the character in the reference image data; performing denoising processing on at least one noise vector representation based on a diffusion model according to the audio sequence data, the visual feature vector representation and the mask feature vector representation to generate at least one frame of image data of a digital human; and generating a digital human animation with voice according to the at least one frame of image data of the digital human and the audio sequence data.
Owner:ULTRAPOWER SOFTWARE

A display control circuit

This utility model relates to the field of IoT intelligent control technology, and in particular to a display control circuit applied to a smart cloud box device. It includes a CPU module, a WIFI / Bluetooth module, a display module, an audio codec module, an audio amplifier module, a transceiver module, a power supply module, and a step-down module. The WIFI / Bluetooth module, display module, audio codec module, audio amplifier module, transceiver module, and step-down module are all electrically connected to the CPU module. The CPU module uses an RK3399 processor chip, the display module uses a TC358772XBG display bridge chip, and the step-down module is also electrically connected to the power supply module. Through the organic combination of multiple modules, this display control circuit comprehensively solves the problem of insufficient display functionality in smart cloud boxes, improving display effect, response speed, and system compatibility.
Owner:SHENZHEN DATAMAX TECHNOLOGY CO LTD

Image defogging method based on hybrid large-scale convolution and attention fusion

ActiveCN118552442BImprove expressivenessquality improvementPattern recognitionComputer vision
The application discloses a method for image defogging based on a mixed large-scale convolution attention mechanism, and the implementation steps are as follows: a large-scale convolution module for extracting multi-scale information of a feature image is constructed, a parallel attention module for extracting shared global information and position-related local information of an original feature is constructed, a mixed large-scale convolution network is built, the mixed large-scale convolution network is trained by using a generated training set, an image to be defogged is input into the trained mixed large-scale convolution network, and a haze-free image is output. The application overcomes the defects of the prior art, such as incomplete defogging of non-uniform fog images, insufficient learning of scene characteristics, and loss of image details and color distortion caused by neglecting multi-scale information of the image. The mixed large-scale convolution network can still achieve a high defogging effect in complex scene defogging.
Owner:XIDIAN UNIV

Face and gesture recognition method and product fusing attention mechanism

The application relates to a face and gesture recognition method and product with a fusion attention mechanism. The method comprises the following steps: inputting a camera image into a CenterNet network to generate a face image region of interest and a hand image region of interest; inputting a face image in a database into a twin neural network with a fusion channel and spatial attention mechanism to generate a face image feature, and inputting the face image region of interest into another twin neural network to generate a face image feature; comparing the face image feature and the face image region of interest feature to recognize a face and generate a face recognition result; inputting the hand image region of interest into a ResNet network based on a multi-scale fusion mechanism to perform semantic segmentation and generate a hand binary image; inputting the hand binary image into a classification network to generate a hand recognition result; and controlling an intelligent wheelchair according to the face recognition result or the hand recognition result. The application can improve the efficiency and precision of face and gesture recognition.
Owner:BEIJING SCI & TECH HUAHUI INTELLIGENT TECH CO LTD

Video generation method and device, electronic equipment and storage medium

The invention provides a video generation method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to technologies such as computer vision, video content understanding, automatic video editing and intelligent media synthesis. According to the specific scheme, the method comprises the following steps: extracting a plurality of candidate video clips containing each target object from an original video based on a source video identifier and reference information of at least two target objects; performing quality screening on the candidate video clips of each target object, and combining the screened clips into a plurality of clip sets according to a preset matching rule; for each segment set, based on the face position information of each video segment in the segment set, respectively performing vertical clipping on each video segment, and synthesizing the clipped video segments according to a preset multi-split-screen layout to obtain a synthesized frame sequence; and combining the synthesized frame sequence with the target audio data to generate a vertical split-screen video.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD