Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2782 results about "Target text" patented technology

A target text (TT) is a translated text written in the intended target language, which is the result of a translation from a given source text. According to Jeremy Munday's definition of translation, "the process of translation between two different written languages involves the changing of an original written text (the source text or ST) in the original verbal language (the source language or SL) into a written text (the target text or TT) in a different verbal language (the target language or TL)". The terms 'source text' and 'target text' are preferred over 'original' and 'translation' because they do not have the same positive vs. negative value judgment.

Model training method and device, equipment, storage medium and product

The invention relates to a model training method and device, equipment, a storage medium and a product. The method comprises the following steps: constructing a training data set according to a target text reasoning chain obtained by converting multi-modal data; according to the training data set, performing supervision fine tuning on the pre-trained multi-modal large language model to obtain a basic reasoning model; performing optimization processing on the basic reasoning model according to reinforcement learning training of long thinking to obtain a target reasoning model; the target reasoning model is used for outputting a target answer containing a reasoning process according to the input multi-modal data. Therefore, the long text constraint can be directly used for reinforcement learning, and the training efficiency is greatly improved; and by adopting long-thinking reinforcement learning training, the model can easily learn a correct thinking process in training, so that the reasoning ability of the multi-modal large language model for processing a complex visual reasoning task is improved, and the correct thinking process is displayed in the reasoning process.
Owner:SHUXING TECH (BEIJING) CO LTD

Cross-modal image-text analysis method for machine vision

The invention relates to the technical field of machine vision, and discloses a machine vision-oriented cross-modal image-text analysis method, which comprises the following steps of: partitioning an input image to generate an image block sequence; inputting the image block sequence into a visual converter for multi-scale feature extraction, and generating target visual features; encoding the input text to generate a target text feature; inputting the target visual features and the target text features into a deep reconstruction bottleneck network for compression alignment, and generating a cross-modal compression vector; and inputting the cross-modal compression vector into a large language model to generate cross-modal decoding information, so that cross-modal redundant information can be effectively filtered, compact shared semantic representation can be learned, the information integrity of the compression process is ensured through bidirectional reconstruction verification, cross-modal semantic alignment is realized, and the method has the advantages of high efficiency and high reliability. Omnibearing cross-modal content generation from the whole to details is achieved, and the requirements of different application scenes are met.
Owner:SHENZHEN YOULIANCHUANG WISDOM TECH CO LTD

Audio analysis method and device based on feature fusion, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of medical health, financial science and technology, culture and art and the like, and discloses a feature fusion-based audio analysis method, which comprises the following steps of: obtaining a target audio in a target field and a target text associated with the target audio, extracting a music feature vector of the target audio, extracting a text semantic feature vector of the target text, fusing the music feature vector and the text semantic feature vector to generate fusion features, constructing a knowledge graph containing knowledge nodes of the target domain, inputting the fusion features into the knowledge graph for semantic analysis, and generating an analysis result. According to the method, deep semantic analysis is performed by fusing audio and text features and combining a knowledge graph, so that multi-modal understanding of audio data is realized, the accuracy and interpretability of an analysis result are improved, and the relevance between the analysis result and industry knowledge is enhanced; therefore, the applicability of the audio analysis technology in the fields of culture and art, medical health, finance and the like is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Knowledge graph construction method and device and medium

The invention discloses a knowledge graph construction method and apparatus, and a medium. The method comprises the steps of collecting a to-be-processed text for constructing a knowledge graph in a specified field; screening out a target text from the to-be-processed text, wherein the correlation between the target text and the specified field meets a preset condition; processing the target text through a pre-constructed target large language model to construct an initial triple; filtering the initial triple to obtain a target triple; and constructing the knowledge graph through the target triple. Therefore, through strict correlation screening conditions, texts irrelevant to the specified field are removed, the influence of noise data on knowledge graph construction is reduced, and the quality and accuracy of the knowledge graph are improved. On the basis of the target texts, the knowledge triples are efficiently and accurately extracted from a large number of target texts by means of the powerful language understanding and generating capacity of the large language model, the labor cost overhead of manual extraction is reduced, and the knowledge graph construction efficiency is improved.
Owner:ZHEJIANG LAB

Question answering method based on large model, and electronic device

The present application relates to the technical field of artificial intelligence, and in particular to a question answering method based on a large model, and an electronic device. The electronic device comprises a communication interface, a memory, and a processor, the memory is configured to store a computer instruction, and the processor is configured to execute a computer program to enable the electronic device to: determine, on the basis of each pre-stored text block, a target text block matched with a question to be answered; obtain, if non-text data including a picture and / or a table is pre-stored for the target text block in advance, stored summary text of the non-text data; and input the question to be answered, the target text block, and the summary text into a first large model on the basis of a preset format to obtain answer information. Since the summary text is text that summarizes all content in the non-text data, the first large model does not need to process the picture and / or the table and still can ensure that the content recorded in the picture and / or the table is taken into consideration during generation of the answer information, thereby improving the model-based question answering accuracy.
Owner:HISENSE GRP HLDG CO LTD

Text error correction method and device, electronic equipment and storage medium

The embodiment of the invention relates to the technical field of natural language processing, and provides a text error correction method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a to-be-corrected target text, and carrying out the semantic boundary recognition, and obtaining all semantic boundaries; segmenting the target text into a plurality of clauses, distributing IDs according to the appearing sequence of each clause, and generating structured data with sequence IDs; inputting the structured data into a language model, and calling a target domain knowledge base corresponding to the target text as a constraint condition to perform semantic analysis, and identifying an error type; adjusting an error correction path based on the error type so as to re-identify error information, the error information including error content, the error type and the error position; and recombining a plurality of clauses based on the sequence ID, the error content, the error type and the error position, and generating an error correction result containing the error position, the error type and the correction suggestion. Therefore, collaborative detection and correction of grammar, semantic and logic errors are realized.
Owner:北京观微科技有限公司

Detecting jailbreak attempts on generative models

A computer-implemented method is provided that detects jailbreak attempts against generative models, which may involve a shift between benign and malicious content. The method includes determining a probability-based metric for each of a plurality of tokens in a target text using a language model, the probability-based metric being based on a probability at least one preceding token. The probability-based metrics are processed to identify a subset of the plurality of tokens having a change in the probability-based metric with respect to others of the plurality of the tokens not within the subset of the plurality of tokens. A jailbreak attempt in the target text is detected in response to identifying the change in the probability-based metric in the subset of the plurality of tokens.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Relation extraction method and system based on graph neural network

The invention discloses a relation extraction method and system based on a graph neural network, and belongs to the technical field of natural language processing. A target text is obtained, word segmentation, part-of-speech tagging and named entity recognition are carried out, and an entity set is extracted; constructing a text graph structure containing multiple edge types based on the entity set; performing feature coding on nodes in the graph to generate an initial feature vector fusing semantic, part-of-speech and position information; inputting the graph into the graph neural network model, and obtaining high-order node representation through multi-layer message passing and aggregation; modeling the entity pair in combination with the structure path and the context information, and inputting a multi-channel classification network to predict the relationship type of the multi-channel classification network; and finally, outputting an entity relationship triple according to a prediction result. The method has stronger semantic modeling ability and structure expression ability in a relation extraction task, and is suitable for scenes such as knowledge graph construction and information extraction systems.
Owner:CHANGCHUN GUANGHUA UNIV

Vertical domain relation extraction method and device, electronic equipment and storage medium

The invention relates to a vertical domain relation extraction method and device, electronic equipment and a storage medium, and the method comprises the steps: constructing a knowledge base of a target vertical domain, carrying out the entity relation extraction of a target text through combining with the knowledge base of the target vertical domain, and obtaining a first entity relation set; performing entity relationship extraction on the target text by utilizing a trained entity relationship extraction model to obtain a second entity relationship set, and performing correction processing on the second entity relationship set by utilizing a knowledge base to obtain a fourth entity relationship set, and screening out a target entity relationship set from the first entity relationship set and the fourth entity relationship set in combination with the knowledge base of the target vertical field. By introducing the vertical field knowledge base, external data support is provided for a large model, the defects of the prior art in the aspects of long-tail relation processing and vertical field reasoning capability are overcome, and meanwhile, the filtering capability of noise samples is enhanced and the extraction quality and efficiency are improved by combining a combined comparison screening mechanism, a knowledge base automatic updating mechanism and a health management mechanism.
Owner:BEIJING HOLARDATA TECH CO LTD

Task planning method and device for large language model, storage medium and equipment

The invention discloses a task planning method and device for a large language model, a storage medium and equipment. The method comprises the steps that firstly, a target field to which a target task text belongs is determined; then, M text blocks similar to the target task text are screened out from the knowledge document of the target domain; inputting the key information of the M text blocks into a preset large language model in combination with the first prompt on the basis of a multi-dimensional prompt text abstract generation mode to obtain a text abstract of the M text blocks, and inputting the text abstract of the M text blocks into the large language model in combination with the second prompt to obtain N subtasks corresponding to the target text. According to a pre-constructed tool description mapping relation, tools for executing the N sub-tasks are called through a large language model, and processing results of the N sub-tasks are obtained by executing the tools so as to form a final processing result of the target task. Therefore, the overall efficiency and quality of executing the target task by the large language model are improved.
Owner:HEFEI IFLY DIGITAL TECH CO LTD

Voice interaction task execution method and device based on large model, equipment and medium

The invention discloses a voice interaction task execution method and device based on a large model, equipment and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: carrying out the voice recognition of voice features, obtaining a text character sequence, and carrying out the optimization of the text character sequence through a preset large language model, and obtaining a target text character sequence; performing entity recognition on the target text character sequence by using a preset large language model to obtain an entity recognition result, determining a relationship type among entities in the target text character sequence according to the entity recognition result, and constructing a knowledge graph according to the relationship type among the entities; generating an initial triple based on the knowledge graph, and optimizing the initial triple by using a preset large language model to obtain a target triple; and fusing the target triple with the initial knowledge graph, and executing a voice interaction task in the target voice interaction scene based on the updated knowledge graph. According to the method and the device, the efficiency and the accuracy of extracting the structured knowledge from the Chinese speech are improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Tibetan language multi-dialect real-time semantic conversion method based on cross-language BERT model

The invention discloses a Tibetan language multi-dialect real-time semantic conversion method based on a cross-language BERT model, and the method comprises the following steps: S1, collecting original text corpora of each dialect of the Tibetan language, and constructing a standardized training corpus set; s2, performing parameter initialization on the mBERT model, and preliminarily training the mBERT model; s3, constructing a semantic modeling model, and performing fine adjustment on the semantic modeling model; s4, receiving to-be-converted text input, and encoding; s5, obtaining an intermediate semantic representation vector of the text through a semantic coding sub-module; s6, inputting the intermediate semantic representation vector into a semantic generation sub-module, and generating target text output; and S7, executing syntactic consistency correction and language fluency correction. According to the method, mBERT modeling and an adversarial optimization mechanism are fused, real-time semantic consistency conversion of multiple dialects of the Tibetan language is achieved, and the method has the advantages of being high in accuracy, high in robustness and low in response delay.
Owner:TIBET MIRAN EDUCATION TECH CO LTD

Depth map generation method and device based on large model, three-dimensional reconstruction method and device, electronic equipment and storage medium

The invention provides a depth map generation method and device based on a large model, a three-dimensional reconstruction method and device, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, can be applied to real-time road scene depth perception, environment three-dimensional reconstruction and obstacle avoidance, and can be applied to real-time road scene depth perception. And virtual and real scene fusion and other scenes can be realized. The specific implementation scheme is as follows: performing visual coding on a monocular image to obtain a coded image; inputting the coded image and the target text into a pre-trained large language model for fusion to obtain fusion features; generating global guide features based on the fusion features, wherein the global guide features comprise joint semantic information of visual features and text features; adding noise to the color image of the monocular image to obtain a noise feature sequence; de-noising the noise feature sequence under the condition of the global guide feature, and generating an implicit feature matched with the joint semantic information; a depth map is generated based on the implicit features.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Method for editing facial image attributes based on text of diffusion model

The invention provides a method for editing facial image attributes based on a text of a diffusion model. The method realizes flexible facial editing with high quality and identity consistency. The method comprises the following steps: constructing a description control face diffusion model comprising a noise prediction network and a variational auto-encoder, wherein the noise prediction network comprises a plurality of text alignment face Transform modules and an optimization residual feature module; the method comprises the following steps: inputting an original image and target text description based on a pre-training model, optimizing a target embedding generated by the input target text description by a text alignment face Transform module to obtain an optimized embedding, and performing model fine tuning after optimization is completed; and performing linear interpolation on the target embedding and the optimized embedding to obtain an initial editing result, introducing ArcFace-Loss as identity loss, extracting face features of an original image and the initial editing result through a pre-training model, calculating feature similarity and minimizing the loss, and ensuring consistency of face identities after editing.
Owner:DALIAN NATIONALITIES UNIVERSITY

Video generation method, electronic device, and computer readable storage medium

Provided are a video generation method, an electronic device, and a computer readable storage medium, relating to the technical fields of computers and video processing. The video generation method comprises: acquiring a target text, wherein the target text is used for describing video content to be generated (S21); and using a target video generation model to perform video generation processing on the target text to obtain a target video, wherein the target video generation model is a model obtained by performing model alignment on an initial video generation model and a preset reward model in a fine-tuning manner (S22). The present invention solves the technical problems in the related art that a video generated by a video generation model obtained by training based on network data has poor quality and does not meet user expectations.
Owner:ALIBABA (CHINA) CO LTD

Machine generated text detection method, terminal, medium and program product

The invention discloses a machine generated text detection method, a terminal, a medium and a program product in the field of natural language processing, and the method comprises the steps: inputting a to-be-detected target text into a trained text detection model, and obtaining a detection result outputted by the text detection model; the text detection model comprises a first semantic feature extraction network, a second semantic feature extraction network and a fusion classification network; the first semantic feature extraction network is used for extracting semantic features of keywords in a target text to generate a first representation vector; the second semantic feature extraction network is used for extracting overall semantic features of the target text and generating a second representation vector; the fusion classification network processes the first representation vector and the second representation vector to generate a detection result about whether the target text is a machine-generated text; the training of the text detection model adopts a multi-task learning strategy. According to the method, multi-task training is introduced, so that the accuracy and robustness of machine generated text detection are effectively improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Electronic archive intelligent retrieval method and system

The invention discloses an intelligent retrieval method and system for electronic archives, and belongs to the field of electric digital data processing.The intelligent retrieval method for the electronic archives comprises the following steps that search content input by a user is received, spelling errors are automatically corrected, and related words are expanded; preliminarily screening the candidate text electronic archive set based on a reverse index mechanism to obtain an initial sorting result, performing correlation resorting and extension recall on the initial sorting result through a semantic similarity model, and outputting a final sorting result; and after determining the target text electronic archive in the final sorting result, according to the uploading information of the target text electronic archive, the method has the beneficial effects that the non-text electronic archives are matched through the determined uploading information of the target text electronic archive, and the non-text electronic archives are sorted according to the association strength, so that the efficiency of sorting the non-text electronic archives is improved. The non-text format content can be identified during keyword retrieval; and the tendency of the user is predicted, so that the use experience of the user in downloading or checking the electronic file is improved.
Owner:SHANDONG YANGMING NETWORK INFORMATION TECHNOLOGY CO LTD

Three-dimensional digital human generation method and system capable of voice interaction

The invention belongs to the technical field of three-dimensional reconstruction, and discloses a three-dimensional digital human generation method and system capable of voice interaction. According to the invention, brand new speaking audios in different languages are automatically generated according to different languages of the input target text and the sampled human voice audios; the sequential stability and detail reduction capability of three-dimensional human motion are guaranteed by using multi-model joint estimation and a sequential loss function, and facial expression details and hand postures in the image can be accurately estimated. After the high-precision three-dimensional human body model is obtained through estimation, human body action and expression generation is carried out based on voice driving, accurate synchronization of actions and expressions generated through voice is achieved, and facial expression movement and body posture movement, namely a whole-body three-dimensional human body model, conforming to brand-new speaking audio are accurately generated; and finally, rendering the whole-body three-dimensional human body model into a real digital human capable of voice interaction by using a three-dimensional neural rendering model. According to the invention, the realization of single person picture input, high-precision three-dimensional digital person generation and voice interaction is facilitated.
Owner:NANJING UNIV OF SCI & TECH

Knowledge graph construction method and device based on multi-source data

The invention relates to a knowledge graph construction method, device and equipment based on multi-source data. The method comprises the following steps: acquiring original data from different data sources; preprocessing the original data to obtain target text data; wherein the preprocessing comprises format conversion, text cleaning and normalization and / or sentence segmentation and segmentation; performing knowledge extraction on the target text data through a pre-optimized large language model to obtain original structured data including an original entity, an original relationship and an original attribute; post-processing the original structured data to obtain target structured data including a target entity, a target relationship and a target attribute; wherein the post-processing comprises format analysis, entity standardization and ambiguity elimination, and relation and attribute verification; and updating nodes and edges of the current knowledge graph according to the target structured data. The method can adapt to multi-source heterogeneous data, and the accuracy and consistency of the knowledge graph are improved.
Owner:CHINESE PEOPLES LIBERATION ARMY UNIT 32802

Intelligent data processing system based on AI large model

InactiveCN120471064ASemantic analysisBiological modelsLinguistic modelCitation frequency
The invention relates to the field of data processing, and discloses an intelligent data processing system based on an AI large model, which comprises the following steps: when detecting that a plurality of candidate entities exist in a text, judging whether the candidate entities need to be subjected to semantic disambiguation processing or not according to context semantic relevancy and entity historical reference frequency; constructing a cross-domain context representation vector, introducing a pre-training language model to encode the context, generating deep semantic representation of candidate entities, and judging whether a clustering result has ambiguity or not; whether the candidate entities have conflicts or not is judged by fusing language model embedding and structuring knowledge graph information; constructing an entity relationship graph based on a graph neural network, performing causality, temporal and attribute dependence reasoning on the relationship between the existing candidate entities, and judging whether the relationship between the candidate entities should be combined or split; and carrying out label replacement on the candidate entities in the target text in combination with the context and the reasoning result. The method has the advantage of improving the accuracy of data processing.
Owner:CHANGSHA DILU DIGITAL TECH

Causal interpretability and illusion suppression method, system and device for text generation

The invention relates to the technical field of financial management, and discloses a causal interpretability and illusion suppression method, system and device for text generation, and the key point of the technical scheme is that the method comprises the following steps: S1, obtaining related data according to a generation target, extracting the causal relationship between entities, and constructing a causal map; s2, extracting a causal chain related to the generated target from the causal atlas, and inputting the causal chain and the input variables into the large language model to obtain a target text; s3, performing anti-fact intervention processing on the input variable, inputting the input variable into the large language model, and recording a logic test result; s4, identifying an entity from the target text, comparing the entity with the fact data related to the generated target, and calculating an entity alignment score; and S5, according to the causal chain, the target text, the logic test result and the entity alignment score, generating an interpretation report and performing structured output, so that the output text has higher logicality and higher credibility.
Owner:JIANGSU SUNING BANK CO LTD +2

Intelligent system conflict point review system based on knowledge graph and large language model

The invention relates to the technical field of electrical digital data processing, and discloses a system conflict point intelligent review system based on a knowledge graph and a large language model, which comprises the following steps: constructing a dual-mode storage space containing an unstructured index and a structured logic graph, analyzing target text extraction features and triggering graph-based generation logic; converting the topological structure of the associated sub-atlas into a natural language instruction sequence to construct a forced logic constraint template, filling the template with a text, and inputting a pre-training language model to generate a verification result; according to the method, the discrete atlas topology is mapped into the linear logic constraint, random divergence of the generative model is restrained on the calculation principle, and precise decoupling and dynamic evolution of unstructured semantics and structured logic are achieved.
Owner:CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Response text generation method and device, computer equipment and storage medium

The embodiment of the invention relates to a reply text generation method and device, computer equipment and a storage medium. After an input text is preprocessed, a target text is obtained; acquiring context information corresponding to the target text, a first prompt word template and a first keyword; retrieving related information corresponding to the input text from a first database according to the input text, the context information, the first cue word template and the first keyword to obtain a retrieval result; generating cue words of a large language model according to the input text, the context information, the first cue word template, the first keyword and the retrieval result; and inputting the cue word into the large language model to output the reply text through the large language model. Therefore, after the input text is combined with the context information, the cue word template of the industry and the keyword to retrieve the related information, the cue word of the large language model is generated according to the retrieval information, the reply text is accurately generated through the large language model, and the question and answer accuracy and the user satisfaction degree are improved.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Computer-Implemented Methods and Systems for Generative Text Painting

A system and method for transforming text within documents using, such as by using large language models (LLMs). Users can select source text from a source document, in response to which a painting configuration is identified or generated based on the source text, such as by providing the source text and a source prompt to a large language model to produce source output, and selecting or generating the painting configuration based on the source output. The user can select destination text, in response to which the painting configuration is applied to the destination text, such as by selecting or generating a destination action definition based on the painting configuration and the destination text, and providing the destination action definition to a large language model to produce destination output. The destination text may be replaced with the destination output, or output derived therefrom. In this way, the system can extract a variety of sophisticated properties, such as style or tone, from user-selected text source text, and apply those properties to user-selected destination text, with minimal user input.
Owner:QUABBIN PATENT HOLDINGS INC

Emotional voice generation method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses an emotional voice generation method, device, equipment and medium, and the method comprises the steps: obtaining a target text of a to-be-generated voice and a target emotional prompt text based on natural language description; inputting the target emotion prompt text into a pre-trained emotion encoder to obtain an emotion embedding vector; inputting the target text into a pre-trained text encoder to obtain a semantic embedding vector; inputting the emotion embedding vector and the semantic embedding vector into a pre-trained joint speech generation model, performing fused speech feature prediction on the emotion embedding vector and the semantic embedding vector, and generating a speech feature with an expected emotion; and decoding the voice features to generate target emotional voice corresponding to the target text. The speech emotion is controlled through the emotion prompt text described by the natural language, and the flexibility and synthesis effect of emotion speech generation interaction are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Text sequence generation method and device of intelligent agent, equipment, medium and product

The embodiment of the invention provides a text sequence generation method of an intelligent agent, which can be applied to the technical field of artificial intelligence so as to remarkably improve the language interaction capability of the intelligent agent. The text sequence generation method of the agent comprises the following steps: generating a candidate draft text corresponding to a received prefix input text based on a partial cache model; generating a target draft text corresponding to the candidate draft text through a complete cache model; and outputting a target text sequence according to the matching degree of the candidate draft text and the target draft text. The embodiment of the invention further provides an agent text sequence generation device and equipment, a storage medium and a program product.
Owner:BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE

Reimbursement approval method, device and equipment and storage medium

The invention discloses a reimbursement approval method and device, equipment and a storage medium. In the scheme, based on the tax identifier of the target enterprise, the electronic invoice data associated with the tax identifier is acquired through the third-party electronic invoice platform interface. And extracting target text information from the electronic invoice data according to a predefined field rule, and generating structured reimbursement data based on the target text information. And generating a reimbursement application form according to the structured reimbursement data, and performing compliance verification on the reimbursement application form to obtain a compliance verification result. And if the compliance verification result is that the compliance verification is passed, performing risk assessment on the reimbursement application form based on the intelligent risk control model, and generating an approval suggestion. And executing a corresponding reimbursement approval process operation according to the approval suggestion and a preset approval rule. According to the technical scheme, the efficiency and accuracy of reimbursement approval are remarkably improved while the financial risk of an enterprise is reduced.
Owner:太保科技有限公司

Large model named entity recognition method and system based on representative sample selection and context enhancement

The invention provides a large model named entity recognition method based on representative sample selection and context enhancement, which comprises a representative sample selection module, an entity knowledge construction module, a dynamic context selection module, a large model calling module and an iterative feedback optimization module, according to representative sample selection, samples with representativeness and information diversity are automatically selected from unlabeled data for labeling through a sample screening strategy based on clustering, entity description integration aims at each entity type, a plurality of high-quality instances are extracted from labeled samples, and standardized entity definition or description prompts are constructed. According to the dynamic context selection, for to-be-recognized text content, a context example most relevant to a target text is dynamically selected from a historical annotation sample or a description set through a semantic similarity retrieval mechanism to serve as auxiliary prompt input, and the adaptability and generalization ability of LLM in a complex or variable scene are improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Spoken language pronunciation training correction system based on intelligent equipment

The invention belongs to the technical field of intelligent voice processing, and particularly relates to a spoken language pronunciation training and correcting system based on intelligent equipment, which acquires a user rhythm feature set including pitch change rate, accent intensity, pause duration and intonation contour by acquiring spoken language audio data sent by a user for a target text, and corrects the spoken language pronunciation training and correcting system. The method comprises the following steps: acquiring spoken language audio data, converting the spoken language audio data into a phoneme sequence aligned with target text time, generating a rhythm deviation degree report according to comparison evaluation of a user rhythm feature set and the phoneme sequence, an execution intonation mode, accent distribution, speech stream sound change and speech speed rhythm, and determining the rhythm deviation degree according to a rhythm problem type in the report in combination with user historical learning data. And forming and outputting a correction scheme including a text prompt, a targeted minimum contrast training unit and a listen-and-read simulation task, solving the problem of weak capability of correcting hyper-phoneme in a second language spoken language of an adult, and improving the authentic and fluency of the spoken language, thereby realizing efficient personalized learning.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Method and system for converting personalized text into voice, and related equipment

The invention provides a personalized text-to-speech conversion method and system and related equipment, and the method comprises the steps: training a deep learning model through a text / audio corpus of a non-standard speaker, and obtaining a non-standard text-to-speech conversion model; obtaining a single speaker reference audio with customized timbre and a to-be-converted target text; inputting the target text into the non-standard text-to-speech model to obtain a sound spectrum representation of the target text; extracting a timbre embedding vector of a target speaker from the single speaker reference audio by using a voiceprint encoder; and fusing the sound spectrum representation and the timbre embedding vector, and inputting the fused sound spectrum representation and timbre embedding vector into a neural vocoder to obtain a personalized voice waveform. According to the method provided by the invention, the synthesis of the tone migration personalized audio can be realized only through the non-standard language data of the single speaker, the scheme realization difficulty is reduced, and the user demand can be better met.
Owner:SHANGHAI JITU SCI & TECH CO LTD