Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

66 results about "Timed text" patented technology

Timed text refers to the presentation of text media in synchrony with other media, such as audio and video.

AI digital human conference proxy method and device under off-line local area network and medium

The invention discloses an AI digital human conference proxy method and device under an offline local area network and a medium, and the method comprises the steps: collecting conference voice in real time, converting the conference voice into a real-time text, and pre-judging a subject set in combination with a localized industry knowledge base and a user historical conference track; if it is detected that the user or the to-be-decided item is mentioned, reply voice is generated in combination with historical corpora of the user; and if the conference enters the pre-judgment topic set, calling the pre-loaded user feature packet to generate reply voice, and controlling the digital person to generate a corresponding audio and video stream. The invention provides an AI digital human conference proxy method and device under an offline local area network and a medium, and aims to realize topic pre-judgment in combination with a local knowledge base and a historical conference track of a user and provide preparation for real-time reply; meanwhile, for different scenes, the user historical corpus and the user feature packet are called respectively to generate the reply, so that the problem that the real-time personalized reply generation of the AI digital person conference agency and the conference issue pre-judgment are difficult to collaboratively realize in an offline local area network environment can be solved.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Public opinion information intelligent processing method and system based on multi-modal data fusion

The invention relates to a public opinion information intelligent processing method and system based on multi-modal data fusion. The method comprises the steps of obtaining real-time text data, image data, video data and social attribute data; performing keyword rule matching and deep learning model analysis on the text data in parallel, comparing confidence coefficient difference values of rule output and model output, and determining a text analysis result; performing cross-modal fusion on the multi-source feature vectors, and dynamically distributing feature weights according to the data quality of each modal to generate weighted fusion features; calculating a social influence value based on the user authority, the propagation path depth and the time attenuation coefficient in the social attribute data, and identifying a key community cluster; further carrying out weighted matching on a pre-constructed domain strategy matrix through multi-modal evidence; and adjusting a rule weight and a model parameter according to the target domain configuration result and the weighted fusion feature. According to the method, the public opinion field can be accurately identified, the processing scheme is intelligently matched, and a visual and understandable decision basis is provided.
Owner:XUZHOU COLLEGE OF INDAL TECH

Method and system for realizing real-time two-way visual interaction of digital human by calling camera

The invention discloses a method and system for realizing real-time two-way visual interaction of a digital human by calling a camera, and belongs to the technical field of multi-modal interaction, and the method comprises the steps: collecting video streams of a user and an environment in real time through a mobile terminal camera, and synchronously obtaining a natural language question of the user; screening related video keys based on question semantics; performing cross-modal semantic fusion analysis on the question text and the key frame image by using a multi-modal model to generate a corresponding answer result; and the virtual digital person outputs an answer through voice synthesis and synchronous anthropomorphic animation, so that real-time image-text question and answer interaction between the digital person and the user is realized. Aiming at the problem that the existing digital human interaction lacks visual situation feeling, the invention provides an end-cloud collaborative visual question-answer interaction scheme, so that the video semantic processing load of a mobile end can be effectively reduced, the end-cloud collaborative real-time image-text question-answer interaction is realized, the perception ability of the digital human to the environment is greatly improved, and the user experience is improved. And the intuition and the naturalness of user interaction experience are improved.
Owner:LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD

A long text matching method combining noise filtering and divide-and-conquer strategy

This invention discloses a long text matching method combining noise filtering and a divide-and-conquer strategy. The method includes: constructing a long text matching model, which comprises a keyword extraction layer, an association extraction layer, and a filtering layer. The keyword extraction layer extracts keywords from the text; the filtering layer filters noise from the text based on sentence similarity to obtain a denoised text sequence; and the association extraction layer further removes keywords from the denoised text sequence to obtain the remaining associated text. The long text matching model is trained with the optimization objective of minimizing a set overall loss function, which reflects the global matching distribution and combines the keyword and association matching distributions. For the target text, real-time text matching is performed using the trained long text matching model. This invention improves the generalization ability and accuracy of text matching.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Natural language-based standard time conversion method and device, equipment and medium

PendingCN122654309AEliminate formatting differencesImprove interaction efficiencyEngineeringTarget text
The application relates to the technical field of natural language processing, financial technology and medical technology, and discloses a standard time conversion method and device based on natural language, equipment and a medium, which comprises the following steps: collecting current session text information in real time; when target text representing a time attribute appears in the text, acquiring corresponding collection time and standardizing the collection time into a specified format reference time; extracting preset time period text containing the target text as context information; reasoning the target text and the context by combining the reference time with a preset natural language model to determine a target date, a target time word and a target number of days; and performing consistency verification on the target time word and the target number of days, and outputting the target date if the verification is passed. The application can be applied to a data processing platform in the fields of financial technology and medical health, improves the standardization conversion efficiency of time text strings, improves the accuracy of large models, and improves the convenience of the corresponding platform.
Owner:PING AN TECH (SHENZHEN) CO LTD

A Text Printing Method and System Based on Speech Recognition

This invention belongs to the field of text printing technology and discloses a text printing method and system based on speech recognition. The method includes the following steps: constructing a text printing template database; constructing a mixed-language speech recognition model; acquiring speech audio data in real time and performing speech recognition; matching several corresponding text printing templates; selecting a text printing template; fusing the real-time speech text data with the selected text printing template; and printing the real-time text printing data. The system includes a database construction unit, a model construction unit, a storage unit, a speech audio acquisition unit, a speech recognition application unit, a template matching unit, a template selection unit, a data fusion unit, and a printer. This invention solves the problems of low intelligence, poor speech recognition effect, low recognition efficiency, and lack of organic integration in existing technologies.
Owner:魏鹏飞

Portable real-time text recognition and translation method and system based on lightweight Transform

The invention discloses a portable real-time text recognition and translation method and system based on a lightweight Transform, and belongs to the technical field of computer vision and natural language processing. The system comprises an image preprocessing module, a text recognition module, a text translation module, a result presentation module, a model optimization engine and a system control module. A lightweight Transform architecture is adopted, and the calculation complexity is reduced through a layered mixed attention mechanism, a shared weight design, dynamic depth adjustment and a compression feedforward network. The text recognition module adopts a CNN and Transform hybrid architecture to realize detection and recognition, and the text translation module shares partial encoder weight and supports incremental decoding. The system integrates compression technologies such as knowledge distillation, quantification and pruning, realizes end-side off-line real-time operation in combination with hardware perception optimization, supports multi-language image-text translation, and has the characteristics of high precision, low delay and low power consumption.
Owner:SHAANXI SCI TECH UNIV

Error sentence generation method and device for model training, computer device and medium

This invention relates to the field of artificial intelligence technology, and more particularly to a method, apparatus, computer device, and medium for generating incorrect sentences for model training. The method maps the number of terms in the text to be processed to an adjustment quantity, samples target terms from the text to be processed, adjusts the target terms to obtain augmented text, occludes the text to be processed, inputs it into a reconstruction model to obtain reconstructed text, trains a generative model based on the text to be processed, the augmented text, and the reconstructed text, inputs real-time text into the trained generative model to obtain generated incorrect sentence text, processes the text to be processed using multiple data augmentation methods to obtain text containing rich error types as labels, trains an end-to-end generative model, allowing real-time text to be directly input into the trained generative model to obtain generated incorrect sentence text without frequent switching of data augmentation methods, greatly improving the efficiency of incorrect sentence generation while ensuring high quality.
Owner:PING AN TECH (SHENZHEN) CO LTD

Conversational analysis system and method

This specification provides a dialogue analysis system and method. The system includes: an acquisition module for acquiring a text sequence, wherein the text sequence is obtained by real-time text conversion of classroom dialogues; a matching module for searching, in a pre-established rule base, using a large language model, whether a target sequence matching the text sequence exists; wherein the rule base includes classification rule sequences for various dialogue types and / or reference text sequences corresponding to each dialogue type; and an analysis module for analyzing, based on the target sequence matched by the matching module, to obtain the target dialogue type corresponding to the classroom dialogue. This disclosure combines a rule base with a large language model to achieve usability and flexibility in classroom dialogue analysis, ensuring both deep integration of the dialogue analysis system with educational theory and the ability to process complex dialogue sequences through a large language model, achieving accurate dialogue classification and analysis.
Owner:TSINGHUA UNIVERSITY

Artificial intelligence-based text generation model construction method and system

The invention belongs to the technical field of natural language processing, and discloses a text generation model construction method and system based on artificial intelligence. The method comprises the following steps: constructing an initial integrated text generation model, and setting an experience playback pool and an elastic loss function; performing optimization training on the initial integrated text generation model according to the training text data to obtain a trained integrated text generation model and a plurality of historical training experiences, and storing the trained integrated text generation model and the historical training experiences in an experience playback pool; extracting a plurality of real-time text generation experiences, and randomly extracting a plurality of historical training experiences from the experience playback pool; and based on the elastic loss function, according to a plurality of historical training experiences and a plurality of real-time text generation experiences, continuously training the trained integrated text generation model to obtain an updated integrated text generation model. According to the method, the problems of unstable generation quality, strong data dependence, model function overflow and insufficient generalization ability in the prior art are solved.
Owner:BEIJING JUNDE INTELLIGENT COMPUTING TECHNOLOGY CO LTD

A text image quality detection method, system, device and medium

The application discloses a kind of text image quality detection method, system, equipment and medium, the method includes: obtaining text image, text image is divided into training set and test set;Text image quality detection model is constructed, and training set is input into text image quality detection model for training;Test set is input into the text image quality detection model that has been trained, obtains detection result, and text image quality detection model is evaluated and optimized.The application detects the quality of text image using the text image quality detection model without reference image, does not need fixed reference image, is more suitable for practical application scene, is more flexible and practical, improves the detection precision of text image quality evaluation, and improves the accuracy and robustness of text image quality detection, realizes accurate and real-time text image quality evaluation.
Owner:DONGFANG (GUANGZHOU) HEAVY MASCH CO LTD

Automated timed text workflow system integrating machine translation, ai-driven tools, and human review for high-volume media processing

A scalable, automated timed text workflow system designed to optimize the generation and refinement of time-synchronized textual content for video is disclosed. Integrated Machine Translation (MT) models and advanced AI-driven tools such as Computer Vision, Generative AI, and Traditional AI automatically generate and refine timed text, including subtitles, closed captions (CC), and SDH (Subtitles for the Deaf and Hard of Hearing). Human review is coordinated through a workflow orchestration system when needed, offering flexibility and scalability for handling high-volume media processing across various platforms and formats.
Owner:PANTOJA PAULETTE

A network information collection method and system based on real-time text analysis

The application discloses a network information collection method and system based on real-time text analysis, comprising the following steps: acquiring real-time text data, preprocessing the real-time text data, acquiring first data and second data according to the preprocessed real-time text data, calculating the similarity of the first data and the similarity of the second data, weighting the similarity of the first data and the similarity of the second data to obtain a classification target, constructing a text classification model according to the classification target, inputting the real-time text data into the text classification model to obtain classification data, and outputting the classification data as collected network information. The method can not only improve the accuracy of network information collection, but also has good interpretability and can be directly applied to a network information collection system.
Owner:CHINA NAT INST OF STANDARDIZATION

Memory-enhanced deep personalized input method system based on end-side large language model

PendingCN122507293APersonalizationUser input
The present application relates to a memory-enhanced deep personalized input method system based on end-side large language model, and belongs to the technical field of artificial intelligence. The present application solves the problems of the prior art, such as the loose coupling of the model and the input core process, the shallow personalization mechanism, the delay and privacy risk caused by cloud deployment, and the insufficient efficiency of the end-side operation that cannot meet the real-time input requirements. The system includes a front-end real-time interaction module, a stylized large model post-training module, a hierarchical memory mechanism module, and an end-side reasoning optimization module. The present application can realize low-delay real-time text generation under the condition of highly limited mobile terminal resources, and at the same time, through continuous modeling of user input behavior, it can complete deep personalization adaptation and ensure the end-side retention and privacy security of user input data throughout the process.
Owner:HARBIN INST OF TECH

Discrimination method for position and type of centralizer based on multi-source data fusion

PendingCN122336657AAlgorithmTimed text
This invention provides a method for determining the placement location and type of a centralizer based on multi-source data fusion, comprising: Step 1, acquiring casing record table data before the casing operation begins; Step 2, acquiring real-time text data and real-time video data at the start of the casing operation; Step 3, determining the target casing information currently being lowered and the type of centralizer to be placed therein based on the real-time text data; Step 4, determining the type of centralizer actually placed on the target casing based on the real-time video data; Step 5, determining the specific time of centralizer placement on the target casing based on the real-time video data; Step 6, determining whether the corresponding type of centralizer was correctly placed on the target casing during the actual placement process. This method for determining the placement location and type of a centralizer based on multi-source data fusion standardizes the casing lowering operation process, effectively replaces traditional manual monitoring methods, and improves on-site safety and supervision efficiency.
Owner:CHINA PETROLEUM & CHEMICAL CORP +1

Enhanced controls for the display of real-time text in calls and meetings

The techniques disclosed herein provide enhanced controls for the display of real-time text (RTT) in calls and meetings. RTT is the ability for someone to send a text message on a character-by-character basis to everybody else in a call or meeting. The system disclosed herein integrates RTT, video, and live captions all in one central experience. This integrated experience allows users to participate equitably by making RTT accessible to users regardless of the operating mode they are in and still concurrently access other meeting content, including video streams, chat messages, live captions, transcripts, and artificial intelligence (AI) tools, such as Copilot. In one embodiment, during an online conference, in response to one of the attendees activating a RTT mode, when at least one user minimizes a meeting stage, such as for the purpose of multitasking while listening, the conference application maintains a display area for displaying RTT.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Real-time online text violation recognition method and system and storage medium

PendingCN121071653ASemantic analysisBiological modelsText recognitionSingle sentence
The invention discloses a real-time online text violation recognition method and system and a storage medium, and belongs to the technical field of text recognition. The method comprises the following steps: collecting user behavior data in real time, and generating a user-level weak tag; comparing and learning the user-level weak tags to obtain a single sentence code and a single sentence risk assessment model; automatically distributing a pseudo label and an attention weight for each sentence by utilizing a pseudo label distribution rule, and training to obtain a lightweight sentence-level model; for an input real-time text, calculating a violation probability through a lightweight sentence-level model to obtain a risk index; according to the risk indexes and a preset risk threshold value, performing grading processing, blocking users whose risk indexes are in a high risk interval, performing secondary reasoning and random sampling on the users whose risk indexes are in a middle and low risk area, and then performing manual auditing; and iteratively updating the user-level weak label according to a user disposal result and newly added report data, and realizing self-evolution of each model. A large number of user-level tags can be obtained at low cost, and the detection rate is improved.
Owner:CHENGDU RENPINGSHENG NETWORK TECHNOLOGY CO LTD

Text processing method and device and storage medium

PendingCN121859902ASolve technical problems with low processing efficiencyeasy to handleNatural language translationTime informationAlgorithm
The invention discloses a text processing method and device and a storage medium. Relates to the technical field of civil aviation information, and comprises the following steps: receiving a to-be-processed first text in a natural language form, and extracting a time text segment in the first text, the time text segment comprising time information; determining a time expression category of the time text segment, the time expression category being a category of an expression mode of the time text segment for the time information; processing the time text segment according to the time expression category to obtain a target time text; and replacing the time text segment in the first text with the target time text to obtain a second text. Through the method and the device, the problem of relatively low efficiency of processing the time text in related technologies is solved.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Real-time text deentanglement-based real image editing

The embodiment of the invention relates to real-time text deentanglement-based real image editing. A method, apparatus, non-transitory computer readable medium and system for image processing includes obtaining an input image depicting a first element, a textual description of the input image, and a modification prompt describing a second element, the second element being different from the first element; generating an intermediate output based on the input image and the textual description, where the intermediate output represents the first element; and generating a composite image based on the intermediate output and the modification cue, where the composite image replaces the first element from the input image with a second element from the modification cue.
Owner:ADOBE INC

Real-time text-based disentangled real image editing

A method, apparatus, non-transitory computer readable medium, and system for image processing include obtaining an input image depicting a first element, a text description of the input image, and a modification prompt describing a second element different from the first element, generating an intermediate output based on the input image and the text description, where the intermediate output represents the first element, and generating a synthetic image based on the intermediate output and the modification prompt, where the synthetic image replaces the first element from the input image with the second element from the modification prompt.
Owner:ADOBE INC

Real-time text-based disentangled real image editing

A method, apparatus, computer readable medium, and system for image processing include: 1105: obtaining an input image depicting a first element, a text description of the input image, and a modificat
Owner:ADOBE INC

Enterprise due diligence method, apparatus, device, and storage medium

The application provides an enterprise due diligence method, device, equipment and storage medium, the method comprises the following steps: obtaining the enterprise information of the client to be due diligence and converting into a label, matching the content of the speech from the historical speech database according to the label, and obtaining the personalized speech list; real-time speech recognition is performed on the due diligence recording collected in the due diligence field to obtain real-time text, the real-time text is subjected to keyword retrieval and comparison with the personalized speech list, respectively obtaining the violation items and the missing items, which are used for generating reminder information and sending to the due diligence personnel; after the due diligence process is completed, the due diligence recording is subjected to speech recognition to obtain the full-amount text, the key information is extracted from the full-amount text and filled into the report template to obtain the due diligence report; the position information of the due diligence field and the voiceprint characteristics of the due diligence recording are collected, the position information and the voiceprint characteristics are verified, and the due diligence report is labeled according to the verification result. The application solves the problem of weak due diligence authenticity verification capability in the prior art.
Owner:SHENGYE INFORMATION TECH SERVICE (SHENZHEN) CO LTD

A method and system for information extraction based on natural language text

This invention belongs to the field of text information extraction technology and discloses a method and system for information extraction based on natural language text. The method includes the following steps: acquiring real-time natural language text and preprocessing it to obtain preprocessed real-time natural language text; using a text analysis model to analyze the preprocessed real-time natural language text to obtain corresponding real-time text analysis results; based on the real-time text analysis results, using a strategy generation model to generate strategies for the preprocessed real-time natural language text to obtain corresponding real-time information extraction strategies; and using an information extraction model to extract information from the preprocessed real-time natural language text according to the real-time information extraction strategies to obtain corresponding real-time key information. This invention solves the problems of slow processing speed, low accuracy, poor flexibility, and insufficient semantic understanding in existing technologies.
Owner:BEIJING JUYUAN RUISI DATA TECHNOLOGY CO LTD

Enhanced controls for the display of real-time text in calls and meetings

The techniques disclosed herein provide enhanced controls for the display of real-time text (RTT) in calls and meetings. RTT is the ability for someone to send a text message on a character-by-character basis to everybody else in a call or meeting. The system disclosed herein integrates RTT, video, and live captions all in one central experience. This integrated experience allows users to participate equitably by making RTT accessible to users regardless of the operating mode they are in and still concurrently access other meeting content, including video streams, chat messages, live captions, transcripts, and artificial intelligence (AI) tools, such as Copilot. In one embodiment, during an online conference, in response to one of the attendees activating a RTT mode, when at least one user minimizes a meeting stage, such as for the purpose of multitasking while listening, the conference application maintains a display area for displaying RTT.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Named entity extraction method based on lexical item combination

The invention relates to the technical field of text lexical item recognition, and particularly discloses a lexical item merging-based named entity extraction method, which comprises the following steps of: sequentially performing sequential combination judgment, full-text frequency statistics and lexical item merging, and extracting a named entity on the premise of not depending on an external dictionary or a training model. According to the method, continuous and high-frequency lexical item combinations can be automatically mined and merged from original texts, named entity extraction is achieved, the method is suitable for open domain, low-resource or real-time text processing scenes, model loading or training is not needed, and the method can efficiently run on terminals or edge devices with limited resources.
Owner:CENT SOUTH UNIV

Method and device for implementing intelligent question and answer service based on large model

The application discloses a method and device for realizing intelligent question and answer service based on a large model, the method comprises the following steps: acquiring real-time text data; inputting the real-time text data into a trained intelligent question and answer network model to perform corresponding answer prediction, and outputting multiple candidate answers according to the answer prediction result; acquiring scores corresponding to the multiple candidate answers, and outputting an optimal answer in the multiple candidate answers according to the sorting result of the scores. The application adds a personalized capability to the large model, solves the problem that the large model is seriously illusory and difficult to achieve ideal effects of problems by relying on knowledge base retrieval, and can effectively and accurately predict and output answers to problems in multiple fields.
Owner:BEIJING CAIZHI TECH CO LTD

Real-time text-to-voice conversion method without Internet connection, equipment and storage medium

The invention discloses a real-time text-to-speech conversion method without Internet connection, equipment and a storage medium. The method comprises the following steps: deploying and configuring a preset speech synthesis engine on local computing equipment; starting a local service to receive a voice synthesis request from a front-end application, and analyzing the voice synthesis request to obtain text content to be synthesized; calling the speech synthesis engine to process the to-be-synthesized text content locally, and generating a corresponding audio data stream; and returning the generated audio data stream to the front-end application, and playing the audio data stream. According to the invention, the preset speech synthesis engine is deployed and configured on the local computing device, so that the whole text-to-speech process can be completed locally without depending on external cloud service, and text-to-speech conversion can still be carried out normally even in an environment without network connection.
Owner:WUHAN HONGXIN TECH SERVICE CO LTD

Real-time discourse correction method and related device, electronic equipment and storage medium

PendingCN122065819ASemantic analysisKnowledge representationTimed textCorrection text
The invention discloses a real-time discourse correction method, a related device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a current recognition text of a real-time audio stream, and obtaining a plurality of historical recognition texts subjected to text correction before the current recognition text and historical correction texts obtained through text correction; based on at least one of the current recognition text and the historical correction text, retrieving in the scene knowledge graph to obtain a reference knowledge text related to the execution of the text correction; wherein the scene knowledge graph is a reference knowledge graph related to an interaction scene where the real-time audio stream is located; and performing text correction on the current recognition text based on the historical correction text and the reference knowledge text to obtain a current correction text of the current recognition text. According to the scheme, the accuracy of real-time discourse correction can be improved, especially when the real-time discourse relates to terminologies or specific domain knowledge.
Owner:IFLYTEK CO LTD

Parallel computing categorization process

To provide a computer-implemented method for parallel computing categorization, a program, and an information processing apparatus.SOLUTION: A method comprises: obtaining real-time text data relating to a matter, the text data comprising a plurality of portions of information; performing a categorization process, the categorization process being configured to run a plurality of threads in parallel, each thread of the plurality of threads acting on one portion of information at a time; obtaining a sentiment score on the basis of the portion of information using a Sentiment Analysis machine learning (ML) model for each thread; assigning a category to the matter on the basis of the sentiment score using a classification ML model trained on historical data; updating a live category on the basis of the category assigned to the matter; and outputting the live category to a user in real-time.SELECTED DRAWING: Figure 8
Owner:FUJITSU LTD