Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

98 results about "Punctuation" patented technology

Punctuation (formerly sometimes called pointing) is the use of spacing, conventional signs and certain typographical devices as aids to the understanding and correct reading of written text whether read silently or aloud. Another description is, "It is the practice action or system of inserting points or other small marks into texts in order to aid interpretation; division of text into sentences, clauses, etc., by means of such marks."

Short video copywriting tone automatic adjusting method driven by hierarchical rhythm mapping

The invention discloses a hierarchical rhythm mapping-driven short video copywriting mood automatic adjustment method, and relates to the technical field of video processing, and the method comprises the steps: 1, receiving a text character string and a language type identifier, and building an occupation column for bearing a tone mark, an accent mark and a duration mark at each level; 2, dividing each sentence into phrase segments based on the hierarchical index table, freezing boundaries by taking the phrase segments as units, presetting sentence end termination styles according to punctuations, determining kernel phrases according to semantic anchor points, initializing trends of the kernel phrases, and performing time sequence elastic alignment and hierarchical backfilling to obtain a sentence end termination pattern; and finally outputting a triple sequence which covers all syllables and is composed of tone marks, accent marks and duration marks as a target rhythm control sequence. And step 3, performing audio generation based on the target rhythm control sequence to obtain new dubbing. According to the method, the tone accuracy and expressive force of short video dubbing are improved, and the time and cost of manual adjustment are remarkably reduced.
Owner:CLOUD ATTACK NETWORK TECH HEBEI CO LTD

Automated segmentation and transcription of unlabeled audio speech corpus

ActiveUS12512100B2Speech recognitionTimestampAudio segmentation
A method includes obtaining initial transcription for input natural speech; performing segmentation of initial transcription into text portions, based on punctuation marks in initial transcription; determining segment-level timestamps for text portions based on the input natural speech; performing audio segmentation on input natural speech, by cutting input natural speech based on segment-level timestamps, to obtain audio chunks; generating transcription portions for each of the audio chunks; merging transcription portions to form re-transcription; determining word-level timestamps for re-transcription, by aligning input natural speech against re-transcription; calculating silence time periods, each corresponding to silence between each two adjacent words of input natural speech, based on word-level timestamps; performing a final segmentation on input natural speech and re-transcription, based on silence time periods, to generate final audio segments and corresponding final transcription portions. The final audio segments and corresponding final transcription portions may be included in training dataset for training a model.
Owner:ORACLE INT CORP

Methods and apparatuses for the condensation of spoken text

A speech condensation processing system and method includes an ASR system for a source language that receives an audio stream with speech and outputs at least one word sequence and time stamps in the language spoken, a memory that stores a condensation program and corresponding data and databases that store training data, which may include manually condensed data, two-way translated data, and aligned subtitle data, and a processor coupled to the ASR system and memory that executes the condensation program to format and condense text by transforming the at least one word sequence from ASR into human-readable text with proper casing and punctuation, and condenses the text based neural training to remove words from the at least one word sequence that are not relevant for meaning.
Owner:APPL TECH APPTEK

Teaching interaction quality evaluation method and system based on large language model

The invention relates to the field of teaching interaction quality evaluation, in particular to a teaching interaction quality evaluation method and system based on a large language model. The method comprises the following steps: audio transcription: converting classroom audio into an original transcription text through voice activity detection, speaker classification, automatic voice recognition and punctuation recovery; transcriptional refining: performing context-based text error correction on the original transcriptional text by using a large language model in combination with a preschool education field knowledge base to generate a refined transcriptional text; a quality evaluation step: based on a preschool education quality evaluation scale, using few sample example guidance and thinking chain reasoning for each scoring point, judging whether a voice segment conforming to the scoring point exists in the refined transcriptional text, performing binary scoring, and determining whether the voice segment conforms to the scoring point; and generating an interactive quality evaluation report containing the standard-reaching rate of each evaluation dimension, teaching bright spot analysis and staged teaching optimization suggestions. The evaluation efficiency is remarkably improved. The method is suitable for teaching interaction quality evaluation.
Owner:THE CHINESE UNIV OF HONG KONG (SHENZHEN)

Large model training method and device, voice recognition text processing method and device, equipment and medium

The invention discloses a large model training method and device, a voice recognition text processing method and device, equipment and a medium, relates to the technical field of communication, and aims to improve the accuracy of text output obtained by a post-processing task. The method comprises the steps that text data sets used for model training are obtained, the text data sets comprise a first text data set and a second text data set, the first text data set comprises a constructed text smoothing task text data set, a text error correction task text data set, a punctuation recovery task text data set and an ITN task text data set, and the second text data set comprises a constructed text smoothing task text data set and a constructed text error correction task text data set; the second text data set is a text data set manually labeled in a real scene; adding a task label for the text data set; and training the first large model by utilizing the text data set added with the task label. According to the embodiment of the invention, the output accuracy of the text obtained by the post-processing task can be improved.
Owner:CHINA MOBILE COMM LTD RES INST +1

Double-path voice stream real-time identification method, system and application

The invention discloses a double-channel voice stream real-time identification method, system and application, and the method comprises the steps: carrying out the preprocessing of collected VOIP call double-channel audio, and maintaining the time sequence synchronization; extracting Mel-frequency cepstral coefficient features and speech spectrogram features of the preprocessed audio, and inputting the spliced features into a Transform deep neural network model for stream speech recognition to obtain a two-way character sequence; generating a unique identifier based on channel identification and voice energy difference, and establishing a corresponding relation with the character sequence; carrying out punctuation prediction and text standardization by utilizing an LSTM-based model, sorting and aligning character sequences according to timestamp fields, and generating a time sequence dialogue stream; and performing anomaly detection and / or storage management on the time sequence dialogue stream to realize real-time quality inspection and agent assistance. According to the invention, synchronous recognition and role distinguishing of double-channel voice are realized, the recognition delay is low, and the recognition accuracy, the detection precision and the real-time performance are high; and high-efficiency management and safety compliance of data are realized by combining distributed encryption storage.
Owner:XUNMENG COMMUNICATION TECHNOLOGY CO LTD

Speech punctuation detection method, apparatus, device, storage medium, and program product

The application discloses a speech punctuation detection method, device, equipment, storage medium and program product. The method comprises the following steps: obtaining first speech basic data of a first speech and second speech basic data of a second speech; determining a second silence duration threshold according to the first speech basic data and the second speech basic data; determining a first punctuation detection result of the second speech according to a silence duration of the second speech and the second silence duration threshold, wherein the first punctuation detection result is used for indicating whether the second speech needs to be punctuated. The second speech basic data of the speech that needs to be punctuated and the first speech basic data of the speech that has been punctuated are used to dynamically adjust the silence duration threshold according to the needs, and then the speech punctuation detection is performed according to the second speech basic data and the adjusted silence duration threshold, so that the speech punctuation detection result is obtained, and the accuracy of the speech punctuation detection is improved.
Owner:MASHANG CONSUMER FINANCE CO LTD

Method for automatically labeling work order types based on agents

PendingCN121434405ADigital data information retrievalSemantic analysisSemantic vectorComputational probability
The invention provides a method for automatically labeling work order types on the basis of agents, which comprises the following steps of: performing punctuation standardization processing on an original work order, and converting spoken and non-standardized work order texts into segmented word segments conforming to field specifications in combination with word segmentation in a power field dictionary; unifying and normalizing the segmented word segments through a preset synonym mapping table to obtain a standardized text sequence; the standardized text sequence is input into a bidirectional encoder expression model to output a semantic vector sequence, and deep semantic understanding of the work order text is achieved; related external information is called to be coded into a feature vector, and then the feature vector is fused with the semantic vector sequence through an attention mechanism to generate an enhanced semantic vector; multi-dimensional label prediction tasks are executed in parallel based on a multi-task learning architecture, and multi-class labels and probability distribution are output; the confidence coefficient is obtained by calculating the maximum value of the probability distribution, and the preset process is executed, so that automation and quality management and control of label generation are realized, and the problem of low efficiency of manual power work order processing in the prior art is solved.
Owner:NORTH CHINA GRID MEASUREMENT CENT

Punctuation prediction method, content display method, device, equipment, medium and product

This application relates to a punctuation prediction method, content display method, apparatus, device, medium, and product. The method includes: acquiring a set of user texts corresponding to a target punctuation mark, wherein the user texts in the set contain the target punctuation mark; extracting usage preferences for the target punctuation mark from the user texts in the set to obtain usage preference features corresponding to the target punctuation mark; filtering target user texts from the set that match the usage preference features based on the usage preference features; and generating text based on the target user texts to obtain target generated text corresponding to the target punctuation mark. The target user texts and the target generated text are used to train a target punctuation prediction model, which is used to predict punctuation marks in text. This method can improve the accuracy of punctuation prediction.
Owner:SHUXING TECH (BEIJING) CO LTD

Real-time punctuation recovery method based on efficient corpus screening

The invention relates to a real-time punctuation recovery method based on efficient corpus screening. According to the method, firstly, a plurality of open-source Chinese error correction corpus data sets are downloaded, data are cleaned, punctuations are removed, and therefore a simulated voice recognition result is constructed; then mixing a plurality of corpora by using different methods to form a plurality of data sets, and performing data weighting; and finally, comparing the accuracy rates of the prediction results of the plurality of data sets, and continuously changing the generation mode of the data sets according to the recovery effect of the model to finely adjust the model. According to the method, the Chinese error correction corpus and the open-source CT-transformer model are effectively utilized, so that a better experimental result is obtained on the task of speech recognition post-processing. Through the data enhancement method that multiple sentences are spliced into one line, different corpora are mixed according to different proportions, and different punctuations are subjected to data weighting, the problem of real-time punctuation recovery of the corpora after real speech recognition is solved, and the punctuation recovery effect is effectively improved.
Owner:KUNMING UNIV OF SCI & TECH

Text punctuation adding method, device, medium and electronic device

The application provides a text punctuation adding method and device, a medium and an electronic equipment. The method comprises the following steps: obtaining a text to be added, performing word segmentation on the text to be added to obtain a plurality of words, obtaining the relationship between the words, obtaining the dependent word of each word and the relationship between each word and its dependent word, determining the relationship vector of each word based on each word, the dependent word of each word and the relationship between each word and its dependent word, obtaining the relationship between the relationship vectors of the plurality of words, and adding punctuation between the plurality of words based on the relationship between the relationship vectors. The relationship between the words in the text to be added and the relationship between the words and the text in the text to be added are considered, and the accuracy of punctuation addition can be improved to a certain extent.
Owner:PING AN TECH (SHENZHEN) CO LTD

Patent document database construction method and device based on technology description

The invention provides a patent document database construction method and device based on technical description. The method comprises the following steps: identifying a main body tag and a description tag of each patent document based on a named entity identification model; carrying out punctuation mark-based feature statement division on the technical description part, and combining all main body tags and / or all description tags belonging to the same feature statement to obtain a technical description combination; and associating the patent number of the patent file with all the technical description combinations to form a patent file database based on the technical description. The patent document is subjected to entity recognition, the feature statements are used as the combination range of the main tags and the description tags, and the technical description of the patent document is combined by using more simplified main information, so that complete coverage of the feature information is realized; and the problem that the retrieval precision is influenced by redundant descriptions or interference words in traditional independent keywords or key sentences is avoided.
Owner:BEIJING AUGUST MELON TECHNOLOGY CO LTD

Graphical user interface for two-document semantic content comparison of electronic devices

1. The name of the design product: the graphical user interface of the semantic content comparison of two documents of an electronic device. 2. The use of the design product: for an electronic device. 3. The design points of the design product: in the graphical user interface. 4. The picture or photo that best indicates the design points: front view. 5. The use of the graphical user interface: for the display interface of the AI intelligent comparison of the differences between two documents. 6. The human-computer interaction mode of the graphical user interface: the front view is the display interface of the comparison of the differences between two documents, and the full-screen view button in the form of four arrows at the top of the interface is clicked to enter the change state diagram, at which time the comparison of the two documents is displayed full screen. 7. Other circumstances that need to be explained: other views are omitted. The "X" in each view represents replaceable or changeable text, numbers, or punctuation marks. The part covered by the gray block in the view belongs to the replaceable or changeable content screen, which does not belong to the content of the design itself.
Owner:JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD

Electronic text analysis for detecting computer-generated interaction in text-based communications

A text analysis processing for detecting computer-generated text is provided. In some cases, a text-based chat interaction may be initiated and analyzed to determine whether the text-based chat generated by a communicating entity is computer-generated. The text of the chat session may be analyzed to evaluate punctuation, use of emojis, spacing, grammar, words, phrases, and the like to determine a further likelihood of whether the text is computer-generated. A duration of the chat session may be used as a scoring factor. The various probabilities and scores may be combined to provide a composite score.
Owner:BANK OF AMERICA CORP

The word of god (WOG): the 1,197,000 letter string of encoded hebrew letters underlying the original bible

A data structure and associated methods for analysis of a continuous 1,197,000-letter unvocalized Hebrew string referred to as the Word of God (WOG). The data structure contains only the twenty-two classical Hebrew letters and their five final forms, with no spacing, punctuation, vowelization, or editorial symbols. Intrinsic placement of the final letters enables deterministic segmentation of the string into 305,490 lexical units and 23,206 verses without external conventions. Fixed letter-number assignments provide a numeric architecture for evaluating substrings, detecting alterations, identifying encoded mathematical correspondences, and performing pattern analysis. The system preserves full semantic range by supporting multiple morphologically valid interpretations of unvocalized Hebrew strings. Methods for segmentation, numeric evaluation, reconstruction, integrity verification, semantic analysis, and mathematical pattern detection are provided thereby providing a reproducible foundation for computational and linguistic research.
Owner:JURAVIN DON KARL

Determining semantic and grammatical correctness of user-expanded sentence using integrated programmatic and specialized guided and constrained artificial intelligence

A system and method guide an Artificial Intelligence engine to determine the semantic and grammatical correctness of a user-expanded sentence in real-time. The sentence validation process involves receiving input from the user, the input includes sentence fragment that the user wishes to expand and user-expanded sentence that the user constructs on the fragment provided. The inputs are broken down into tokens. The word-level tokenization algorithm is used, which identifies tokens by splitting the text into spaces, punctuation marks, and other delimiters. Further, a token comparison algorithm is used to assess the relationship between the sentence fragment and the user-expanded sentence to analyze order and placement. Once the token comparison is complete, a prompt is generated using prompt generator to evaluate grammatical and semantic evaluation of the user-expanded sentence. Real-time feedback is provided to the user based on grammatical and semantic evaluation.
Owner:2HR LEARNING INC

Electronic Text Analysis for Detecting Computer-Generated Interaction in Text-Based Communications

ActiveUS20260010717A1Semantic analysisE-textData science
A text analysis processing for detecting computer-generated text is provided. In some cases, a text-based chat interaction may be initiated and analyzed to determine whether the text-based chat generated by a communicating entity is computer-generated. The text of the chat session may be analyzed to evaluate punctuation, use of emojis, spacing, grammar, words, phrases, and the like to determine a further likelihood of whether the text is computer-generated. A duration of the chat session may be used as a scoring factor. The various probabilities and scores may be combined to provide a composite score.
Owner:BANK OF AMERICA CORP

A method for regulatory speech segmentation based on speech recognition and end-point detection

ActiveCN117238279BAutomatic segmentationSpeech segmentation
The application provides a regulation voice segmentation method based on speech recognition and endpoint detection, which is applied to air traffic control voice audio stream segmentation, and comprises the following steps: step 1, constructing a punctuation model based on speech recognition and a speech endpoint detection model; step 2, using the punctuation model based on speech recognition to recognize the audio data stream of the regulation voice, and outputting the corresponding text and sentence end identifier of the audio data stream; step 3, using the speech endpoint detection model to judge the speech starting point and ending point contained in the audio data stream of the regulation voice; step 4, segmenting the audio data stream of the regulation voice into audio segments; and step 5, applying the audio segments as data materials to the speech recognition process of the air traffic control system. Through the combination of speech recognition and endpoint detection, the application realizes the automatic segmentation of the air traffic control voice audio stream, and improves the accuracy and efficiency of the segmentation.
Owner:THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP

A method for intelligent punctuation compression and line layout for East Asian text typesetting

This invention discloses an intelligent punctuation compression and line layout method for East Asian text typesetting. East Asian full-width punctuation marks are categorized into six types based on their typesetting function: start punctuation, end punctuation, sentence-end punctuation, sentence-in-sentence punctuation, exclamation / interrogative punctuation, and center punctuation. Each type defines an independent compressibility and compression direction. The categorized punctuation marks are modeled as composite typesetting elements carrying both width flexibility parameters and line break cost parameters. The line break cost and remaining compression capacity are correlated within the same element, and the composite typesetting element is incorporated into the cost optimization process of the line layout algorithm. The selection of line break points and punctuation spacing allocation are jointly determined in the same optimization calculation. A power-law cost function and hierarchical spacing allocation priority (punctuation compression takes precedence over word spacing adjustment, which in turn takes precedence over character spacing adjustment) are employed, and rules are implemented through hard constraints. Vertical layout mode, vertical-center-horizontal unit modeling, and cross-language configurable rules are supported. This invention solves the problem of suboptimal global typesetting quality caused by the decoupling of punctuation compression and line break algorithms in existing technologies.
Owner:BEIJING ADVANCED OPEN SOURCE TECHNOLOGY CO LTD

Character display and voice broadcast synchronization method and device, computer equipment, readable storage medium and program product

The invention relates to a text display and voice broadcast synchronization method and device, computer equipment, a readable storage medium and a program product. The method comprises the following steps: acquiring an original character string needing to be displayed on a current page of a client and / or an original character string already displayed on the current page of the client; punctuation marks in the original character string are filtered, and a pure character string corresponding to the original character string can be obtained; obtaining the word number of a read character string in the pure character string; then, according to the word number of the pure character string and the word number of the read character string, a broadcast demarcation point used for segmenting the read character string and an unread character string in the pure character string can be determined; and based on the document object model, positioning to the broadcast demarcation point in real time, and displaying the original character string corresponding to the broadcast demarcation point on the current page so as to realize synchronous display of character display and voice broadcast of the original character string.
Owner:CHINA LIFE INSURANCE CO LTD

Decoder Tool, System and Method for Deriving Divine Messaging

A decoder tool, system and method helps a user derive secondary meaning from an original Hebrew bible letter string. A primary sequence of letters is arranged devoid of spaces and punctuation to form the bible letter string. Each letter is associated with a numerical letter value. Reference words or phrases are also identified within the bible string, each of which correspond to a primary word or phrase value. The primary word value is correlated with at least one other secondary word value, which secondary word value is equated to the primary word value. The primary meaning of the reference word or phrase can then be interpreted in view of the at least one secondary word value, and the at least one secondary word value is associated with at least one secondary meaning for bolstering an understanding of the primary meaning.
Owner:ORIGINAL BIBLE FOUNDATION & CODE2GOD

Large language model prompt generation method based on knowledge graph optimization

The invention discloses a big language model prompt generation method based on knowledge graph optimization, and relates to the technical field of artificial intelligence natural language processing. The problems existing in an existing retrieval enhancement generation technology are solved. The method specifically comprises the following steps: preprocessing a text; extracting subject terms; performing context retrieval based on the knowledge graph; context pruning: pruning the extracted context by selecting a semantically most relevant context that can be used to answer a given prompt; generating retrieval enhancement; the text preprocessing comprises text cleaning, text marking, stop word deletion, word form and word drying, dictionary mapping and word bag construction; the content of text cleaning specifically comprises the following aspects: removing all characters and noise irrelevant to semantics; removing the URL link; carrying out line feed and redundant blank treatment; the characters and the noise comprise punctuations and special symbols. According to the method, the burden of the model is reduced, and the accuracy and correlation of the language model generation content are remarkably improved.
Owner:贵州省通信产业服务有限公司

Text sentence breaking method, system and related device based on combination of acoustics and semantics

The application provides a text punctuation method and system based on the combination of acoustics and semantics, and related equipment. The method comprises: obtaining audio data containing voice instructions; processing the audio data based on a preset voice recognition engine and a semantic model, identifying at least one candidate punctuation point in the text corresponding to the audio data, and obtaining a semantic punctuation probability of each candidate punctuation point; for each candidate punctuation point, extracting an acoustic feature set corresponding to the candidate punctuation point in the audio data; calculating a fusion punctuation score for the candidate punctuation point according to the semantic punctuation probability and the acoustic feature set; based on the fusion punctuation score, determining whether to perform a punctuation operation at the candidate punctuation point, and outputting the final text punctuation result of the audio data. The application fuses and calibrates pure text semantic analysis by introducing acoustic information, reduces the ambiguity of punctuation, and improves the accuracy of complex voice instruction punctuation.
Owner:SHENZHEN TONGXINGZHE TECH

Intelligent display voice-to-text method and system, and medium

The application discloses an intelligent display voice-to-text method and system and a medium, and the method comprises the following steps: when an audio and video file is captured, the audio and video file is converted into text; each sentence of text is segmented according to a pre-agreed punctuation mark to obtain segmented sentence text information; a pre-established dictionary table is searched; if historical data corresponding to the segmented sentence text information is found in the dictionary table, a corresponding previously combined text paragraph is obtained from the dictionary table; the text paragraph is rendered; single characters in the rendered text paragraph content are processed according to sensitive words, forbidden words or search words; and the processed single characters are synchronized with the text paragraph on a page for display. Through the application, a user can quickly find out where a violation is located, and can also quickly jump to a corresponding progress for manual auditing to confirm whether a problem exists, thereby greatly reducing the workload of the user for compliance processing.
Owner:SHENZHEN CRAFTSMAN NETWORK TECH CO LTD

A small language-oriented audio and video subtitle optimization generation method and system

PendingCN122313949ASoutheast asiaSpoken language
This invention relates to the field of artificial intelligence technology and discloses a method and system for optimizing and generating audio and video subtitles for less commonly spoken languages. The method includes the following steps: Step 1: Audio extraction and preprocessing; Step 2: Language recognition for the less commonly spoken language; Step 3: Language-aware punctuation restoration module; Step 4: Subtitle readability-driven segmentation module; Step 5: Subtitle format encapsulation and output. This application establishes a complete technical process encompassing audio preprocessing, language-aware speech recognition, structured text restoration, subtitle segmentation optimization, and time alignment. This method significantly improves the accuracy of speech-to-text conversion in less commonly spoken language videos and enhances the structural integrity and readability of subtitles. It is particularly suitable for practical application scenarios involving less commonly spoken languages ​​in regions such as Southeast Asia, where data is scarce and languages ​​are diverse. It has broad application value and practical significance in media dissemination, educational videos, and government services.
Owner:XINGZHOU DIGITAL TECH (ZHUHAI) CO LTD

Radio and television historical program voice translation system

The invention discloses a broadcast television historical program voice translation system, relates to the technical field of artificial intelligence, and solves the problem that an existing large model is difficult to be widely applied due to high price and high cost. A program history recording file is played back on a computer, an interface circuit is utilized to send an audio signal to an artificial intelligence system to be translated into characters, and then the characters are sent back to the computer to be stored through the interface circuit. The system supports a computer to simultaneously translate a plurality of sets of historical program recording files, reproduces a time axis of historical broadcast of each program on the computer, segments text contents according to punctuation marks, prints timestamps on the text contents, and stores the text contents in the system. Compared with video and audio content retrieval, the method has the advantages of intuition and high efficiency when the retrieval operation is executed by taking characters as carriers. The interface circuit further has the function of simulating manual operation of intelligent equipment, and accidental faults of third-party software can be automatically eliminated.
Owner:JILIN UNIVERSITY

Subtitle generation method, intelligent playing device, storage medium and computer program

The embodiment of the invention provides a subtitle generation method, intelligent playing equipment, a storage medium and a computer program, and the method comprises the steps: inputting audio information in audio and video contents into an ASR model for voice recognition, and obtaining text information outputted by the ASR model in a streaming manner; wherein the ASR model is obtained based on punctuation-removed corpus training, and the text information does not contain punctuation marks; determining audio time information corresponding to each piece of text information output by the ASR model in a streaming manner; determining target text information of voice pause and voice pause duration after the target text information according to the audio time information of the adjacent text information; and determining a corresponding punctuation mark according to the voice pause time length, and inserting the determined punctuation mark after the target text information to form a subtitle content containing the text information and the punctuation mark. According to the embodiment of the invention, the voice recognition accuracy can be ensured, the caption readability is considered, the real-time performance of caption generation is improved, and the audio and video experience of a user is improved.
Owner:JINGCHEN SEMICON SHENZHEN CO LTD

Text auxiliary analysis and evaluation method

The invention relates to a text auxiliary analysis and evaluation method. The method comprises the following steps: S1, data preparation: preprocessing a text, including stop word removal, punctuation mark processing, special character processing and case and small case conversion; s2, establishing a bag-of-words model, converting a text into a word frequency vector, and weighting by using TF-IDF; s3, word embedding: converting each word into a high-dimensional vector by respectively using Word2Vec and GloVe, and then performing result comparison; s4, performing feature selection, and performing screening according to importance or correlation of features to reduce noise and overfitting; s5, selecting a learning model, and importing the features for training; and S6, predicting new text data by using the trained model. According to the method, preprocessing, feature extraction, model training and evaluation are carried out on text data through a series of steps, and finally a model capable of being used for text auxiliary analysis and evaluation is obtained.
Owner:SUZHOU AEROSPACE INFORMATION RES INST

Method and apparatus for processing video data

Embodiments of the present disclosure provide a video data processing method and device, relating to the technical field of computer, which solves the problem of high error rate caused by current manual acceptance of video. The method comprises: obtaining video data to be put, wherein the video data comprises oral broadcast data and picture data; performing segmentation on the oral broadcast data through an asr interface and an open-source punctuation sentence segmentation algorithm to obtain segmented asr text; filtering invalid information in the picture data through an ocr interface and preset invalid information to obtain filtered ocr text; and obtaining a matching result of the video data according to the segmented asr text, the filtered ocr text and a preset word matching rule. The embodiments of the present disclosure are suitable for the acceptance process of brand parties for the video to be put.
Owner:特赞(上海)信息科技有限公司

Generating unified text using speech recognition models for conversational ai systems and applications

In various examples, generating unified text using speech recognition models for AI systems and applications is described herein. Systems and methods are disclosed that use a machine learning model that is trained to generate unified text associated with user speech, where the unified text includes punction marks, capitalizations of words, inverse text normalization formatting, end of sentence (EOS) detections, and / or end of utterance (EOU) detections. For instance, the machine learning model may receive audio data representing speech as input. The machine learning model may then process the audio data and, based at least on the processing, generate output data associated with the speech. In some examples, the output data may represent tokens, such as tokens associated with automatic speech recognition processing, punctuation and capitalization processing, EOS and / or EOU processing, and / or inverse text normalization processing. In such examples, the tokens may then be processed to generate the unified text.
Owner:NVIDIA CORP