Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

228 results about "Lexicon" patented technology

A lexicon, word-hoard, wordbook, or word-stock is the vocabulary of a person, language, or branch of knowledge (such as nautical or medical). In linguistics, a lexicon is a language's inventory of lexemes. The word "lexicon" derives from the Greek λεξικόν (lexicon), neuter of λεξικός (lexikos) meaning "of or for words."

Large language model generation content security test system and method in black box scene

The invention discloses a big language model generation content security test system and method in a black box scene. The system comprises a jailbreak prompt word library module used for storing jailbreak prompt words for performing security test on a big language model; the violation question and answer pair module is used for storing violation question and answer pairs covering different types; the response acquisition module is used for obtaining a data request packet according to query content formed by the jailbreak prompt word and the query request; the security analysis module is used for calculating the similarity between response data corresponding to the query request and an expected violation answer, taking the similarity as a security score, and inputting the security score into the adaptive optimization module; and the adaptive optimization module is used for optimizing the jailbreak prompt words output by the jailbreak prompt word bank module by using a genetic algorithm according to the security score output by the security analysis module. According to the method and the device, the security of the large language model generation content can be effectively tested.
Owner:CHINA ELECTRONICS TECH CYBER SECURITY CO LTD +1

Component identifier matching method and system based on NLP semantic segmentation and multi-level word bank

The invention discloses a component identifier matching method and system based on NLP semantic segmentation and a multi-level word library, and relates to the field of constructional engineering informatization, and the method comprises the steps: building a standard main word library and a preset compensation word library, carrying out the semantic segmentation of an input component identifier, and obtaining a segmented lexical element sequence, calculating semantic similarity and morphological similarity between the segmented lexical element sequence and entries in a standard main word bank, and performing standard term mapping based on the semantic similarity and the morphological similarity to obtain a standard term mapping result; carrying out structured data conversion on the segmented lexical element sequence which does not complete the standard term mapping by utilizing a regular expression mode to obtain structured data, and carrying out AI extension library matching on the segmented lexical element sequence which does not complete the structured data conversion to obtain a synonym matching result; and combining the standard term mapping result, the structured data and the synonym matching result to obtain component identification description. By implementing the method, the matching accuracy of the component identifier names in the engineering project can be improved.
Owner:CHINA CONSTR THIRD ENG BUREAU GRP CO LTD

Marketing mobile terminal dynamic behavior compliance monitoring method and system

The invention discloses a marketing mobile terminal dynamic behavior compliance monitoring method and system, belongs to the technical field of information transmission monitoring, and aims to solve the problems that existing monitoring is insufficient in non-text information processing, redundant in sensitive word judgment, incomplete in violation expression variant coverage and low in data reuse rate. The method comprises the steps that dynamic behavior information is captured, a composite retrieval code retrieval history library is constructed, and whether sensitive word recognition is conducted or not is judged; if not, generating a standardized word segmentation sequence; comparing the segmented words to determine matched segmented words and unmatched segmented words, calling a historical synonym set of the matched segmented words, positioning core semantics of the unmatched segmented words, screening evading expressions to generate a synonym set of the unmatched segmented words, and constructing a synonym data set; calling a matched word segmentation historical judgment result, establishing a synonym sequence calculation probability, and judging a sensitive word in combination with a significant proportion; executing a processing flow and updating the history library and the word segmentation library; according to the invention, efficient and accurate compliance monitoring is realized, the data reuse rate is improved, and the effectiveness of a monitoring closed loop is guaranteed.
Owner:JIANGSU ELECTRIC POWER INFORMATION TECH

Method, device and equipment for constructing dynamic expansion word bank and medium

The invention relates to the technical field of passenger service, and discloses a construction method and device of a dynamic extension word bank, equipment and a medium. The method comprises the following steps: acquiring multi-source heterogeneous original knowledge data related to civil aviation passenger service, and converting the multi-source heterogeneous original knowledge data into a standard format text; processing the unstructured data in the standard format text based on the adaptive word segmentation model and the stop word list to obtain a plurality of candidate words; performing multi-dimensional corpus feature quantitative analysis on each candidate word to determine candidate keywords; and performing similarity calculation based on the entity word segmentation and the candidate keywords, determining effective candidate keywords, and adding the effective candidate keywords into the dynamic expansion word bank. By means of the method and device, the technical problems that in the prior art, a word bank used by a civil aviation service system is usually based on a static vocabulary or depends on manual compiling and updating, the static word bank cannot be updated in real time, the characteristics of the civil aviation field cannot be flexibly handled, and the maintenance cost is high due to manual updating and maintenance are solved.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Intelligent labeling multi-dimensional retrieval system for education case library

The invention relates to the technical field of education informatization, and discloses an intelligent annotation multi-dimensional retrieval system for an education case library, which comprises a case storage server used for storing education texts, images and audio and video case resources; the intelligent labeling engine is connected with the case storage server, and performs three-order structured labeling on case resources through an element layering analysis model: extracting labels of a basic entity layer, a political theory layer and a value guiding layer; the knowledge graph construction module is used for constructing a multi-modal case knowledge graph which takes an element-teaching scene-childbearing target as an association relationship on the basis of the hierarchical label output by the intelligent labeling engine; the intelligent retrieval terminal comprises a semantic analysis unit and a multi-dimensional feedback unit. According to the method, an intelligent annotation engine pre-trains a theoretical word bank and an emotion analysis model through a BERT-BiLSTM-CRF model, and automatically completes three-order structured annotation to form a structured data system supporting advanced retrieval.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

Table instruction conversion system based on electric energy meter communication protocol

The invention discloses a table instruction conversion system based on an electric energy meter communication protocol, and the system comprises a data collection module which obtains and processes an electric energy meter communication protocol file and a historical manual processing instruction data set; the basic information reading module is used for extracting an input field and a data format and mapping a power grid instruction format to generate a compliance instruction; the instruction missing comparison module is used for comparing the compliance instruction with the historical data set and judging field missing content; the lexicon association and expansion module is used for extracting initial keywords and generating virtual expansion words covering missing scenes; the evaluation storage module is used for verifying the virtual expansion words and storing the virtual expansion words into an expansion associated word bank in a classified manner; the table content information extraction module is used for analyzing the file through an index strategy, automatically extracting content to fill the intelligent table and recording a source; the method has the advantages that electric energy meter communication protocol deep analysis, dynamic lexicon expansion and automatic table instruction generation can be carried out, and electric energy meter data processing refinement and real-time performance are improved.
Owner:ZHEJIANG REALLIN ELECTRON CO LTD

Data processing method and device, equipment, storage medium and program product

The invention provides a data processing method and device, equipment, a storage medium and a program product, and relates to the technical field of computers. The method comprises the steps that first question information input by a user is subjected to word segmentation processing and then subjected to semantic matching with a pre-configured lexicon, and second question information is determined; the second question information is normalized and vectorized and then subjected to similarity calculation with a pre-configured table library, an entity meeting a similarity threshold value and a standardized query template are determined, the pre-configured table library comprises a plurality of tables, and the standardized query template comprises a structure query language and a skeleton question; inputting the first question information, the entity and the standardized query template into a first large language model, and determining multi-table query information; and inputting the first question information and the multi-table query information into a second large language model, and determining response information. Based on the large language model, the cross-multi-table question can be effectively answered by deeply analyzing the user question and combining with the table library generated by multi-table preprocessing, and the problem that cross-multi-table answers cannot be effectively processed due to the fact that table question answering driven by the large language model mainly aims at a single table is solved.
Owner:CHINA TELECOM CORP LTD

Video Conferencing Network System and Method Using Natural Language Processing

PendingUS20260254670A1Networked systemProcessing
Systems and methods for performing non-statistical text analysis are provided. Essentially unconstrained natural language text may be segmented into individual terms, which are assigned a semantic category with the aid of a category lexicon and category association table. Categorized terms may be grouped into expressions if they commonly refer to a concept. Terms, as well as information about their categories and groupings into expressions, may be organized into information blocks, which each may contain a discrete piece of information about a text. Information blocks may be used to populate a database about the text. Information blocks may be queried to answer a question, be compared to identify inconsistencies in the text, be used to identify functions to execute from a user's command, or how to respond to a user during a dialog.
Owner:LAIKE INC

Teacher classroom speech recognition method and system based on hot word guidance, and readable storage medium

PendingCN121838769ASemantic analysisBiological modelsSpeech recognition performanceSpeech sound
The invention relates to the technical field of semantic recognition, in particular to a teacher classroom speech recognition method and system based on hot word guidance and a readable storage medium. According to the method, the hot word bank strongly related to the teaching scene is constructed, the hot words in the hot word bank are subjected to data processing to obtain the fusion features with prominent features, and the fusion features are input into the large language model for speech recognition, so that the model can improve the recognition precision by using the hot word information. The objective of the invention is to improve the classroom speech recognition performance of teachers.
Owner:SOUTHWEST FORESTRY UNIVERSITY

Electronic device lexicon word list scene word graphical user interface

1. The name of the design product: electronic device's word book word list scene word graph user interface. 2. The use of the design product: for an electronic device. 3. The design points of the design product: the interface content of the graphical user interface in the screen. 4. The picture or photo that best indicates the design points: front view. 5. The product carrier is the conventional design, and the rear view, left view, right view, top view, and bottom view are omitted. 6. The use of the graphical user interface: for the user to switch 10 scene word books, and manage the words of each word book through the word list. 7. The human-computer interaction mode of the graphical user interface: the front view is the initial interface. Click the expand word list details button in the front view to enter the interface change state diagram.
Owner:FOSHAN FUTURE CLASSROOM INFORMATION TECHNOLOGY CO LTD

Advanced technology maturity grade interval discriminant analysis method and device

The invention provides a leading-edge technology maturity level interval discriminant analysis method and device, and the method comprises the steps: obtaining a technology text which is pre-collected for a to-be-judged technology, and constructing an original data set; preprocessing the original data set to obtain an updated data set; inputting each piece of data in the updated data set into a preset data extraction model, wherein the data extraction model extracts triple data consisting of a technical direction, a technical dynamic state and a technical index in each piece of data in the updated data set; determining a transition level of the to-be-judged technology based on matching of a preset transition keyword library and the updated data set; technical information including the triple data and the transition level is supplemented into a preset structured JSON cue word template, the filled structured JSON cue word template is input into a preset large language model, and the large language model outputs a maturity judgment level.
Owner:MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY

Low-orbit satellite field term automatic discovery method based on local semantic heat map

The application provides a low-orbit satellite field term automatic discovery method based on a local semantic heat map, comprising the following steps: S1. corpus acquisition and preprocessing; S2. local semantic heat map construction; S3. convolution discrimination; S4. term screening and merging; and S5. vocabulary closed-loop updating and service. According to the application, the statistical feature image of a candidate fragment is converted into a local semantic heat map, and a convolutional neural network (CNN) is used for pixel-level discrimination, so that unregistered terms and combined terms are dynamically identified; on this basis, the incremental updating and closed-loop management of the term vocabulary are realized, so that the field adaptation capability and the explainability of downstream tasks such as word segmentation, NER, retrieval and RAG are improved.
Owner:SHANGHAI LINGSHU INTELLIGENT TECH CO LTD +2

Deep learning-based illegal advertising board detection method, device, equipment and medium

The application discloses a method, device and equipment for detecting illegal advertising boards based on deep learning and a medium, wherein the method comprises the following steps: constructing a sample set and a test set; inputting sample images in the sample set into a YOLOv7 model for detection, and outputting detected targets; performing character recognition on the detected targets, comparing all words in the recognized text with a pre-created word library, searching for words with the highest similarity in the word library as keywords and corresponding probability values; using a loss function to optimize the model training process; using sample images in the test set to test the optimized YOLOv7 model, and detecting illegal advertising boards for to-be-detected images after the test is completed. The application enhances the feature expression capability, uses keywords obtained through character recognition to assist in training, is beneficial to distinguishing illegal advertising boards from regular advertising boards, and improves the detection accuracy.
Owner:SHENZHEN ALL THINGS CLOUD TECH CO LTD

A text segmentation and sensitive word detection method based on matrix multiplication

ActiveCN115757721BEnergy efficient computingText database indexingAlgorithmDeterministic finite automaton
A text segmentation and sensitive word detection method based on matrix multiplication, comprising the following steps: obtaining an original text string and a sensitive word library; constructing a deterministic finite state automaton tree diagram of sensitive words according to the sensitive word library; converting the original text string into a text character two-dimensional matrix and recording the length of the text character two-dimensional matrix; constructing a matching two-dimensional matrix according to a horizontal matching rule, a vertical matching rule, an oblique matching rule and an inverse oblique matching rule, the length of the matching two-dimensional matrix being the same as that of the text character two-dimensional matrix; performing dot multiplication processing on the text character two-dimensional matrix and the matching two-dimensional matrix to obtain a corresponding result matrix; generating a corresponding matching text string according to the result matrix, and matching with the deterministic finite state automaton tree diagram to determine whether there is a sensitive word. The application supports interval text character information detection, improves the sensitive word detection accuracy and detection efficiency.
Owner:XUANCAI INTERACTIVE NETWORK SCI & TECH

A text-guided face spoofing detection method and system

This application provides a text-guided method and system for detecting face forgery. The method includes obtaining a face image to be detected (whether it is genuine or fake), inputting the face image into a face detection model to obtain the face authenticity detection result. The face detection model includes: constructing a text prompt lexicon covering multiple granularities and generating multi-dimensional text prototypes; extracting visual features, optimizing the feature distribution of visual features to obtain global visual features; performing feature separation and enhancement on the global visual features; mapping the global visual features after feature separation and enhancement to predicted text features; applying similarity constraints to obtain cross-modal prototype matching results; applying discriminative constraints on different-dimensional text prototypes to obtain feature measurement learning results; and outputting the detection result based on the cross-modal prototype matching results and feature measurement learning results. This method solves the problem of poor generalization in face forgery detection and improves detection accuracy and generalization ability.
Owner:NANJING UNIV OF POSTS & TELECOMM

Parcel stack address full word association method and device, equipment and storage medium

The present application relates to the technical field of intelligent logistics, and particularly relates to a parcel pile address full word association method and device, equipment and a storage medium, the method first carries out data preprocessing to the original address data set obtained, obtains a plurality of address full word entries, respectively allocates a unique address full word identifier to each address full word entry, and is one-to-one associated and bound with the delivery code region number set obtained, obtains a standardized address full word library, determines a parcel pile in response to the operation instruction of the courier, obtains at least one address full word entry and the unique address full word identifier corresponding to the address full word entry from the standardized address full word library, and constructs an association relationship database based on the corresponding unique address full word identifier and the number of the parcel pile, matches the obtained address full word to be matched based on the standardized address full word library, if the matching is successful, determines the parcel pile information corresponding to the address full word to be matched based on the association relationship database, and aims to improve the distribution efficiency.
Owner:SHANGHAI DONGPU INFORMATION TECH CO LTD

Multi-scene-oriented privacy extraction method and device and medium

The invention discloses a privacy extraction method and device for multiple scenes and a medium, and relates to the technical field of privacy information processing. The method comprises the steps of obtaining to-be-processed data input by a user, and preprocessing the to-be-processed data to obtain standard to-be-processed data; inputting the standard to-be-processed data into a preset named entity recognition model, and recognizing semantic entities in the standard to-be-processed data in combination with a preset multi-scene sensitive word library; inputting the semantic entities into a preset relation extraction model to extract a relation between the semantic entities; and calculating a total weight value of the entity initial sensitivity, the relation association degree and the scene adaptation degree by using a preset initial weight coefficient, comparing the total weight value with a preset privacy extraction threshold value, and performing privacy extraction on the standard to-be-processed data. According to the method, the exclusive sensitive words and rules are constructed for various scenes, the incidence relation between the entities is extracted through the relation extraction model, and the compliance and accuracy of privacy extraction are ensured by quantifying the weight.
Owner:DIGITAL INTELLIGENCE CLOUD ALLIANCE (SHANDONG) DIGITAL TECHNOLOGY CO LTD

Conversation generation method and device, computer equipment, storage medium and program product

The invention relates to a dialogue generation method and device, computer equipment, a storage medium and a program product, and the method comprises the steps: obtaining a plurality of target person feature words corresponding to the role description information of a digital person according to a preset person feature word library; combining the feature words set by each target person to obtain a plurality of feature word groups; matching the plurality of feature phrases with a preset dialogue scene library to determine a feature phrase set; under the condition that the current output dialogue is the first section dialogue, determining a first feature phrase; generating a current output dialogue based on the first feature phrase through a dialogue generation model; under the condition that the current output dialogue is any non-first-section dialogue, determining a second feature word group; and through the dialogue generation model, based on the second feature phrase and the output dialogue before the current output dialogue, the consistency of long-range dialogue role features is maintained.
Owner:GUANGZHOU QUWAN NETWORK TECH CO LTD

Physiological index self-adaptive health assessment method and device driven by user feedback

PendingCN121862406ASolve health risk identification challengesImprove early abnormality detection rateHealth-index calculationSensorsContact sensorMedical emergency
The invention relates to the technical field of physiological index data processing, in particular to a user feedback driven physiological index adaptive health assessment method and device. The method comprises the following steps: acquiring a BCG signal of a user based on a non-contact sensor, and acquiring index monitoring data of a plurality of indexes based on the BCG signal; the user inputs a user feedback text representing the abnormal state of the body, matching is conducted on the basis of the user feedback text and the disease words in the lexicon, and if the disease words are matched, indexes corresponding to the disease words serve as final monitoring indexes; if the disease word cannot be matched, screening out an initial monitoring index based on the user feedback text, and determining a final monitoring index based on the initial monitoring index; and obtaining daily index monitoring data based on the final monitoring index. In this way, mapping from coarse-grained subjective perception to refined objective indexes can be realized, and monitoring blind areas and misjudgment risks are remarkably reduced while the early-stage anomaly detection rate is increased.
Owner:ZHEJIANG QISHENG DATA SERVICE CO LTD

A method for generating a pinyin bucket word library of a four-level index

PendingCN122452553ADigital dataData integrity
The present application relates to the technical field of electronic digital data processing, and discloses a four-level index pinyin bucket word library generation method, acquires a Chinese word library text, extracts the first Chinese character and the first pinyin of each line of words as a classification key, and generates a triple; the triple is classified into a corresponding pinyin bucket according to the pinyin, the words in the bucket are sorted in descending order of word frequency, an independent Chinese character block is generated for each unique Chinese character, and the words are solidified according to the word length by using a first character multiplexing mechanism; a four-level static index of an initial letter statistical area and a pinyin index area and a Chinese character word index area and a word storage area is sequentially constructed, each index area uses fixed-length entries and absolute offset addressing; a metadata area containing a file length double-semantic field is constructed, a high-frequency-first deterministic truncation is performed on the word library according to a preset threshold, each data area is spliced and a data integrity check value is appended, and an embedded binary word library is generated. The problems of high storage redundancy and uncontrollable memory are solved, and the purposes of deterministic analysis, resource adaptation and high security are achieved.
Owner:SICHUAN HAIGE HENGTONG PRIVATE NETWORK TECH CO LTD

Lyrics translation module for player page word translation presentation and interactive graphical user interface for electronic devices

1. Name of the product in this design: Lyrics translation module for word translation display and interactive graphical user interface on player page of electronic device. 2. Purpose of this design: An electronic device. 3. The key design features of this product are: the parts in the graphical user interface, and the parts not marked with dotted lines are the parts for which protection is required. 4. The picture or photo that best illustrates the key design points: Design 1 front view. 5. Design 1 is designated as the basic design. 6. Purpose of the graphical user interface: The overall appearance design of the interface is used to display the song playback page, display lyrics, control song playback, and display word translations based on the dictionary (words marked in Chinese lyrics will be switched to English words with Chinese translations, while words marked in foreign language lyrics will directly display the Chinese translations); the partial appearance design of the interface is used to display lyrics, control song playback, and display word translations based on the dictionary (words marked in Chinese lyrics will be switched to English words with Chinese translations, while words marked in foreign language lyrics will directly display the Chinese translations). 7. Human-computer interaction method of graphical user interface: In the main view of Design 1, the foreign language lyrics are displayed in the middle of the interface. The corresponding Chinese translation is displayed below the words marked by the dictionary in the lyrics. Users can click on the lyrics area to bring up the word list module. Design 2 uses the same human-computer interaction method as Design 1. In the main view of Design 3, the Chinese lyrics are displayed in the center of the interface. The words marked in the lyrics are switched to English words with Chinese translations. Users can click on the lyrics area to bring up the word list module. Design 4's main view is a landscape playback page, with foreign language lyrics displayed in the center of the interface. The corresponding Chinese translations are displayed below the words marked in the lyrics according to the dictionary. Users can click on the lyrics area to bring up the word list module. In the main view of Design 5, when the user clicks on the blank area in the middle of the interface, the lyrics are displayed in an immersive state, showing the interface changes from the main view of Design 5 to the interface change diagram of Design 5.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

A traditional Chinese medicine auxiliary syndrome differentiation system based on semantic analysis

PendingCN122388177ASemantic gapMedicine
The application belongs to the technical field of data retrieval, and particularly relates to a traditional Chinese medicine auxiliary syndrome differentiation system based on semantic analysis. The system comprises a traditional Chinese medicine symptom preprocessing module, a dynamic inverted index construction module and a semantic retrieval matching module. The preprocessing module extracts symptom entities and degree modifiers of the inquiry text through a traditional Chinese medicine field dependency syntax analysis tree, maps them to a standard symptom word library and converts them into weight coefficients. The dynamic inverted index construction module establishes an inverted list with the standard symptom words as index keys, and appends the bias values of the basic syndrome vectors and the document syndrome vectors after the document identification. The semantic retrieval matching module generates a query vector with weight coefficients according to a symptom query request, calculates the inner product of the query vector and the bias values in the inverted list, eliminates the document identification below the preset threshold, and outputs the syndrome conclusion corresponding to the remaining documents. The application solves the semantic gap and the synonym omission problem caused by the non-standardized expression of traditional Chinese medicine.
Owner:BEIJING JIANHUI SMART MEDICAL TECHNOLOGY CO LTD

Information retrieval prompting method and device, electronic equipment and storage medium

ActiveCN116402044BImprove multi-dimensional promptsImprove the ability of accurate searchDigital data information retrievalNatural language data processingThe InternetEngineering
The application provides an information retrieval prompting method and device, electronic equipment and a storage medium, the method comprising: preprocessing device retrieval information input by a user to determine a retrieval keyword; sequentially matching the retrieval keyword with a prompt word library, a synonym library and a pinyin word library to obtain multiple prompt information corresponding to the device retrieval information; the prompt word library is constructed based on attribute information of various devices, the synonym library is constructed based on synonymous words corresponding to the attribute information of various devices, and the pinyin word library is determined based on the prompt word library; and the multiple prompt information is sent to a front end for display. The application can provide accurate retrieval prompt information for a user, help the user quickly and effectively retrieve attribute information of an Internet of Things device of interest, improve the multi-dimensional prompting and accurate retrieval capability of the Internet of Things device, and provide a good user experience.
Owner:INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES

Domain large model harmful cue word generation method based on knowledge graph

The invention discloses a field large model harmful cue word generation method based on a knowledge graph. The method comprises the following steps: screening a universal harmful cue word data set based on a constructed risk knowledge graph to obtain a seed harmful cue word library; processing the field corpus to obtain an embedded context; generating candidate harmful cue words through a synthesis model based on seed harmful cue words, risk entities, embedded contexts and examples, cleaning and enhancing toxicity indexes to obtain high-risk harmful cue words, and adding the high-risk harmful cue words into a seed harmful cue word library; and based on the semantic relevancy between the seed harmful cue words and the risk entities and the toxicity scores of the seed harmful cue words, screening out cue word input of a next round, and constructing an iteratively updated field harmful cue word data set. The method has the characteristics of high automation, multi-dimensional evaluation, controllable generation and the like, multi-round man-machine collaborative prompt word construction can be realized, and the red team drilling efficiency and the safety test quality of a large language model in a specific application field are remarkably improved.
Owner:ZHEJIANG UNIV +1

A professional word segmentation method applied to nuclear power industry

ActiveCN116306611BSolve the problem of inaccurate word segmentationreduce investmentData processing applicationsSemantic analysisProcess engineeringChinese word
The application belongs to the field of natural language processing in the nuclear power industry, and particularly relates to a professional word segmentation method applied to nuclear power industry corpus, which comprises nuclear power professional vocabulary construction, nuclear power stop word vocabulary construction, nuclear power synonym vocabulary construction, nuclear power same pronoun vocabulary construction, nuclear power field new word recognition, nuclear power field entity automatic recognition, nuclear power field synonym automatic recognition, nuclear power corpus Chinese accurate word segmentation and the like operations. The application has the beneficial effect of completely solving the problem of inaccurate Chinese word segmentation of nuclear power industry corpus, laying a solid foundation for the application of subsequent big data and artificial intelligence in the field of nuclear power natural language processing, and reducing the investment of other practitioners in the field of nuclear power in natural language processing.
Owner:CHINA NUCLEAR POWER OPERATION TECH CORP

A data transmission system based on artificial intelligence services

The application discloses a data transmission system based on artificial intelligence service, and relates to the technical field of data transmission.The system comprises a data acquisition module, a word library construction module, a word segmentation comparison module, a transmission analysis module, a conflict analysis module and a transmission optimization module.The data acquisition module performs word segmentation on an artificial intelligence service data packet to construct a word segmentation data group.The model construction module constructs a key word library containing weights according to a scene.The word segmentation comparison module calculates word frequency and weight ratio to obtain a word group priority value to determine a data packet transmission sequence.The transmission analysis module compares word frequency of data packets with network maximum byte values to determine conflicts.The conflict analysis module calculates vector dot products and continuous invalid word frequency Wx to determine the priority sequence of equivalent data packets.The transmission optimization module optimizes data packets and sends the data packets to a network.The application can optimize data packet transmission, improve efficiency and provide a basis for data transmission optimization.
Owner:王振

Railway signal fault text enhancement method, system and equipment based on improved EDA

The invention discloses a railway signal fault text enhancement method, system and device based on improved EDA, and the method comprises the following steps: constructing a railway field entity dictionary, carrying out the proper noun recognition of a preprocessed to-be-enhanced railway signal fault text, and outputting a structured text; on the basis of a railway signal fault related text data set, a similar word bank of railway signal fault text related vocabularies is generated by adopting a Word2Vec model and used for replacing non-professional vocabularies in railway signal fault text statements, and enhancement operation is performed on the basis of a structured text and the similar word bank of the railway signal fault text related vocabularies by adopting an EDA improvement method. Enhancing the railway signal fault text to be enhanced; and screening the enhanced railway signal fault text. According to the method, the integrity and accuracy of domain terms are ensured through an entity fixing mechanism, an enhanced sample conforming to a real scene is generated in combination with a semantic constraint editing strategy, and the problem of unbalanced data distribution is effectively solved.
Owner:YANSHAN UNIV

Voice AI quality inspection method and platform

The invention relates to the technical field of voice AI quality inspection, and discloses a voice AI quality inspection method and platform, and the method comprises the steps: carrying out the Whisper voice recognition transcription of a first voice signal, obtaining a transcription text, carrying out the acoustic emotion feature extraction of the first voice signal, and obtaining an emotion acoustic parameter; inputting the transcriptional text into a large language model for context semantic analysis and sensitive lexicon matching to obtain text semantic data; generating a voice quality inspection result according to the emotional acoustic parameters and the text semantic data, and determining a second voice signal according to the voice quality inspection result; and performing Laplacian noise addition and differential privacy on the second voice signal in combination with the voice quality inspection result to generate a voice quality inspection report, the method can more comprehensively capture key information in customer service dialogues, improve the accuracy and integrity of quality inspection analysis, improve the organization efficiency and query convenience of the quality inspection result, and improve the user experience. And an efficient data access and analysis tool is provided for quality inspection management personnel.
Owner:GUANGDONG ICAR GUARD INFORMATION TECH

Method for constructing a word dictionary based on a double array tree and a word query method

The application provides a double-array tree-based dictionary construction method and word query method, which comprises the following steps: generating a sensitive word tree based on the split text characters of each sensitive word and the arrangement order of the text characters, wherein each path in the sensitive word tree corresponds to a sensitive word; performing dictionary coding on the split text characters of the sensitive word after deduplication to obtain the character coding corresponding to each text character; determining the base coding of each character node in the sensitive word tree in layers according to at least one of random selection trial calculation, self-increment selection trial calculation, overall query selection trial calculation and preset data selection trial calculation; determining the storage location coding of the text character of each character node in the storage space based on the character coding corresponding to the text character of each character node and the base coding of the parent node of each character node, and taking the storage location coding of the parent node of each character node as the check coding of each character node, so as to construct a sensitive word dictionary.
Owner:SHANGHAI XIYU JIZHI TECH CO LTD