Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

392 results about "Lexicon" patented technology

A lexicon, word-hoard, wordbook, or word-stock is the vocabulary of a person, language, or branch of knowledge (such as nautical or medical). In linguistics, a lexicon is a language's inventory of lexemes. The word "lexicon" derives from the Greek λεξικόν (lexicon), neuter of λεξικός (lexikos) meaning "of or for words."

Voice data interaction feedback control processing method based on large language model

The invention relates to the technical field of large language models, and discloses a voice data interaction feedback control processing method based on a large language model. The method comprises the following steps: acquiring an industrial voice instruction through an AMR main controller, and obtaining a standardized vector through industrial lexicon matching and intention classification; inputting an industrial large language model for reasoning processing, and generating an AMR execution scheme; carrying out distributed coordination and task allocation on the AMR cluster to form a control instruction sequence; and the motion controller executes monitoring, processes exceptions and outputs a feedback strategy. Accurate classification and semantic understanding of complex industrial instructions are achieved, the processing capacity of professional knowledge in the industrial field is improved, and meanwhile high-precision real-time state monitoring is achieved.
Owner:TIANJIN HONGHUANG TECH CO LTD

Context-aware, artificial intelligence-based system for increasing employee engagement and automating the integration of business processes.

A system for improving employee engagement and automating the integration of business processes. The system includes: a user interaction module configured to receive multimodal inputs, including natural language text and voice requests from employees via enterprise portals, mobile applications, voice-activated devices, and collaboration platforms, and to generate real-time responses via conversational output channels; a context processor that is operationally coupled with the user interaction module, wherein the context processor stores the interaction history, employee preferences, organizational role metadata, and task results in both short-term and long-term memory and dynamically derives context for controlling responses and initiating workflows; A natural language and intent processor that is communicatively linked to the context processor. The intent engine consists of large language models trained on company-specific lexicons to perform intent detection, entity extraction, ambiguity resolution, and sentiment or urgency classification from employee queries; a workflow orchestration layer in communication with the intent engine and the context processor, wherein the workflow orchestration layer includes a rule-based and AI-powered process execution engine configured to automatically initiate, route, and complete cross-functional business tasks, including approvals, escalations, and compliance checks, with workflows defined using modular templates that include conditional logic and time-based triggers; Webhooks and authentication protocols provide secure interoperability between the workflow orchestration layer and external enterprise platforms; a feedback and learning module that is operationally coupled with the workflow orchestration layer and the natural language understanding engine, wherein the feedback and learning module is configured to analyze task completion rates, response accuracy, latency metrics, and user feedback signals, and to retrain underlying language and workflow models for adaptive improvement in real time; and an administration console with role-based access controls, compliance dashboards, audit trails and interfaces for workflow configuration, whereby the administration console enables authorized personnel to monitor system operations, adjust interaction rules and enforce data protection restrictions across departments.
Owner:KURAPATI SURESH KHAMMAM +3

Retrieval method and device based on static word embedding, computer equipment and medium

The invention relates to a retrieval method and device based on static word embedding, computer equipment, a computer readable storage medium and a computer program product. The method comprises the steps of performing word segmentation on an original text to obtain a first word segmentation result, training a static word embedding model by utilizing the first word segmentation result to obtain word vectors, and generating a synonym word library; the method comprises the following steps: establishing a full-text inverted index by utilizing an original text, and expanding query words by utilizing a synonym library in a retrieval stage; encoding the original text into a semantic vector by using a semantic generation model, and constructing a vector index based on the semantic vector; based on a to-be-queried text in a user query request, performing retrieval by using the full-text inverted index to obtain a first candidate document, and performing retrieval by using the vector index to obtain a second candidate document; performing fusion processing on the first candidate document and the second candidate document to obtain a target candidate document; and inputting the target candidate document into the text generation model to obtain a retrieval result. By adopting the method, the accuracy of text retrieval can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Large language model generation content security test system and method in black box scene

The invention discloses a big language model generation content security test system and method in a black box scene. The system comprises a jailbreak prompt word library module used for storing jailbreak prompt words for performing security test on a big language model; the violation question and answer pair module is used for storing violation question and answer pairs covering different types; the response acquisition module is used for obtaining a data request packet according to query content formed by the jailbreak prompt word and the query request; the security analysis module is used for calculating the similarity between response data corresponding to the query request and an expected violation answer, taking the similarity as a security score, and inputting the security score into the adaptive optimization module; and the adaptive optimization module is used for optimizing the jailbreak prompt words output by the jailbreak prompt word bank module by using a genetic algorithm according to the security score output by the security analysis module. According to the method and the device, the security of the large language model generation content can be effectively tested.
Owner:CHINA ELECTRONICS TECH CYBER SECURITY CO LTD +1

Steel bridge disease detection and identification method based on large language model

The invention relates to a steel bridge disease detection and identification method based on a large language model, and belongs to the technical field of artificial intelligence and civil engineering crossing. According to the method, a cross-modal feature alignment mechanism is constructed through a pre-trained multi-modal large language model by fusing a steel bridge image and a field customized text prompt, and a cascade detection process of'component identification-disease classification-region segmentation 'is realized. Comprising the following steps: designing a structured text prompt word bank to enhance semantic consistency, and dynamically fusing general knowledge and instance features in combination with a mixed prompt mechanism; a multi-level cross-modal alignment strategy is adopted to generate an anomaly graph, and a disease area is accurately positioned; a visual prompt enhancement module is introduced to improve the multi-scale feature discrimination ability, and the robustness in a complex environment is adjusted and optimized through data self-adaption. Under the condition of few samples or even zero samples, high-sensitivity detection and pixel-level segmentation of steel bridge cracks, corrosion and other diseases are achieved, and the problems that a traditional method is low in efficiency, poor in generalization, high in labor cost and the like are effectively solved.
Owner:HEBEI UNIV OF TECH

A / B test question and answer method and device based on large model agent

The invention relates to an A / B test question answering method and device based on a large model agent, and relates to the technical field of large models, agents and computers. The method comprises the following steps: displaying an intelligent interaction page associated with an intelligent agent; displaying a plurality of prompt word options in response to a trigger operation on a prompt word bank in the intelligent interaction page; in response to a trigger operation on a target cue word option in the plurality of cue word options, at least displaying a target cue word template corresponding to the target cue word option; in response to a content editing operation on the target cue word template, filling a target content corresponding to the content editing operation into the target cue word template to obtain a target cue word; and in response to a sending operation on the target cue word, displaying a target answer of the intelligent agent to the target cue word through the large model on the intelligent interaction page. The expression logic of the user for the intention can be optimized through the cue word, the content quality of questions is improved, then the accuracy of answers generated by the large model is improved, and the use experience of the user is improved.
Owner:BEIJING VOLCANO ENGINE TECH CO LTD

Component identifier matching method and system based on NLP semantic segmentation and multi-level word bank

The invention discloses a component identifier matching method and system based on NLP semantic segmentation and a multi-level word library, and relates to the field of constructional engineering informatization, and the method comprises the steps: building a standard main word library and a preset compensation word library, carrying out the semantic segmentation of an input component identifier, and obtaining a segmented lexical element sequence, calculating semantic similarity and morphological similarity between the segmented lexical element sequence and entries in a standard main word bank, and performing standard term mapping based on the semantic similarity and the morphological similarity to obtain a standard term mapping result; carrying out structured data conversion on the segmented lexical element sequence which does not complete the standard term mapping by utilizing a regular expression mode to obtain structured data, and carrying out AI extension library matching on the segmented lexical element sequence which does not complete the structured data conversion to obtain a synonym matching result; and combining the standard term mapping result, the structured data and the synonym matching result to obtain component identification description. By implementing the method, the matching accuracy of the component identifier names in the engineering project can be improved.
Owner:CHINA CONSTR THIRD ENG BUREAU GRP CO LTD

Multi-scale sea wave height prediction method and system based on large language model

The invention belongs to the technical field of sea wave height prediction, and discloses a multi-scale sea wave height prediction method and system based on a large language model, and the method comprises the steps: S1, carrying out the decoupling interaction based on an amplitude phase, and obtaining enhanced sea wave height data; s2, extracting based on multi-scale semantic information: constructing a feature pyramid structure for the enhanced sea wave height data, outputting sea wave self-encoding features of different scales, performing self-encoding on a lexicon of a large language model, outputting word vector features, and obtaining the multi-scale semantic information by using a cross-modal attention mechanism and a dynamic fusion mechanism; and S3, a step of large-scale language model prediction based on sea wave height domain cue word driving: inputting the multi-scale semantic information and the domain specific semantic features into a large-scale language model, and outputting prediction features. According to the method, sea wave complex spatial-temporal change information is effectively mined, and accurate prediction of the sea wave height is realized.
Owner:OCEAN UNIV OF CHINA

Marketing mobile terminal dynamic behavior compliance monitoring method and system

The invention discloses a marketing mobile terminal dynamic behavior compliance monitoring method and system, belongs to the technical field of information transmission monitoring, and aims to solve the problems that existing monitoring is insufficient in non-text information processing, redundant in sensitive word judgment, incomplete in violation expression variant coverage and low in data reuse rate. The method comprises the steps that dynamic behavior information is captured, a composite retrieval code retrieval history library is constructed, and whether sensitive word recognition is conducted or not is judged; if not, generating a standardized word segmentation sequence; comparing the segmented words to determine matched segmented words and unmatched segmented words, calling a historical synonym set of the matched segmented words, positioning core semantics of the unmatched segmented words, screening evading expressions to generate a synonym set of the unmatched segmented words, and constructing a synonym data set; calling a matched word segmentation historical judgment result, establishing a synonym sequence calculation probability, and judging a sensitive word in combination with a significant proportion; executing a processing flow and updating the history library and the word segmentation library; according to the invention, efficient and accurate compliance monitoring is realized, the data reuse rate is improved, and the effectiveness of a monitoring closed loop is guaranteed.
Owner:JIANGSU ELECTRIC POWER INFORMATION TECH

Large model output content security test method and device

The invention relates to the field of large model security testing, and particularly provides a large model output content security testing method and device, and the method comprises the following steps: S1, preparing and managing a test set, a sensitive word library and a regular expression which are required by testing; s2, reading a test set, and obtaining a large model output result according to the test set and the large model interface information; s3, judging whether the output content of the large model is safe or not according to the sensitive lexicon and the regular expression; s4, extracting semantic risk features according to the output content of the large model by using the large model and the oriented Prompt, and automatically storing the semantic risk features after confidence verification; and S5, storing the information result of each request in a file. Compared with the prior art, the test time can be shortened, and the evaluation efficiency can be improved; and the security of the output content of the large model can be effectively evaluated by using a method for dynamically constructing the sensitive word bank by using the output result of the large model.
Owner:INSPUR QILU SOFTWARE IND

Voice navigation model training method and device, navigation method and vehicle-mounted navigation system

The invention discloses a voice navigation model training method and device, a navigation method and a vehicle-mounted navigation system. The method comprises the following steps: constructing a language database comprising a point-of-interest core word library, an easily confused word mapping list and a corrected sentence pattern template library for correcting navigation points of interest; calling a specified model to generate an annotation data set based on a pre-constructed cue word project and a language database; inputting the labeled data set into a pre-constructed navigation model for iterative training to obtain a target navigation model; the target navigation model is used for performing initial interest point recognition and initial interest point correction on a voice navigation instruction output by a user and determining a target interest point. Therefore, the core word library provides a point-of-interest vocabulary basis, and the easy-to-confuse word mapping list effectively solves the problem of recognition errors caused by non-standard pronunciation of a user and the like. Meanwhile, due to the introduction of the corrected sentence pattern template library, the interest points can be corrected, the navigation failure caused by the input error of the user is reduced, and the navigation precision is improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Three-dimensional point cloud recognition method based on visual angle specific cue word and related equipment

The invention discloses a three-dimensional point cloud recognition method and related equipment based on a visual angle specific cue word, and the method comprises the steps: obtaining point cloud data, carrying out the projection of a point cloud through a plurality of preset visual angles, and obtaining a multi-visual-angle image; inputting the multi-view image into a contrast language-image pre-training image encoder to obtain a multi-view feature; obtaining view angle specific cue words from a preset multi-source view angle specific cue word library, wherein each view angle corresponds to a group of view angle specific cue words; inputting the visual angle specific prompt words into a contrast language-image pre-training text encoder to obtain text features; and performing similarity calculation according to the multi-view features and the text features, and obtaining an identification result according to the calculated similarity. According to the method, the point cloud is projected into the multi-view image, and geometric information is reserved; and distributing a special cue word for each view angle, and carrying out similarity calculation on the point cloud multi-view-angle features and the multi-view-angle semantic features to realize high-precision recognition under the condition of zero samples or few samples.
Owner:SOUTH CHINA UNIV OF TECH

Data query method and device based on text conversion, medium and equipment

The invention discloses a data query method and device based on text conversion, a medium and equipment, relates to the technical field of big data and financial science and technology, and mainly aims at solving the problem that existing SQL statement query is poor in effectiveness. Comprising the steps of obtaining to-be-queried text content; the text content is labeled on the basis of a labeling model after model training is completed, text segmented words with labels are obtained, the labeling model is obtained by training a large language model on the basis of labeling samples, and the labeling samples comprise sample data with intention labels, condition labels, table and / or field labels and function labels; determining query logic matched with the tag according to a preset query statement conversion relationship, and determining a query object matched with the text segmented word based on a preset query statement word bank; and combining the query logic with the query object to obtain a query statement, and querying a query result matched with the text content based on the query statement.
Owner:PING AN PAY ELECTRONIC PAYMENT CO LTD

Methods and systems for generating textual outputs from images

Embodiments of the present disclosure provide systems and methods for performing text extraction from an image including textual data. The method performed by a processor includes extracting machine-readable textual data from the image. The machine-readable textual data includes one or more words. The method includes comparing each of the one or more words with a dataset including a domain lexicon database and a language dictionary database to determine a first set of words and a second set of words. The first set of words is words successfully matching with words available in the dataset, and the second set of words is words with no successful match with words available in the dataset. Further, the method includes splitting at least one word of the second set of words into two or more words to determine a third set of words and generating a textual output associated with the image.
Owner:MAERSK AS

Method, device and equipment for constructing dynamic expansion word bank and medium

The invention relates to the technical field of passenger service, and discloses a construction method and device of a dynamic extension word bank, equipment and a medium. The method comprises the following steps: acquiring multi-source heterogeneous original knowledge data related to civil aviation passenger service, and converting the multi-source heterogeneous original knowledge data into a standard format text; processing the unstructured data in the standard format text based on the adaptive word segmentation model and the stop word list to obtain a plurality of candidate words; performing multi-dimensional corpus feature quantitative analysis on each candidate word to determine candidate keywords; and performing similarity calculation based on the entity word segmentation and the candidate keywords, determining effective candidate keywords, and adding the effective candidate keywords into the dynamic expansion word bank. By means of the method and device, the technical problems that in the prior art, a word bank used by a civil aviation service system is usually based on a static vocabulary or depends on manual compiling and updating, the static word bank cannot be updated in real time, the characteristics of the civil aviation field cannot be flexibly handled, and the maintenance cost is high due to manual updating and maintenance are solved.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Intelligent report generation method, system and equipment based on NL2SQL (Non-Layer 2Structured Query Language) and medium

The invention provides an intelligent report generation method, system and device based on NL2SQL and a medium, and belongs to the technical field of data processing and report generation. The method comprises the steps that data of all business systems of an enterprise are obtained, scene tables are constructed according to five dimensions of personnel, equipment, products, quality and production, and the corresponding relation between fields of each scene table and business system fields is determined; regularly extracting data from a business system in an incremental extraction mode, processing the data, and storing the processed data in a corresponding scene table; in response to a natural language query request of a user, identifying a query intention, matching and retrieving a pre-constructed scene SQL template library and a scene prompt word library through the query intention, and determining an associated scene table; and generating an SQL code based on a matching retrieval result, performing auditing, executing the SQL code passing the auditing to the associated scene table to generate a report, and displaying the generated report to the user in a front-end adaptation mode. According to the invention, enterprise data integration and report automatic generation are realized, and efficiency and accuracy are improved.
Owner:山东浪潮智能生产技术有限公司

Literature retrieval method based on large language model

The invention discloses a literature retrieval method based on a large language model, and belongs to the technical field of large language models. The method comprises the steps that the large language model obtains a natural language text from a user; according to the natural language text, a domain knowledge base of the target domain is constructed, and the domain knowledge base comprises a core definition corresponding to the target domain, a multi-dimensional lexicon, a theme retrieval formula, an ambiguity lexicon and a constraint rule base; performing literature retrieval according to a theme retrieval formula in the domain knowledge base to obtain a first literature set of the target domain; based on a core definition, an ambiguous word library and a constraint rule in the domain knowledge base, screening literatures in the first literature set to obtain a second literature set; and generating a literature analysis report based on the second literature set. According to the method, large-range and high-accuracy literature retrieval can be realized.
Owner:WUHAN UNIV

Header field intelligent benchmarking method, system and device based on semantic index segmentation

The invention discloses a header field intelligent benchmarking method and system based on semantic index segmentation, relates to the technical field of data management, and aims to solve the problems that in the prior art, header field standardization depends on manpower, efficiency is low, and semantic fuzzy and non-Chinese fields are difficult to process. An intelligent benchmarking scheme fusing semantic comprehension and vector matching is provided. According to the method, header fields are decomposed through a large language model to generate an initial word bank, and a historical benchmarking database is constructed to realize rapid matching; semantic segmentation is carried out by adopting a fine tuning BERT model, and precise benchmarking is completed in combination with vector library retrieval and Top-K recommendation; code set recognition is achieved by applying regular rules and semantic logic for non-Chinese fields. The system comprises a lexicon construction module, a history matching module, a semantic segmentation module, a vector matching module and the like. The device comprises hardware units such as a processor and a memory. According to the invention, automatic standardized processing of header fields is realized, and the data governance efficiency is improved.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Network chat sensitive word auditing method and system based on multi-dimensional recognition

ActiveCN120745646ASemantic analysisWord identificationEngineering
The invention belongs to the technical field of network chat auditing, and provides a network chat sensitive word auditing method and system based on multi-dimensional recognition, and the method comprises the following steps: collecting multi-source data in real time, extracting a core keyword, obtaining a text with the core keyword and interaction data corresponding to the text, the hot events are identified by analyzing the interaction data, and the temporary sensitive words are extracted based on the hot events. According to the method, multi-source data are collected in real time, temporary sensitive words related to hot events are dynamically extracted, the problem that a traditional static sensitive word bank lags in response to emerging sensitive words generated by hot events is solved, the relevance score is calculated by means of semantic similarity and the co-occurrence frequency, dominant related words can be covered, and the probability of occurrence of the hot events is lowered. And hidden anaphora vocabularies can be identified, so that missed judgment caused by variable vocabulary forms is reduced, and the timeliness and comprehensiveness of sensitive word identification are improved.
Owner:CO OP LAND CORP

Intelligent labeling multi-dimensional retrieval system for education case library

The invention relates to the technical field of education informatization, and discloses an intelligent annotation multi-dimensional retrieval system for an education case library, which comprises a case storage server used for storing education texts, images and audio and video case resources; the intelligent labeling engine is connected with the case storage server, and performs three-order structured labeling on case resources through an element layering analysis model: extracting labels of a basic entity layer, a political theory layer and a value guiding layer; the knowledge graph construction module is used for constructing a multi-modal case knowledge graph which takes an element-teaching scene-childbearing target as an association relationship on the basis of the hierarchical label output by the intelligent labeling engine; the intelligent retrieval terminal comprises a semantic analysis unit and a multi-dimensional feedback unit. According to the method, an intelligent annotation engine pre-trains a theoretical word bank and an emotion analysis model through a BERT-BiLSTM-CRF model, and automatically completes three-order structured annotation to form a structured data system supporting advanced retrieval.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

Cross-lingual speech recognition

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for cross-lingual speech recognition are disclosed. In one aspect, a method includes the actions of determining a context of a second computing device. The actions further include identifying, by a first computing device, an additional pronunciation for a term of multiple terms. The actions further include including the additional pronunciation for the term in the lexicon. The actions further include receiving audio data of an utterance. The actions further include generating a transcription of the utterance by using the lexicon that includes the multiple terms and the pronunciation for each of the multiple terms and the additional pronunciation for the term. The actions further include after generating the transcription of the utterance, removing the additional pronunciation for the term from the lexicon. The actions further include providing, for output, the transcription.
Owner:GOOGLE LLC

Table instruction conversion system based on electric energy meter communication protocol

The invention discloses a table instruction conversion system based on an electric energy meter communication protocol, and the system comprises a data collection module which obtains and processes an electric energy meter communication protocol file and a historical manual processing instruction data set; the basic information reading module is used for extracting an input field and a data format and mapping a power grid instruction format to generate a compliance instruction; the instruction missing comparison module is used for comparing the compliance instruction with the historical data set and judging field missing content; the lexicon association and expansion module is used for extracting initial keywords and generating virtual expansion words covering missing scenes; the evaluation storage module is used for verifying the virtual expansion words and storing the virtual expansion words into an expansion associated word bank in a classified manner; the table content information extraction module is used for analyzing the file through an index strategy, automatically extracting content to fill the intelligent table and recording a source; the method has the advantages that electric energy meter communication protocol deep analysis, dynamic lexicon expansion and automatic table instruction generation can be carried out, and electric energy meter data processing refinement and real-time performance are improved.
Owner:ZHEJIANG REALLIN ELECTRON CO LTD

Adaptive masking method and system, and device and medium

The present invention relates to the technical field of data security, and in particular to an adaptive masking method and system, and a device and a medium. The method comprises: firstly, acquiring a keyword of a current file to be masked of a sender user; then, on the basis of the keyword, performing addition, deletion and modification on a current sensitive word library to obtain a new sensitive word library; and finally, generating a regular expression on the basis of the new sensitive word library, determining the positions of sensitive words on the basis of the regular expression, so as to obtain the sensitive words, and performing masking to obtain a masked file. Therefore, masking of a plurality of types of data is realized, and data types before and after masking remain unchanged, thereby ensuring the readability of a masked file while guaranteeing the security of masked data, and further increasing the masking speed; on the basis of different roles of receiver and sender users, the masking intensity is adaptively selected, thereby preventing important information from being leaked to untrusted persons, and overcoming the defect of an existing masking algorithm failing to resist collusion attacks and brute-force enumeration attacks; and multiple threads are used to concurrently process sensitive word retrieval and sensitive-word masking operations, thereby greatly increasing the masking speed.
Owner:CHENGDU AIRCRAFT INDUSTRY GROUP

Strong association control method for long text streaming conversion

The invention provides a strong association control method for long text streaming conversion, belongs to the field of long text processing, and solves the problems of uncontrollable association, semantic deviation and key information loss in long text conversion. The method comprises the following steps: receiving a streaming input fragment, and calculating a multi-dimensional dynamic strong association relationship containing a semantic retention degree, a structure correspondence degree, a key information extraction degree and style consistency; generating key cue words according to the key cue words, and forcibly outputting focusing core elements through dynamic weight adjustment and large model generation probability correction; and when the association is insufficient, a feedback closed loop is formed through weight enhancement and keyword supplementation. Cue words are screened based on associated weak dimensions, and each dimension is quantitatively evaluated in combination with a pre-training model and a domain word library. According to the scheme, accurate input and output association control is achieved, the method is suitable for scenes such as abstract generation, structured extraction and style migration, and the semantic retention precision, the structure alignment degree and the multi-task adaptation capacity are effectively improved.
Owner:CLP CLOUD BRAIN (TIANJIN) TECH CO LTD

Power big data dictionary construction system based on improved SO-PMI algorithm

The invention relates to the technical field of power big data processing, and discloses a power big data dictionary construction system based on an improved SO-PMI algorithm. The system comprises a corpus preprocessing module, an algorithm adaptation module, a word frequency correlation analysis module, a dictionary hierarchy construction module and a domain lexicon fusion module. The corpus preprocessing module is used for acquiring a field text, extracting basic lexical elements, marking positions and frequencies of terminologies, screening high-frequency core lexical elements and constructing a basic lexical element library; an algorithm adaptation module assigns values to lexical element weights, adjusts a co-occurrence window and an association threshold value of the SO-PMI algorithm, and constructs an adaptation parameter set; the word frequency association analysis module extracts a co-occurrence sequence, compares association strength, screens associated word tuples, and obtains an association relationship set; the dictionary hierarchy construction module collects hierarchy labels and obtains a hierarchy affiliation label group after matching; and the domain lexicon fusion module divides the lexical element association groups to corresponding classification nodes to generate a domain dictionary.
Owner:WEIHAI POWER SUPPLY COMPANY OF STATE GRID SHANDONG ELECTRIC POWER COMPANY +1

Sensitive data detection method, related device and medium

The embodiment of the invention provides a sensitive data detection method, a related device and a medium. The method comprises the following steps: acquiring to-be-detected data; a double-encoder rule generation model is called to convert the user query instruction into first vector data, text information of a knowledge source library is converted into second vector data, target knowledge text information is determined according to the first vector data and the second vector data, and the knowledge source library at least comprises one of a law and regulation library, an industry word library and an enterprise database; calling a dual-encoder rule generation model to generate a first detection rule according to the templated instruction and the target knowledge text information; calling a large language model to carry out adaptive optimization on the first detection rule to obtain an optimized second detection rule; and calling the multi-mode joint detection model to perform sensitive data detection on the to-be-detected data according to a second detection rule to obtain a sensitive data detection report. The embodiment of the invention aims to improve the accuracy and timeliness of sensitive data detection.
Owner:PENG CHENG LAB

Automatic speech recognition result optimization system and method, medium and equipment

The invention provides an automatic speech recognition result optimization system and method, a medium and equipment, and the system comprises an ASR preprocessing module which carries out the initial ASR recognition of an original audio, obtains a preliminary recognition text, carries out the named entity recognition of the preliminary recognition text, enables a recognition result to be matched with a hot word with similar pronunciation in a hot word library, and carries out the recognition of the hot word; inputting the hot words into an ASR model for secondary identification; the ASR text post-processing module is used for carrying out semantic error correction on the text output by the ASR model by utilizing a large language model and verifying the reasonability of the text; and the ASR speaker recognition module is used for segmenting and numbering the audio by using a speaker separation technology, and mapping the number with a specific name. According to the invention, by adopting a mode of entity word extraction and hot word matching, the problem that a speech recognition system cannot screen related hot words in advance is solved, and the recognition effect of an ASR system is improved.
Owner:SHANGHAI SHENGHEKUN INFORMATION TECH CO LTD

Problem shunting method, device and equipment

The embodiment of the invention discloses a problem shunting method, device and equipment. In the embodiment of the invention, user question information is obtained; inputting the user question information into a preset high-recall-rate word library, and determining a comparison result; responding to the comparison result that the user question information comprises the same vocabulary in the high-recall-rate word bank, and determining the user question information as risk question information; inputting the risk problem information into a preset dichotomy model, and outputting a classification result; and in response to the fact that the classification result is high risk, determining that the risk problem information is high-risk problem information, and intercepting the high-risk problem information. According to the method, question shunting can be accurately and efficiently carried out, and the user experience is improved.
Owner:ALIBABA (CHINA) CO LTD

Multi-language AI (artificial intelligence)-based personalized ascending plan dynamic generation system

The invention discloses a multi-language AI-based personalized ascending plan dynamic generation system, which comprises a multi-language semantic interaction module, a dynamic gradient prediction engine, a real-time policy adaptation system and a 1 + N double-track service module, and is characterized in that the multi-language semantic interaction module is used for processing Uyghur language and Kazakh language voice / text input; cross-language intention recognition is realized through a bidirectional LSTM neural network, and a Uighur-Chinese bilingual contrast policy lexicon covering college entrance examination, examination and research and doctor stages is constructed to realize multi-language policy semantic alignment; the dynamic gradient prediction engine is used for generating a volunteer scheme including three levels of sprint-adaptation-stability for college entrance examination, generating a sprint-adaptation-stability target scheme for the college entrance examination, and generating a core target-alternative target-bottom guarantee target scheme for the blog by fusing characteristic parameters of different learning stages on the basis of an XGBoost algorithm, and generating a college scheme including sprint-adaptation-stability for the college entrance examination, the college entrance examination, the college entrance examination and the college entrance examination and the college entrance examination and the college entrance examination and the college entrance examination and the college entrance examination. The real-time policy adapts the system.
Owner:YINING KAIMENG TECHNOLOGY INFORMATION CONSULTING CO LTD

Virtual digital human estrus sharing ability enhancement method, device, equipment, medium and product

The invention discloses a method, a device, equipment, a medium and a product for enhancing the estrus sharing ability of a virtual digital human, and relates to the technical field of human-computer interaction and virtual digital humans. The method comprises the following steps: firstly, based on dialogue data of a user, extracting to obtain multi-modal feature information for reflecting a dialogue understanding result of the user, voice emotion of the user and facial emotion of the user, and then importing the multi-modal feature information into a user real intention recognition model obtained by pre-training based on a machine learning algorithm; and finally, according to the big language model prompt word bank, combining the recognition result and the dialogue content, generating a response prompt word, and according to the prompt word, controlling the virtual digital person to carry out big language model-based common situation response to the user, so that the response content of the digital person is accurate and appropriate, and the user experience is improved. And the mood, expression and the like of the digital human can be kept consistent with those of the user, so that more real common-situation interaction experience is realized.
Owner:BEIJING SITU CHANGJING DATA TECH SERVICE CO LTD