Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

158 results about "Lexical frequency" patented technology

Description Lexical usage frequency is known to influence the application rate of some variable processes. Specifically, variable lenition processes typically affect frequent lexical items more often than infrequent lexical items. For instance, variable t/d-deletion in English is more likely to apply to a frequent word (just)...

Deep learning large model-based medical record information extraction and analysis method and system

The invention provides a medical record information extraction and analysis method and system based on a deep learning large model, and the method comprises the steps: carrying out the semantic understanding of a medical record text through a deep learning model, and recognizing key clinical features; extracting the key clinical features, and carrying out vectorization processing on the extracted key clinical features; performing sentence segmentation and word segmentation on the case text based on a rule engine and the key clinical features; adjusting the segmented words based on a symptom classification rule; and carrying out word frequency statistics on the adjusted segmented words. According to the method, accurate recognition and structured conversion of key clinical features can be ensured through a semantic understanding technology, knowledge graph construction and machine learning model optimization are supported, and a new view angle and a new tool are provided for exploration and research of diseases.
Owner:SHANDONG MENTAL HEALTH CENT

Electric bicycle form design method and system, electronic equipment and storage medium

The invention belongs to the field of vehicle industry design, and discloses an electric bicycle form design method and system, electronic equipment and a storage medium, and the method comprises the steps: mining user emotion vocabularies through online comments, constructing an emotion lexicon in combination with an improved word frequency-inverse document frequency algorithm and a D-S evidence theory, and screening out key perceptual vocabularies; constructing a convolutional long-short-term memory neural network model optimized by a snake swarm algorithm to realize a mapping model between customer sensibility and product morphological characteristics; carrying out subjective and objective comprehensive evaluation on the design scheme in combination with an eye movement experiment and subjective evaluation so as to screen out an optimal design scheme; an optimal scheme is selected and input into a generative AI platform for multi-angle visual rendering, ergonomic modeling and aerodynamics simulation are assisted, and the structural feasibility and performance are verified. The method provides systematic technical support for emotional value improvement and design optimization of industrial products.
Owner:NANCHANG UNIV

Government affair knowledge retrieval method and device based on large model, equipment and medium

The invention discloses a large model-based government affair knowledge retrieval method, apparatus and device, and a medium, and relates to the technical field of natural language processing, and the method comprises the steps of performing information extraction on a to-be-processed government affair work order through a preset government affair large model to obtain target structured information; based on a word processing fusion algorithm constructed by a word frequency-inverse document frequency algorithm and a graph sorting algorithm, performing word processing on the target structured information to obtain work order keywords; converting the work order keyword into a target semantic vector; performing retrieval operation on a preset government affair knowledge base by utilizing a preset retrieval condition and the target semantic vector to obtain a target government affair knowledge document; the preset retrieval condition is a retrieval condition constructed based on time, space and the work order emergency degree. Therefore, the information accuracy and the processing efficiency can be improved through word processing; and in combination with retrieval conditions constructed based on time, space and work order emergency degrees, accurate retrieval can be provided, and the processing effect on various complex government affair work orders is improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Lightweight knowledge graph rapid construction method and system based on NLP technology

The invention discloses a lightweight knowledge graph rapid construction method and system based on an NLP technology. The method comprises the steps of preprocessing an unstructured text and segmenting the unstructured text into semantic segments; unsupervised clustering is combined with the contour coefficient to determine the optimal clustering number, and a semantic association fragment cluster is obtained; performing word segmentation, part-of-speech tagging, NER and entity linking on the fragment cluster, extracting an entity and initial relationship, and fusing semantic similarity, TF-IDF word frequency collaboration degree and co-occurrence frequency to calculate a relationship edge weight; constructing a lightweight knowledge graph; in the question and answer stage, questions are disassembled through a process engine, related sub-graphs are retrieved, and answers with reasoning links are generated. The system correspondingly comprises a text preprocessing module, a clustering module, an entity relation processing module, a graph construction module, a question and answer reasoning module and a storage module. According to the method, the construction cost is reduced, the interpretability and the module coupling degree are improved, multiple scenes such as government and enterprise public opinions and medical assistance are adapted, and the problems of weak generalization, poor real-time performance and'black box 'in the traditional technology are solved.
Owner:XIAMEN MEIYA PICO INFORMATION CO LTD +1

Method and system for generating intelligent insight report based on AI large model

The invention relates to the field of intelligent report generation, in particular to an intelligent insight report generation method and system based on an AI large model, and the method comprises the steps: inputting an insight demand, and generating an insight data package comprising insight contents, associated data and industry labels; extracting a basic keyword set of the insight data packet to form a mixed feature code; after mixed feature coding preprocessing, weight distribution is carried out; and after weight distribution of the AI large model, injecting a high-weight feature vector into a semantic understanding core layer, injecting a low-weight feature vector into a logical reasoning layer, and outputting analysis data to form an intelligent insight report. According to the method, word embedding parameters are optimized according to field semantic characteristics, a parameter verification mechanism is introduced, field text characteristics are adapted by adjusting vector dimensions, context windows and low-frequency vocabulary filtering threshold values, window parameter validity is verified through cosine similarity, word frequency threshold value reasonability is verified through standard deviation, and field text characteristic matching is achieved. And it is ensured that the feature vectors can accurately capture domain-specific semantics.
Owner:SUZHOU YINGTIANDI INFORMATION TECH CO LTD

Retrieval optimization method based on dynamic mixed retrieval and adjacent paragraph introduction

The invention discloses a retrieval optimization method based on dynamic mixed retrieval and adjacent paragraph introduction. The method comprises the following steps: taking related documents as a corpus of a retriever; performing retrieval by using a BM25 algorithm and a HyDE algorithm to generate corresponding preliminary retrieval document paragraphs; calculating a query specificity value for user query by using TF-IDF; determining weights of a BM25 algorithm and a HyDE algorithm, and respectively performing weighted fusion on a document paragraph correlation score based on word frequency statistical calculation in the BM25 algorithm and a document paragraph correlation score based on semantic embedding similarity calculation in the HyDE algorithm to generate candidate document paragraphs and comprehensive scores corresponding to the candidate document paragraphs; the document paragraphs ranked in the top are selected from the preliminary retrieval document paragraphs, adjacent document paragraphs are obtained, weighted scores of the adjacent document paragraphs are calculated, the adjacent document paragraphs are added into a candidate document paragraph list, and the document paragraphs with the top scores are selected from the list to serve as retrieval results.
Owner:SOUTH CHINA UNIV OF TECH

Resume label generation method and device and medium

The invention discloses a resume label generation method and device and a medium, and relates to the field of natural language processing, and the method comprises the steps: carrying out the preprocessing of an original resume text, and carrying out the word segmentation of original resume data into a plurality of text words; counting word frequencies corresponding to the text words, and endowing the text words with corresponding importance weights according to the word frequencies; vectorizing the text words to obtain corresponding word vectors, and aggregating the word vectors according to the importance weight to obtain resume text vectors corresponding to the original resume text; based on a preset hierarchical label library, performing multi-label classification on the text semantic vector through a nonlinear classification algorithm, and outputting a prediction label of the original resume text; and screening the label prediction values based on a preset threshold to obtain a final resume label set. By fusing word frequency weighted semantic representation and a nonlinear classification algorithm, the accuracy and adaptability of resume label generation are remarkably improved.
Owner:SHENZHEN INSPUR HAIYUE HUMAN RESOURCES TECHNOLOGY CO LTD

Semantic understanding-based completion environmental protection acceptance survey report auxiliary auditing system

The invention relates to the technical field of natural language processing, in particular to a semantic understanding-based completion environmental protection acceptance survey report auxiliary auditing system, which comprises a structure recognition module, a semantic chain construction module, a semantic reconstruction module, an image-text coordinate fusion module and a spatial offset judgment module. In the method, through extracting subjects, terms and facility words in facility paragraphs and establishing a combination relationship, ordered labeling of semantic elements is realized, statement structure definition is enhanced, a semantic chain group is constructed and facility classification is associated, semantic consistency comparison is supported, semantic sequences and word frequency distribution in design and reply contents are compared, and description consistency is verified; facility coordinates in an image are extracted, character paragraphs are matched, image-text content anchoring, orientation, terrain and distance word extraction and conversion into space parameters to be compared with an image path are achieved, space consistency verification is completed, and the intelligent level of auditing on semantic comprehension, image-text fusion and space logic is overall improved.
Owner:SHANGHAI RUIDUN INFORMATION TECHNOLOGY CO LTD +1

Extension query method and system based on large language model entity extraction

The invention discloses an extended query method and system based on big language model entity extraction, key entities in original queries are extracted, the original queries are repeatedly spliced, the spliced original queries and key entities are fused, extended queries are obtained, entity interpretation is used as an extension basis, and compared with a traditional word frequency statistics or synonym replacement method, the extended query method and system have the advantages that the extension efficiency is improved; the large language model can generate extended fragments with semantic consistency and information integrity in combination with the context, it is guaranteed that the finally generated extended query is consistent with the original intention of a user, semantic drift and invalid extension are reduced, the correlation and precision of extended content are improved, meanwhile, the spliced original query and key entities are fused, and the user experience is improved. According to the method, the original query is locally enhanced by adopting an entity-level semantic interpretation mode, higher interpretability and controllability are prepared, and a complicated model structure and a training process are not needed, so that the deployability and universality of the method are remarkably improved, the expansion effect is more stable, and diversified retrieval requirements can be met.
Owner:PEOPLE CN CO LTD +1

Speed Up Methods and Systems for Large Language Model Training

A method initializes and accelerates training of neural network based large language model, including by: (i) accessing a corpora for training a neural-network based large language model having word embeddings and word projections in respective word embedding and word projection layers and at least one hidden layer; (ii) counting raw token frequencies associated with content within the corpora; (iii) smoothing the raw token frequencies into a series of vector norms based on log or scaled log functions parameterized by maximum norm information; and (iv) injecting vector norm information into word embeddings and / or word projections based on norm-angle reparameterization to prepare the large language model for training.
Owner:APPL TECH APPTEK

Data mining system and method for mail data

The invention discloses a data mining system and method for mail data, relates to the technical field of data retrieval processing, solves the technical problem of inaccurate retrieval matching caused by single semantic analysis retrieval, and improves data relevance by unifying multi-source data formats and combining an enterprise address book to supplement metadata. Structured data standardization and unstructured text vectorization are carried out, a computable semantic space is constructed, keywords are screened through a matching proportion, a word frequency weight is calculated in combination with a BM25 algorithm, the keyword matching accuracy is improved, a semantic vector is generated by utilizing a pre-training language model, and semantic correlation is calculated through cosine similarity. Keyword matching scores and semantic similarity scores are dynamically fused, accuracy and comprehensiveness are balanced, accuracy and recall rate are calculated in real time, retrieval deviation types are automatically recognized, keyword weights and similarity threshold values are adjusted in a targeted mode or synonym lists are expanded, and self-adaptive optimization of the system is achieved.
Owner:SHANGZHONG ONLINE TECH CO LTD

Description generation method based on continuous zero sample

The invention provides a description generation method based on continuous zero samples. The description generation method comprises the following steps: sequencing plain text training corpora according to the difficulty of continuous learning; setting that each training step t only can access the t-th corpus, and for the t-th learning step, generating a synthetic image for an input text of the t-th corpus by using a diffusion model; extracting image embedding of the synthesized image by using a CLIP image encoder; retrieving a group of text description embedding with similar semantics in a text corpus by using image embedding; image embedding and text description embedding are input into a fusion module network, and the output of a fusion module is defined as a soft prompt of a large language model; for an input text, key entities are extracted from the text through word frequency statistics so as to construct a soft prompt and a hard prompt of the large language model; hard prompts and soft prompts are input into a large language model network, and key parts in input are concerned through an attention mechanism.
Owner:JIANGSU VOCATIONAL & TECHNICAL UNIVERSITY OF ARCHITECTURE

Question and answer evaluation method and system for autism intervention

The invention discloses a question and answer evaluation method and system for autism intervention, and relates to the field of medical electronic systems.The method comprises the steps that historical data are extracted, noise reduction processing is carried out to obtain noise reduction data after noise reduction processing, and a database is built according to the noise reduction data, a word frequency-inverse document frequency algorithm and a BERT model; extracting question and answer information of a plurality of different autism stages in the database, and constructing a judgment model according to the question and answer information and the database; obtaining current question and answer data, and obtaining a judgment result of the child according to the current question and answer data and the judgment model; the teacher resume in the database is extracted, multiple recommendation teachers are determined according to the teacher resume and a judgment result, the two are combined to form a multi-level semantic feature matrix, local word frequency information is reserved, global context understanding is fused, and therefore the problem that feature extraction of a single model is one-sided is solved; accurate recognition capability of autism core symptoms (such as social obstacles and communication defects) is remarkably improved, and a more reliable behavior characteristic basis is provided for an intervention scheme.
Owner:高班超

Method and device for improving RAG recall rate in voice question and answer scene

The invention provides an RAG recall rate improving method and device in a voice question and answer scene, and relates to the technical field of data processing.The method comprises the steps that semantic cleaning processing is conducted on an original corpus containing a voice recognition result, semantic compression is conducted on the cleaned original corpus, and the original corpus is obtained; utilizing the plurality of candidate embedded vector generation models to respectively execute vector generation operation, and outputting word vectors; for each word vector, calculating a semantic fidelity score; evaluating the plurality of semantic fidelity scores, and selecting a target embedding vector generation model with the optimal semantic fidelity score from the plurality of candidate embedding vector generation models; calculating a word frequency value and an inverse document frequency value of each word according to data input, judging whether the words are professional hot words or not, and screening out the professional hot words to construct a hot word list; and jointly inputting an embedded vector output by the target embedded vector generation model and the hot word list into a question and answer module, and outputting a target answer text. According to the invention, the RAG recall rate in the voice question and answer scene can be improved.
Owner:ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD

Power big data dictionary construction system based on improved SO-PMI algorithm

The invention relates to the technical field of power big data processing, and discloses a power big data dictionary construction system based on an improved SO-PMI algorithm. The system comprises a corpus preprocessing module, an algorithm adaptation module, a word frequency correlation analysis module, a dictionary hierarchy construction module and a domain lexicon fusion module. The corpus preprocessing module is used for acquiring a field text, extracting basic lexical elements, marking positions and frequencies of terminologies, screening high-frequency core lexical elements and constructing a basic lexical element library; an algorithm adaptation module assigns values to lexical element weights, adjusts a co-occurrence window and an association threshold value of the SO-PMI algorithm, and constructs an adaptation parameter set; the word frequency association analysis module extracts a co-occurrence sequence, compares association strength, screens associated word tuples, and obtains an association relationship set; the dictionary hierarchy construction module collects hierarchy labels and obtains a hierarchy affiliation label group after matching; and the domain lexicon fusion module divides the lexical element association groups to corresponding classification nodes to generate a domain dictionary.
Owner:WEIHAI POWER SUPPLY COMPANY OF STATE GRID SHANDONG ELECTRIC POWER COMPANY +1

Analysis method and device for credit granting approval process

The invention provides a credit granting approval process analysis method and device. Comprising the steps of performing word segmentation on text content in a credit granting approval document set, and combining the obtained segmented words to obtain combined phrases; calculating weighted word frequency-inverse document word frequency of the combined phrases; target phrases with weighted word frequency-inverse document word frequency larger than a word frequency threshold value are screened out from the combined phrases, and an amplified word list is generated according to the target phrases and the original word list; training a target word segmentation device based on the amplified word list; finely adjusting the basic vector model according to the training corpus to obtain a term vector model; analyzing the auditing opinion information based on a large language model to obtain an initial risk analysis rule containing a preset dimension; based on a similarity algorithm, a target word segmentation device and a term vector model, grouping and integrating the initial risk analysis rules to obtain risk analysis rules; and processing the information of each examination and approval stage by using a big language model and adopting a risk analysis rule to obtain examination and approval suggestion information.
Owner:MINSHENG BANKING CORP

Automatic identification, classification and development trend analysis method of net red villages based on multi-source data fusion and natural language processing

The method for automatic identification, classification and development trend analysis of net red villages based on multi-source data fusion and natural language processing comprises the following steps: UGC data is crawled from Xiaohongshu and Douyin through a distributed master-slave architecture, de-duplicated based on SimHash, and normalized in time and coding format; a text semantic fingerprint is generated, and multi-level semantic cache fingerprint matching is performed; for unassigned text, its complexity is calculated, and a large language model API is adaptively called to automatically complete and extract five-level administrative divisions; weights are determined based on the analytic hierarchy process, interaction indicators such as likes, comments, collections and forwards are integrated, and a comprehensive network heat index of the village is obtained; an external text mining tool is connected, and batch word frequency analysis, semantic network analysis and sentiment tendency evaluation are performed; a document-term matrix is constructed, TF-IDF weighting is performed, and unsupervised clustering algorithm is used for clustering analysis of village characteristics; cross-dimension analysis is performed on the clustering results, and a development portrait, advantage mining and operation suggestion warning are automatically generated in combination with the SWOT model.
Owner:ZHEJIANG UNIV OF TECH

Word-tag-based language system for sentence acceptability judgment

Methods and systems for sentence acceptability judgment. A word frequency distribution for a predetermined textual data set is obtained. A replacement rate is determined based on an obtained word frequency distribution for a predetermined textual data set and every occurrence of a word having a frequency lower than the replacement rate is replaced in text of a training data set with a corresponding tag to generate revised text of the training data set (the training data set comprising at least a portion of the predetermined textual data set). A plurality of language models are trained with the revised text and a best performing trained language model of the plurality of trained language models is selected. An acceptability of each sentence of a plurality of candidate sentences is rated using the selected trained language model and a best sentence of the plurality of candidate sentences is selected based on the ratings.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

A domain knowledge triple extraction method, system, medium and device

The application discloses a domain knowledge triple extraction method, system, medium and equipment, and relates to the technical field of knowledge engineering.The domain knowledge triple extraction method and system provided by the application can provide high-quality and efficient data support for triple extraction by analyzing and converting the obtained domain professional text into a structured vector knowledge base which can be efficiently searched, then the domain professional text is subjected to quantitative calculation and fusion screening, the initial entities of domain core high-frequency words which have global high frequency and cross-text universality are mined, the problem of "local optimum and global deviation" caused by traditional single word frequency is avoided, non-core term interference is effectively avoided, and the stability, relevance and efficiency of iterative extraction are improved, and further, the high-frequency word initial entities are subjected to closed-loop iteration through iterative RAG algorithm combined with the structured vector knowledge base, accurate extraction of the large model and iterative correlation entities, the explicit and implicit semantic correlation triples in the text can be mined layer by layer, the strong relevance and structural integrity of the extracted triples are ensured, and the professionalism and accuracy of the model in extracting triples are improved.
Owner:XIAN UNIV OF TECH

Data classification management method and device and computing equipment

The data classification management method comprises the steps that a to-be-classified text is input into a first classification model, and the first classification model determines key information according to the to-be-classified text; obtaining semantic features according to the key information; the semantic features comprise word vectors of a plurality of key words and feature vectors of word frequency features; attribute information of the to-be-classified text is determined according to the semantic features, the attribute information comprises a first attribute, and the first attribute is a preset category. Therefore, key information highly related to semantics in the text is reserved, redundant invalid information is deleted, the semantic features of the original text can be reserved while the length of the text is shortened, and the classification speed of the long text is greatly increased on the premise that the classification precision is not reduced.
Owner:HUAWEI TECH CO LTD

Method and device for fine-tuning bert model, equipment and storage medium

Embodiments of the present application provide a BERT model fine-tuning method and device, equipment and a storage medium, and relate to the technical field of computers. The method comprises: generating an original token list according to a text corpus, and obtaining word frequency statistical information of the original token list; determining a mask generation strategy based on the word frequency statistical information, and performing mask processing on the original token list according to the mask generation strategy to obtain a mask processing result; and performing iterative learning on a pre-trained BERT model based on the mask processing result until a preset condition is met, to obtain a fine-tuned BERT model. The present application statistically obtains word frequency information corresponding to a text corpus, and adaptively adjusts a mask generation probability according to the word frequency, so that the mask distribution generated by tokens of different frequencies is more uniform, and the semantic information of the text corpus is better preserved, thereby improving the learning ability of the BERT model for masked corpus.
Owner:BEIJING TOPSEC NETWORK SECURITY TECH +2

A multi-directional text alignment method

The application provides a multi-directional text comparison method, comprising: exporting a packaging design drawing into a PDF format and parsing text content and corresponding position information; splitting the text parsed from the PDF according to intervals, and judging whether the split text is consistent; calculating 2-gram word frequency by using a large amount of Chinese corpus, calculating the probability of text in a normal order and a reverse order, and taking the probability as a basis to judge whether the text is in a reverse order; processing all the text according to position coordinates, judging the direction of the text block according to the normal order and the reverse order of the text, then sorting and merging the text in the text block; comparing the PDF text content with the standard text content for examination, matching similar lines, and marking the differences of the similar lines. The processing result can greatly reduce the structured difference between the parsed text and the actual text, improve the detection accuracy of the parsed text and the actual text, reduce the workload of manual intervention, thereby reducing the packaging design cost and precision.
Owner:PU HUA KE JI YOU XIAN GONG SI

Watermark generation method and system based on large language model

The application belongs to the field of artificial intelligence, and particularly relates to a watermark generation method and system based on a large language model, which comprises the following steps: obtaining a training data set, preprocessing data in the training data set; presetting a watermark encoding rule; training a large prophecy model according to the watermark encoding rule and the preprocessed data; obtaining text data to be processed; cleaning and standardizing the text data; calculating the word frequency and inverse document frequency of the standardized text data to generate an initial vector; processing the initial vector by using a Word2Vec model to obtain a high-dimensional word vector; inputting the high-dimensional word vector into a trained neural network model to extract deep features; classifying the text according to the deep features; inputting the classified problem into the trained large prophecy model to generate a watermark corresponding to the text type; and the application designs a watermark encoding rule and generates a corresponding watermark for the classified text, thereby improving the security of the text.
Owner:CAIQIMAO (GUANGZHOU) INFORMATION TECHNOLOGY CO LTD

Topic model generation method and device, storage medium and program product

The embodiment of the invention provides a topic model generation method and device, a storage medium and a program product, and relates to the technical field of artificial intelligence. The method comprises the following steps: preprocessing an obtained user comment text and comment metadata, and correspondingly obtaining a word frequency vector and a text label; inputting the word frequency vector into a preset topic model, converting the word frequency vector into potential space representation, and performing normalization processing to obtain topic distribution corresponding to a user comment text output by the topic model; inputting the user comment text into a preset attention mechanism model to obtain a text embedding vector corresponding to the user comment text; and based on the topic distribution, the text embedding vector and the text label, carrying out joint training, and determining a trained target model. The attention mechanism and the comment metadata are combined in the modeling process, the trained target model is determined, and the problems that in the prior art, a topic model is insufficient in real scene data processing capacity and low in efficiency are solved.
Owner:CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1

A method for generating a pinyin bucket word library of a four-level index

PendingCN122452553ADigital dataData integrity
The present application relates to the technical field of electronic digital data processing, and discloses a four-level index pinyin bucket word library generation method, acquires a Chinese word library text, extracts the first Chinese character and the first pinyin of each line of words as a classification key, and generates a triple; the triple is classified into a corresponding pinyin bucket according to the pinyin, the words in the bucket are sorted in descending order of word frequency, an independent Chinese character block is generated for each unique Chinese character, and the words are solidified according to the word length by using a first character multiplexing mechanism; a four-level static index of an initial letter statistical area and a pinyin index area and a Chinese character word index area and a word storage area is sequentially constructed, each index area uses fixed-length entries and absolute offset addressing; a metadata area containing a file length double-semantic field is constructed, a high-frequency-first deterministic truncation is performed on the word library according to a preset threshold, each data area is spliced and a data integrity check value is appended, and an embedded binary word library is generated. The problems of high storage redundancy and uncontrollable memory are solved, and the purposes of deterministic analysis, resource adaptation and high security are achieved.
Owner:SICHUAN HAIGE HENGTONG PRIVATE NETWORK TECH CO LTD

An intelligent customer service system, method, medium and device based on emotion recognition

The present disclosure relates to an intelligent customer service system and method based on emotion recognition, equipment and medium, the system comprises: a voice data extraction module for extracting voice data of a user based on a conversation between the system and the user and storing; a word segmentation processing module for performing word segmentation processing on the voice data to obtain word vector data; an entity recognition module for extracting entity keywords in the word vector data and performing first sorting on the word frequency of the entity keywords; an emotion recognition module for extracting emotion keywords in the word vector data and performing second sorting according to the word frequency of the emotion keywords; a decision module for selecting a corresponding response strategy based on a decision algorithm according to the first sorting and the second sorting; a voice module for generating a reply voice statement based on the response strategy. The system and method of the present disclosure recognize user emotions, assist in decision-making, and ultimately improve user service quality.
Owner:PING AN PAY ELECTRONIC PAYMENT CO LTD

A language processing-based word vector representation method, device and terminal equipment

The application is suitable for the field of artificial intelligence technology, and provides a word vector representation method and device based on language processing and a terminal device, the method comprising: constructing an input word table; setting word frequency information of each word in the input word table; classifying the words into N clusters according to the word frequency information, words with the same word frequency information being classified into the same set to obtain a clustering set with N subsets; defining a first word vector dimension of the words in the first cluster through an experience factor, and obtaining a second word vector dimension of the words in each subset according to the subset where the words are located and the first word vector dimension; defining an nth linear mapping matrix for the nth cluster; and processing a target word vector through parameters shared by the input word table as an input of a natural language processing model. Through the application, the problem that the calculation amount is large, the efficiency is low and the word distribution features cannot be considered when converting words into word vectors in natural language processing can be solved.
Owner:SICHUAN LAN-BRIDGE INFORMATION TECHNOLOGY CO LTD

A semantic recognition model based on NLP and a training method thereof

The application discloses a semantic recognition model based on NPL and a training method thereof, relates to the technical field of semantic recognition model training, and aims to solve the problem of low recognition accuracy of a semantic recognition model. The current whole sentence is searched, multiple segmented sentences are screened out, words in the segmented sentences are recognized, the part-of-speech information of the words is collected, the part-of-speech weight under the corresponding part-of-speech information is called, the number of repeated occurrences of each word in the current whole sentence is counted and the repeated word frequency is calculated, the processing level of each word is calculated by comprehensively considering the repeated word frequency and the part-of-speech weight, the target words are screened out and recognized, the semantics of the target words are acquired and the number of semantics is counted under the current interactive window, the interactive duration of the target words is analyzed by calling historical data, the number of word reference sources is set as a reference, the number of semantics of the target words and the interactive duration are comprehensively considered to adjust the number of word reference sources of each target word, and the semantic analysis accuracy of each target word in the interactive sentence is improved.
Owner:SHANGHAI WUKE INTERNET TECHNOLOGY CO LTD

Document retrieval method and device based on AI large model

The invention provides a document retrieval method and device based on an AI large model, and belongs to the technical field of artificial intelligence. The document retrieval method comprises the steps that semantic extension is conducted on keywords in a query statement, and a plurality of extension words are obtained; obtaining a word frequency and an inverse document frequency of each keyword relative to each candidate document, and a word frequency and an inverse document frequency of each expansion word relative to each candidate document; determining a correlation score of each candidate document according to each word frequency and each inverse document frequency; and outputting a plurality of candidate documents of which the correlation scores are ranked in the top. According to the method, semantic enhanced sparse retrieval is realized, the method has the advantage of high efficiency of traditional sparse retrieval, and the retrieval accuracy is also improved.
Owner:BEIJING XIN LI FANG TECH INC