Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

217 results about "Lexical frequency" patented technology

Description Lexical usage frequency is known to influence the application rate of some variable processes. Specifically, variable lenition processes typically affect frequent lexical items more often than infrequent lexical items. For instance, variable t/d-deletion in English is more likely to apply to a frequent word (just)...

TF-IDF and cross entropy-based cue word compression method and system

The invention discloses a cue word compression method and system based on TF-IDF and cross entropy, belongs to the technical field of large model cue word compression, and aims to solve the problems that redundant information is introduced into long cue words, the model efficiency is reduced and the cost is increased. To-be-compressed content is divided into sentences at the sentence level and then converted into embedded vectors, and the Euclidean distance is calculated in combination with problem vectors so as to screen related sentences; calculating a TF-IDF value at the word level through a word frequency and an inverse document frequency to extract keywords and recombine sentences; and selecting a reference model and a basic model at the Token level, identifying the key Token based on a cross entropy loss difference value, and splicing the key Token in sequence to generate a compressed cue word. According to the method, a complex calculation structure is avoided, the inference efficiency is improved while the semantic integrity is maintained, and the resource consumption is reduced.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Deep learning large model-based medical record information extraction and analysis method and system

The invention provides a medical record information extraction and analysis method and system based on a deep learning large model, and the method comprises the steps: carrying out the semantic understanding of a medical record text through a deep learning model, and recognizing key clinical features; extracting the key clinical features, and carrying out vectorization processing on the extracted key clinical features; performing sentence segmentation and word segmentation on the case text based on a rule engine and the key clinical features; adjusting the segmented words based on a symptom classification rule; and carrying out word frequency statistics on the adjusted segmented words. According to the method, accurate recognition and structured conversion of key clinical features can be ensured through a semantic understanding technology, knowledge graph construction and machine learning model optimization are supported, and a new view angle and a new tool are provided for exploration and research of diseases.
Owner:SHANDONG MENTAL HEALTH CENT

Electric bicycle form design method and system, electronic equipment and storage medium

The invention belongs to the field of vehicle industry design, and discloses an electric bicycle form design method and system, electronic equipment and a storage medium, and the method comprises the steps: mining user emotion vocabularies through online comments, constructing an emotion lexicon in combination with an improved word frequency-inverse document frequency algorithm and a D-S evidence theory, and screening out key perceptual vocabularies; constructing a convolutional long-short-term memory neural network model optimized by a snake swarm algorithm to realize a mapping model between customer sensibility and product morphological characteristics; carrying out subjective and objective comprehensive evaluation on the design scheme in combination with an eye movement experiment and subjective evaluation so as to screen out an optimal design scheme; an optimal scheme is selected and input into a generative AI platform for multi-angle visual rendering, ergonomic modeling and aerodynamics simulation are assisted, and the structural feasibility and performance are verified. The method provides systematic technical support for emotional value improvement and design optimization of industrial products.
Owner:NANCHANG UNIV

Government affair knowledge retrieval method and device based on large model, equipment and medium

The invention discloses a large model-based government affair knowledge retrieval method, apparatus and device, and a medium, and relates to the technical field of natural language processing, and the method comprises the steps of performing information extraction on a to-be-processed government affair work order through a preset government affair large model to obtain target structured information; based on a word processing fusion algorithm constructed by a word frequency-inverse document frequency algorithm and a graph sorting algorithm, performing word processing on the target structured information to obtain work order keywords; converting the work order keyword into a target semantic vector; performing retrieval operation on a preset government affair knowledge base by utilizing a preset retrieval condition and the target semantic vector to obtain a target government affair knowledge document; the preset retrieval condition is a retrieval condition constructed based on time, space and the work order emergency degree. Therefore, the information accuracy and the processing efficiency can be improved through word processing; and in combination with retrieval conditions constructed based on time, space and work order emergency degrees, accurate retrieval can be provided, and the processing effect on various complex government affair work orders is improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Large language model training method based on knowledge graph

The invention discloses a big language model training method based on a knowledge graph, which comprises the following steps: collecting professional text data in a specific field, extracting entities and relationships from the professional text data, constructing a field knowledge graph, and constructing the knowledge graph based on data of different data sources on the basis of an external network; and introducing the domain knowledge graph into an intranet, storing the domain knowledge graph into a knowledge base in the intranet, extracting terminologies from the domain knowledge graph, analyzing word frequencies of the terminologies in the professional text data, and selecting the terminologies with the word frequencies higher than a set threshold value as newly added high-frequency terminologies. According to the method, the professional knowledge graph of the target field is constructed, the big language model is finely adjusted based on the professional knowledge graph, and the professional knowledge big language model is generated.
Owner:CHONGQING NORMAL UNIVERSITY

Lightweight knowledge graph rapid construction method and system based on NLP technology

The invention discloses a lightweight knowledge graph rapid construction method and system based on an NLP technology. The method comprises the steps of preprocessing an unstructured text and segmenting the unstructured text into semantic segments; unsupervised clustering is combined with the contour coefficient to determine the optimal clustering number, and a semantic association fragment cluster is obtained; performing word segmentation, part-of-speech tagging, NER and entity linking on the fragment cluster, extracting an entity and initial relationship, and fusing semantic similarity, TF-IDF word frequency collaboration degree and co-occurrence frequency to calculate a relationship edge weight; constructing a lightweight knowledge graph; in the question and answer stage, questions are disassembled through a process engine, related sub-graphs are retrieved, and answers with reasoning links are generated. The system correspondingly comprises a text preprocessing module, a clustering module, an entity relation processing module, a graph construction module, a question and answer reasoning module and a storage module. According to the method, the construction cost is reduced, the interpretability and the module coupling degree are improved, multiple scenes such as government and enterprise public opinions and medical assistance are adapted, and the problems of weak generalization, poor real-time performance and'black box 'in the traditional technology are solved.
Owner:XIAMEN MEIYA PICO INFORMATION CO LTD +1

Distributed aggregation retrieval method for law and regulation text database

The invention relates to the technical field of aggregation retrieval, in particular to a distributed aggregation retrieval method for a law and regulation text database, which comprises the following steps: collecting and integrating law and regulation text data, preprocessing the law and regulation text data, removing useless punctuations and blank characters, and generating a preprocessed text set. According to the method, data integration and preprocessing are carried out on the regulation text, useless punctuations and redundant blank characters are effectively removed, the purity of text data is guaranteed, and the accuracy of subsequent word segmentation and word frequency statistics is improved; the implementation of the word segmentation and word frequency statistical process is helpful for identifying key feature vocabularies in the regulatory text, so that the accuracy of the subsequent inverted index construction stage is enhanced; by preliminarily constructing the inverted index and implementing similar item merging and low-frequency vocabulary removing operation, the index scale and redundant interference are reduced, and the index query efficiency and accuracy are improved.
Owner:NANJING XIAOZHUANG UNIV

Retrieval enhancement generation optimization method and device based on adaptive classification and autonomous reasoning

The invention discloses a retrieval enhancement generation optimization method and device based on adaptive classification and autonomous reasoning. The method comprises the following steps: constructing a domain knowledge vector database; inputting a proprietary domain problem, and finding a plurality of matched knowledge documents with the highest score in the vector database from two dimensions of semantics and word frequency; inputting the matched knowledge document into an adaptive classification module to obtain a semantic attribute category of the document; the document and the corresponding semantic attribute category are input into an autonomous reasoning module, uncertain items and error items are eliminated, and enhanced knowledge is obtained; the enhanced knowledge is sent to the intelligent question and answer module to generate a response answer; the response answers are sent to the self-adaptive classification module again, and semantic attribute categories of the response answers are checked; if the answer is not judged to be correct, the autonomous reasoning module and the subsequent steps are executed again on the response answer; and if the answer is judged to be correct, displaying the answer to the user. The problem that in the prior art, an external knowledge source has negative effects on a generation result can be solved.
Owner:ZHEJIANG UNIV +2

Multimodal Large Model Evaluation Method and Device for Urban Governance Based on Multiple Layers of Text

The present invention discloses a method and device for evaluating a multi-modal large model for urban governance based on multi-level texts, comprising the following steps: obtaining the true label text of a sample picture and the output text of the multi-modal large model for urban governance; using an urban management thesaurus to segment the text into independent words and calculating the similarity score of lexical units; using the urban management thesaurus to create a bag-of-words model to convert the text data into a word frequency matrix and calculating the low-level word frequency similarity score; using a trained language model to extract semantic features with context in the text and converting them into vector representations, and calculating the semantic similarity score; using a general large language model with a formatted request with business guiding words and evaluation requirements to obtain the high-level semantic similarity score corresponding to the text; and combining the similarity scores at each level to calculate the comprehensive evaluation score of the multi-modal large model. The present invention fully considers the similarity measurement in the multi-level space of texts, improving the accuracy of evaluating the text output of the urban governance large model.
Owner:CHENGDU ZHIHUI HENENG CITY TECHNOLOGY CO LTD

Method and system for generating intelligent insight report based on AI large model

The invention relates to the field of intelligent report generation, in particular to an intelligent insight report generation method and system based on an AI large model, and the method comprises the steps: inputting an insight demand, and generating an insight data package comprising insight contents, associated data and industry labels; extracting a basic keyword set of the insight data packet to form a mixed feature code; after mixed feature coding preprocessing, weight distribution is carried out; and after weight distribution of the AI large model, injecting a high-weight feature vector into a semantic understanding core layer, injecting a low-weight feature vector into a logical reasoning layer, and outputting analysis data to form an intelligent insight report. According to the method, word embedding parameters are optimized according to field semantic characteristics, a parameter verification mechanism is introduced, field text characteristics are adapted by adjusting vector dimensions, context windows and low-frequency vocabulary filtering threshold values, window parameter validity is verified through cosine similarity, word frequency threshold value reasonability is verified through standard deviation, and field text characteristic matching is achieved. And it is ensured that the feature vectors can accurately capture domain-specific semantics.
Owner:SUZHOU YINGTIANDI INFORMATION TECH CO LTD

Method and system for realizing data classification and grading based on large language model technology

The invention discloses a method and system for realizing data classification and grading based on a large language model technology, and the method comprises the steps: collecting related data from a plurality of data sources, and carrying out the preprocessing, thereby forming a standardized preprocessing data set; obtaining text data from the text data, and extracting features of word frequency, TF-IDF values and semantic vectors to obtain a comprehensive feature vector; randomly extracting a part of the standardized preprocessed data set, and dividing and labeling the standardized preprocessed data set to obtain labeled training set data and test set data; training a classification and grading model; an optimal classification and grading model is selected; inputting the data to be classified and graded into the optimal classification and grading model to obtain a classification and grading result; the system comprises a data acquisition module, a data processing module, a feature extraction module, a data labeling module, a model training module, a model evaluation module and a data classification and grading module. According to the method, the scientificity, accuracy and efficiency of classification and grading are improved, and the requirements of enterprises and organizations for data classification and grading management are met.
Owner:XIAN AMAI XINKE TECH CO LTD

Retrieval optimization method based on dynamic mixed retrieval and adjacent paragraph introduction

The invention discloses a retrieval optimization method based on dynamic mixed retrieval and adjacent paragraph introduction. The method comprises the following steps: taking related documents as a corpus of a retriever; performing retrieval by using a BM25 algorithm and a HyDE algorithm to generate corresponding preliminary retrieval document paragraphs; calculating a query specificity value for user query by using TF-IDF; determining weights of a BM25 algorithm and a HyDE algorithm, and respectively performing weighted fusion on a document paragraph correlation score based on word frequency statistical calculation in the BM25 algorithm and a document paragraph correlation score based on semantic embedding similarity calculation in the HyDE algorithm to generate candidate document paragraphs and comprehensive scores corresponding to the candidate document paragraphs; the document paragraphs ranked in the top are selected from the preliminary retrieval document paragraphs, adjacent document paragraphs are obtained, weighted scores of the adjacent document paragraphs are calculated, the adjacent document paragraphs are added into a candidate document paragraph list, and the document paragraphs with the top scores are selected from the list to serve as retrieval results.
Owner:SOUTH CHINA UNIV OF TECH

Resume label generation method and device and medium

The invention discloses a resume label generation method and device and a medium, and relates to the field of natural language processing, and the method comprises the steps: carrying out the preprocessing of an original resume text, and carrying out the word segmentation of original resume data into a plurality of text words; counting word frequencies corresponding to the text words, and endowing the text words with corresponding importance weights according to the word frequencies; vectorizing the text words to obtain corresponding word vectors, and aggregating the word vectors according to the importance weight to obtain resume text vectors corresponding to the original resume text; based on a preset hierarchical label library, performing multi-label classification on the text semantic vector through a nonlinear classification algorithm, and outputting a prediction label of the original resume text; and screening the label prediction values based on a preset threshold to obtain a final resume label set. By fusing word frequency weighted semantic representation and a nonlinear classification algorithm, the accuracy and adaptability of resume label generation are remarkably improved.
Owner:SHENZHEN INSPUR HAIYUE HUMAN RESOURCES TECHNOLOGY CO LTD

AI accompanying system and method based on general artificial intelligence

The invention discloses an AI accompanying system and method based on general artificial intelligence, and the method comprises the following steps: S1, collecting and preprocessing user data, and generating a standardized data set; s2, performing dimension reduction by adopting diffusion mapping, extracting voice and behavior characteristics, and retaining a state change trend; s3, constructing a behavior knowledge graph based on symbol logical reasoning, and recording symbol state changes; s4, similar states and coping strategies are retrieved by using sparse distributed memory storage symbol representation; s5, combining manifold embedding, historical matching and symbolic reasoning, updating a user state and generating a personalized accompanying strategy; s6, analyzing the voice rhythm and the word frequency, constructing an emotion evaluation model, and generating a personalized emotion intervention scheme; s7, optimizing the knowledge graph by adopting incremental memory, and improving the companion strategy adaptability; and S8, recording feedback, continuously optimizing emotion intervention, and realizing intelligent accompanying. The AI accompanying system improves intelligence, personalization and long-term adaptability of AI accompanying, and is suitable for the fields of old-age care, mental health management and the like.
Owner:BEIJING GUANGRONG INNOVATION TECHNOLOGY CO LTD

Semantic understanding-based completion environmental protection acceptance survey report auxiliary auditing system

The invention relates to the technical field of natural language processing, in particular to a semantic understanding-based completion environmental protection acceptance survey report auxiliary auditing system, which comprises a structure recognition module, a semantic chain construction module, a semantic reconstruction module, an image-text coordinate fusion module and a spatial offset judgment module. In the method, through extracting subjects, terms and facility words in facility paragraphs and establishing a combination relationship, ordered labeling of semantic elements is realized, statement structure definition is enhanced, a semantic chain group is constructed and facility classification is associated, semantic consistency comparison is supported, semantic sequences and word frequency distribution in design and reply contents are compared, and description consistency is verified; facility coordinates in an image are extracted, character paragraphs are matched, image-text content anchoring, orientation, terrain and distance word extraction and conversion into space parameters to be compared with an image path are achieved, space consistency verification is completed, and the intelligent level of auditing on semantic comprehension, image-text fusion and space logic is overall improved.
Owner:SHANGHAI RUIDUN INFORMATION TECHNOLOGY CO LTD +1

Internet-oriented illegal advertisement identification method, device and system

The invention relates to the technical field of text processing, in particular to an internet-oriented illegal advertisement recognition method, device and system. According to the method, text information data carried by an advertisement is specifically analyzed, the high distribution consistency condition of part-of-speech and word frequency of each segmented word in a local text before and after special symbol processing and the similarity condition of an overall text are compared and analyzed, and the necessary processing degree of a special symbol is obtained in combination with the connection association condition of words before and after the special symbol processing; through the processed text data and the abnormal possibility of the image and the release time, the violation score of the advertisement is evaluated and identified, and a more accurate evaluation result is obtained. According to the method, by comparing semantic change conditions represented by coherent text contents before and after special symbol processing, interference special symbols carried in the advertisement text are processed, the accuracy of advertisement violation information identification is improved, and violation advertisement identification is more comprehensive and reliable.
Owner:BEIJING MEISHU INFORMATION TECH

Text data parallel semantic deduplication method, system and equipment and medium

The invention provides a text data parallel semantic deduplication method, system and device and a medium, and belongs to the technical field of text processing. The method comprises the steps of collecting large-scale text data to be subjected to duplicate removal, performing preprocessing to generate sample data, performing word segmentation and word frequency statistics by using a BERT word list, and calculating a word frequency vector of each piece of sample data; using a pre-trained BERT model to extract semantic features of the sample data, calculating word semantic weights of words in the sample data according to the word frequency vectors, and calculating similarity of the sample data based on the semantic features and a simhash algorithm; and according to the semantic similarity of the sample data, maintaining a pre-deletion dictionary and a global deduplication graph, and executing a deletion operation of the sample data based on the global deduplication graph to form a deduplicated data set. By combining the BERT model, the simhash algorithm and the global deduplication graph, the efficiency and accuracy of parallel semantic deduplication of large-scale text data are effectively improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Extension query method and system based on large language model entity extraction

The invention discloses an extended query method and system based on big language model entity extraction, key entities in original queries are extracted, the original queries are repeatedly spliced, the spliced original queries and key entities are fused, extended queries are obtained, entity interpretation is used as an extension basis, and compared with a traditional word frequency statistics or synonym replacement method, the extended query method and system have the advantages that the extension efficiency is improved; the large language model can generate extended fragments with semantic consistency and information integrity in combination with the context, it is guaranteed that the finally generated extended query is consistent with the original intention of a user, semantic drift and invalid extension are reduced, the correlation and precision of extended content are improved, meanwhile, the spliced original query and key entities are fused, and the user experience is improved. According to the method, the original query is locally enhanced by adopting an entity-level semantic interpretation mode, higher interpretability and controllability are prepared, and a complicated model structure and a training process are not needed, so that the deployability and universality of the method are remarkably improved, the expansion effect is more stable, and diversified retrieval requirements can be met.
Owner:PEOPLE CN CO LTD +1

Speed Up Methods and Systems for Large Language Model Training

A method initializes and accelerates training of neural network based large language model, including by: (i) accessing a corpora for training a neural-network based large language model having word embeddings and word projections in respective word embedding and word projection layers and at least one hidden layer; (ii) counting raw token frequencies associated with content within the corpora; (iii) smoothing the raw token frequencies into a series of vector norms based on log or scaled log functions parameterized by maximum norm information; and (iv) injecting vector norm information into word embeddings and / or word projections based on norm-angle reparameterization to prepare the large language model for training.
Owner:APPL TECH APPTEK

Data mining system and method for mail data

The invention discloses a data mining system and method for mail data, relates to the technical field of data retrieval processing, solves the technical problem of inaccurate retrieval matching caused by single semantic analysis retrieval, and improves data relevance by unifying multi-source data formats and combining an enterprise address book to supplement metadata. Structured data standardization and unstructured text vectorization are carried out, a computable semantic space is constructed, keywords are screened through a matching proportion, a word frequency weight is calculated in combination with a BM25 algorithm, the keyword matching accuracy is improved, a semantic vector is generated by utilizing a pre-training language model, and semantic correlation is calculated through cosine similarity. Keyword matching scores and semantic similarity scores are dynamically fused, accuracy and comprehensiveness are balanced, accuracy and recall rate are calculated in real time, retrieval deviation types are automatically recognized, keyword weights and similarity threshold values are adjusted in a targeted mode or synonym lists are expanded, and self-adaptive optimization of the system is achieved.
Owner:SHANGZHONG ONLINE TECH CO LTD

Description generation method based on continuous zero sample

The invention provides a description generation method based on continuous zero samples. The description generation method comprises the following steps: sequencing plain text training corpora according to the difficulty of continuous learning; setting that each training step t only can access the t-th corpus, and for the t-th learning step, generating a synthetic image for an input text of the t-th corpus by using a diffusion model; extracting image embedding of the synthesized image by using a CLIP image encoder; retrieving a group of text description embedding with similar semantics in a text corpus by using image embedding; image embedding and text description embedding are input into a fusion module network, and the output of a fusion module is defined as a soft prompt of a large language model; for an input text, key entities are extracted from the text through word frequency statistics so as to construct a soft prompt and a hard prompt of the large language model; hard prompts and soft prompts are input into a large language model network, and key parts in input are concerned through an attention mechanism.
Owner:JIANGSU VOCATIONAL & TECHNICAL UNIVERSITY OF ARCHITECTURE

Question and answer evaluation method and system for autism intervention

The invention discloses a question and answer evaluation method and system for autism intervention, and relates to the field of medical electronic systems.The method comprises the steps that historical data are extracted, noise reduction processing is carried out to obtain noise reduction data after noise reduction processing, and a database is built according to the noise reduction data, a word frequency-inverse document frequency algorithm and a BERT model; extracting question and answer information of a plurality of different autism stages in the database, and constructing a judgment model according to the question and answer information and the database; obtaining current question and answer data, and obtaining a judgment result of the child according to the current question and answer data and the judgment model; the teacher resume in the database is extracted, multiple recommendation teachers are determined according to the teacher resume and a judgment result, the two are combined to form a multi-level semantic feature matrix, local word frequency information is reserved, global context understanding is fused, and therefore the problem that feature extraction of a single model is one-sided is solved; accurate recognition capability of autism core symptoms (such as social obstacles and communication defects) is remarkably improved, and a more reliable behavior characteristic basis is provided for an intervention scheme.
Owner:高班超

Method and device for improving RAG recall rate in voice question and answer scene

The invention provides an RAG recall rate improving method and device in a voice question and answer scene, and relates to the technical field of data processing.The method comprises the steps that semantic cleaning processing is conducted on an original corpus containing a voice recognition result, semantic compression is conducted on the cleaned original corpus, and the original corpus is obtained; utilizing the plurality of candidate embedded vector generation models to respectively execute vector generation operation, and outputting word vectors; for each word vector, calculating a semantic fidelity score; evaluating the plurality of semantic fidelity scores, and selecting a target embedding vector generation model with the optimal semantic fidelity score from the plurality of candidate embedding vector generation models; calculating a word frequency value and an inverse document frequency value of each word according to data input, judging whether the words are professional hot words or not, and screening out the professional hot words to construct a hot word list; and jointly inputting an embedded vector output by the target embedded vector generation model and the hot word list into a question and answer module, and outputting a target answer text. According to the invention, the RAG recall rate in the voice question and answer scene can be improved.
Owner:ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD

Retrieval sorting method and device based on large language model

The embodiment of the invention relates to a retrieval sorting method and device based on a large language model. The method comprises the following steps: selecting a large language model which is realized based on a coder-decoder structure and is pre-trained as a first LLM (Logical Language Model); constructing a document scoring / sorting model based on the first LLM model; training a document scoring / sorting model on the basis of a preset score / data set under the condition that model parameters of the first LLM model are unchanged; after training is finished, receiving a first query text input by the user, and taking a document library specified by the user as a first document library; performing first-stage document screening on the first document library according to the first query text in a word frequency scoring mode; performing two-stage document screening on the primarily screened document sequence by using a document scoring model according to the first query text; and performing three-stage document reordering on the secondary screening document sequence by using a document ordering model according to the first query text to obtain a recommended document sequence and feeding back the recommended document sequence to the user. According to the invention, the retrieval precision can be improved.
Owner:BEIJING DP TECH CO LTD

Power big data dictionary construction system based on improved SO-PMI algorithm

The invention relates to the technical field of power big data processing, and discloses a power big data dictionary construction system based on an improved SO-PMI algorithm. The system comprises a corpus preprocessing module, an algorithm adaptation module, a word frequency correlation analysis module, a dictionary hierarchy construction module and a domain lexicon fusion module. The corpus preprocessing module is used for acquiring a field text, extracting basic lexical elements, marking positions and frequencies of terminologies, screening high-frequency core lexical elements and constructing a basic lexical element library; an algorithm adaptation module assigns values to lexical element weights, adjusts a co-occurrence window and an association threshold value of the SO-PMI algorithm, and constructs an adaptation parameter set; the word frequency association analysis module extracts a co-occurrence sequence, compares association strength, screens associated word tuples, and obtains an association relationship set; the dictionary hierarchy construction module collects hierarchy labels and obtains a hierarchy affiliation label group after matching; and the domain lexicon fusion module divides the lexical element association groups to corresponding classification nodes to generate a domain dictionary.
Owner:WEIHAI POWER SUPPLY COMPANY OF STATE GRID SHANDONG ELECTRIC POWER COMPANY +1

Technical achievement classification and retrieval method and system

The invention relates to the technical field of information retrieval, in particular to a technical achievement classification and retrieval method and system.The technical achievement classification and retrieval method comprises the following steps that on the basis of metadata of literatures in a scientific and technical literature database, word frequency analysis is executed, keyword frequency distribution is recorded, the keyword frequency distribution is associated with the literature metadata, and a metadata association analysis result is generated; and according to the metadata association analysis result, mapping the keyword and the classification label to obtain a keyword field mapping table. According to the method, through word frequency analysis and keyword frequency recording of literature metadata, accurate mapping from keywords to field tags is achieved, the accuracy of literature classification is improved, high efficiency of results is ensured by optimizing classification parameters, and the accuracy of literature classification is improved by updating a literature index framework and adopting user feedback optimization classification and index parameters. According to the method, the retrieval flexibility and the user satisfaction degree are improved, more accurate and quicker retrieval is realized through logic programming of keyword and field combination, the time of researchers is greatly saved, and research and development resource configuration is optimized.
Owner:SHENZHEN WOCHENG NETWORK TECH CO LTD +1

Analysis method and device for credit granting approval process

The invention provides a credit granting approval process analysis method and device. Comprising the steps of performing word segmentation on text content in a credit granting approval document set, and combining the obtained segmented words to obtain combined phrases; calculating weighted word frequency-inverse document word frequency of the combined phrases; target phrases with weighted word frequency-inverse document word frequency larger than a word frequency threshold value are screened out from the combined phrases, and an amplified word list is generated according to the target phrases and the original word list; training a target word segmentation device based on the amplified word list; finely adjusting the basic vector model according to the training corpus to obtain a term vector model; analyzing the auditing opinion information based on a large language model to obtain an initial risk analysis rule containing a preset dimension; based on a similarity algorithm, a target word segmentation device and a term vector model, grouping and integrating the initial risk analysis rules to obtain risk analysis rules; and processing the information of each examination and approval stage by using a big language model and adopting a risk analysis rule to obtain examination and approval suggestion information.
Owner:MINSHENG BANKING CORP

Automatic identification, classification and development trend analysis method of net red villages based on multi-source data fusion and natural language processing

The method for automatic identification, classification and development trend analysis of net red villages based on multi-source data fusion and natural language processing comprises the following steps: UGC data is crawled from Xiaohongshu and Douyin through a distributed master-slave architecture, de-duplicated based on SimHash, and normalized in time and coding format; a text semantic fingerprint is generated, and multi-level semantic cache fingerprint matching is performed; for unassigned text, its complexity is calculated, and a large language model API is adaptively called to automatically complete and extract five-level administrative divisions; weights are determined based on the analytic hierarchy process, interaction indicators such as likes, comments, collections and forwards are integrated, and a comprehensive network heat index of the village is obtained; an external text mining tool is connected, and batch word frequency analysis, semantic network analysis and sentiment tendency evaluation are performed; a document-term matrix is constructed, TF-IDF weighting is performed, and unsupervised clustering algorithm is used for clustering analysis of village characteristics; cross-dimension analysis is performed on the clustering results, and a development portrait, advantage mining and operation suggestion warning are automatically generated in combination with the SWOT model.
Owner:ZHEJIANG UNIV OF TECH

Word-tag-based language system for sentence acceptability judgment

Methods and systems for sentence acceptability judgment. A word frequency distribution for a predetermined textual data set is obtained. A replacement rate is determined based on an obtained word frequency distribution for a predetermined textual data set and every occurrence of a word having a frequency lower than the replacement rate is replaced in text of a training data set with a corresponding tag to generate revised text of the training data set (the training data set comprising at least a portion of the predetermined textual data set). A plurality of language models are trained with the revised text and a best performing trained language model of the plurality of trained language models is selected. An acceptability of each sentence of a plurality of candidate sentences is rated using the selected trained language model and a best sentence of the plurality of candidate sentences is selected based on the ratings.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION