Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

403 results about "Word list" patented technology

Voice generation method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a voice generation method which comprises the following steps: constructing a multi-language voice synthesis model, obtaining plain text data and paired voice text data, and constructing an expansion vocabulary; updating a language perception embedding layer and model parameters, and converting an input text into a mark sequence; and the encoder extracts context semantic features, extracts pronunciation rule features, and the decoder fuses the features to generate an acoustic feature sequence, and converts the acoustic feature sequence into target voice data. According to the invention, the multi-language speech synthesis model is combined with the language perception embedding layer, so that the speech generation capability of a low-resource language is improved; the text conversion accuracy is improved by expanding the vocabulary, the target language learning ability is enhanced by unsupervised training, the low data environment adaptability is optimized by supervised training, and the speech naturalness and fluency are improved by feature fusion.
Owner:PING AN TECH (SHENZHEN) CO LTD

Financial fraud detection method based on large language model

The invention provides a financial fraud detection method based on a large language model. The method comprises the steps of obtaining a to-be-recognized text; performing word segmentation on the to-be-recognized text through the target word segmentation tool and the financial fraud dictionary, and determining a fraud sensitive word list; calculating the weight of each sensitive word in the sensitive word list according to a target algorithm to obtain a sensitive word weight feature vector; inputting the sensitive word weight feature vector and a to-be-recognized text into a large language model, and determining a context semantic vector of the sensitive word in combination with a word embedding technology; determining the similarity between the context semantic vector of the sensitive word and a preset financial fraud type semantic vector; and determining a financial fraud type according to the similarity. Through the implementation of the method, the generalization ability of the pre-training model is utilized to capture text deep semantics, priori knowledge is injected in combination with a sensitive word weight mechanism, the model is guided to focus high-risk vocabularies, the defect of a traditional method in semantic comprehension is overcome, the financial fraud recognition rate is increased, and the omission ratio is reduced.
Owner:CHONGQING UNIV OF TECH

Methods and systems for speech emotion retrieval via natural language prompts

Methods and systems for generating training data for training a contrastive language-audio machine-learning model. A plurality of audio segments are retrieved from a speech emotion recognition (SER) database along with metadata associated with the audio segments. The metadata of each audio segment includes an emotion class. Words or terms associated with emotions are retrieved from a lexicon. A large language model (LLM) is executed on (i) the classes of emotion associated with the audio segments and (ii) the words or terms from the lexicon. This generates a plurality of text captions associated with emotion, which are stored in a caption pool. For each audio segment retrieved from the SER database, that audio segment is paired with one or more of the text captions from the caption pool that were generated based on the emotion class associated with that audio segment. This yields audio-text pairs for training a contrastive learning model.
Owner:ROBERT BOSCH GMBH

Three-dimensional model intelligent generation method and system based on natural language

The invention discloses a three-dimensional model intelligent generation method and system based on a natural language, and the method comprises the steps: receiving and describing the natural language inputted by a user in real time, based on the three-dimensional model data structure, the de-noising symbol set and the stop word list, key information data are extracted according to natural language description and mapped into a three-dimensional model parameter set; according to the scene type described by the natural language and the three-dimensional model parameter set, calling a three-dimensional model generation network corresponding to the scene type described by the natural language to generate an initial three-dimensional model; and receiving and analyzing a model optimization instruction input by a user in real time, updating the geometric structure, detail features and physical attributes of the initial three-dimensional model according to an analysis result, generating a dynamic interaction model, and outputting a three-dimensional model file of the dynamic interaction model. By applying the method and the system provided by the invention, the efficiency, the accuracy, the flexibility, the usability and the user experience of three-dimensional model modeling are improved.
Owner:BEIJING DONGFANG AIDIPU DIGITAL TECH CO LTD

Speech recognition method and system for constructing small language based on whispertoken

The invention provides a voice recognition method and system for constructing a small language based on whispertoken, and relates to the technical field of natural language processing and voice recognition, and the method comprises the steps: extracting all tokens related to a target small language in a whispertoken, and forming an initial candidate set; matching and analyzing the tokens in the initial candidate set and the collected training text corpus of the target small language, and counting the occurrence frequency of the tokens in the corpus; and screening high-frequency tokens according to a frequency statistical result, and supplementing low-frequency tokens to construct a dynamic vocabulary. According to the method, the vocabulary quality is improved, the model training efficiency is optimized, the speech recognition accuracy is enhanced, the model generalization ability is improved, and the model construction process is simplified, so that an efficient, accurate and easy-to-implement solution is provided for the field of minority language speech recognition.
Owner:BEIJING RUI KELUN INTELLIGENT TECH CO LTD

Language model training method and device, text processing method and device, equipment and medium

The invention provides a language model training method and device, a text processing method and device, equipment and a medium, and relates to the technical field of natural language processing, and the method comprises the steps: predicting a first probability matrix corresponding to each data unit in a sample text based on a teacher model; the first probability matrix comprises a probability value of each data unit belonging to each lexical element in the first word list; according to the numerical value of each probability value in the first probability matrix, compressing the first probability matrix to obtain a second probability matrix corresponding to each data unit; performing alignment operation on the second word list according to lexical elements corresponding to probability values in the second probability matrix to obtain a third word list; and according to the third word list and the second probability matrix, carrying out distillation training on the student model to obtain a target language model, thereby reducing the storage cost, improving the distillation training efficiency, and enabling the target language model trained according to the method to better adapt to different model architectures and text processing scenes while keeping high performance.
Owner:IFLYTEK CO LTD

Electronic archive information extraction method and extraction system

The invention relates to the technical field of information extraction, in particular to an information extraction method and system for electronic archives. The method comprises the following steps: obtaining a to-be-processed electronic file and carrying out OCR identification to generate initial text data; error detection is carried out on the initial text data, and OCR error recognition candidate items in the initial text data are recognized; for each OCR misrecognition candidate item, generating a first data name according to context semantics and a layout structure of the candidate item; extracting low-level features of the first data name, and performing named entity recognition on each OCR misrecognition candidate item by utilizing a preset field word list in combination with the first data name; through a four-in-one process of ''misrecognition detection + named entity recognition + semantic error correction + templated extraction'', the core technology bottlenecks of inaccurate recognition, poor error correction capability, low information extraction intelligence and the like in the prior art are solved, and the accuracy, stability and intelligent level of electronic archive information extraction are remarkably improved.
Owner:INNER MONGOLIA FINANCE AND ECONOMICS UNIVERSITY

Dynamic vocabularies for conditioning a language model for transforming natural language to a logical form

Techniques are disclosed herein for generating dynamic vocabularies for conditioning a language model. A dynamic vocabulary is constructed from an input prompt, database schema information for a database to be queried, and programming language information for a programming language to be used for querying the database to condition the language model to predict an output statement in the programming language. The dynamic vocabulary can be included in prompt information that is provided to the language model. The number of tokens in the dynamic vocabulary can be different than a number of tokens included in a vocabulary of the language model. By utilizing a dynamic vocabulary, the language model can be conditioned to predict tokens for the output statement that are contextually consistent with the tokens included the dynamic vocabulary.
Owner:ORACLE INT CORP

Drug-target interaction prediction method based on pre-training language model

According to the pre-training language model-based drug-target interaction prediction method designed by the invention, natural language processing and graph neural network technologies are fused, context semantic features are automatically extracted from drug molecule SMILES character strings and protein sequences, and by constructing a graph structure taking drug-target pairs as nodes, the drug-target interaction is predicted. The weight of an edge is defined according to the similarity between embedded vectors, and a simplified graph convolutional network is adopted to carry out graph structure modeling to realize complex relation learning, so that the accuracy, generalization and interpretability of prediction are improved, the limitation of a traditional method on the problems of sparse feature expression, mutual information loss and'words outside a vocabulary 'is overcome, and the prediction accuracy, generalization and interpretability are improved. And finally, the accuracy of predicting the drug-target interaction relationship is improved.
Owner:SHANGHAI JIAOTONG UNIV

Intelligent personalized interaction system and construction method thereof

The invention provides a construction method of an intelligent personalized interaction system. The construction method comprises the following steps: constructing a database; preparing a training set and a test set with problem texts, constructing a vocabulary of high-frequency words for the training set and the test set, and performing text coding; constructing a workflow scheduler for counting high-frequency problem types; the user portrait building module is used for obtaining a user portrait according to historical dialogues, user questions and high-frequency question types of the database; the intelligent personalized interaction system comprises a user question input interface, a question classification module and a database which are connected in sequence, and the database is connected with a workflow scheduler and a user portrait construction module. And outputting the user portrait and the high-frequency question type to a pre-trained large language model to generate a recommendation question. The invention further provides a corresponding system. According to the method, deep understanding, accurate analysis and personalized interaction of user questions are realized, so that the user experience is remarkably improved, and diversified requirements of users in different scenes are met.
Owner:EAST CHINA UNIV OF SCI & TECH

Electronic medical record content quality control method and system based on large language model

The invention provides an electronic medical record connotation quality control method and system based on a large language model, and the method comprises the steps: carrying out the modal classification of a whole-process medical record of a patient, and uniformly converting the whole-process medical record into a corresponding text modal; performing hierarchical division on the whole-process medical record of the text mode to form a medical record content tree; according to different nodes of the medical record content tree, retrieving a quality control rule set corresponding to the nodes of the medical record content tree, and retrieving a corresponding diagnosis and treatment knowledge set; for each quality control rule in the quality control rule set, if a plurality of medical record content tree nodes are needed, establishing a corresponding context set, constructing a quality control cue word list, inputting cue words into the connotation quality control big language model, judging whether a quality control problem exists in the cue words or not by the connotation quality control big language model, and proposing a modification suggestion; and checking and revising the output result to obtain a final quality control result. According to the method, key point missing, description conflicts and deep diagnosis and treatment logic contradictions in medical records can be recognized and corrected, and the accuracy and stability of quality control results are improved.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Data privacy protection method, system and device for large language model application

The embodiment of the invention is suitable for the technical field of artificial intelligence, and provides a data privacy protection method, system and device for large language model application, the method is applied to client equipment, and the client equipment is deployed with an input layer and an output layer of a large language model. The method comprises the steps that text input data of the large language model is coded according to a vocabulary, a mark list represented by integers is obtained, and the vocabulary comes from a server; processing the mark list through the input layer to obtain intermediate input data; the intermediate input data is sent to the server; receiving intermediate output data returned by the server; and processing the intermediate output data based on the output layer and the vocabulary to obtain text output data. Through the method, the privacy protection of the input data can be realized while the output data is automatically obtained by using the large language model.
Owner:NATIONAL UNIVERSITY OF SINGAPORE +1

Large language model high-quality text data set construction method and system

The invention discloses a large language model high-quality text data set construction method and system, belongs to the technical field of deep learning, and aims to solve the technical problems of how to construct a high-quality text data set and how to reduce noise data. Comprising the following steps: collecting related literatures from industry data to obtain a text corpus; performing data preprocessing on the collected text corpus to obtain a pre-labeled text corpus; constructing a label system based on the industry terminologies and the indexes, and labeling the relationship between the entities in the pre-labeled text corpus based on the label system; the method comprises the following steps: updating an original vocabulary of a pre-trained large language model based on an industry dictionary, carrying out unsupervised training on the pre-trained large language model, taking a masked text corpus and corresponding label information as input attributes, carrying out upper and lower semantic analysis on the input attributes through the large model subjected to unsupervised training, predicting words masked by the text corpus, and obtaining a word masked by the text corpus. And outputting the enhanced text data, and performing data testing on the text data.
Owner:INSPUR SOFTWARE TECH CO LTD

Automatic deep thinking model selection training method and system for large language model

The invention relates to the technical field of artificial intelligence, in particular to an automatic deep thinking model selection training method for a large language model. Comprising the following steps: expanding a word segmentation device vocabulary, and adding a special mark lt; carrying out think gt; and lt; / (ginkgt); the identification module is used for identifying a content structure of the deep thinking mode; a dialogue template is designed, input and output formats of a common mode and a deep thinking mode are distinguished, and the deep thinking mode comprises lt; carrying out think gt; marking a guided reasoning process; a common mode training data set and a deep thinking mode training data set with the same scale are constructed, and joint training is performed on the model, so that the model has response capabilities of the two modes at the same time. The method comprises the following steps: adding a special mark (lt; carrying out think gt; and lt; / (ginkgt); and a dialogue template is customized, semantic distinguishing between a common mode and a deep thinking mode is achieved, and the model can accurately switch response strategies according to requirements.
Owner:SUZHOU CHANGYUXING TECHNOLOGY CO LTD

Unmanned vehicle inspection small target detection method based on efficient attention mechanism

The invention discloses an unmanned vehicle inspection small target detection method based on an efficient attention mechanism. The method comprises four stages of image feature extraction, vocabulary embedding extraction, efficient attention coding and cross attention decoding. And image feature extraction: performing feature extraction on the input image by using the backbone network to generate a multi-scale feature map. And vocabulary embedding extraction: generating vocabulary embedding in the defined category vocabulary through a CLIP text encoder. And efficient attention coding: performing deep feature interaction, space attention guidance and multi-scale feature aggregation processing on the multi-scale feature map of the picture to obtain an image feature map fusing the visual context and the multi-scale information. And cross attention decoding: embedding the aggregated feature map and vocabulary, and outputting a final detection result through cross attention fusion, IoU perception query and regional text comparison processing. Compared with the prior art, the method has the advantages of good prediction effect, good practicability and the like.
Owner:HOHAI UNIV +1

Encrypted network traffic classification method based on pre-trained large language model

The invention relates to an encrypted network traffic classification method based on a pre-trained large language model, and belongs to the technical field of encrypted network traffic classification. The method mainly comprises three stages: a pre-training stage: converting original encrypted network traffic data into a double-byte hexadecimal format through preprocessing, generating and optimizing a basic vocabulary by using a byte pair coding algorithm, then constructing a large language model, and obtaining a pre-training model through distributed training; in the retraining stage, the data is subjected to head byte shuffling processing, and the pre-training model is quickly retrained to improve the generalization ability. In the fine tuning stage, to-be-classified data is preprocessed to generate hexadecimal double-byte data with labels, and classification task training is performed by using the retrained model to obtain an encrypted network traffic classification fine tuning model and classification accuracy. Through the combination of pre-training and retraining and the fine tuning of the pre-training model by means of data classification, the efficient processing and accurate classification of the complex encrypted network traffic are realized.
Owner:CHONGQING UNIV

Text information extraction method and system based on large language model

The invention discloses a text information extraction method and system based on a large language model, and belongs to the technical field of natural language processing. Inputting the obtained text information data into the trained text information extraction model, and performing text information processing; wherein the training of the text information extraction model comprises the following steps: constructing a generative and discriminant text information extraction model based on a large language model; constructing a dialogue strategy gradient discrete prompt optimization framework to enhance the prompt stability; and constructing a noise channel matrix, and performing data enhancement and stability enhancement. According to the method, the stability of the model and the accuracy of the extraction result are effectively improved by introducing the judgment component, the balance between generation and copying is optimized by designing the dynamic generation probability calculation module and the extension word list mechanism, the stability and accuracy of text information extraction are effectively improved, and the user experience is improved. And efficient and reliable technical support is provided for scenes such as public opinion analysis and knowledge graph construction.
Owner:WUHAN UNIV

Sensitive data sharing method and system based on data consanguinity

The invention relates to a sensitive data sharing method and system based on data consanguinity, and belongs to the technical field of data management and sharing. The system comprises four core components including a data consanguinity module, a sensitive word list recognition module, a data desensitization module and a data sharing module, wherein the data consanguinity module dynamically collects and analyzes a full life cycle flow path of data from a source end to an application end, and constructs a consanguinity map of structured storage; the sensitive word list identification module identifies sensitive data according to the sensitive word list and the blood relationship map, and executes risk grading, dynamic monitoring and compliance auditing; the data desensitization module configures a desensitization rule to deform or convert sensitive data; and the data sharing module performs intelligent approval, dynamic authorization and full-link audit on the data sharing request based on the sensitive word list and the desensitization rule. According to the method, the data security sharing efficiency is remarkably improved, and the security goals of data availability and invisibility and no risk in sharing are achieved.
Owner:STATE GRID FUJIAN ELECTRIC POWER CO LTD

Power grid equipment intelligent question and answer optimization method and system based on large language model

The invention relates to the technical field of intelligent questioning and answering workers, in particular to a power grid equipment intelligent questioning and answering optimization method and system based on a large language model. The method comprises the following steps: acquiring problem data of power grid equipment, and creating a dynamically updated power grid equipment vocabulary to store the problem data; a part-of-speech tagging model of a conditional random field is used for tagging part-of-speech of question data, power grid equipment knowledge retrieval and large language model building are performed based on the question data of tagging and semantic roles, parameter updating is performed on the built large language model, and answer optimization is performed according to a processing result of the large language model. According to the method, a position coding formula combined with semantic information is introduced, so that a large language model can better understand the relationship between a word order and semantics in a text in the field of power grid equipment, semantic logic in the text can be more deeply mastered, the ability of understanding the operation process and the equipment relationship is improved, and answers more conforming to professional logic are output.
Owner:ZHONGWEI POWER SUPPLY COMPANY OF STATE GRID NINGXIA ELECTRIC POWER

Trusted time sequence prediction method based on big language model fusion knowledge graph

The invention provides a credible time sequence prediction method based on a big language model fusion knowledge graph, and the method comprises the steps: carrying out the reversible normalization processing of original time sequence data, and decomposing the normalized time sequence data into trend, season and residual components; constructing a time sequence knowledge graph with rich semantics by utilizing a vocabulary of the pre-trained large language model; screening the time sequence knowledge graph to obtain a prefix prompt sequence, and splicing the prefix prompt sequence with the time sequence embedding to obtain an enhanced time sequence embedding; and embedding and inputting the enhanced time sequence into a pre-trained large language model for processing, and performing inverse standardization processing to obtain a predicted value. According to the method, diversity of time sequence data components can be effectively captured, high-quality context prompts are provided by means of the time sequence knowledge graph, and the prediction performance of the model is remarkably improved; and meanwhile, cognitive uncertainty and accidental uncertainty of a prediction result can be quantified, so that credible time sequence prediction is realized.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Using semantic hierarchy trees to increase the robustness of open-vocabulary object detection and vocabulary adapter

An object identification system includes: a category module configured to, for a category of a vocabulary of objects, retrieve a hierarchy including at least: a sub-category that is more specific than the category; and a super-category that is less specific than the category; a sentence module configured to generate a set of sentences for the category that describe the hierarchical relationship between sub-category, super-category, and the category; an encoder module configured to encode the sentences into encodings, respectively, for the category; an aggregator module configured to generate an aggregated encoding for the category by aggregating the encodings of the category; and an identification module configured to selectively identify an object included in a region of interest of an input image as being in the category based on a comparison of (a) an encoding of the region of interest and (b) the aggregated encoding for the category.
Owner:NAVER CORP

Method, device and equipment for constructing dynamic expansion word bank and medium

The invention relates to the technical field of passenger service, and discloses a construction method and device of a dynamic extension word bank, equipment and a medium. The method comprises the following steps: acquiring multi-source heterogeneous original knowledge data related to civil aviation passenger service, and converting the multi-source heterogeneous original knowledge data into a standard format text; processing the unstructured data in the standard format text based on the adaptive word segmentation model and the stop word list to obtain a plurality of candidate words; performing multi-dimensional corpus feature quantitative analysis on each candidate word to determine candidate keywords; and performing similarity calculation based on the entity word segmentation and the candidate keywords, determining effective candidate keywords, and adding the effective candidate keywords into the dynamic expansion word bank. By means of the method and device, the technical problems that in the prior art, a word bank used by a civil aviation service system is usually based on a static vocabulary or depends on manual compiling and updating, the static word bank cannot be updated in real time, the characteristics of the civil aviation field cannot be flexibly handled, and the maintenance cost is high due to manual updating and maintenance are solved.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Malicious URL detection method based on character-level language model and structural feature fusion

The invention discloses a malicious URL detection method based on character-level language model and structural feature fusion. The method comprises the steps that character-level URL semantic features are obtained through a character-level language model; the URL semantic features are sent to expansion pyramid attention, semantic enhancement is carried out, and enhanced character-level URL semantic features are obtained; extracting structural features of the URL character string to obtain URL structural features; and carrying out dynamic weighted fusion on the URL structural features and the enhanced character-level URL semantic features, and outputting a malicious URL judgment result through a classifier. Through character-level semantic understanding, a predefined word list library is separated, the adaptive capacity of random characters appearing in the URL is improved, meanwhile, the structural features of the URL are fused, and the recognition performance of the model for the malicious URL is further improved.
Owner:CHANGSHU INSTITUTE OF TECHNOLOGY

File anti-desensitization self-learning recognition system and method based on information entropy

The invention discloses a file anti-desensitization self-learning recognition system and method based on information entropy, belongs to the technical field of intersection of natural language processing and content security recognition, and is applied to document screening and risk recognition in a multi-task scene. The implementation method comprises the following steps of: 1, performing character recognition and noise reduction processing on an original file to form a data set; 2, training labeled sample data through small samples, respectively adopting probability distribution of a data sliding window and information entropy to carry out maximum and minimum normalization screening, and further utilizing a fitted linear regression model to form an anti-desensitization word list; 3, screening the anti-desensitization degrees of the sentence segments of the data set by adopting a dictionary tree Trie structure to form an anti-desensitization sentence segment table; 4, marking the chapter-level anti-desensitization degree data set text fragments by using the large model; 5, generating an anti-desensitization report according to the anti-desensitization word and the anti-desensitization degree of the marked anti-desensitization file; compared with the prior art, the anti-desensitization file screening method and device have the advantage that the anti-desensitization file screening accuracy is improved.
Owner:BEIJING INST OF TECH

Open-Vocabulary Object Detection Based on Frozen Vision and Language Models

An example method of training a detector head for object detection of a training object category based on a frozen vision and language model (VLM) is provided. The method includes receiving the frozen VLM pre-trained on a plurality of image-text pairs. The method includes determining, for an image embedding generated by a pre-trained image encoder of the frozen VLM and by the detector head, a detection region embedding indicative of one or more regions of interest in an image. The method includes generating, by a pre-trained text encoder of the frozen VLM, a text embedding of the training object category. The method includes predicting, by the detector head and based on the detection region embedding and the text embedding of the training object category, an object from a target object vocabulary associated with the training object category. The method includes providing the pre-trained frozen VLM and the trained detector head.
Owner:GOOGLE LLC

Bid invitation document qualification examination processing method and system

The invention provides a bid invitation file qualification review processing method and system. The method comprises the following steps: constructing a qualification performance vocabulary; designing a query template to fit user input; constructing a sample vector data set; filtering user input based on the qualification performance vocabulary; similarity comparison is carried out by calculating the cosine similarity of keywords input by a user and the query templates in the sample vector data set, and the query template corresponding to the highest similarity is taken to obtain a recombined query statement; performing preliminary matching on the recombined query statement in a vector database; according to the existence of the text additional field in the screened result and the repetition rate of the text additional field content, screening to obtain a corresponding qualification performance term input by the user; and combining the corresponding qualification performance regulations input by the user with the keywords input by the user, and then inputting the combination into the LLM model to generate answers. According to the method, the accuracy of qualification legitimacy and compliance review of the bid inviting document can be remarkably improved.
Owner:CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Efficient and effective system and method to build multi-lingual large language models

A method of enabling a language model trained in a first language to support a second language includes extending an existing vocabulary of the language model to include additional tokens for text in the second language. The method also includes initializing the additional tokens for text in the second language based on subtokens of tokens for text in the first language from the existing vocabulary. The method further includes training the language model using a mixed language dataset that includes a first language corpus and a second language corpus. In addition, the method includes performing instruction tuning using a dataset that includes (i) instruction and response pairs involving the first language and (ii) instruction and response pairs involving the second language.
Owner:SAMSUNG ELECTRONICS CO LTD

Large model cue word generation method and device, storage medium and program product

The embodiment of the invention provides a large model cue word generation method and device, a storage medium and a program product. In the method, quality index analysis can be performed on the initial cue word based on multiple task dimensions, so that iterative optimization is performed on the initial cue word accurately and efficiently based on the analysis result, the obtained reference cue word is matched with the task target of the model task on multiple task dimensions, and the effect is better. The reference cue word is derived, a cue word list of the target large model can be obtained, the cue word list comprises a plurality of candidate cue words, and different candidate cue words perform the same task guidance on the target large model through different description styles and / or cue word contents. According to the method, the user can select the target cue word as required when calling the target large model, so that the diversity of the cue word is enriched, the cue word list adapts to different description styles and content requirements, and the flexibility and applicability of calling the large model by the user are improved.
Owner:BEIJING 58 INFORMATION TTECH CO LTD

Key technology identification and IPC classification method and system based on large language model

The invention relates to the field of natural language processing, and provides a key technology identification and IPC classification method and system based on a large language model, and the method comprises the steps: S1, collecting domain data, and carrying out the data preprocessing, so as to construct a structured domain corpus; s2, reconstructing the mixed word segmentation device and generating a corresponding field vocabulary; performing incremental pre-training on the Qwen2.5 base model by using the pre-training data corpus and the field vocabulary and adopting a low-rank adaptation technology to inject field data; a supervision fine tuning training data set oriented to a specified task is constructed based on domain data, data preprocessing is carried out to convert the data into an input format suitable for an SFT unit, and supervision fine tuning is carried out on the SFT unit through a low-rank adaptation technology; according to the method, the key core technology and the IPC classification number of any specified field can be efficiently output, and work such as expert interview, data annotation and model training fitting does not need to be repeatedly carried out on the specified field like an existing method.
Owner:UNIV OF SCI & TECH OF CHINA +1

Dynamic scene 4D semantic map generation method and device and processing equipment

The invention provides a dynamic scene 4D semantic map generation method, a dynamic scene 4D semantic map generation device and processing equipment, and aims to realize geometric perception and semantic alignment combined processing in a single framework by designing a first feedforward framework for 4D semantic map generation. The framework comprises two core components, namely a streaming visual geometric converter for capturing space-time geometric features of a dynamic scene and a semantic bridging decoder for mapping the space-time geometric features to language aligned semantic spaces, so that the structural integrity is kept, and the semantic interpretability is improved. Different from a traditional method depending on time-consuming scene-level optimization, the method can effectively support multi-dynamic scene merging training, can be directly applied during reasoning, and is high in calculation efficiency and generalization ability. According to the design, the practicability of large-scale deployment is remarkably improved, a new thought of open vocabulary 4D scene understanding is developed, good data support can be provided for scene understanding tasks of applications such as intelligence, meta universe and digital twinning, and the method has good application prospects.
Owner:JIANGHAN UNIVERSITY