Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

262 results about "Word list" patented technology

Electronic archive information extraction method and extraction system

The invention relates to the technical field of information extraction, in particular to an information extraction method and system for electronic archives. The method comprises the following steps: obtaining a to-be-processed electronic file and carrying out OCR identification to generate initial text data; error detection is carried out on the initial text data, and OCR error recognition candidate items in the initial text data are recognized; for each OCR misrecognition candidate item, generating a first data name according to context semantics and a layout structure of the candidate item; extracting low-level features of the first data name, and performing named entity recognition on each OCR misrecognition candidate item by utilizing a preset field word list in combination with the first data name; through a four-in-one process of ''misrecognition detection + named entity recognition + semantic error correction + templated extraction'', the core technology bottlenecks of inaccurate recognition, poor error correction capability, low information extraction intelligence and the like in the prior art are solved, and the accuracy, stability and intelligent level of electronic archive information extraction are remarkably improved.
Owner:INNER MONGOLIA FINANCE AND ECONOMICS UNIVERSITY

Data privacy protection method, system and device for large language model application

The embodiment of the invention is suitable for the technical field of artificial intelligence, and provides a data privacy protection method, system and device for large language model application, the method is applied to client equipment, and the client equipment is deployed with an input layer and an output layer of a large language model. The method comprises the steps that text input data of the large language model is coded according to a vocabulary, a mark list represented by integers is obtained, and the vocabulary comes from a server; processing the mark list through the input layer to obtain intermediate input data; the intermediate input data is sent to the server; receiving intermediate output data returned by the server; and processing the intermediate output data based on the output layer and the vocabulary to obtain text output data. Through the method, the privacy protection of the input data can be realized while the output data is automatically obtained by using the large language model.
Owner:NATIONAL UNIVERSITY OF SINGAPORE +1

Unmanned vehicle inspection small target detection method based on efficient attention mechanism

The invention discloses an unmanned vehicle inspection small target detection method based on an efficient attention mechanism. The method comprises four stages of image feature extraction, vocabulary embedding extraction, efficient attention coding and cross attention decoding. And image feature extraction: performing feature extraction on the input image by using the backbone network to generate a multi-scale feature map. And vocabulary embedding extraction: generating vocabulary embedding in the defined category vocabulary through a CLIP text encoder. And efficient attention coding: performing deep feature interaction, space attention guidance and multi-scale feature aggregation processing on the multi-scale feature map of the picture to obtain an image feature map fusing the visual context and the multi-scale information. And cross attention decoding: embedding the aggregated feature map and vocabulary, and outputting a final detection result through cross attention fusion, IoU perception query and regional text comparison processing. Compared with the prior art, the method has the advantages of good prediction effect, good practicability and the like.
Owner:HOHAI UNIV +1

Sensitive data sharing method and system based on data consanguinity

The invention relates to a sensitive data sharing method and system based on data consanguinity, and belongs to the technical field of data management and sharing. The system comprises four core components including a data consanguinity module, a sensitive word list recognition module, a data desensitization module and a data sharing module, wherein the data consanguinity module dynamically collects and analyzes a full life cycle flow path of data from a source end to an application end, and constructs a consanguinity map of structured storage; the sensitive word list identification module identifies sensitive data according to the sensitive word list and the blood relationship map, and executes risk grading, dynamic monitoring and compliance auditing; the data desensitization module configures a desensitization rule to deform or convert sensitive data; and the data sharing module performs intelligent approval, dynamic authorization and full-link audit on the data sharing request based on the sensitive word list and the desensitization rule. According to the method, the data security sharing efficiency is remarkably improved, and the security goals of data availability and invisibility and no risk in sharing are achieved.
Owner:STATE GRID FUJIAN ELECTRIC POWER CO LTD

Trusted time sequence prediction method based on big language model fusion knowledge graph

The invention provides a credible time sequence prediction method based on a big language model fusion knowledge graph, and the method comprises the steps: carrying out the reversible normalization processing of original time sequence data, and decomposing the normalized time sequence data into trend, season and residual components; constructing a time sequence knowledge graph with rich semantics by utilizing a vocabulary of the pre-trained large language model; screening the time sequence knowledge graph to obtain a prefix prompt sequence, and splicing the prefix prompt sequence with the time sequence embedding to obtain an enhanced time sequence embedding; and embedding and inputting the enhanced time sequence into a pre-trained large language model for processing, and performing inverse standardization processing to obtain a predicted value. According to the method, diversity of time sequence data components can be effectively captured, high-quality context prompts are provided by means of the time sequence knowledge graph, and the prediction performance of the model is remarkably improved; and meanwhile, cognitive uncertainty and accidental uncertainty of a prediction result can be quantified, so that credible time sequence prediction is realized.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Using semantic hierarchy trees to increase the robustness of open-vocabulary object detection and vocabulary adapter

An object identification system includes: a category module configured to, for a category of a vocabulary of objects, retrieve a hierarchy including at least: a sub-category that is more specific than the category; and a super-category that is less specific than the category; a sentence module configured to generate a set of sentences for the category that describe the hierarchical relationship between sub-category, super-category, and the category; an encoder module configured to encode the sentences into encodings, respectively, for the category; an aggregator module configured to generate an aggregated encoding for the category by aggregating the encodings of the category; and an identification module configured to selectively identify an object included in a region of interest of an input image as being in the category based on a comparison of (a) an encoding of the region of interest and (b) the aggregated encoding for the category.
Owner:NAVER CORP

Method, device and equipment for constructing dynamic expansion word bank and medium

The invention relates to the technical field of passenger service, and discloses a construction method and device of a dynamic extension word bank, equipment and a medium. The method comprises the following steps: acquiring multi-source heterogeneous original knowledge data related to civil aviation passenger service, and converting the multi-source heterogeneous original knowledge data into a standard format text; processing the unstructured data in the standard format text based on the adaptive word segmentation model and the stop word list to obtain a plurality of candidate words; performing multi-dimensional corpus feature quantitative analysis on each candidate word to determine candidate keywords; and performing similarity calculation based on the entity word segmentation and the candidate keywords, determining effective candidate keywords, and adding the effective candidate keywords into the dynamic expansion word bank. By means of the method and device, the technical problems that in the prior art, a word bank used by a civil aviation service system is usually based on a static vocabulary or depends on manual compiling and updating, the static word bank cannot be updated in real time, the characteristics of the civil aviation field cannot be flexibly handled, and the maintenance cost is high due to manual updating and maintenance are solved.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Open-Vocabulary Object Detection Based on Frozen Vision and Language Models

An example method of training a detector head for object detection of a training object category based on a frozen vision and language model (VLM) is provided. The method includes receiving the frozen VLM pre-trained on a plurality of image-text pairs. The method includes determining, for an image embedding generated by a pre-trained image encoder of the frozen VLM and by the detector head, a detection region embedding indicative of one or more regions of interest in an image. The method includes generating, by a pre-trained text encoder of the frozen VLM, a text embedding of the training object category. The method includes predicting, by the detector head and based on the detection region embedding and the text embedding of the training object category, an object from a target object vocabulary associated with the training object category. The method includes providing the pre-trained frozen VLM and the trained detector head.
Owner:GOOGLE LLC

Dynamic scene 4D semantic map generation method and device and processing equipment

The invention provides a dynamic scene 4D semantic map generation method, a dynamic scene 4D semantic map generation device and processing equipment, and aims to realize geometric perception and semantic alignment combined processing in a single framework by designing a first feedforward framework for 4D semantic map generation. The framework comprises two core components, namely a streaming visual geometric converter for capturing space-time geometric features of a dynamic scene and a semantic bridging decoder for mapping the space-time geometric features to language aligned semantic spaces, so that the structural integrity is kept, and the semantic interpretability is improved. Different from a traditional method depending on time-consuming scene-level optimization, the method can effectively support multi-dynamic scene merging training, can be directly applied during reasoning, and is high in calculation efficiency and generalization ability. According to the design, the practicability of large-scale deployment is remarkably improved, a new thought of open vocabulary 4D scene understanding is developed, good data support can be provided for scene understanding tasks of applications such as intelligence, meta universe and digital twinning, and the method has good application prospects.
Owner:JIANGHAN UNIVERSITY

Time sequence prediction method and system based on large language model semantic prompt enhancement

The invention discloses a time sequence prediction method and system based on large language model semantic prompt enhancement. The adopted model comprises a semantic prompt enhancement module, a pre-training large language model and a self-adaptive mapping network; the semantic prompt enhancement module maps a vocabulary of the large language model into a sparse subset, and encodes the sparse subset of the vocabulary into a key matrix and a value matrix; embedding the time sequence into a vector code to form a query matrix, calculating the sparsity similarity between the query vector and a key matrix, and selecting the query vector to form a sparse query matrix; calculating sparse attention, and splicing the sparse attention with the time sequence embedded vector to obtain an enhanced time sequence embedded vector; the pre-trained large language model is used for embedding a vector according to the enhanced time sequence to obtain a time sequence preliminary prediction representation; and the adaptive mapping network adopts a B-spline function to carry out nonlinear mapping on the preliminary prediction representation of the time sequence. The modeling capability of the model on long-term dependence and complex sequence modes is enhanced, and the prediction precision is improved.
Owner:HEBEI UNIV OF TECH

Method for bidirectional translation between sign language and text using ai, deep learning, and dictionary search techniques

The present invention facilitates communication between sign language users and machines by translating sign language and text using AI models, deep learning computer vision, and word embeddings. Users interact via sign language, captured and processed through deep learning and NLP modules. The system converts sign language videos into text, constructs coherent sentences, and generates contextually appropriate responses using a Retrieve and Generate (RAG) model. Responses are translated back into sign language videos, spelling out words not found in the dictionary. If requested, a human agent can respond. Key features include high-accuracy recognition, context-aware response generation, dynamic vocabulary updates, and optional human interaction. The method ensures efficient processing with LLM, embedding techniques, and deep learning, optimizing translation accuracy and user experience. The system adapts to multiple languages and dialects by training on specific sign languages, making it applicable globally.
Owner:MAHGOUB AHMED

Zero-Hallucination Specialized Large Language Model Search

Systems and methods are disclosed herein for receiving, based on user interaction with a user interface, a user input of a natural language search query for identifying a cybersecurity threat by way of a search interface, the natural language search query requesting a specialized search of a threat database. An application generates a search vocabulary based on the natural language search query. The application performs a query lookup using the threat database, the query lookup returning a plurality of files that at least partially match the search query. The application prompts a large language model to generate an answer to the natural language search query using the plurality of files that at least partially match the search query, and outputs for display the answer using the user interface.
Owner:ANOMALI INC

Data processing method and device, equipment and storage medium

The embodiment of the invention provides a data processing method and device, equipment and a storage medium, and the method comprises the steps: predicting N first prediction probability lists corresponding to N lexical element positions through a sub-language model based on cue word information, and predicting N first prediction probability lists corresponding to N lexical element positions through a main language model based on the cue word information and the N first prediction probability lists, and predicting N second prediction probability lists corresponding to the N lexical element positions. And determining probability distribution characteristics corresponding to the prediction probabilities in the N first prediction probability lists, and adjusting the prediction probabilities in the N first prediction probability lists according to the probability distribution characteristics corresponding to the N first prediction probability lists to obtain N first adjustment probability lists. And similarly, obtaining N second adjustment probability lists based on the N second prediction probability lists. And according to the N first adjustment probability lists, the N second adjustment probability lists and the lexical element list, determining text content meeting the prompt word information. Through the method, the generation efficiency of the text content can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

System and method for combining unsupervised machine learning and information extraction models for topic modeling

A data processing system and method include receiving a set of documents from which to generate at least one topic, creating at least one information extraction model for the set of documents, executing the at least one information extraction model to extract a plurality of text segments from the set of documents by applying one or more rule sets defined for at least one entity, word list, or grammatical pattern and extracting a plurality of text portions from the set of documents based on the one or more rule sets as the plurality of text segments, inputting at least a subset of the plurality of text segments into an unsupervised machine learning model, and executing the unsupervised machine learning model to output the at least one topic for the set of the documents.
Owner:SAS INSTITUTE INC

End-cloud cooperative training method, system and device

The embodiment of the invention provides an end-cloud cooperative training method, system and device in the field of artificial intelligence, and can be used for stripping an embedding layer to terminal training in an end-cloud cooperative training process so as to improve the privacy security of end-side data. The method comprises the steps that a client uses training data as input of a representation layer to obtain a representation vector, the representation layer is used for obtaining a vector corresponding to the input data from a representation word list stored in the client, and the representation word list is stored in the client; the client sends the representation vector to the cloud platform, so that the cloud platform takes the representation vector as the input of a language model deployed at the cloud platform side to obtain an output feature; the client receives an output feature sent by the cloud platform, wherein the output feature is obtained by inputting a representation vector into a language model by the cloud platform; and the client calculates a loss value by using the output feature, updates the representation layer according to the loss value, and sends the loss value to the cloud platform, and the loss value is used for the cloud platform to update the language model.
Owner:HUAWEI TECH CO LTD

Large language model watermarking method based on triple word list division strategy

The invention relates to the technical field of large language model security, and particularly provides a large language model watermarking method based on a triple word list division strategy. The method comprises a watermark embedding step and a watermark detection step, wherein in the watermark embedding step, a vocabulary is dynamically divided into three mutually exclusive subsets of green, yellow and red based on context information at each time step; dynamically determining a gating value according to the context entropy value; applying a positive bias to the green subset vocabulary and a negative bias to the yellow subset vocabulary based on a gating value, and prohibiting selection of the red subset vocabulary, thereby adjusting an output probability distribution and embedding a watermark; in the watermark detection step, a subset corresponding to each word is reproduced, the hit rate of green and red in a sliding window is calculated, statistical significance is calculated based on Poisson-binomial distribution, and the existence of the watermark is judged through Fisher merge test. According to the method, the problems of low detection success rate, high false alarm rate, text quality reduction and insufficient robustness of the existing watermark technology are solved.
Owner:NATIONAL INTERDISCIPLINARY RESEARCH CENTER FOR ENGINEERING PHYSICS

Instant messaging-based proper noun speech recognition processing method and computer device

The invention discloses a proper noun speech recognition processing method based on instant messaging and a computer device. The method comprises the following steps: firstly, constructing annotation data sets of different scenes and a user-specific custom word list, and dynamically obtaining related data and a hot word list according to a to-be-recognized voice scene; afterwards, a training voice recognition model is subjected to fine tuning by using the annotation data set to obtain a first model, and after a recognition instruction is received, a hot word list is loaded for recognition to obtain a voice initial recognition text; and after the initial recognition text is obtained, dynamic optimization is carried out by utilizing a user-specific custom word list, and recognition errors are corrected. According to the method, matching data can be dynamically obtained, the model is combined with a scene to understand proper nouns, and the recognition difficulty caused by pronunciation and meaning complexity is reduced; in combination with targeted data and hot word information training, the proper noun recognition accuracy is improved, the defects of an existing model error correction mechanism are overcome, the accuracy of a final recognition result is remarkably improved, and powerful support is provided for speech recognition and subsequent application.
Owner:BEIJING VRV SOFTWARE CO LTD

Word segmentation labeling method and device, electronic equipment and storage medium

The invention discloses a word segmentation labeling method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining a text to be subjected to word segmentation; performing a word segmentation strategy based on a preset word list on the to-be-segmented text to obtain a first word segmentation path; performing a rule-based word segmentation strategy on the text to be subjected to word segmentation to obtain a second word segmentation path; performing a statistics-based word segmentation strategy on the text to be subjected to word segmentation to obtain a third word segmentation path; generating a plurality of candidate paths based on a preset dictionary and the text to be subjected to word segmentation; determining a plurality of word segmentation paths to be scored in the first word segmentation path, the second word segmentation path and the third word segmentation path based on the plurality of candidate paths; determining multi-dimensional scores of the plurality of word segmentation paths to be scored; according to the method, the target word segmentation path is determined in the multiple word segmentation paths to be scored based on the multi-dimensional score, and the optimal word segmentation annotation result is selected by performing score comparison on the screened word segmentation annotation results, so that the method is closer to actual requirements and semantics, and can adapt to more diversified text contents and application scenes.
Owner:郭鹏

Domain-specific retrieval language models

Various examples, systems, and methods are disclosed relating to domain-specific document retrieval that incorporates custom vocabulary integration and embedding model updates. A computing system can extract multiple segments from a collection of documents and generate queries that correspond to at least one segment. The computing system can identify terms that satisfy a uniqueness criterion and input the terms into a tokenizer to create a vocabulary dataset. The vocabulary dataset, the document segments, and the queries can be used to update an embedding model to support retrieval and semantic alignment within private documents.
Owner:NVIDIA CORP

Intelligent rail transit passenger flow volume prediction method and system based on AI large model

The invention particularly relates to a smart rail transit passenger flow volume prediction method and system based on an AI large model. According to the rail transit passenger flow volume prediction method and system, semantic features in text data are extracted through an AI large model, event influence quantification is carried out, an event influence matrix is generated, a field self-adaptive pre-training method is adopted, and customized word list construction is carried out; processing the spatio-temporal data by adopting a neural network algorithm and a GNN-LSTM hybrid network, and outputting a preliminary predicted value; a linear regression model with a forgetting factor is utilized to dynamically calibrate the preliminary predicted value, and a linear regression coefficient and a confidence interval are dynamically displayed through an interpretable output module. According to the rail transit passenger flow volume prediction method and system, the safety, expandability, flexibility and intelligence of data application in the field of intelligent rail transit are improved, and a whole set of accurate intelligent rail multi-scene prediction data is provided for the rail industry; personnel and vehicle scheduling tasks can be completed and coordinated timely and accurately in festivals and holidays or when emergency events occur.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Client-cloud collaborative training method, system and apparatus

Provided in the embodiments of the present application are a client-cloud collaborative training method, a system and an apparatus in the field of artificial intelligence, which can be used for offloading an embedding layer to a terminal for training during client-cloud collaborative training, thereby improving the privacy security of client side data. The method comprises: a client uses training data as an input of a representation layer, so as to obtain a representation vector, the representation layer being used for obtaining from a representation word list stored in the client a vector corresponding to the input data, and the representation word list being stored in the client; then, the client sends the representation vector to a cloud platform, such that the cloud platform uses the representation vector as an input of a language model deployed at the cloud platform side, so as to obtain an output feature; the client receives the output feature sent by the cloud platform, the output feature being obtained by inputting the representation vector into the language model by the cloud platform; and the client uses the output feature to calculate a loss value, updates the representation layer on the basis of the loss value, and sends the loss value to the cloud platform, the loss value being used for the cloud platform to update the language model.
Owner:HUAWEI TECH CO LTD

RAG financial credit decision-making method based on policy mask constraint decoding

The invention belongs to the technical field of financial science and technology, and provides an RAG financial credit decision-making method based on policy mask constraint decoding. The method comprises the steps of obtaining a structured financial credit granting policy codebook, screening credit granting parameter values which can be coded into single lexical units by a pre-training word segmentation device, and constructing an effective policy lexical unit ID set; adding an original logits vector and a mask vector during decoding through the mask vector matched with the vocabulary of the large language model, and forcibly constraining an output lexical element in an effective set to realize one-time decoding compliance; and meanwhile, a dynamic policy updating mechanism, a parameter type sub-mask mechanism, a semantic offset error recovery mechanism and a policy codebook structured verification mechanism are additionally arranged. The method guarantees the compliance of the credit decision text from the source, improves the decision efficiency, shortens the policy adjustment response time, and reduces the decision interruption rate.
Owner:SHENZHEN MAGIC NUMBER INTELLIGENT ARTIFICIAL INTELLIGENCE CO LTD

Watermark detection method and device based on Bayesian detection

The invention provides a watermark detection method and device based on Bayesian detection. The method comprises the steps of obtaining a target text to be subjected to large model watermark detection and a cue word corresponding to the target text; processing a token sequence spliced by the cue word and the target text through a language model to obtain probability distribution output by the language model for each token position; for the jth token in the T tokens of the target text, selecting the first k tokens before the jth token to be input into the hash function, and obtaining a random seed of the jth token; dividing a word list of the large model into a preferential selection set and a non-preferential selection set based on random seeds; on the basis of preset watermark bias, preferentially selecting the set and the probability distribution, and generating probability distribution after disturbance of the jth token; calculating a log-likelihood ratio by using the probability distribution before and after disturbance, and accumulating the log-likelihood ratio to a detection score of the current text; and if the detection score is higher than the threshold value, judging that the target text has the watermark of the large model.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Time series prediction method based on multi-modal enhanced large language model

The invention discloses a time series prediction method based on a multi-modal enhanced large language model, which comprises the following steps: acquiring historical time series data, constructing a standardized input matrix and setting core task parameters; dividing data blocks through a sliding window mechanism, and converting the data blocks into uniform dimension features through linear embedding; a semantic prototype is generated based on a large language model vocabulary, and time sequence features and semantic features are fused through multi-head cross-attention; alternately splicing the data blocks and the corresponding semantic information, and constructing a self-multi-modal input sequence; designing a three-level structured prompt including context, task target and modal guidance, and fusing the three-level structured prompt with a multi-modal sequence; and training the lightweight model by adopting a frozen training strategy, and outputting a prediction result of a specified time step in the future. According to the method, through single-source data enhancement and prompt guidance, the time sequence reasoning capability of a small-parameter large language model is activated, high-precision prediction is guaranteed, efficient deployment is achieved, and the method adapts to long-term and short-term prediction tasks in the fields of electric power, traffic, meteorology and the like.
Owner:HANGZHOU DIANZI UNIV

A method and system for constructing a small language voice recognition based on a whisper token

The application provides a speech recognition method and system for constructing a small language based on a whisper token, and relates to the technical field of natural language processing and speech recognition. The method comprises the following steps: extracting all tokens related to a target small language in a whisper tokenizer to form an initial candidate set; matching and analyzing the tokens in the initial candidate set with collected training text corpus of the target small language, and counting the frequency of the tokens in the corpus; and screening high-frequency tokens and supplementing low-frequency tokens according to the frequency counting result to construct a dynamic vocabulary. The application improves the vocabulary quality, optimizes the model training efficiency, enhances the speech recognition accuracy, improves the model generalization ability, and simplifies the model construction process, thereby providing an efficient, accurate and easy-to-implement solution for the field of small language speech recognition.
Owner:BEIJING RUI KELUN INTELLIGENT TECH CO LTD

Display apparatus and voice control method

The invention provides a display device and a voice control method. The display device comprises a display, a sound collector and a controller connected with the display and the sound collector. Wherein the display is configured to display an image picture and a user interface; the sound collector is configured to collect a voice control instruction of a user, and the controller is configured to determine at least one scrolling control contained in a current display interface; constructing a voice scrolling control word list of a scrolling control in the current display interface; the voice scrolling control word list is used for representing a corresponding relationship between the scrolling direction of the scrolling control and the semantic control word; and in response to a voice control instruction of a user, controlling a target scrolling control in the current display interface to execute scrolling operation based on the voice scrolling control word list and a voice control text corresponding to the voice control instruction. Therefore, the voice control of the scroll control is realized, the flexibility and convenience of the control mode of the display equipment are improved, and the user experience is improved.
Owner:HISENSE VISUAL TECH CO LTD

Tongue diagnosis information intelligent acquisition method and device based on cloud platform, and medium

The invention discloses a tongue diagnosis information intelligent collection method and device based on a cloud platform and a medium, and relates to the technical field of information collection, and the method comprises the steps: constructing a tongue diagnosis situation formula library, registering tongue diagnosis collection resources and business labels of all medical institutions, and forming a tongue diagnosis business knowledge base; based on the tongue diagnosis service knowledge base, receiving a tongue diagnosis service request and matching a tongue diagnosis situation formula, and generating a tongue diagnosis collection situation instance and a corresponding tongue diagnosis collection script; and based on the tongue diagnosis acquisition session packet, the corresponding tongue diagnosis situation formula and the tongue diagnosis acquisition situation instance, generating a tongue diagnosis acquisition business file. According to the method, the structured bidirectional association between the tongue diagnosis acquisition actions and the tongue diagnosis acquisition resources is realized by constructing the tongue diagnosis situation formula library, establishing the tongue diagnosis acquisition resource capability word list and the tongue diagnosis situation formula resource demand word list, calculating the scene adaptation level and forming the tongue diagnosis business knowledge base.
Owner:DONGGUAN SANYAN BIOTECHNOLOGY DEV CO LTD

Method and device for improving word learning efficiency of middle and primary school students

A method and device for improving word learning efficiency of middle and primary school students relates to the technical field of recommendation learning, and comprises the following steps: S1, defining feature vectors of words and a word list of learning materials; s2, establishing feature vectors of all words in the English learning textbook, initializing a word recommendation value dictionary and performing learning material storage; s3, dynamically updating the word feature vectors and the recommendation value dictionary according to the expressions of the students in different learning types in the learning process; s4, screening out alternative learning materials, calculating recommendation values of the alternative learning materials through a word dictionary and a word recommendation value dictionary of the alternative learning materials, and pushing the learning material with the highest recommendation value to the student; the problem that traditional word learning lacks personalized dynamic recommendation and is difficult to adapt to multi-modal learning scenes and individual memory attenuation laws is solved, dynamic feature engineering is combined with multi-scene recommendation, visual closed-loop optimization of long-term learning effects is realized, and word learning efficiency is effectively improved.
Owner:读书郎教育科技有限公司

Method and device for detecting privacy disclosure of applet, electronic equipment and storage medium

The invention discloses a privacy leakage detection method and device for an applet, electronic equipment and a storage medium, the method and device are applied to the electronic equipment and are used for carrying out privacy leakage detection on the applet, specifically, an API in the applet is recognized according to a developer document, the API involving user personal data reading serves as a taint source, and the taint source is used for detecting privacy leakage of the applet. An API related to user sensitive information transmission is used as a taint convergence position; performing static taint analysis on the program code of the to-be-analyzed applet to obtain a behavior set of the applet for processing the user sensitive information; constructing a cue word list of the applet based on the behavior set, wherein cue words comprise a data type and an operation type of data accessed by the applet; and based on a chain thinking prompt strategy and a multi-query mechanism, performing compliance judgment on the prompt words in the prompt word list by utilizing a large language model in combination with an applet privacy policy to obtain judgment results, and forming a detection report based on all the judgment results. Whether the privacy data use behavior of the applet is consistent with the privacy policy of the applet can be confirmed through the detection report, so that whether the applet has privacy leakage or not is confirmed.
Owner:广州市公安局网络安全保卫支队

System and method for combining unsupervised machine learning and information extraction models for topic modeling

A data processing system and method include receiving a set of documents from which to generate at least one topic, creating at least one information extraction model for the set of documents, executing the at least one information extraction model to extract a plurality of text segments from the set of documents by applying one or more rule sets defined for at least one entity, word list, or grammatical pattern and extracting a plurality of text portions from the set of documents based on the one or more rule sets as the plurality of text segments, inputting at least a subset of the plurality of text segments into an unsupervised machine learning model, and executing the unsupervised machine learning model to output the at least one topic for the set of the documents.
Owner:SAS INSTITUTE INC