Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

464 results about "Confusion" patented technology

A feeling that you can't think clearly, focus or make decisions.

Remote education data processing system

The invention relates to a remote education data processing system which comprises the following steps: under a remote teaching task, pre-defining a task intention and a data expectation point; a semantic timestamp and a task binding label are printed on each data fragment; mapping the collected confusion data including silence, eye movement drift and prediction into a unified learning semantic vector; a micro-expression + interactive behavior + time sequence decision path ternary modeling mode is introduced, and a potential cognitive intention corresponding to the feature combination is recognized; teaching context information is fused; constructing a cognitive state mapping model; reasoning a current cognitive state label from multi-modal sensing data; searching intervention track VS effect feedback data in a historical database; generating a predicted intervention behavior sequence by using a sequence modeling algorithm; a dynamic combination suggestion chain including light prompt, content reconstruction, personalized practice and tutoring invitation is adopted; and superposing the cognitive state sequences of all students into a group cognitive trajectory map.
Owner:SHENZHEN ZHONGJING EDUCATION TECH CO LTD

Complex task-based high-quality pseudo-annotation data set construction method

The invention discloses a high-quality pseudo-annotation data set construction method based on a complex task, and relates to the technical field of multi-modal learning, and the method comprises the steps: constructing a cross-modal causal graph based on multi-modal original data, loading a domain knowledge graph, recognizing an inter-modal confusion variable, and generating an initial pseudo-tag; an anti-fact sample is generated by forcibly cutting off a non-causal path in the cross-modal causal graph, and a cross-modal depolarization pseudo-label is generated by comparing the pseudo-label difference between the original sample and the anti-fact sample; and in combination with the cross-modal depolarization pseudo-labels and the semantic consistency pseudo-labels, a standardized pseudo-annotation data set with multi-modal alignment, clear entity relationship and semantic consistency is generated through fusion. According to the method, an anti-fact intervention framework is adopted, non-causal path influence in cross-modal interaction is identified and eliminated by analyzing probability distribution difference, and false association is effectively inhibited.
Owner:CHINA NAT INST OF STANDARDIZATION

Zip-MoE model grouping mixed expert layer-based Chinese and English speech recognition method and system

The invention provides a Zip-MoE model grouping hybrid expert layer-based Chinese and English speech recognition method and system, a Zip-MoE model comprises six encoder blocks, a Bypass module is included between every two encoder blocks, and the weights of the output of the previous encoder block and the output of the current encoder block are learned; the first three encoder blocks are of a standard Zipform structure; the last three encoder blocks adopt a Zipform-MoE structure containing a grouped hybrid expert layer, and the grouped hybrid expert layer is used for replacing the last feed-forward network of the Zipform structure; the grouping mixed expert layer comprises a Chinese expert group, an English expert group and a language router, and each expert group is composed of a plurality of expert networks and is provided with an unsupervised router. The problem of language confusion is relieved, different time delay streaming scenes can be adapted, the number of experts is flexibly expanded, pre-training is not needed, and the Chinese and English recognition efficiency is greatly improved.
Owner:XIAMEN UNIV

Active defense system and method based on multi-protocol dynamic simulation and distributed trapping

The invention provides an active defense system and method based on multi-protocol dynamic simulation and distributed trapping. The active defense method based on multi-protocol dynamic simulation and distributed trapping comprises the following sub-steps: S1, constructing a multi-protocol dynamic simulation environment; s2, deploying distributed trapping nodes; s3, deep trapping of attack behaviors; s4, attack chain reconstruction and behavior analysis; s5, performing adaptive confusion and adversarial enhancement; s6, automatic threat intelligence production and feedback; by loading the protocol template library and initializing the state machine, the response can be dynamically generated according to the real-time session context, and dynamic simulation of various service protocols is adopted, so that the detection capability on network attacks is improved, potential threats can be captured more quickly, and the risks of missing report and false report are reduced; and through an automatic threat intelligence generation and feedback mechanism, in combination with IOC index identification, structured output and real-time response, a defense strategy can be quickly responded and adjusted.
Owner:CHINA LIFE INSURANCE CO LTD

Test case generation method and system based on multi-agent efficient collaboration

The invention relates to a test case generation method and system based on multi-agent efficient collaboration, belongs to the technical field of software testing, and solves the problems of incomplete scene coverage, logic disorder and the like when a single large model processes a complex task. The method comprises the following steps: constructing a multi-modal professional field knowledge base; the task planning agent module generates a test case generation task based on a project development document and a software source code file of a to-be-tested project and distributes the test case generation task to the test demand analysis agent module; based on the test case generation task, extracting a test demand point related to the tested software configuration item, describing the test demand point, and labeling a corresponding tracking relationship between a test demand point ID and a code snippet or a function in the software source code file to construct a test demand-code snippet / function set; and generating a test case and a test description document based on the test demand-code snippet / function set and the multi-modal professional domain knowledge base. And the ability of generating the test case of the large model is enhanced through organic and efficient cooperation of multiple agents.
Owner:BEIJING JINGHANG COMPUTING & COMM RES INST

Business English scene knowledge graph generation and intelligent retrieval method

The invention discloses a business English scene knowledge graph generation and intelligent retrieval method, and relates to the technical field of natural language processing, and the method comprises the steps: constructing a document time sequence index and anaphora alignment mapping table, and registering an evidence anchor point index; extracting event units under the constraint and merging the event units into event clusters in a cross-document manner; arranging an event chain according to unified time and generating a process graph, and binding evidence anchor points to relation edges; analyzing the query into a query intention expression and a path constraint, and generating an evidence presentation sub-plan and an image-text combined execution plan; and positioning answers in the atlas, integrally returning the answers, paths and evidences, and meanwhile, forming replayable audit tracks. The method solves the problems of cross-channel time dislocation and confusion of references, and supports interpretable and auditable retrieval. The method is suitable for scenes such as mails, conference summary, instant communication and the like, guarantees locatable sources, path replay and evidence verification, and is convenient for internal control review.
Owner:LANZHOU INST OF TECH

Large model dynamic knowledge distillation method and system

The embodiment of the invention provides a large-model dynamic knowledge distillation method and system, and belongs to the field of large-model dynamic knowledge distillation. The method comprises the following steps: acquiring a question-answer pair data set, and encoding a training input question-answer pair by using a student model to generate a first soft label; performing distillation, and using the student model again to encode the training input question and answer pair to generate a second soft label; calculating a predicted uncertainty value and a distinguishing capability of the student model; dynamically adjusting the distillation intensity weight and the distillation temperature value; and iteratively executing the process of training the student model until the performance index of the student model reaches a preset threshold, and completing dynamic knowledge distillation. According to the method, by dynamically adjusting the distillation parameters, the model output stability is improved, the noise of fuzzy samples is reduced, the class boundary confusion rate in class task interpretation is reduced, and the generalization ability of complex semantics is improved.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Method and system for constructing hydrological survey knowledge model based on large model

The invention relates to the technical field of hydrological survey skill training, in particular to a construction method and system of a hydrological survey knowledge model based on a large model. The method comprises the following steps: acquiring multi-source hydrological data, and constructing a hydrological knowledge graph; generating a confusion sample corresponding to the hydrological knowledge graph, and performing adversarial training on the confusion sample; according to the confrontation training result and historical test questions in the multi-source hydrological data, obtaining bidirectional mapping test questions; inputting the bidirectional mapping test questions into a question setting interface of a hydrological survey learning platform; collecting answer data and student physiological data corresponding to the bidirectional mapping test questions in real time, and updating a student knowledge state according to the answer data and the student physiological data; constructing a hydrological survey knowledge model based on the updated student knowledge state; and generating a student recommendation learning scheme by using the hydrological survey knowledge model. According to the invention, the intellectualization and practicability of hydrological survey education can be improved.
Owner:BUREAU OF HYDROLOGY CHANGJIANG WATER RESOURCES COMMISSION

Code generation method and device based on large model, medium and equipment

The embodiment of the invention discloses a code generation method based on a large model, and the method comprises the steps: carrying out the grammatical analysis of a business logic process described by a natural language, and obtaining atomic task units and a logic relation between the atomic task units; and then re-editing according to the logical relationship to obtain a secondary expression text, and generating a code through LLM. And the problem of subsequent code generation through LLM due to logic relation chaos caused by natural language description for a business logic process is avoided. In the syntactic analysis process, the simplest expression corresponding to the natural language can be determined, ambiguity of description of steps to be executed in the business logic process is reduced, secondary expression texts are obtained through re-editing, the influence of logic chaos possibly occurring in the business logic process can be avoided, and the business logic processing efficiency is improved. Therefore, the normalization of the input LLM text is not affected regardless of the writing degree of the business logic process, and the accuracy and reliability of code generation are improved.
Owner:ZHEJIANG ANT MISUAN TECHNOLOGY CO LTD

Automatic address governance method and device based on intelligent agent and medium

The invention relates to the field of geographic information systems, and discloses an agent-based automatic address governance method and device and a medium, and the method comprises the steps: obtaining a to-be-governed address through a data perception and preprocessing agent; inputting the address to be governed into a master control agent; the main control agent analyzes the address to be governed and gives an analysis result; according to the analysis result, selecting a tool execution agent to process the address to be treated; and outputting the governed address, and storing all data by adopting a memory and interaction agent. According to the method, the analysis accuracy, the semantic understanding depth and the robustness of the LLM in processing complex and non-standard addresses containing confusion information are greatly improved, meanwhile, the requirement for manual intervention is greatly reduced, the efficiency of processing mass address data is remarkably improved, and the real-time response speed is remarkably increased.
Owner:WUDA GEOINFORMATICS CO LTD

Phishing document de-obfuscation and feature extraction method and application thereof in attack detection

The invention discloses a phishing document de-obfuscation and feature extraction method and an application thereof in attack detection. The de-obfuscation comprises the steps of obtaining an obfuscation macro code of a phishing document, constructing a prompt engineering template by utilizing a pre-training language model, analyzing an obfuscation logic structure and generating a de-obfuscation rule and a reduction strategy; the method comprises the following steps: structuring a confused macro code into an abstract syntax tree through an analysis tool, matching a typical confusion mode based on a regular expression, and performing simulation and cell reference analysis in combination with a function to realize structure preliminary reduction, control flow semantic reduction and operation path construction; according to the unmixing rule and the abstract syntax tree, the confusion structure is converted into a readable macro statement, a macro code instruction sequence with confusion semantics removed is generated, and a semantic sequence obtained after unmixing is output. The feature extraction comprises word feature extraction, Token feature extraction, abstract syntax tree feature extraction and relation feature extraction. According to the method, the bottleneck that traditional phishing attack detection and confusion documents are difficult to recognize is broken through, and the detection accuracy of phishing document attacks is improved.
Owner:GUIZHOU UNIV

Text segmentation method and related equipment

PendingCN121328561AMathematical modelsSemantic analysisSemantic changeSemantic variation
The invention provides a text segmentation method and related equipment. The method comprises the steps of obtaining a to-be-processed text; segmenting the to-be-processed text into ordered statement sequences to obtain an initial statement set of the to-be-processed text; wherein the ordered statement sequence comprises a plurality of statements; the semantic variation, the confusion degree variation and the information entropy variation of a first target text block are calculated when a to-be-decided statement in an initial statement set of the to-be-processed text is added into the first target text block, and the first target text block is a set of multiple statements meeting a merging condition; determining the collaboration degree of the statement to be decided and the first target text block based on the semantic variable quantity, the confusion variable quantity and the information entropy variable quantity; and partitioning the to-be-processed text based on the collaboration degree of the to-be-decided statement and the first target text block to obtain a target partitioned text of the to-be-processed text. The text segmentation quality can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Multi-modal data desensitization method, system, equipment and medium

The invention provides a multi-modal data desensitization method, system and device and a medium, and belongs to the technical field of information security. The method comprises the following steps: firstly, acquiring multi-modal original data based on preset access information and acquisition frequency, and after standardized preprocessing, positioning sensitive information by using a classification recognition algorithm and generating an analysis report. Then, according to a report, matching a strategy from a desensitization strategy library bound with the sensitivity level, adjusting a dynamic confusion algorithm parameter to desensitize, integrating and generating desensitized multi-modal data, and recording and storing a mapping relationship between the original data and the desensitized data; and if a restoration request is received, after authorization verification is passed, reverse desensitization restoration data is carried out by using the mapping relation. Besides, full-process feedback information is collected, the desensitization effect is quantitatively evaluated according to a preset index, the desensitization strategy and algorithm are iteratively optimized accordingly, and effective desensitization and safety management of multi-modal data are achieved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Legal decision prediction method and device based on large language model and logic enhancement

PendingCN120373535AForecastingBiological modelsFirst-order logicLinguistic model
The invention provides a law decision prediction method and device based on a large language model and logic enhancement, which can realize adaptive adjustment of law decision rules and improve the accuracy of decision prediction, and comprises the following steps: generating decision rules in the form of a first-order logic language by adopting an LLM model in combination with a law article original text and a historical case, describing key elements in legal provisions and historical cases through first-order logic symbols to form a consequent of a judgment rule, and taking a judgment label as a consequent of the judgment rule; calculating statement similarity of the cases based on a BGE vector model, and screening historical cases with high similarity to construct a confusion-prone case verification set; based on the confusion-prone case verification set, using a confusion perception comparative learning method to optimize a judgment rule; and for a to-be-predicted target case, generating a candidate tag of the target case based on a pre-trained domain small model, judging whether the target case accords with the candidate tag through the LLM model in combination with the thinking chain and the optimized judgment rule, and outputting a predicted judgment result.
Owner:NAT UNIV OF DEFENSE TECH

Method and system for realizing role playing dimension

The invention discloses a role playing dimension implementation method and system, and relates to the technical field of data processing. Comprising the steps that a meta-model of a dimension model related to a role playing dimension is created, the meta-model comprises five types of data tables, a role playing dimension name is selected from a role playing dimension name source according to requirements, when the role playing dimension name is different from a physical dimension table name, it is considered that the dimension corresponding to the role playing dimension name is the role playing dimension, and the role playing dimension is selected from a physical dimension table name; when the related dimension of one index is displayed, the role playing dimension name is used for replacing the physical dimension name to display the actual business meaning of the dimension, when the dimension is selected to summarize the indexes, the role playing dimension name is used for constructing a dynamic SQL, and when the composite index is analyzed, the actual business meaning of the dimension is displayed. When multiple indexes are analyzed in parallel, if the facts where the polyatomic indexes forming the composite indexes are located are all associated with the same physical dimension, the composite indexes are summarized through the physical dimension, and when the multiple indexes are analyzed in parallel, if the facts where all the indexes participating in analysis are located are all associated with the same physical dimension, the multiple indexes are summarized through the physical dimension; and the associated dimension list is displayed by using the physical dimension name, so that confusion caused by different role names is avoided.
Owner:INSPUR SOFTWARE TECH CO LTD

Mathematical classroom real-time participation degree and cognitive state intelligent perception analysis system

The invention discloses an intelligent perception analysis system for the real-time participation degree and cognitive state of a mathematics classroom. The intelligent perception analysis system comprises a multi-source data acquisition module, a fusion analysis engine module, a real-time feedback module and an offline optimization module. The method has the beneficial effects that a dynamic participation index is innovatively designed, and an attention attenuation resetting mechanism is introduced; the real-time feedback module generates a cognitive thermodynamic diagram of HSL color mapping to position group obstacle points, and triggers self-adaptive question pushing; and the off-line optimization module dynamically updates the edge weight of the knowledge graph through error co-occurrence analysis, and early warns a cognitive confusion relationship without textbook association. According to the method, the limitation of single-mode perception is broken through, the problem solving step quality visual diagnosis and the cross-cycle cognitive impairment prediction are realized, a teaching closed loop of'multi-source perception-hierarchical quantification-real-time intervention-knowledge evolution 'is formed, and the mathematical classroom cognitive state analysis precision and the teaching intervention timeliness are remarkably improved.
Owner:SHIHEZI UNIVERSITY

AI intelligent education tutoring system based on thinking chain and retrieval enhancement generation technology

The invention relates to the technical field of artificial intelligence education, and discloses an AI intelligent education tutoring system based on a thinking chain and a retrieval enhancement generation technology. Transparent generation and dynamic verification of a problem solving logic chain are realized through a CoT-RAG fusion architecture, and the interdisciplinary confusion rate is reduced in combination with a knowledge graph technology of subject boundary isolation; the non-question-bank-dependent generation engine creates questions in real time based on semantic association, and the coverage rate is increased; federated learning and differential privacy technologies are integrated, and the safety of user data is ensured while the generalization ability of the model is improved; the learning behavior analysis unit dynamically optimizes a knowledge recommendation path through time sequence modeling, and cooperates with the interactive logic correction module to form a self-evolution closed loop. Finally, an intelligent education system which is high in interpretability, high in professional precision and compliant in privacy is constructed, and the learning efficiency and the interdisciplinary problem solving capability are remarkably improved.
Owner:HUAZHONG NORMAL UNIV

Online course learning management method based on knowledge graph

The invention relates to the technical field of online education, and discloses an online course learning management method based on a knowledge graph. The method comprises the following steps: acquiring multi-modal learning behavior data of a learner, and extracting a deep learning state vector reflecting knowledge understanding depth, learning input degree and cognitive confusion through semantic fusion; and dynamically calculating and updating the logical relationship strength among the knowledge points in the course knowledge graph by using the vector, so that the knowledge structure can adaptively evolve along with the actual cognitive state of the learning group. And generating a real-time personalized learning path based on the updated knowledge graph and the current state vector of the learner. Meanwhile, according to cognitive confusion features in the state vector, intervention measures such as pushing of remedial resources, adjusting of content sequence or starting of self-adaptive testing are triggered in real time. According to the method, the dynamic optimization of the knowledge graph and the accurate and immediate response of learning intervention are realized, and the adaptability and management efficiency of online learning are improved.
Owner:SHENYANG UNIV

A Method for Enhancing NL2SQL Questions Driven by Multivariate Knowledge Linkage

The present invention belongs to the technical field of power natural language data question answering, and specifically relates to a method for enhancing NL2SQL questions driven by multi-source knowledge links. The method includes: constructing a power data schema using database table information and sorting out power domain knowledge; constructing a question parsing Prompt template and using a large language model to analyze the structure of the original question to extract key entities from the original question; retrieving power domain knowledge in a hybrid similarity retrieval manner based on the sorted out power domain knowledge and the key entities extracted from the original question; obtaining database tables and data schemas related to the original question through a multi-level schema linking method; standardizing knowledge and designing a question enhancement Prompt template based on the retrieved power domain knowledge and the obtained data schema, and using a large language model to reconstruct and enhance the original question to eliminate confusion and interference factors and improve the accuracy of question answering.
Owner:YANTAI HAIYI SOFTWARE

Government affair application platform based on artificial intelligence model

The invention relates to the technical field of government affair intelligence, in particular to a government affair application platform based on an artificial intelligence model, and the platform comprises a semantic recognition module, a tag affiliation module, a rule screening module, a path generation module and a path verification module. According to the method, the subject-called structure for recognizing noun phrases and approval verbs in the government affair approval text is adopted, semantic boundaries are defined, the structure recognition accuracy is enhanced, keyword and field item comparison adopts a field-level consistency mode, the tag affiliation precision is improved, field value comparison logic is introduced in rule node screening, and the accuracy of tag affiliation is improved. In path generation, a path structure with clear logic is established through sorting index and field dependence identification, field conflicts and path chaos are avoided, field value consistency verification ensures that path content is highly matched with an approval text, and it is guaranteed that the path is effective and available; automatic processing from text analysis and field affiliation to path verification is achieved, and the efficiency of the government affair approval process is improved.
Owner:SHENZHEN YUNHENG INTELLIGENT CO LTD

Digital human student driving method based on large language model agent

The invention discloses a digital human student driving method based on a large language model agent, and the method specifically comprises the steps: initializing digital human student configuration information, and loading a knowledge system and background setting; during teaching, voice of a teacher is converted into characters in real time, and accurate transmission of information is ensured; teaching content is accurately analyzed by means of an intention recognition module, and teaching intentions and stages are recognized, so that digital human students can respond in time; during question asking, the digital person student feeds back according to knowledge reserve and state, such as hand raising or doubt; through short-time memory updating and dynamic knowledge base adjustment, continuous learning adapts to the teaching progress, and diversified learning behaviors and emotional responses are shown. According to the digital human student driving method, the teaching process is divided into a plurality of stages, and input and output of stage intelligent agents are subjected to formatted constraint, so that digital human students can simulate real classroom interaction and show different learning behaviors, emotional responses and classroom performance, and the method is used for teaching simulation practice of normal teachers.
Owner:EAST CHINA NORMAL UNIV

Empirical knowledge graph construction and question answering method based on heuristic self-question answering

The invention provides an experience knowledge graph construction and question answering method based on heuristic self-question answering, and the method comprises the steps: randomly dividing a text corpus into small batches, inputting the small batches into a large language model, and obtaining a heuristic rule set through induction processing and quality evaluation; constructing question and answer pairs by using the heuristic rule set and the text content cues, converting the question and answer pairs, and then carrying out confidence calculation to obtain a structured experience triple with confidence; constructing a knowledge graph by using the structured experience triad with confidence; extracting the core fragment of the original narrative query by using the heuristic rule set and generating a refined query text; obtaining an empirical path by using the embedded vector of the refined query text and the knowledge graph; and inputting the original narrative query and the experience path into the large language model to generate a final answer. According to the method, logic confusion and content contradiction possibly caused by a traditional RAG method are effectively avoided, the generated answer is more reliable, and the reasoning process is more interpretable.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Weak supervision full-view digital slice classification method and device based on causal intervention

The invention discloses a weak supervision full-view digital slice classification method and device based on causal intervention, and the method comprises the steps: segmenting a full-view digital slice (WSI) into a plurality of image blocks, forming a package containing a plurality of instances, and obtaining a package level label of the package; performing feature extraction on each image block by using a feature extractor to obtain instance features; processing the instance features by adopting a multi-instance learning (MIL) model to obtain packet level features; constructing a confusion set, and representing confusion features frequently appearing in the core mode in the training process through the confusion set; performing causal intervention on the packet level features based on the confusion set to obtain intervened packet level features; and performing classification based on the intervened packet level features to obtain a classification result of the WSI. According to the method, the problems of inaccurate classification and poor robustness caused by attention deviation and false correlation interference in traditional weak supervision full-view digital slice classification are solved.
Owner:HUNAN UNIV

Method and device for multi-party joint fine tuning of language model based on data protection

A multi-party joint fine-tuning language model method and device based on data protection, multiple parties comprise a first party holding a target language model and a second party providing computing power, the target language model comprises an embedding table and a plurality of serially arranged network layers, and the first party determines a plurality of confusion layers and a confusion embedding table; at least sending the multi-layer confusion layer to a second party; the confusion embedding table is obtained by performing second confusion on the embedding table, the confusion layer is obtained by performing first confusion on a parameter matrix in the corresponding network layer, and the first confusion mode corresponds to the second confusion mode; the second party obtains a training sample comprising a target confusion word embedding sequence and a label thereof, wherein the training sample is determined based on the text and the confusion embedding table; based on the target confusion word embedding sequence, obtaining a first output result through multiple confusion layers; and based on the difference between the first output result and the tag, adjusting a multi-layer confusion layer to realize a multi-party joint fine tuning model on the premise of protecting model parameters and data privacy.
Owner:ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD

Reverse confusion resisting method and system for deep learning model of end-side equipment

The invention relates to an anti-reverse confusion method and system for an end-side equipment deep learning model, and the core process comprises the steps: firstly inputting an original model obtained through the training of a model training frame into a deep learning compiling frame, and extracting three types of key information, namely, a model operator, a topological structure, parameters and dimensions, through the characteristics of a compiler; then constructing a feature analysis module to evaluate model features, dynamically matching a confusion scheme from a strategy library, and balancing safety and performance; the confusion module is embedded into a plurality of different levels such as a computational graph level, an operator template level and tensor intermediate expression through a hierarchical compiling mechanism, and a complex scheme can be jointly implemented across multiple levels; and finally, the compiler synchronously completes confusion reinforcement when generating the target code. According to the method, hardware adaptation is not needed, low-overhead confusion is achieved through a native pass mechanism of a compiler, fine-grained customized protection is supported, a model structure, parameters and computational logic can be effectively hidden, reverse engineering attacks can be resisted, and the method is particularly suitable for end-side equipment scenes with limited computing power.
Owner:WUHAN UNIV

LLMs pre-training data set optimization method and device and storage medium

The invention discloses an LLMs pre-training dataset optimization method, which comprises the following steps of: selecting internal data of a dataset, and performing hidden Markov model confusion degree calculation on a text segment by segment by adopting a sliding window; and based on a hidden Markov model confusion degree calculation result, obtaining a comprehensive confusion degree score of the whole text through confusion degree weighted average calculation, and using the comprehensive confusion degree score to screen sentences with semantic chaos in the data set. According to the method, a sliding window technology and a weighted average strategy are introduced on the basis of traditional puzzle calculation, so that a comprehensive quality evaluation index can be provided for the whole text, the language model adaptation degree of the text can be evaluated more accurately, and compared with single puzzle calculation, the quality condition of the text can be reflected more comprehensively; and meanwhile, sentences with disordered semantics in the data set can be effectively screened out, a network low-quality corpus text obtained for free is converted into a high-quality and valuable corpus, the LLMs pre-training cost is effectively saved, and the model capability is improved.
Owner:XIAOVO TECH

Semantic matching-based field name standard dropping mark mapping method and semantic matching-based field name standard dropping mark mapping system

The invention discloses a semantic matching-based field name standard abbreviation mapping method and system, and particularly relates to the technical field of data governance, a source field set to be subjected to abbreviation and a standard field set in a standard field library are acquired, field names and field descriptions are normalized, and field text representation is generated; using the pre-training language model to generate a semantic vector and recalling a candidate mapping pair set; calculating character string similarity, synonym expansion matching score, data type consistency score and sample value pattern consistency score for the candidate mapping pairs; generating a neighbor confusion risk index based on semantic neighbor density, entity category distribution deviation, synonym expansion and context anti-verification, performing risk gating correction, and judging a score in combination with a difficult case; and applying table-level consistency constraints in the same service table or entity domain to carry out joint consistency solution to obtain a table-level optimal mapping scheme, generating confirmation confidence, and outputting an automatic confirmation or manual recheck mapping result.
Owner:SHANGHAI INTERNET SOFTWARE

Course dynamic optimization method and system, electronic equipment and storage medium

The invention provides a course dynamic optimization method and system, an electronic device and a storage medium, and relates to the technical field of online education, and the method achieves the quantitative association of a course unit-knowledge fragment through a soft mapping matrix, enables the perplexity to be bidirectionally transmitted between the course unit and the knowledge fragment, and achieves the quantitative traceability. The comprehensive confusion degree is obtained by fusing the confusion degrees of the course unit layer and the knowledge fragment layer, so that the high-confusion course unit can be positioned more accurately and explainably, and fine optimization of the course can be realized; through a course optimization decision with cost constraint, automatic generation of a course optimization scheme with controllable cost is realized; besides, personalized learning path optimization of a course structure layer is realized, and a course optimization strategy continuously acts on subsequent learners, so that dynamic evolution of course contents and personalized learning paths is realized.
Owner:BEISEN CLOUD COMPUTING CO LTD

Class increment image classification method and system based on multi-modal pre-training model

The invention discloses a class increment image classification method and system based on a multi-modal pre-training model, and the method comprises the steps: obtaining a class increment learning data set, and for each task, model training comprises two stages: an intra-task training stage and a cross-task fine tuning stage; in-task training stage: for the data of the current task, performing fine adjustment on the pre-trained visual language model by adopting a task-specific adapter to realize classification among categories in the task; a cross-task fine tuning stage: introducing a mapping module for image feature expression, mapping features of a specific feature space of a task to a feature space shared by the task, and realizing cross-task category separability; during reasoning, a reasoning strategy based on prediction uncertainty is adopted for image classification. According to the method, the problem of category confusion existing across tasks can be solved, and the accuracy of selecting the output features is improved.
Owner:SUN YAT SEN UNIV

Multi-channel APP automatic packaging system

The invention relates to the technical field of mobile application continuous integration automation, and discloses a multi-channel APP automatic packaging system which comprises the steps that byte code files in a released mobile terminal APP installation package are acquired, and an occupied symbol space set is generated; generating an available confusion namespace by using set complementary operation; extracting method-level code change by using an abstract syntax tree difference algorithm, and generating a change method set; generating a patch special confusion rule for the change method set by using a namespace allocation algorithm; compiling and obfuscating the change method set by applying a patch special obfuscating rule to generate a safe obfuscating patch byte code; performing packaging and signature operation on the safe obfuscation patch byte code, and outputting a minimized obfuscation safe hot repair patch pack; according to the method, the technical problems that class definition conflicts or method calling errors occur after the hot repair patch is loaded, and the downloading success rate of a user is affected due to the overlarge size of the hot repair patch are solved.
Owner:ANHUI DUOXIAO INFORMATION TECH CO LTD