Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

154 results about "Text query" patented technology

Text Query Applications. The purpose of a text query application is to enable users to find text that contains one or more search terms. The text is usually a collection of documents. A good application can index and search common document formats such as plain text, HTML, XML, or Microsoft Word.

Method and system for constructing vector databases used for converting free text queries to cyber language queries

A system and method for querying data sources for cybersecurity analysis is presented. The system and method include: receiving security logs from at least one data source, wherein the security logs lack pre-defined schema; generating a schema of the security logs based on at least a type of data of the security logs, wherein the generated schema includes fields of the security logs and values of the fields; embedding field vectors, wherein each field vector is a vector representation of a value of each respective field; embedding value vectors, wherein each value vector is a vector representation of a natural language description of each value in each respective field; and generating a query in a cyber language query, using an AI system, for execution on at least one target data source based, in part, on the generated schema, the embedded field vectors, and the embedded value vectors.
Owner:VEGA CYBER SOLUTIONS LTD

Weak supervision online video moment positioning method and system based on memory perception

The invention relates to a weak supervision online video moment positioning method and system based on memory perception, and belongs to the technical field of artificial intelligence, and the method comprises the steps: carrying out the multi-modal feature fusion of a given video and a text query thereof, and obtaining the unified frame level representation at each stage; inputting the fused features into an offline module and an online module in an integral and frame-by-frame manner by using an offline guide online model architecture; in the off-line module, generating a Gaussian mask to reconstruct query of a covered part of words, and obtaining a proposal of an action starting moment; in the on-line module, the long-term historical memory in the window is used for enhancing the score, the attention weight of the score in the window is dynamically generated, and the score of the current frame is calculated in a weighted mode; taking the proposal obtained by the offline module as a pseudo tag, and providing supervision information for the score sequence of the online module; and high-performance weak supervision on-line moment positioning can be completed only by independently deducing the on-line module. The expansion capability and the application value of the model are remarkably improved.
Owner:SHANDONG UNIV

Visual positioning method based on semantic comprehension and attribute distinguishing enhancement

The invention belongs to the technical field of visual positioning, and relates to a visual positioning method based on semantic comprehension and attribute distinguishing enhancement. The framework mainly comprises a feature coding module, a semantic sensitive data enhancement module, a fine-grained attribute guiding module and a multi-stage cross-modal decoder module. Specifically, the feature coding module performs feature coding on an input graph. The semantic sensitive data enhancement module generates a plurality of queries which are consistent with long text semantics by keeping the consistency of spatial relation words in combination with a large language model, so that a training data set for the long text is expanded. The fine-grained attribute guiding module extracts attribute prior information from a text query in combination with a text graph model and an image encoder, constructs a visual feature representation with higher discrimination by using the information guiding model, and generates a target query with attribute difference at the same time.
Owner:DALIAN UNIV OF TECH

Target-level three-dimensional point cloud cross-modal semantic retrieval method and device and electronic equipment

The invention discloses a target-level three-dimensional point cloud cross-modal semantic retrieval method and device and electronic equipment. The target-level three-dimensional point cloud cross-modal semantic retrieval method comprises the steps that independent three-dimensional target point clouds are segmented from three-dimensional scene point clouds, and unique identifiers are given to the independent three-dimensional target point clouds; projecting each target point cloud to three orthogonal two-dimensional observation planes to generate a multi-view composite image; a pre-trained vision-language basic model is adopted to code and fuse the synthesized image, and a unified multi-modal feature vector is generated; a vector database in which the feature vectors are associated with their identifiers is constructed. And encoding a natural language query text into a text query vector by using the model, calculating the semantic similarity between the text query vector and the feature vector in the database, and positioning and returning the corresponding three-dimensional target point cloud. According to the method, accurate and efficient cross-modal retrieval from a natural language to a three-dimensional point cloud target is realized, the problem that a traditional method is difficult to support semantic fine-grained retrieval of the three-dimensional target point cloud is solved, and a key technical support is provided for training data management in the fields of intelligence and the like.
Owner:MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Dynamic document annotation system

A set of locations in a data object to be annotated is identified as corresponding to metadata of the data object. A natural language text query is generated using the metadata of a data object. A set of scores is generated for the set of locations using a generative neural network, and the set of scores indicate whether individual candidate locations satisfy the natural language text query. Based on the set of scores, a location in the data object is annotated to generate an annotated location as corresponding to the metadata.
Owner:CITIGROUP

Systems and methods for orchestrating interaction with an artificial intelligence application

Systems and methods for orchestrating interaction with an artificial intelligence (AI) application in a contact center environment receive, via an AI agent, a voice message from a user; convert the message from voice to text; generate an initial computational inference process based on the text message; determine whether or not all information required to execute the initial computational inference process is available to the processor; when a determination is made that all information required is available: execute the initial computational inference process; generate a text reply based on the initial computational inference process; convert the text reply to a voice reply; and send the voice reply to the user via the AI agent; when a determination is made that information is unavailable: generate a text query requesting the information; convert the text query to a voice query; and send the voice query to the user via the AI agent.
Owner:THE BANK OF NEW YORK MELLON

Image text alignment method based on multi-modal large language model

PendingCN121542772AText alignmentFeature vector
The invention relates to the technical field of image texts, in particular to an image text alignment method based on a multi-modal large language model, which comprises the following steps of: segmenting an image into local areas and encoding the local areas into visual feature vectors by adopting a pre-trained visual encoder; visual and text features are mapped to the same dimension space through a learning projection layer, an InfoNCE loss function is adopted to maximize the similarity of a positive sample pair, dynamic interaction between text query and an image area is achieved through a cross attention mechanism, language module parameters are frozen to pre-train a visual encoder, an overall module is combined to be fine-tuned, and text query is completed. An error case is analyzed, and a verified and improved alignment module is obtained through a data enhancement optimization module. The method solves the problems that a traditional image text alignment method is incomplete in manual feature extraction and low in precision, an existing deep learning model is limited in cross-modal semantic understanding ability, and the generalization ability is poor when facing long-tail data.
Owner:GUANGDONG UNIVERSITY OF FOREIGN STUDIES

Information processing

To-be-queried information is received by a server device. The to-be-queried information is transmitted from a terminal device. A multimodal information search is performed by the server device based on the to-be-queried information to obtain multimodal search results. A content digest extraction of a target text in the multimodal search results is performed to obtain one or more content digest fragments. Based on the one or more content digest fragments, a text query result is generated. When a first preset number of search results in the multimodal search results lack a rich media query result, key information extraction on the text query result is performed to obtain key description information. A supplemental rich media query result is acquired according to the key description information. The text query result and the supplemental rich media query result are fused into a target query result that is transmitted to the terminal device.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Generating video descriptions using a machine learning model

The present disclosure describes techniques for generating video descriptions using a machine learning model. A plurality of sets of visual tokens corresponding to a plurality of frames of a video is generated. A first type of tokens is generated by implementing temporal pooling on the plurality of sets of visual tokens corresponding to the plurality of frames. A second type of tokens is generated by compressing each of the plurality of sets of visual tokens corresponding to each of the plurality of frames. A third type of tokens is generated by applying cross-attention between each of the plurality of sets of visual tokens and a fourth type of tokens including text tokens generated based on an input text query. A text description of the video is generated based on the first type of tokens, the second type of tokens, the third type of tokens, and the fourth type of tokens.
Owner:LEMON INC(GB)

System for realizing intelligent office guidance of office staff based on government affair element universe digital person

The invention provides a government affair element universe digital person-based system for realizing intelligent guide of affair handling personnel, and belongs to the technical field of government affair service intellectualization. The system deeply integrates element universe, intelligent interaction, big data and artificial intelligence technologies, realizes natural interaction with the affair handling personnel by means of an intelligent virtual digital person, and realizes intelligent guide of the affair handling personnel. And the government affair service process can be smoothly completed. A local intelligent knowledge base system and a knowledge base exclusive to multiple government affair departments are built, and a DeepSeek large model is applied. Voice or text query of business knowledge, handling processes, policies and regulations is supported, and contents cover latest policies of tax payment services, meta universe, offline business handling places and modes and the like. The digital human has abundant functions, provides business consultation, hall navigation guidance and intelligent explanation, can also perform daily question and answer interaction, and creates a relaxed communication atmosphere. And after logging in the meta universe cloud hall, the accompanying real-time efficient interactive guide service can be obtained by clicking the guide digital person, so that the government affair service efficiency and the user satisfaction are remarkably improved.
Owner:SHANDONG INSPUR COMML SYST CO LTD +1

Dialog agents with two-sided modeling

A central learning model is deployed as a user model and as an assistant model. Sensitive information utterances from a corpus of previously stored conversation language corresponding to user queries and chat agent responses thereto are used to train the user model to become an updated user model and to train the assistant model to become an updated assistant model, respectively. The user model provides user contexts corresponding to user queries to the assistant model and the assistant model provides assistant contexts corresponding to chat agent responses to the user model. During training, the user model does not provide plain-text queries to the assistant model and the assistant model does not provide plain-text responses to the user model. The updated assistant model may facilitate a federated training process produce an updated central model. An updated central model may be used to provide real-time chat agent responses to live user queries.
Owner:THE HONG KONG UNIV OF SCI & TECH

Multi-interface collaborative natural language data query method and system

The invention provides a multi-interface collaborative natural language data query method and system, and relates to the technical field of natural language processing.The method comprises the steps that text query data transmitted by a text input interface, voice query data transmitted by a voice input interface and touch instruction data transmitted by a touch selection interface are obtained; performing optimization processing on the voice query data and the touch instruction data to obtain a multi-interface fusion text set; analyzing the annotation data set based on preset natural language semantics to obtain a semantic analysis result; generating a unified query instruction through a digital information transmission technology; determining a data source corresponding to each piece of query field information in the semantic analysis result; according to the method, the initial query result is obtained according to the preset retrieval rule of each data source, so that accurate analysis and cross-interface unified execution of the natural language query intention are realized, and the accuracy, robustness and interactive experience of data query in a complex scene are improved.
Owner:BEIJING ALL VIEW CLOUD DATA TECH CO LTD

Retrieving digital images in response to search queries for search-driven image editing

Systems, methods, and non-transitory computer-readable media implements related image search and image modification processes using various search engines and a consolidated graphical user interface. For instance, one or more embodiments involve receiving an input digital image and search input and further modify the input digital image using the image search results retrieved in response to the search input. In some cases, the search input includes a multi-modal search input having multiple queries (e.g., an image query and a text query), and one or more embodiments involve retrieving the image search results utilizing a weighted combination of the queries. Some implementations involve generating an input embedding for the search input (e.g., the multi-modal search input) and retrieving the image search results using the input embedding.
Owner:ADOBE INC

Training method of video time positioning model, video time positioning method, equipment and medium

The invention relates to the technical field of computer vision, particularly provides a training method of a video time positioning model, a video time positioning method, equipment and a medium, and aims to solve the problem of large video time positioning error. In order to achieve the purpose, the model training method comprises the steps that multiple frames of images are sampled from a training video to serve as training data, the training data and a first preset query text are coded to obtain visual features and text query features, and a video time positioning model is trained based on the visual features and the text query features, obtaining a plurality of candidate answers, respectively calculating the relative advantage value of each candidate answer based on the real answer of the first preset query text and the plurality of candidate answers, adjusting the parameters of the video time positioning model based on each relative advantage value, and continuing to execute the step of sampling multiple frames of images from the training video. Therefore, the performance and accuracy of video time positioning can be improved.
Owner:PEKING UNIV +1

Machine-learned models for processing structural representations of therapeutic materials

PCT designated stageWO2025235348A1BiostatisticsCharacter and pattern recognitionStructural representationSequence processing
An example computer-implemented method for training a machine-learned sequence processing model to process queries over material structures includes obtaining a plurality of training inputs comprising a plurality of input sequences respectively corresponding to a plurality of representational schemas for representing material structures of a plurality of different materials, wherein each respective input sequence of the plurality of input sequences comprises a respective textual query and a respective sequence-based structural representation of a respective material, wherein each different respective sequence-based structural representation is characterized by a different respective representational schema; generating, using a machine-learned sequence processing model, a plurality of training outputs that respectively correspond to the plurality of training inputs; evaluating the plurality of training outputs; and updating, based on the evaluation of the plurality of training outputs, one or more parameters of the machine-learned sequence processing model.
Owner:GOOGLE LLC

Recursive deep text query method, system and equipment and storage medium

The invention discloses a recursive deep text query method, system and device and a storage medium, and the method comprises the steps: firstly creating an initial query and initializing a hierarchical depth, executing a first-layer query, carrying out evolution processing, and then carrying out retrieval; if the information density is insufficient, sub-query is generated to execute second-layer query, and retrieval is performed after second evolution processing; if the second-layer retrieval is contradictory, third-layer query is executed, retrieval is performed after timeline query and information source tracing verification are passed, and a third-layer retrieval result is finally output; through multi-level query and targeted evolution processing, query can be optimized step by step, more information dimensions can be covered, information gaps can be filled up, and retrieval comprehensiveness and accuracy can be improved; and a third-layer verification mechanism can effectively check contradictions by means of timelines and source tracing, so that the accuracy and reliability of results are guaranteed, and better retrieval results are provided for users.
Owner:ZHILU CLOUD (SHENZHEN) ARTIFICIAL INTELLIGENCE CO LTD

Multi-modal data processing and retrieval method

The invention particularly relates to a multi-modal data processing and retrieving method. The multi-modal data processing and retrieval method comprises the following steps: respectively carrying out depth feature extraction on image data and text data to generate an image vector and a text vector; carrying out interaction on the image vector and the text vector, and capturing semantic association information among multiple modes; mapping the vectors after interaction to a unified semantic space to realize semantic alignment among different modes; dividing the unified semantic space vector into fragments, storing the fragments in distributed nodes, and establishing a vector index; and generating a query vector by using an image or text queried by a user, carrying out parallel calculation on the similarity with a storage vector, and returning a retrieval result according to the similarity. According to the multi-modal data processing and retrieval method, the problems that the semantic difference between different modal data such as images and texts is large, the retrieval efficiency is low and storage is difficult to expand are solved, the accuracy and efficiency of multi-modal data retrieval are remarkably improved, and the method has remarkable technical advantages and wide application scenes.
Owner:INSPUR QILU SOFTWARE IND

API document recommendation method based on large language model

The invention discloses an API document recommendation method based on a large language model. The method comprises the steps of obtaining a plurality of text queries and a plurality of API documents, and presetting correlation values, so as to construct a reordering data set; training the constructed reordering model according to the reordering data set to obtain a trained reordering model; obtaining a to-be-tested text query, preprocessing all the API documents, and obtaining an initial sorting API document recommendation sequence of the to-be-tested text query according to the to-be-tested text query and the preprocessed API documents; and processing the initially sorted API document recommendation sequence and the to-be-tested text query by adopting the trained resorting model to obtain a resorted API document recommendation sequence of the to-be-tested text query. According to the method, the API document related to the development task is automatically recommended by understanding the query intention of developers, so that the consulting time of the developers is effectively shortened, and the development efficiency is improved.
Owner:ZHEJIANG UNIV

A method for real-time deduplication clustering of long texts

The application discloses a method for real-time deduplication and clustering of long texts, comprising the following steps: step one, infrastructure construction: a data structure capable of quickly fuzzy comparison of feature vectors is established in a central service storage; a text list A with feature values as primary keys is established in a central database; a text list B with feature values as primary keys is established in the central database; step two, coarse classification, real-time calculation when each text in the text list A and the text list B enters the system; step three, timed background modeling; step four, fine classification, timed and quantitative calculation is performed on the text list A, and all similar text pairs thus judged are marked with the minimum feature value in the set as the associated feature value in the table A; and step five, real-time query of terminal users: single text query and time period query, and the application relates to the technical field of data analysis. According to the application, the real-time data processing is ensured, and high accuracy is also considered.
Owner:OEFFECT INFORMATION TECH CO LTD

Text-guided human ear three-dimensional point cloud anaphora segmentation method and system

The invention provides a text-guided human ear three-dimensional point cloud anaphora segmentation method and system, and the method comprises the steps: inputting a to-be-segmented human ear three-dimensional point cloud and a corresponding text description into a human ear point cloud anaphora segmentation model which comprises a text encoder, a point cloud clustering module, a point cloud encoder and a text query feature decoding module; the text encoder encodes the text description into a text feature vector L; the point cloud clustering module is used for dividing the point cloud into different clusters and carrying out dimension transformation on the central point of each cluster to obtain a point cloud clustering feature Fd; the point cloud encoder uses L and Fd to guide point cloud feature extraction to obtain a point cloud encoding feature Fp '; and the text query feature decoding module is used for decoding through a fusion query decoder to obtain a mask preliminary vector Y4, and screening the Y4 through the text guide point cloud mask output module to obtain a mask most conforming to a text description related region as a final output mask. According to the invention, anaphora segmentation can be carried out on the human ear three-dimensional point cloud.
Owner:UNIV OF SCI & TECH BEIJING

Memory perception based weakly supervised online video temporal instance localization method and system

The present application relates to a memory-aware based weakly supervised online video moment localization method and system, belonging to the field of artificial intelligence technology, comprising: multi-modal feature fusion on a given video and its text query to obtain unified frame-level representation at each stage; using an offline-guided online model architecture, the fused features are input into the offline and online modules in the form of the whole and frame by frame; in the offline module, a Gaussian mask is generated to reconstruct the hidden part of the query, obtaining the proposal of the action starting moment; in the online module, the long-term historical memory in the window is used for enhancement, and the attention weight in the window is dynamically generated, and the score of the current frame is calculated by weighting; the proposal obtained by the offline module is used as a pseudo label to provide supervision information for the score sequence of the online module; only the online module needs to be inferred separately, and the weakly supervised online moment localization with high performance is completed. The present application significantly improves the expansion capability and application value of the model.
Owner:SHANDONG UNIV

Intelligent query method and device based on thematic map

The invention provides an intelligent query method and device based on a thematic map, and relates to the technical field of artificial intelligence. And inputting the thematic map and the text query instruction into the multi-modal expert model to obtain a target query result output by the multi-modal expert model. According to the intelligent query method based on the thematic map, on the basis that the thematic map is accurately analyzed, comprehensive perception and analysis of the space environment are improved, and efficient and accurate thematic map query is achieved.
Owner:AEROSPACE INFORMATION RES INST CAS

A multi-interface cooperative natural language data query method and system

The application provides a multi-interface cooperative natural language data query method and system, relates to the technical field of natural language processing, and obtains text query data transmitted by a character input interface, voice query data transmitted by a voice input interface and touch instruction data transmitted by a touch selection interface; the voice query data and the touch instruction data are respectively optimized to obtain a multi-interface fusion text set; based on a preset natural language semantic analysis annotation data set, a semantic analysis result is obtained; a unified query instruction is generated through digital information transmission technology; data sources corresponding to each query field information in the semantic analysis result are determined; and an initial query result is obtained according to preset retrieval rules of each data source, so that accurate analysis and cross-interface unified execution of a natural language query intention are realized, and the accuracy, robustness and interactive experience of data query in a complex scene are improved.
Owner:BEIJING ALL VIEW CLOUD DATA TECH CO LTD

Multimodal context selection for large language model based resolutions addressing technical issues

A method for technical issue resolution. The method includes: receiving, from a user, a text query concerning a technical issue; obtaining query-related context relevant to the text query; and processing, through a large language model (LLM), the text query and the query-related context to produce a multimodal query response used by the user to address the technical issue. More specifically, embodiments described herein utilize text topic and zero shot classification models to translate multimodal technical documentation (e.g., including text and images) into topic relevant metadata; and process queries, pertaining to technical issues, using a multimodal LLM provided with query-related text and image context derived from said topic relevant metadata.
Owner:DELL PROD LP

A cross-modal representation learning and retrieval method and system for grain production

This application discloses a cross-modal representation learning and retrieval method and system for grain production, relating to the field of agricultural informatization. The method includes: performing multi-granularity semantic alignment on image-text pairs in the grain production process based on a bidirectional guided image-text fusion network to obtain semantically segmented images; performing image spatial decoupling and temporal enhancement on video-text pairs in the grain production process based on global semantic guidance to obtain structured semantic image features; constructing a text feature library and an image feature library; determining a transmission plan matrix based on the modality of the data to be retrieved; generating query features for the data to be retrieved based on the transmission plan matrix; and outputting text query results or image query results using a similarity measurement method based on the query features, text feature library, and image feature library. This application enables deep fusion of cross-modal features, improves the accuracy of image-text semantic matching, and achieves fast and accurate matching and retrieval between images and text.
Owner:AGRI INFORMATION INST OF CHINESE ACAD OF AGRI SCI

Video time retrieval method based on bidirectional semantic enhancement

A video time retrieval method based on bidirectional semantic enhancement comprises the following steps: firstly, respectively extracting video features and querying text features by using a pre-trained video encoder and a pre-trained text encoder; secondly, respectively inputting the extracted features into a feature alignment module to obtain aligned video features and text features; then, designing a TGVM module, dynamically enhancing video features according to global features of a text mode, and reducing irrelevant information in the video features; thirdly, designing a Feedback Decoder module and feeding back comparison loss, shortening the distance between decoded text features and original video features, and enhancing text feature representation through the original video features; then, using a cross attention mechanism to obtain joint features after interaction of the video and the text; finally, a corresponding time slice in the video is queried through a decoder prediction text, and a model is trained under the supervision of time retrieval joint loss; according to the method, the multi-modal reasoning capability of the multi-modal interaction part is enhanced, so that video and text information can be better aligned.
Owner:XIDIAN UNIV

Method and controller configured to execute agent for adapting LLM response based on user preferences

A method for an agent based on a large-language model LLM comprises receiving a text query from a user, transforming to an LLM query and sending to the LLM. Thereafter, receiving an LLM response and evaluating the LLM response by determining one or more features of the LLM response, determining one or more targets of the text query, determining a distance between the one or more features of the LLM response and the one or more targets of the text query, determining that the LLM response does not meet an expected target when the distance falls below a threshold acceptance distance. If the LLM response does not meet an expected target, then generating an improved LLM response by re-parsing the text query to generate an improved LLM query and sending the improved LLM query to the LLM. Finally, the method includes providing the improved response to the user.
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD +1

Method and system for protecting the authenticity of machine learning model-generated images

The invention relates to the field of information technology, and more particularly to digital data protection means for confirming the authenticity of digital data. A method for protecting the authenticity of images generated by a machine learning model in response to a text query comprises the steps of: receiving user input data for the generation of an image; generating an image with the aid of a machine learning model; registering the type of user input data; extracting from the image pixels of at least one colour component thereof; separating each colour component into blocks of pixels with a set size; forming a first sequence of blocks; calculating an identifier of the image; generating an insertion for protecting the generated image; generating an encrypted sequence; generating a sequence to be embedded; embedding the sequence; and presenting the protected image to the user. The technical result is that of providing more efficient protection of AI-generated images in order to determine the authenticity thereof.
Owner:PUBLICHNOE AKTSIONERNOE OBSHCHESTVO SBERBANK ROSSII (PAO SBERBANK)

Multi-modal large model illusion relieving method based on multi-view multi-path reasoning

The invention discloses a multi-modal large model illusion relieving method based on multi-view multi-path reasoning. The method comprises the following steps: acquiring an input image and a text query of a user; the image is sent into a visual encoder of a visual language model LVLM to be processed, and a visual mark is obtained; processing the text query into a text mark through a word segmentation device of the LVLM; obtaining multi-view information; converting the multi-view description into a mark, and combining the mark with an original visual mark and a text mark to form an enhanced input sequence; 5) inputting the enhanced input sequence into the LVLM for decoding, activating a multi-path reasoning mechanism, and distributing an independent decoding path for information of each view angle; 6, the overall deterministic scores of all the potential answers are ranked, and the answer with the highest deterministic score is selected as the final output. According to the LVLM text output generation method, through cooperative work of the two strategies of multi-view information acquisition and multi-path reasoning, illusion of the LVLM during text output generation can be effectively reduced.
Owner:HUAZHONG UNIV OF SCI & TECH

A learner concentration detection method and device based on MRI images and a medium

The application discloses a learner concentration detection method and device based on MRI images and a medium, relates to the technical field of computer vision, and extracts features related to cognitive activities in learner fMRI signals through a local-global feature aggregator (LGA) to generate text query features and brain subtitles; in visual stimulus reconstruction, a three-condition comprehensive guidance condition diffusion model is introduced into a reverse generation process, brain subtitles, stimulus images and fMRI potential features are combined, an image generation result is optimized through a cross-attention mechanism, and learner concentration is evaluated according to the similarity between the reconstructed image and the original stimulus; through multi-modal feature fusion and weight distribution, the objectivity and accuracy of the concentration detection are significantly improved, and a new paradigm is provided for cognitive ability evaluation.
Owner:ZHEJIANG NORMAL UNIV