Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

496results about "Multimedia data clustering/classification" patented technology

Automatic construction method of end-to-end agent based on graph structure semantic fusion

The invention relates to the technical field of artificial intelligence, in particular to an automatic construction method of an end-to-end agent based on graph structure semantic fusion. The method comprises the following steps: receiving business demand data input by a user; business target and demand constraint condition analysis is carried out on the business demand data, and a core workflow framework of the intelligent agent is generated; performing end-to-end execution path analysis on the core workflow framework of the intelligent agent to obtain an end-to-end workflow; constructing a dynamic evolution semantic map; and constructing an end-to-end call chain execution strategy based on the end-to-end workflow, and performing agent instance packaging and agent instance reinforcement learning enhancement processing according to the dynamic evolution semantic map, thereby automatically constructing an end-to-end agent. According to the invention, by fusing the graph structure knowledge and the generation capability of the large language model, an efficient, accurate and extensible agent automatic construction scheme is provided for various complex business scenes.
Owner:BEIJING ZHONGSHURUIZHI TECH CO LTD

Foundation generative artificial intelligence (AI) model with transformer architecture for environmental, social, and governance (ESG) impact

PendingUS20250299059A1Multimedia data clustering/classificationBiological modelsSustainability reportingEngineering
An ESG-specific multimodal AI foundation model is disclosed, featuring a Transformer-based architecture with approximately 30 billion parameters, designed explicitly for environmental, social, and governance (ESG) domain applications. This model uniquely supports extremely long context windows (up to 128,000 tokens), critical for comprehensive ESG analyses of lengthy documents such as sustainability reports and policies. It integrates textual and visual data through gated cross-attention and a Mixture-of-Experts (MoE) architecture, achieving precise multimodal context comprehension. The invention employs Group Relative Policy Optimization (GRPO) reinforcement learning strategy, refining model outputs based on group-relative advantages computed from multiple candidate generations, thus significantly enhancing ESG-specific reasoning and output quality.
Owner:ECORATINGS SOFTWARE SOLUTIONS PTE LTD

Visual scene method, system and device based on digital twinning and medium

The invention discloses a visual scene method, system and equipment based on digital twinning and a medium, and relates to the technical field of digital twinning and visual modeling, and the method comprises the steps: collecting multi-format source data, carrying out the field standard mapping, and carrying out the consistency verification of the multi-format source data after the field standard mapping; writing the multi-format source data after consistency verification into a to-be-fused buffer area, selecting model component data in the to-be-fused buffer area to execute coordinate reference conversion, and performing association binding on the converted model component data and the structured graphic and text information by constructing an identification field; and synchronously loading the associated and bound model component data and the structured graphic and text information to a predefined template container, generating a scene configuration file according to a container structure, and calling a release engine to register to a multi-terminal rendering service channel. According to the method, accurate alignment of multi-format model components in a unified space coordinate system is realized by constructing an affine transformation coordinate reference conversion scheme.
Owner:中亿丰数字科技集团股份有限公司

Retrieval enhancement method based on multi-modal data fusion and modal perception

The invention relates to the technical field of information retrieval and generation, in particular to a retrieval enhancement method based on multi-modal data fusion and modal perception. According to the method, firstly, a dual-channel architecture is adopted to perform feature extraction and coding on a text and an image respectively, and mutually independent embedded representation spaces are constructed, so that high-quality collaboration and matching of cross-modal representation are realized; and a pseudo-pairing generation mechanism is introduced to effectively mine and reconstruct the existing non-paired data in the knowledge base. And designing a query modal perception and dynamic weighting mechanism for accurately controlling the fusion proportion of the image-text bimodal information in the retrieval stage so as to match the modal demand difference of different query contents. And further executing aggregation retrieval and reordering of the cross-modal information by using dynamic weighted fusion retrieval to generate a candidate set of multi-modal responses. According to the method, accurate matching and dynamic weight adjustment of the image-text content are realized, and the accuracy and expression integrity of the generated content are improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Multi-user data storage docking and secure transmission method based on AI

The invention provides an AI-based multi-user data storage docking and secure transmission method, which comprises the following steps of: acquiring a propagation path of a hot topic according to a multi-modal content association graph, performing fragmentation storage on the propagation path by adopting a clustering algorithm, and generating a fragmentation storage index table; when the relevance of different modal contents is higher than a preset threshold value, determining that the sensitivity grading result is consistent with the multi-modal relevance, performing dynamic rule adjustment by adopting a deep learning model, generating a dynamic auditing rule set, and obtaining boundary judgment parameters of the rule set; and distributing the dynamic auditing rule set and the decision transparency index to each service line by adopting a cross-platform data synchronization protocol according to the auditing decision log, generating a cross-platform consistency auditing standard, and obtaining a synchronization log of standard execution.
Owner:SHENZHEN MINGHUI INTELLIGENT TECH CO LTD

Intelligent multi-mode patrol sensing method, system and equipment and storage medium

The invention belongs to the field of multi-mode sensing fusion, and relates to an intelligent multi-mode patrol sensing method, system and equipment and a storage medium, which are used for carrying out automatic patrol, real-time sensing and abnormity warning in a complex environment. Comprising the steps of natural language instruction analysis, task decomposition and distribution, intelligent agent scheduling and execution, multi-modal information collection and fusion, anomaly recognition and event response, result feedback and task closed loop, and through the integration of an OWL framework and a DeepSeek-R1 large model, a multi-modal data fusion technology and an NL-SLAM autonomous navigation algorithm, the multi-modal data fusion algorithm and the multi-modal data fusion technology are integrated. The problems that in the prior art, dynamic task understanding based on natural languages cannot be achieved, a sensing system lacks efficient multi-modal data fusion, so that the anomaly recognition precision is low, and complex tasks or emergencies are difficult to deal with are solved, various abnormal conditions in a complex environment are effectively dealt with, and the method is suitable for popularization and application. The stability and expansibility of task execution are improved, and the navigation success rate and the task completion rate of the system in an unstructured environment are enhanced.
Owner:QINGDAO UNIV

Knowledge question and answer rapid processing method and system based on artificial intelligence

The invention provides a knowledge question and answer rapid processing method and system based on artificial intelligence, and relates to the field of artificial intelligence. A multi-modal knowledge graph is constructed, collected multi-source teaching data is fused through a mixed retrieval strategy, and the mixed retrieval strategy comprises semantic retrieval, vector retrieval and metadata retrieval; multi-level question and answer processing is executed based on an RAG enhancement framework, a multi-modal input intention is analyzed, cross-library joint retrieval is performed, and an optimization answer is generated in combination with a teaching scene; distilling the global model to a lightweight TinyBERT architecture, dynamically optimizing question and answer quality through a cognitive reinforcement learning framework, positioning a key document from a comprehensive retrieval list, evaluating an optimized answer, and reconstructing an answer with a key document verification score; according to the invention, the professional skill level of teachers and students in the fields of artificial intelligence and large model application can be improved, and the personalized requirements of teachers and students in teaching, scientific research and innovation courses can be met.
Owner:RONGKE LIANCHUANG (TIANJIN) INFORMATION TECH CO LTD

System and method for using artificial intelligence (AI) to analyze social media content

Systems and methods for reducing the search space by processing media content to refine search parameters. A computing device may obtain the media content in response to receiving a request for inclusion of the media content in a media content knowledge repository, extract an audio component, a video component, and a text component of the media content, and determine attributes within the extracted components. The computing device may determine segment attributes based on a result of correlating the determined audio, video, and text attributes, integrate the segment attributes into the media content knowledge repository, and / or perform any of a variety of responsive actions.
Owner:SOCIAL VOICE LTD

Legal service method and system based on large language model and related equipment

The invention provides a legal service method and system based on a large language model and related equipment, and the method comprises the steps: obtaining user request information, and recognizing the class information and field information of the user request information; according to the category information and the domain information, a preset agent meeting the category information and the domain information is selected from a preset agent library, the preset agent meeting the category information and the domain information serves as a target agent, and the target agent comprises a preset large language model meeting the domain information and a preset function matrix meeting the category information; inputting the user request information into the target agent to obtain an initial result; on the basis of the user request information, enhanced information is inquired in a preset vector knowledge base through an RAG retrieval technology; and generating a target result for the user request information based on the enhanced information and the initial result, and displaying the target result through a visualization technology. And different legal problems can be flexibly handled according to the diversity and individuation requirements of legal services.
Owner:EAST CHINA NORMAL UNIV

Digital collection full-life-cycle traceability tracking system based on block chain

The invention discloses a digital collection full-life-cycle traceability tracking system based on a block chain, and the system comprises a digital identification module, a verification module, a packaging transfer module, a verification module, an integration module, and a system construction module, and achieves the full-life-cycle traceability tracking of a digital collection through the technologies of multi-dimensional feature extraction, watermark embedding, distributed storage, and zero-knowledge proof. And tracking and verification of the whole process from creation to transaction of the digital collections are realized. According to the system, the transaction authorization model is encapsulated by adopting the intelligent contract, so that safe transfer of ownership is ensured; a tamper-proof verification mechanism is constructed through multi-level integrity verification; different platform data are integrated by using a cross-chain technology to form a comprehensive traceability graph, so that key problems of authenticity verification, copyright protection, value evaluation and the like faced by a digital collection market are effectively solved. In addition, the system also establishes a scientific authenticity evaluation and value dynamic evaluation system, provides an objective basis for market pricing, and significantly improves the credibility and safety of the digital collections.
Owner:HANGYING (JIANGSU) INFORMATION TECH CO LTD

Electronic device using personal ai model, and operation method thereof

An electronic device includes a memory storing one or more instructions; a communication circuit; and a processor operatively coupled to the memory and the communication circuit, in which the one or more instructions, when executed by the processor, cause the electronic device to: identify a use pattern of specified media items from a plurality of media items stored in the memory, determine, based on the use pattern, a score of each of the specified media items, extract, using a main AI model stored in the memory, a feature corresponding to a characteristic of each of the specified media items, acquire a first AI model trained based on the score and the feature, based on the first AI model, determine a first preference of each of first media items from the plurality of media items, and based on the first preference, perform a function related to the first media items.
Owner:SAMSUNG ELECTRONICS CO LTD

Multi-modal data identifier generation method and system based on semantic hash

The invention discloses a multi-modal information identifier generation method and system based on semantic hash. The method comprises the following steps: carrying out data preprocessing and multi-modal feature extraction on multi-modal data to obtain a unified representation containing rich semantic information; mapping the features to a shared semantic space by adopting an independent alignment projection network of each mode, and introducing cross-mode contrast learning and label supervision to realize semantic alignment among different modes; compressing the high-dimensional features after multi-modal data alignment to generate a Hash code with a fixed length; generating a unified semantic hash code containing semantic information of at least two modals; searching the number of times that the Hash code appears in a database through the unified semantic Hash code, generating a redundant code, and splicing the redundant code with the unified semantic Hash code to form an identifier of the sample; on the basis of the unified semantic hash codes, the hash codes related to the to-be-queried category in the test set are queried in the training set, and unified identification and efficient retrieval of different modes are achieved.
Owner:BEIHANG UNIV

Active real-time interaction system based on LLM

The invention discloses an active real-time interaction system based on LLM. The active real-time interaction system comprises a self-cognition module, a real-time perception module, an active planning module, a situation decision module, a dynamic execution module, an emotion engine module, an adaptive interaction module and a safety management and control module. The system adopts a multi-dimensional vector to represent role attributes, and constructs a hierarchical memory storage structure to record interaction experience; generating an environment state vector through a multi-modal data processing technology; executing the target task decomposition algorithm to generate a multi-path execution plan; making a real-time decision based on the multi-dimensional decision factor; mapping the abstract decision into an instruction sequence and monitoring an execution process; simulating a system emotional state and influencing decision expression; dynamically adjusting an interaction strategy according to the user characteristics; evaluating decision rationality and starting a corresponding intervention mechanism. All the modules form a closed-loop workflow through a standardized data interface, the problems that a traditional AI system is passive in response, lacks self-cognition, is limited in environmental perception, is single in planning capability, lacks flexibility in decision making and the like are solved, and active service, self-evolution and safety controllability of the system are achieved.
Owner:BEIJING ZHIMING ERXING NETWORK TECHNOLOGY CO LTD

System And Method For Using Artificial Intelligence (AI) To Analyze Social Media Content

Systems and methods for reducing the search space by processing media content to refine search parameters. A computing device may obtain the media content in response to receiving a request for inclusion of the media content in a media content knowledge repository, extract an audio component, a video component, and a text component of the media content, and determine attributes within the extracted components. The computing device may determine segment attributes based on a result of correlating the determined audio, video, and text attributes, integrate the segment attributes into the media content knowledge repository, and / or perform any of a variety of responsive actions.
Owner:SOCIAL VOICE LTD

Deep learning-based teaching corpus construction method and system, and medium

The invention discloses a teaching corpus construction method and system based on deep learning, and a medium, and belongs to the cross technical field of medical education, artificial intelligence and multi-modal data processing. The method comprises the following steps: a data acquisition and annotation stage: acquiring medical inquiry video data of different scenes; in the model training stage, a multi-modal deep learning model architecture is used for carrying out cross-modal alignment on multi-modal data; then semantic understanding and question classification are carried out based on the pre-training language model after fine tuning; in the corpus construction and optimization stage, multi-modal data are integrated, a corpus management system is built, dynamic updating and self-adaptive optimization are carried out on a corpus, meanwhile, a typical non-language posture library is constructed for a regional culture background, and teaching is carried out in combination with the corpus. According to the invention, in combination with medical professional knowledge, an efficient and intelligent medical inquiry video teaching corpus is constructed through deep fusion of natural language processing, computer vision, speech recognition and multi-modal data fusion technologies.
Owner:CHONGQING MEDICAL UNIVERSITY

Personalized teaching content generation method and system based on digital portraits

The invention discloses a personalized teaching content generation method and system based on a digital portrait, and the method comprises the steps: collecting the multi-dimensional learning data of a student in real time, analyzing the data keyword of the student, constructing the portrait of the student, generating a teaching target based on the portrait of the student, and enabling the teaching target to comprise a teaching theme, exercises related to the theme, and knowledge points related to the theme. Converting the teaching target into an executable instruction of a large language model through a structured prompt engineering technology; the large language model generates initial teaching content according to the teaching instruction; verifying and optimizing the generated initial teaching content to ensure that the generated teaching content conforms to a teaching target and can adapt to the current cognitive level and learning requirements of students; and outputting the verified and optimized teaching content to the students and automatically adapting to the content presentation form according to the learning style preference of the students.
Owner:SHAANXI NORMAL UNIV +1

Meeting content intelligent generation processing method and system based on multi-modal large model

The invention discloses a conference content intelligent generation processing method and system based on a multi-modal large model, and the method comprises the steps: collecting the original data of a conference, and completing the standardization preprocessing; inputting a multi-modal large model, extracting multi-modal features and carrying out semantic alignment; executing cross-modal hash coding, generating binary codes and establishing an index database; performing hash retrieval on the related fragments, and constructing a conference content directed graph; based on a conference content directed graph structure, searching an optimized path by adopting a Monte Carlo tree; and generating structured conference content, and outputting a summary, an abstract and an action item. According to the method, efficient extraction, accurate retrieval and structured intelligent generation of the conference content are realized by fusing a multi-modal large model, cross-modal Hash coding and Monte Carlo tree search.
Owner:NANJING WEITEXI NETWORK SCI & TECH

Cross-platform content generation and distribution method based on multi-modal AI

The invention discloses a cross-platform content generation and distribution method based on a multi-modal AI, and belongs to the technical field of cross-platform content generation and distribution, and the method comprises the steps: carrying out the content analysis and feature extraction of an original material based on a multi-modal AI model, and generating a structured content label and a semantic vector. By combining a vector matching degree formula of target portrait features and platform features, an adaptation strategy is dynamically generated, it is ensured that content not only conforms to platform rules, but also can accurately reach a target group, the conversion rate and user viscosity are finally improved, visual, text and semantic vectors are aligned through a Transform multi-modal fusion model, a cross-modal joint representation vector is generated, and the user experience is improved. And in combination with an adversarial generative network, differentiated variants of the same theme are generated in batches, and a content diversity score mechanism ensures that generated contents are balanced between creativity and compliance.
Owner:QUZHOU TIMES ENGINE NETWORK TECHNOLOGY CO LTD

Element fusion-based cultural and creative design auxiliary method and system

The invention provides a cultural creative design assisting method and system based on element fusion. Belongs to the technical field of creative design. The method comprises the steps that design elements are acquired and classified; on the basis of a deep learning algorithm, a semantic network between elements is constructed, the semantic network is converted into a low-dimensional vector space through a network embedding technology, and user preference, emotional tendency and future trend are analyzed in combination with social media data; and generating a fusion scheme, and carrying out continuous iterative optimization on the design. The semantic relation between the design elements is analyzed, so that the internal relation between the design elements can be deeply understood; through a complex network theory, a language network model between design elements is constructed, the relationship between different elements can be visualized, and high efficiency and systematicness of design decision are facilitated.
Owner:HANGZHOU WUSHI WUJI CULTURE TECH CO LTD +1

Multi-source data fusion law enforcement record document generation method

The invention discloses a law enforcement record document generation method based on multi-source data fusion, and belongs to the technical field of data processing, and the method comprises the following steps: S1, obtaining multi-source data of an alarm case, classifying texts, images and audios in the multi-source data, and obtaining a structured data set; s2, extracting case related entities from the structured data set, and generating an event description table; and S3, determining a corresponding event element field in the event description table, and generating a law enforcement record document according to a preset document specification template. The multi-source data fusion law enforcement record document generation method solves the problem of low reliability of a law enforcement document in practical application due to low efficiency and inconsistency of an existing law enforcement record document generation mode.
Owner:GUANGDONG POLICE COLLEGE (GUANGDONG PROVINCIAL PUBLIC SECURITY JUDICIAL MANAGEMENT CADRE COLLEGE)

Outdoor intelligent art display system based on multimedia interaction

The invention discloses an outdoor intelligent art display system based on multimedia interaction, and relates to the technical field of digital art display and intelligent control, and the system comprises an audience behavior collection module which is used for collecting audience staying time and behavior response in real time, obtaining a content attraction initial evaluation result, calculating an attraction index of each content unit, and outputting the attraction index of each content unit; the content analysis and screening module is used for extracting alternative multimedia content units from a preset content library, judging the theme relevance between alternative contents and the currently displayed contents, and obtaining a candidate content set meeting replacement conditions; according to the outdoor intelligent art display system based on multimedia interaction, intelligent replacement and updating of multimedia content are realized on the premise of not interrupting display, content combination is automatically optimized through audience feedback data, content with low attraction is eliminated, and welcome elements are enhanced.
Owner:GUANGZHOU ACADEMY OF FINE ARTS

Tracking concepts within content in content management systems and adaptive learning systems

An example for presenting educational content including converting a multimedia document into a text representation of the multimedia document, partitioning the text representation into multiple portions of text based on a text characteristic of the text representation; determining educational concepts associated with portions of text. Then, generating at least a first cluster and a second cluster where the first and second clusters include portions of text of the multiple portions of text, and each portion of respective text in the first cluster is associated with a first educational concept of the one or more educational concepts and each portion of respective text in the second cluster is associated with a second educational concept of the one or more educational concepts. From the clusters, generating a first educational content item based on the first cluster and a second educational content item based on the second cluster.
Owner:OBRIZUM GRP LTD

Image video retrieval method based on domain fine-tuning large language model

The invention provides an image video retrieval method based on a domain fine-tuning large language model, which comprises the following steps: performing fine-tuning on a pre-training model to obtain a fine-tuning pre-training model for intention classification and keyword extraction; performing dynamic iteration screening on an optimal prompt template through Monte Carlo tree search in combination with a hidden Markov model (HMM); performing noise filtering on the keyword list, and predicting category labels of the filtered keywords through a conditional random field model to obtain a keyword enhancement set; combining with the user intention to generate a query condition, and obtaining a candidate resource set; and according to the similarity between the user query text and the candidate resource set, and in combination with the optimal prompt template, obtaining the resource path with the highest matching score between the user query and the candidate resource, and obtaining the retrieved image or video, so that the identification deviation possibly occurring when a general model processes proper nouns and terminologies can be effectively solved, and the user experience is improved. And the retrieval accuracy and response speed are improved, so that the retrieval accuracy and professional adaptability are improved.
Owner:HUBEI ZHONGKE NETWORK ENG

Digital media content element accurate screening method based on artificial intelligence image recognition

The invention discloses a digital media content element accurate screening method based on artificial intelligence image recognition, and relates to the technical field of digital media, and the method comprises the steps: building a distributed capture network to form a dynamic content pool, and building a metadata index database; calling a multi-modal image perception engine to generate a double-layer characteristic spectrum containing dominant and recessive elements; constructing a distributed recognition model cluster based on federated learning; converting the user demand into a screening parameter set and generating a decision tree; screening and secondarily verifying an output result through a double-path matching mechanism; and constructing a reinforcement learning reward function based on user behaviors, and driving the model and the decision tree to co-evolve. Multi-source heterogeneous content full-dimension analysis is achieved, the recognition comprehensiveness and depth are improved, knowledge barriers and privacy risks are solved, the screening accuracy and flexibility are improved, the system is endowed with the continuous optimization capacity, and the method is suitable for efficient and accurate digital media content screening scenes.
Owner:XIAMEN HUAXIA UNIV

Automated Response Engine Implementing a Universal Data Space Based onCommunication Interactions Via an Omnichannel Electronic Data Channel

Various embodiments relate generally to data science and data analysis, computer software and systems, and control systems to provide a platform to implement automated responses to data representing electronic messages, among other things, and, more specifically, to an automated predictive response computing system independent of electronic communication channel of an electronic message payload, the automated predictive response computing system being configured to, for example, implement a universal data space based on, at least in part, conversational data flows, which may be classified and used to provide a predictive response to assist resolution, such as assisting an agent among other things. In an example, a method may include augmenting at least a subset of one or more portions of communication data, implementing augmented communication portion data to determine a predicted response, and generating data to facilitate the predicted response based on the subset of inbound electronic messages.
Owner:KHOROS LLC

Hybrid artificial intelligence classifier

System, methods, apparatuses, and computer program products are disclosed for generating and using a hybrid artificial intelligence classifier for classifying input into one or more nodes of a taxonomy. Training data is received for at least a first portion of the taxonomy and used to train a supervised machine learning (ML) model to classify input into the first portion of the taxonomy having training data. A large language model (LLM) taxonomy is determined for at least a second portion of the taxonomy. The hybrid AI classifier classifies input based on a first classification obtained by providing the input to the supervised ML, and a second classification obtained by providing at least the input and the LLM taxonomy to a pre-trained LLM.
Owner:CAMELOT UK BIDCO LTD

RAG-based pdf intelligent retrieval and generation method and system

The application discloses a kind of PDF intelligent retrieval and generation method and system based on RAG, by obtaining the document data of input, using the classification model established in advance to parse document data, extract text content and image content to form first data set;Using deep learning model to the image content in first data set carries out feature extraction, while the text content in first data set applies natural language processing technology to carry out semantic analysis, obtains multimodal feature set;According to multimodal feature set, application information integration algorithm is uniformly encoded and is handled to generate second data set, if detecting the integrity of fusion feature vector in second data set is lower than preset threshold value, then supplementary context semantic analysis fills in missing information;Using preset index construction mechanism to the clustering processing of fusion feature vector in second data set, generates the retrieval index library containing classification index structure.The application improves the accuracy and comprehensiveness of document retrieval.
Owner:HUNAN ZHIXUE YOUKE INFORMATION TECHNOLOGY CO LTD +1

Method and system for ai-based evaluation of game animals

A system for an automated evaluation of a game animal based on sensory animal-related data including a processor of an animal evaluation server (AES) node configured to host a machine learning (ML) module and connected to at least one user-entity node over a network and a memory on which are stored machine-readable instructions that when executed by the processor, cause the processor to: receive an evaluation request including animal profile sensory data from the at least one user-entity node; derive the animal profile sensory data from the evaluation request; parse the animal profile sensory data to derive a plurality of key classifying features; query a local animal evaluation database to retrieve local historical animal evaluations'-related data based on the plurality of key classifying features; generate at least one classifier feature vector based on the plurality of key classifying features and the local historical animal evaluations'-related data; and provide the at least one classifier feature vector to the ML module configured to generate an animal evaluation predictive model for producing at least one animal scoring parameter; and generate animal scoring data for the at least one user-entity node based on the at least one animal scoring parameter.
Owner:WARD JUSTIN

Dynamic multimedia data hash retrieval method and system based on extensible increment

The invention discloses a dynamic multimedia data hash retrieval method and system based on extensible increment, and relates to the technical field of multimedia data retrieval. The method comprises the following steps: acquiring dynamic multimedia data to be retrieved; a pre-trained Hash retrieval large model is constructed, the Hash retrieval large model is trained by taking a bit extensible Hash center as global supervision information and taking tag cosine similarity as local supervision information, and specifically, generalization feature representation of new multimedia data maintaining new and old class discrimination is obtained through forward propagation; constructing a linear mapping relationship between the generalization feature representation and the hash code, and introducing an auxiliary variable to continuously update the hash function without playback; and utilizing the trained Hash retrieval large model to generate a query Hash code for the dynamic multimedia data to be retrieved, and utilizing the query Hash code to retrieve. According to the method, low memory occupation, high updating efficiency and non-forgetting retrieval of the dynamic multimedia data stream in the open environment are realized.
Owner:SHANDONG JIANZHU UNIV

Multi-modal content based automated feature recognition

A system includes a computing platform having processing hardware, and a memory storing software code and a machine learning (ML) model-based feature classifier. When executed, the software code receives media content including a first media component corresponding to a first media mode and a second media component corresponding to a second media mode, encodes the first media component using a first encoder to generate multiple first embedding vectors, and encodes the second media component using a second encoder to generate multiple second embedding vectors. The software code further combines the first embedding vectors and the second embedding vectors to provide an input data structure for a neural network mixer, process, using the neural network mixer, the input data structure to provide feature data corresponding to a feature of the media content, and predict, using the ML model-based feature classifier and the feature data, a classification of the feature.
Owner:DISNEY ENTERPRISES INC