Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

553results about "Metadata multimedia retrieval" patented technology

Semantic comprehension driven cross-modal information fusion and retrieval method and system

The invention discloses a cross-modal information fusion and retrieval method and system driven by semantic comprehension, and the method comprises the steps: obtaining text, image and audio original data, and extracting an initial feature set of each modal through a deep neural network; dynamically distributing each modal weight coefficient based on an attention mechanism, and performing weighted fusion on the initial feature set to obtain cross-modal fusion feature representation; through a cross-modal semantic association analysis model, high-dimensional semantic association features are extracted from the fusion feature representation, and semantic enhancement feature vectors are generated; constructing a cross-modal semantic graph network based on the vector, complementing missing modal features, and generating an optimized multi-modal feature set; and inputting the optimized feature set and the query sample into a contrast learning model, calculating a semantic similarity score, and generating a cross-modal retrieval result sorting list according to the score.
Owner:SHANGHAI CIVIL AVIATION VOCATIONAL & TECH COLLEGE

Collaborative personalized learning system and method based on large model

The invention discloses a collaborative personalized learning system and method based on a large model, and relates to the technical field of collaborative learning, and the system collects and fuses multi-modal data, generates a student state vector, plans a personalized learning path based on the state and a knowledge graph, and generates multi-modal explanation content according with the student style. The method comprises the following steps: extracting a student reasoning path, identifying thinking deviation, analyzing task performance, generating a path and content adjustment suggestion, identifying a motivation state, triggering an intervention strategy, receiving a teacher strategy, and realizing path and style co-construction. The multi-modal content of the matched style is generated, adaptive intervention is achieved in combination with inference analysis and emotion adjustment, and the learning efficiency and the teaching response intelligent level are improved.
Owner:SMART DYNAMICS CO LTD

Visual scene method, system and device based on digital twinning and medium

The invention discloses a visual scene method, system and equipment based on digital twinning and a medium, and relates to the technical field of digital twinning and visual modeling, and the method comprises the steps: collecting multi-format source data, carrying out the field standard mapping, and carrying out the consistency verification of the multi-format source data after the field standard mapping; writing the multi-format source data after consistency verification into a to-be-fused buffer area, selecting model component data in the to-be-fused buffer area to execute coordinate reference conversion, and performing association binding on the converted model component data and the structured graphic and text information by constructing an identification field; and synchronously loading the associated and bound model component data and the structured graphic and text information to a predefined template container, generating a scene configuration file according to a container structure, and calling a release engine to register to a multi-terminal rendering service channel. According to the method, accurate alignment of multi-format model components in a unified space coordinate system is realized by constructing an affine transformation coordinate reference conversion scheme.
Owner:中亿丰数字科技集团股份有限公司

Digital collection full-life-cycle traceability tracking system based on block chain

The invention discloses a digital collection full-life-cycle traceability tracking system based on a block chain, and the system comprises a digital identification module, a verification module, a packaging transfer module, a verification module, an integration module, and a system construction module, and achieves the full-life-cycle traceability tracking of a digital collection through the technologies of multi-dimensional feature extraction, watermark embedding, distributed storage, and zero-knowledge proof. And tracking and verification of the whole process from creation to transaction of the digital collections are realized. According to the system, the transaction authorization model is encapsulated by adopting the intelligent contract, so that safe transfer of ownership is ensured; a tamper-proof verification mechanism is constructed through multi-level integrity verification; different platform data are integrated by using a cross-chain technology to form a comprehensive traceability graph, so that key problems of authenticity verification, copyright protection, value evaluation and the like faced by a digital collection market are effectively solved. In addition, the system also establishes a scientific authenticity evaluation and value dynamic evaluation system, provides an objective basis for market pricing, and significantly improves the credibility and safety of the digital collections.
Owner:HANGYING (JIANGSU) INFORMATION TECH CO LTD

Multi-modal data identifier generation method and system based on semantic hash

The invention discloses a multi-modal information identifier generation method and system based on semantic hash. The method comprises the following steps: carrying out data preprocessing and multi-modal feature extraction on multi-modal data to obtain a unified representation containing rich semantic information; mapping the features to a shared semantic space by adopting an independent alignment projection network of each mode, and introducing cross-mode contrast learning and label supervision to realize semantic alignment among different modes; compressing the high-dimensional features after multi-modal data alignment to generate a Hash code with a fixed length; generating a unified semantic hash code containing semantic information of at least two modals; searching the number of times that the Hash code appears in a database through the unified semantic Hash code, generating a redundant code, and splicing the redundant code with the unified semantic Hash code to form an identifier of the sample; on the basis of the unified semantic hash codes, the hash codes related to the to-be-queried category in the test set are queried in the training set, and unified identification and efficient retrieval of different modes are achieved.
Owner:BEIHANG UNIV

Active real-time interaction system based on LLM

The invention discloses an active real-time interaction system based on LLM. The active real-time interaction system comprises a self-cognition module, a real-time perception module, an active planning module, a situation decision module, a dynamic execution module, an emotion engine module, an adaptive interaction module and a safety management and control module. The system adopts a multi-dimensional vector to represent role attributes, and constructs a hierarchical memory storage structure to record interaction experience; generating an environment state vector through a multi-modal data processing technology; executing the target task decomposition algorithm to generate a multi-path execution plan; making a real-time decision based on the multi-dimensional decision factor; mapping the abstract decision into an instruction sequence and monitoring an execution process; simulating a system emotional state and influencing decision expression; dynamically adjusting an interaction strategy according to the user characteristics; evaluating decision rationality and starting a corresponding intervention mechanism. All the modules form a closed-loop workflow through a standardized data interface, the problems that a traditional AI system is passive in response, lacks self-cognition, is limited in environmental perception, is single in planning capability, lacks flexibility in decision making and the like are solved, and active service, self-evolution and safety controllability of the system are achieved.
Owner:BEIJING ZHIMING ERXING NETWORK TECHNOLOGY CO LTD

Personalized teaching content generation method and system based on digital portraits

The invention discloses a personalized teaching content generation method and system based on a digital portrait, and the method comprises the steps: collecting the multi-dimensional learning data of a student in real time, analyzing the data keyword of the student, constructing the portrait of the student, generating a teaching target based on the portrait of the student, and enabling the teaching target to comprise a teaching theme, exercises related to the theme, and knowledge points related to the theme. Converting the teaching target into an executable instruction of a large language model through a structured prompt engineering technology; the large language model generates initial teaching content according to the teaching instruction; verifying and optimizing the generated initial teaching content to ensure that the generated teaching content conforms to a teaching target and can adapt to the current cognitive level and learning requirements of students; and outputting the verified and optimized teaching content to the students and automatically adapting to the content presentation form according to the learning style preference of the students.
Owner:SHAANXI NORMAL UNIV +1

Systems and methods for generating playlists by applying search prompts to a model configured to generate structured queries

An electronic device associated with a media-providing service stores, in a vector space, a plurality of respective vector representations for respective media content items. The electronic device receives a user input, including a text string. The electronic device generates, using a neural network, a structured query based on the text string. The electronic device determines, based on the structured query, whether to generate a vector representation of a portion of the text string. When the electronic device determines to generate the vector representation of the portion of the text string, it generates the vector representation of the portion of the text string, wherein the vector representation is embedded in the vector space, and identifies a set of media items using the vector representation of the portion of the text string. And the electronic device provides one or more select media items from the set of media items to a user.
Owner:SPOTIFY

Cross-modal image-text retrieval method and device based on pulse fusion

The invention discloses a cross-modal image-text retrieval method based on pulse fusion, and belongs to the technical field of image-text retrieval. The method comprises the following steps: inputting a target image set into a target detection network to obtain an image floating point code, and inputting a target text set into a word segmentation device to obtain a word floating point code; inputting the image floating point code and the word floating point code into a pulse encoder to obtain a first image pulse code and a first text pulse code; inputting the first image pulse code and the first text pulse code into a pulse cross attention fusion module to obtain a second image pulse code and a second text pulse code; respectively carrying out weighted accumulation and average pooling on the second image pulse code and the second text pulse code to obtain an image floating point feature vector set and a text floating point feature vector set, and calculating cosine similarity to obtain an image-text alignment result; and the text retrieval result of each image in the target image set is obtained based on the image-text alignment result, so that the accuracy and efficiency of image-text retrieval are improved.
Owner:WUHAN UNIV OF TECH

Meeting content intelligent generation processing method and system based on multi-modal large model

The invention discloses a conference content intelligent generation processing method and system based on a multi-modal large model, and the method comprises the steps: collecting the original data of a conference, and completing the standardization preprocessing; inputting a multi-modal large model, extracting multi-modal features and carrying out semantic alignment; executing cross-modal hash coding, generating binary codes and establishing an index database; performing hash retrieval on the related fragments, and constructing a conference content directed graph; based on a conference content directed graph structure, searching an optimized path by adopting a Monte Carlo tree; and generating structured conference content, and outputting a summary, an abstract and an action item. According to the method, efficient extraction, accurate retrieval and structured intelligent generation of the conference content are realized by fusing a multi-modal large model, cross-modal Hash coding and Monte Carlo tree search.
Owner:NANJING WEITEXI NETWORK SCI & TECH

Data auditing method and device, electronic equipment and nonvolatile storage medium

The invention discloses a data auditing method and device, electronic equipment and a nonvolatile storage medium. The method comprises the following steps: acquiring cross-modal data to be audited; a verification rule set corresponding to the multi-modal knowledge base is determined, first verification processing is carried out on the cross-modal data based on the verification rule set, a first verification result is obtained, and the first verification processing is used for verifying whether the cross-modal data has exceptions which do not conform to rules in the verification rule set or not; the large auditing model is adopted, second verification processing is conducted on the cross-modal data, a second verification result is obtained, and the second verification processing is used for detecting whether logic contradictions exist in the cross-modal data or not based on the semantic reasoning ability of the large model; and determining an exception report corresponding to the cross-modal data according to the first verification result and the second verification result. According to the method and the apparatus, the technical problem of poor content quality inspection effect of cross-modal data in related technologies is solved.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Method, appartus, device and storage medium for media item generation

Embodiment of the disclosure relates to a method, apparatus, device, and storage medium for media item generation. The method provided herein includes: displaying a item search page comprising a search control, the item search page being configured to provide an item search result corresponding to received input information; receiving a search term inputted in the search control; providing, in the item search page, a generation entry in association with the search control; and based on a selection of the generation entry, displaying, in the item search page, a set of media items generated based on the search term. In this way, the embodiments of the disclosure may provide users with more diversified media items to meet user specific needs for media items.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Method, apparatus, electronic device, and storage medium for collecting media content

Embodiments of the present disclosure provide a method, apparatus, electronic device, and storage medium for collecting media content. The method includes: in response to a collection operation on media content, adding the media content to the media content collection list of the current user, and displaying a favorites list of the current user, where the favorites list is used to display first identifiers of at least some of the current user's favorites; in response to a trigger operation on the first identifier, adding the media content to the favorite corresponding to the first identifier on which the trigger operation acts. By adopting the above technical solutions, the embodiments of the present disclosure can enrich the collection methods of media content.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Photo content clustering for digital picture frame display and automated frame storytelling

A method and system for automated routing of pictures taken on mobile electronic devices to a digital picture frame including a camera, microphone, and speaker integrated with the frame, and a network connection module allowing the frame for direct contact and upload of photos from electronic devices or from photo collections of community members. Clustering photos by content is used to improve display and to respond to photo viewer desires. Trends or patterns can be detected from the photo collections and that information used for various purposes beyond photo display. The frame includes a conversational intelligence that provides a verbal communication with a viewer, such as for determining an identity or preferences of the frame viewer, determining photos to display for the viewer, discussing displayed photos with the viewer, or telling stories or life histories to the viewer based upon photo content.
Owner:PUSHD INC

Federal cross-modal retrieval method and system based on interaction prompt

The invention provides a federal cross-modal retrieval method and system based on interactive prompts, and relates to the field of cross-modal information retrieve.The federal cross-modal retrieval method and system based on interactive prompts retrieve similar second modal data from a second modal data set by using first modal data based on the feature similarity of two modal data comprises the steps that initial features of the two modal data are extracted respectively; performing multi-layer bidirectional interaction between the initial features of the two modals by using a cross-modal interaction network obtained by federal learning and taking a prompt vector as an intermediary to obtain final features of the two modals after cross-modal interaction; calculating feature similarity based on the final features of the two modalities, and screening similar second modal data; according to the method, federal learning and prompt learning are combined, so that the effectiveness and universality of a cross-modal retrieval technology are improved, and the problems of privacy protection and performance optimization in cross-modal retrieval are solved.
Owner:SHANDONG UNIV

Electronic government enterprise service intelligent response method based on AI cue word engineering

The invention relates to the crossing field of artificial intelligence technology and e-government, and discloses an e-government enterprise service intelligent response method based on AI cue word engineering, comprising: constructing and maintaining a government knowledge graph and a scenarized cue word template library; receiving and analyzing a service request input by an enterprise, and extracting an enterprise feature tag and a business intention; selecting a target cue word template based on the matching degree of the enterprise feature tag and the scene tag of the template in the scene cue word template library, and generating an adapted cue word template; retrieving associated knowledge data from the government affair knowledge graph according to the service request, and filling the adapted cue word template with the retrieved associated knowledge data to generate a structured cue word; inputting the structured cue word into an AI model to drive the AI model to generate a government affair service response; and collecting multi-dimensional evaluation feedback of government affair service response, and carrying out iterative optimization. The demand understanding precision is improved, and the risk of core information misjudgment is reduced.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

Tracking concepts within content in content management systems and adaptive learning systems

An example for presenting educational content including converting a multimedia document into a text representation of the multimedia document, partitioning the text representation into multiple portions of text based on a text characteristic of the text representation; determining educational concepts associated with portions of text. Then, generating at least a first cluster and a second cluster where the first and second clusters include portions of text of the multiple portions of text, and each portion of respective text in the first cluster is associated with a first educational concept of the one or more educational concepts and each portion of respective text in the second cluster is associated with a second educational concept of the one or more educational concepts. From the clusters, generating a first educational content item based on the first cluster and a second educational content item based on the second cluster.
Owner:OBRIZUM GRP LTD

Tobacco marketing hotspot event real-time analysis method based on big data and AI

The invention provides a tobacco marketing hotspot event real-time analysis method based on big data and AI, and relates to the technical field of big data and artificial intelligence, and the method comprises the steps: collecting multi-source heterogeneous data through a distributed crawler, and achieving the structured processing of unstructured data through the semantic analysis and multi-modal fusion technology; hot event identification and early warning are carried out by combining deep learning and a propagation dynamics model; further fusing the knowledge graph, NLP and space-time analysis to generate brand specification popularity ranking and trend prediction; and finally, an evaluation model is constructed based on historical and real-time data, new product research and development, brand promotion and supply chain optimization strategies are output, and full-process intelligent decision support is realized.
Owner:SHANDONG INSPUR DIGITAL BUSINESS TECHNOLOGY CO LTD

Intelligent traffic accident liability affirmation method and system based on multi-agent cooperation mechanism

The invention relates to the technical field of artificial intelligence and intelligent traffic, in particular to a traffic accident liability intelligent affirmation method and system based on a multi-agent cooperation mechanism and a large language model. The whole process of credible information verification, responsibility affirmation reasoning and standard document generation is simulated. Wherein the legal expert agent adopts a competing mechanism, and the fairness and the accuracy of affirmation are improved through preliminary affirmation, double-party defense and final judgment. The reasoning ability of a large language model and accurate knowledge retrieval of a retrieval enhancement generation technology are fused, the real-time performance and accuracy of legal clause quotation are ensured, and a road traffic accident identification document conforming to specifications is automatically generated. The method effectively solves the problems that traditional manual identification is low in efficiency and high in subjectivity, and an existing intelligent method lacks interpretability and legal accuracy.
Owner:SICHUAN POLICE COLLEGE +1

Advertisement creativity matching method based on multi-modal content generation

The invention discloses an advertisement creativity matching method based on multi-modal content generation, and relates to the technical field of digital media content generation, and the method comprises the following steps: building a cross-modal time anchoring belt facing advertisement creativity matching, carrying out metaphor level decomposition on input text information, marking a symbol axis for image information, and carrying out data processing on the image information; obtaining an initial semantic boundary list; and constructing a culture fingerprint database according to the initial semantic boundary list, and mapping the territory taboo information and the brand symbol information into constraint tags to obtain a semantic guardrail set. According to the method, through cross-modal time anchoring and semantic boundary control, accurate correspondence of the text and the image in time and semantic levels is achieved, and it is ensured that generated content is clear in semantic meaning and adaptive in culture. In combination with breathing type phase traction and cultural fingerprint dynamic adjustment, multi-modal content rhythm and emotion are coordinated and unified, brand expression is kept stable, and the overall consistency and propagation effect of advertisement creativity are improved.
Owner:大根控股股份有限公司

Multi-modal data retrieval, generation and synthesis method and system based on artificial intelligence driving

The invention relates to the technical field of artificial intelligence, in particular to a multi-modal data retrieval, generation and synthesis method and system based on artificial intelligence driving, and the method comprises the steps of multi-modal feature extraction, cross-modal alignment, feature fusion, multi-modal retrieval and sorting and multi-modal generation. Different deep learning models are adopted to carry out feature extraction on multi-modal data and convert the multi-modal data into vectors, through cross-modal alignment, different modal feature vectors are mapped to the same vector space, through feature fusion, correlation weights are calculated through an attention mechanism, weighted summation is carried out on fusion features, and through multi-modal retrieval and sorting, multi-modal data are obtained. According to the method, the similarity between a query vector and candidate data is calculated and sorted, and finally, through multi-modal generation, retrieval knowledge is used as external knowledge to be fused into a generation model, and a multi-modal synthesis answer is generated, so that information in multi-modal data can be effectively integrated and utilized.
Owner:BEIJING SGITG ACCENTURE INFORMATION TECH CO LTD +1

Multi-modal intelligent question answering system and method based on large model

The invention discloses a multi-modal intelligent question answering system and method based on a large model. The system comprises a signaling processing module, a document preprocessing module, a voice preprocessing module, a retrieval recall module, an RAG processing chain module and a knowledge base module. The method comprises the following steps: (1) a knowledge base offline construction step; (2) a knowledge base online construction step; (3) a real-time question and answer step; according to the system and the method provided by the invention, personalized intelligent question and answer service can be provided for audiences by combining materials which are uploaded by a speaker in advance and contain rich contents such as characters, pictures and charts and real-time presentation contents and utilizing multi-mode presentation information such as characters, pictures and voice, so that instant questions of the audiences can be answered, and the audiences can be more interested. Extensible knowledge supplementation can be provided according to the presentation content and materials.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Task processing method and intelligent device with body

The invention discloses a task processing method and an intelligent device, and relates to the technical field of intelligent devices, and the method comprises the steps: obtaining multi-modal data which comprises task information of a to-be-processed task in a real-time interaction scene of a physical entity and an environment; performing feature extraction and feature alignment processing on different modal data in the multi-modal data to obtain a first feature after feature alignment; performing cross-modal data retrieval on a knowledge base based on the first feature to obtain target retrieval data; performing feature fusion on different modal data features in the first features to obtain second features; and making a decision based on the target retrieval data and the second feature by using a decision model to obtain an action instruction sequence, and executing the to-be-processed task based on the action instruction sequence.
Owner:LENOVO (BEIJING) LTD

Learning resource recommendation method and system based on knowledge tracking and retrieval enhancement generation

The invention provides a learning resource recommendation method based on knowledge tracking and retrieval enhancement generation, and belongs to the technical field of education. The method comprises the steps of obtaining data of a user learning platform and / or learning content uploaded by a user, and constructing a background knowledge base related to a current learning task of the user; acquiring learning interaction behavior data of the user, constructing a knowledge state tracking model, and forming a current knowledge mastering state of the user; wherein the learning interaction behavior data of the user comprises knowledge points, exercises, videos and texts; and outputting personalized learning resource recommendation and question and answer response by utilizing a retrieval enhancement generation model in combination with the background knowledge base and the knowledge mastering state. Therefore, dynamic, personalized and knowledge-accurate learning content recommendation can be realized.
Owner:GUANGZHOU PANYU POLYTECHNIC

Dynamic multimedia data hash retrieval method and system based on extensible increment

The invention discloses a dynamic multimedia data hash retrieval method and system based on extensible increment, and relates to the technical field of multimedia data retrieval. The method comprises the following steps: acquiring dynamic multimedia data to be retrieved; a pre-trained Hash retrieval large model is constructed, the Hash retrieval large model is trained by taking a bit extensible Hash center as global supervision information and taking tag cosine similarity as local supervision information, and specifically, generalization feature representation of new multimedia data maintaining new and old class discrimination is obtained through forward propagation; constructing a linear mapping relationship between the generalization feature representation and the hash code, and introducing an auxiliary variable to continuously update the hash function without playback; and utilizing the trained Hash retrieval large model to generate a query Hash code for the dynamic multimedia data to be retrieved, and utilizing the query Hash code to retrieve. According to the method, low memory occupation, high updating efficiency and non-forgetting retrieval of the dynamic multimedia data stream in the open environment are realized.
Owner:SHANDONG JIANZHU UNIV

System to correlate video data and contextual data

In some embodiments, a method of processing image data may include receiving environmental data and associated capture time data from a sensor of a mobile computing device, the capture time data reflecting capture time of the environmental data; processing the environmental data to generate metadata; time stamping the metadata using the capture time data; receiving video data and video time data at a processor; correlating the metadata to the video data using the capture time data and the video time data; receiving a search query; and / or identifying a frame within the video data by performing a search of the metadata using the search criterion.
Owner:SNAP INC