Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

410 results about "Content analytics" patented technology

Content analytics is the act of applying business intelligence (BI) and business analytics (BA) practices to digital content. Companies use content analytics software to provide visibility into the amount of content that is being created, the nature of that content and how it is used.

System and methods for cross platform engagement oriented artificial intelligence enhanced programming

A platform for dynamically generating application experiences. The platform comprises a design management system, an agent orchestration system, an analytics system, a model management system, a user management system, and databases for storing design elements and templates. The design management system provides a portal for application owners / designers to create UX / UI designs, allowing them to select design elements from a set of categories or templates. The platform gathers existing websites / applications to identify common design patterns, stored in a design catalogue database, and suggests historical interfaces for design exploration. It enables the generation of templated applications that integrate with legacy systems. The agent orchestration system parses user specifications, selects generative AI systems, and generates UX / UI content based on the specifications. The analytics system collects and analyzes data to provide insights for improving UX / UI design and optimizing website performance. The model management system trains and maintains generative AI models used for content generation.
Owner:QOMPLX INC

Hypertext markup language (HTML) content analysis using machine learning

HyperText Markup Language (HTML) content analysis (HCA) using machine learning is described. A feature vector schema may be generated based on domain names corresponding to HTML webpages and corresponding indications of a status of the HTML webpage. The schema may map each position in a feature vector of a given HTML webpage to a resource identifier. Information may be processed using the schema to generate respective feature vectors. The feature vectors may be used to train a model to generate risk indicators for HTML webpages. A potentially parked domain webpage or a potentially malicious domain webpage may be received. A feature vector for the webpage may be generated and inputted to the model. The model may generate a risk indicator for the webpage. The risk indicator may be output and may cause responsive actions. The model may be updated based on a determination indicating whether the webpage was a parked domain webpage or a malicious domain webpage.
Owner:CENTRIPETAL NETWORKS INC

Multi-agent automatic picture retouching system based on content analysis

The invention discloses a multi-agent automatic image retouching system based on content analysis, and the method comprises the steps: 1, receiving the input of an original image, calling a finely-adjusted multi-mode large model to carry out the combined semantic-visual analysis of the image, extracting key semantic elements in the image, and carrying out the recognition of the key semantic elements; comprising but not limited to background coordination degree, illumination condition, figure hair style definition, weather quality and composition layout rationality content information. Based on these information, the system performs comprehensive image quality scoring on the image, which combines image style, aesthetic level, definition and subjective quality multi-dimensional evaluation criteria. Meanwhile, based on a multi-dimensional image quality scoring result and understanding of image semantic content, the system automatically generates a repair and optimization strategy set for specific defects so as to guide a subsequent image quality enhancement process. The invention relates to the field of multi-modal model and multi-agent cooperation, and can meet the increasing requirements of rapid optimization and high-quality propagation of image contents.
Owner:信华信(大连)软件服务股份有限公司

Intelligent self-adaptive eye protection display system

The invention relates to the technical field of display, and particularly discloses an intelligent self-adaptive eye protection display system which sequentially comprises a content analysis module, an ambient light detection module, an eye protection parameter calculation module, a display parameter adjustment module and a user feedback module. The system identifies the screen content type and the dynamic degree in real time, synchronously detects the ambient light intensity and the color temperature, obtains the optimal combination of the blue light proportion, the brightness, the contrast ratio, the color temperature and the sharpness through a multi-target optimization model, and smoothly adjusts the display through driving; a user can score, finely adjust and mark a scene, and feedback data enters a self-learning engine to iteratively update a parameter weight so as to form a content-environment-user closed loop. The system can significantly reduce blue light harm and improve visual comfort without external hardware.
Owner:JIANGSU YUANTAI PRECISION INSTRUMENT CO LTD

Al-based video content analysis method and system

The invention relates to the technical field of content recognition, in particular to an Al-based video content analysis method and system, and the method comprises the following steps: carrying out the image segmentation learning based on a video sequence through employing a U-Net convolutional neural network, analyzing scenes and elements in a video frame, recognizing and isolating key visual elements in a video through network learning, and carrying out the recognition of the key visual elements in the video. Comprising objects and figures. According to the method, image segmentation learning is carried out by adopting the U-Net convolutional neural network, key visual elements in the video can be identified and isolated more accurately, a clearer basis is provided for follow-up scene change and key event tracking, video content analysis is carried out by applying the graph neural network, the identification capability of dynamic scenes and events is enhanced, and the identification efficiency is improved. Deeper structured understanding is provided for video content indexing, the video quality is analyzed through a structural similarity index evaluation method, and the reason of visual quality reduction can be accurately recognized and improved.
Owner:CHENDA (GUANGZHOU) NETWORK TECH CO LTD +1

Semantic perception black box large language model training data auditing method and system

The invention relates to the technical field of training data auditing, in particular to a semantic perception black box large language model training data auditing method and system.The method comprises the following steps that content is returned based on a text generation interface, lexical elements and candidate content are analyzed for multi-round sampling distribution, and a convergence index is determined; and calculating semantic paths and weights of the real content and the candidate content, and analyzing differences between the real content and the candidate content and the reference content to obtain an affiliation adaptation judgment result. According to the method, through multi-round sampling behavior fluctuation tracking, section feature dynamic extraction, stability change judgment and sequence-level weight aggregation, fine separation of training data attribution signals is achieved, a multi-level judgment system is constructed for complex expression and diversified output, judgment accuracy is optimized through a tension comparison signal group and an attribution adaptation judgment mechanism, and the judgment accuracy is improved. And a sensitive and adaptive training data member auditing strategy is formed, so that the risk hidden danger omission is effectively prevented, and the model data security boundary control is enhanced.
Owner:NANKAI UNIV

Voice segmentation intelligent editing system based on deep learning

PendingCN121260170ASpeech recognitionSpeech segmentationInformation density
The invention relates to the technical field of voice signal processing, and discloses a voice segmentation intelligent editing system based on deep learning. The system comprises a voice feature extraction module, a segmentation boundary detection module, a semantic content analysis module, an editing strategy generation module and a real-time quality evaluation module. The voice feature extraction module collects multi-dimensional voice features and timestamp information, and verifies feature integrity and timeliness; the segmentation boundary detection module identifies voice pause intervals and semantic turning nodes and divides segmentation units and boundary types; a semantic content analysis module extracts text content and emotion features of each segment, and analyzes semantic topic relevance and information density; an editing strategy generation module formulates a segmentation retention rule and a sequence adjustment scheme, and matches user preferences and scene demands; the real-time quality evaluation module monitors voice fluency and information integrity in the editing process and analyzes splicing errors and user feedback. According to the system, intelligent processing of the whole voice editing process is realized.
Owner:SHENZHEN JYEOO NETWORK TECH CO LTD

Automatic test case generation method and system based on AI

The invention provides an AI-based automatic test case generation method and system, and relates to the technical field of software testing, and the method comprises the steps: converting an original demand document uploaded by a user into an image format document, and calling a multi-modal large model to recognize text features and non-text features in the image format document, integrating to form an analyzed demand document; dynamically cutting the analyzed demand document to obtain a plurality of logically coherent content blocks; calling a large language model to generate a corresponding test case title for each content block, and forming a test case title set; and inputting each test case title in the test case title set and the corresponding content block into the large language model to generate a complete test case corresponding to the original demand document. The method has the beneficial effects that through multi-modal content analysis and intelligent workflow arrangement, automatic analysis and understanding can be performed on image-text features of a demand document, a high-quality test case is automatically generated, and the test period is greatly shortened.
Owner:SHANGHAI YOUKA NETWORK TECH CO LTD

Zero-configuration knowledge graph construction method, equipment and medium

The invention discloses a zero-configuration knowledge graph construction method, equipment and a medium, belongs to the technical field of knowledge graphs, and aims to solve the technical problem of how to overcome the hard coding defect in traditional knowledge graph construction, realize a more flexible construction mode with lower difficulty in use, further realize data source entity association and semantic fusion and improve the knowledge graph construction efficiency. According to the technical scheme, field name semantic analysis and data content analysis are conducted through a dynamic semantic analysis algorithm, and intelligent database deep analysis is achieved; the same entities in the multiple tables are automatically recognized and merged through a cross-table entity intelligent merging algorithm, and multi-table collaborative intelligent search is achieved; constructing an intelligent knowledge graph: automatically selecting an optimal construction strategy through a multi-strategy relation construction engine according to keyword role features, and independently constructing a local knowledge graph for each data table; and optimizing and intelligently exporting the knowledge graph.
Owner:INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD

Cross-platform content generation and distribution method based on multi-modal AI

The invention discloses a cross-platform content generation and distribution method based on a multi-modal AI, and belongs to the technical field of cross-platform content generation and distribution, and the method comprises the steps: carrying out the content analysis and feature extraction of an original material based on a multi-modal AI model, and generating a structured content label and a semantic vector. By combining a vector matching degree formula of target portrait features and platform features, an adaptation strategy is dynamically generated, it is ensured that content not only conforms to platform rules, but also can accurately reach a target group, the conversion rate and user viscosity are finally improved, visual, text and semantic vectors are aligned through a Transform multi-modal fusion model, a cross-modal joint representation vector is generated, and the user experience is improved. And in combination with an adversarial generative network, differentiated variants of the same theme are generated in batches, and a content diversity score mechanism ensures that generated contents are balanced between creativity and compliance.
Owner:QUZHOU TIMES ENGINE NETWORK TECHNOLOGY CO LTD

Large model-based phishing mail detection system and method

The invention discloses a phishing mail detection method and system based on a large model, and relates to the technical field of network security, the system comprises a real-time detection system and an offline analysis system, and a mail content analysis module receives real-time mail data to form structured data; the detection rule center matches the structured data with a detection rule, directly intercepts a mail hitting a high-confidence blacklist rule, and transmits the structured data of a mail hitting a low-confidence blacklist rule into a phishing mail detection agent; the phishing mail detection agent sequentially performs mail header suspicious feature analysis, mail body semantic analysis, mail body structure analysis and attachment content analysis on the structured data to obtain a mail intention, a link and an attachment file; and performing corresponding analysis by combining a detection tool, and judging whether the mail is a phishing mail or not according to an analysis result. A large-model-driven phishing mail detection intelligent agent is utilized, and a detection path is dynamically planned to cope with endless phishing mail attack means.
Owner:山东省大数据中心

Methods, systems, and devices for capturing video content associated with performing an athletic skill and determining biomechanic adjustments for performing the athletic skill

PendingUS20250312680A1Gymnastic exercisingBall sportsComputer graphics (images)Video content analysis
Aspects of the subject disclosure may include, for example, obtaining current video content of a player repeatedly performing a physical skill, analyzing the current video content based on previous video content, the previous video content comprises other video content of the player repeatedly performing the physical skill, determining biomechanic metrics of the player performing the physical skill based on the analysis, and determining each biomechanic metric of a portion of the biomechanic metrics does not satisfy a respective biomechanic metric success rate. Further embodiments include generating a first image of the player performing the physical skill from the current video content, generating a second image of the player performing the physical skill from the previous video content, and presenting the first image and the second image simultaneously and indicating the portion of the biomechanics that did not satisfy the respective biomechanic metric success rate. Other embodiments are disclosed.
Owner:ATHLETIQ LLC

Contextual advertising through multimodal content analysis

A system and method for contextual advertising that analyzes video content through multimodal examination of visual, audio, and textual elements to create detailed contextual understanding of individual scenes. The system segments video content into discrete scenes and simultaneously processes each scene to extract contextual characteristics including objects, settings, dialogue, music, and emotional tone. These characteristics are classified according to advertising industry taxonomies and converted into numerical embeddings that enable semantic similarity matching. During video playback, when advertisement opportunities occur, the system identifies the current scene context, analyzes available advertisements using similar techniques, computes similarity scores between scene and advertisement characteristics, and selects contextually appropriate advertisements for seamless integration. This approach enables privacy-compliant advertising that matches advertisement content with scene context rather than relying solely on user behavioral data, improving advertisement relevance and viewer experience.
Owner:TUBI INC

Content optimization method and system for enhancing search engine optimization (SEO) of a website

The present disclosure provides a search engine optimization (SEO) system comprising a remote server. The remote server comprises a memory with a set of executable routines and a search engine database with multiple fields of applications, each associated with multiple URLs indexed with a user engagement matrix, written data, and a search engine ranking. A processor acquires a web link from a computing device, extracts textual content, analyzes relevancy, and retrieves relevant URLs. The processor evaluates a thematic score, a readability score, and an emotional tone data using NLP techniques, analyzes written data of each URL to determine a topic weight, a legibility weight, and a sentiment tone data, develops a machine learning model, applies the model to recommend content alterations, and renders the alterations at the computing device.
Owner:ORIGINALITY AI INC

Value-oriented content analysis and intervention method in ideological and political education

InactiveCN120764804AForecastingKnowledge representationFeature vectorEducation intervention
The invention provides a value-oriented content analysis and intervention method in ideological and political education, and belongs to the technical field of education and teaching. A pre-trained multi-modal value understanding model is combined with a recurrent neural network to extract dominant value features and capture context dependence, implicit value expressions are recognized through an attention mechanism and comparative learning, and a tendency index is generated based on the similarity between feature vector calculation and core value. A reinforcement learning training content intervention agent is utilized to generate an education intervention strategy, and finally anti-fact reasoning is adopted to evaluate different intervention strategy effects and optimize an intervention scheme, so that a complete technical closed loop from content analysis to intervention implementation to effect evaluation is formed. The technical problem that the implicit value in the ideological and political education content is difficult to quantify to accurately identify and implement effective intervention is solved.
Owner:QINGDAO HUANGHAI UNIV

Video scene content label determination method based on knowledge graph

The invention relates to the technical field of video content analysis, and discloses a video scene content label determination method based on a knowledge graph. The method comprises the following steps: inputting a multi-modal feature sequence formed by visual, audio and text features extracted from a video stream into a knowledge graph inference engine comprising an entity relationship network and a semantic association rule base; performing node mapping through an engine to generate an initial scene entity set; performing hierarchical reasoning on the set based on a semantic association rule base to obtain a scene semantic topological structure; screening core entities according to entity weight distribution in the topological structure, and generating a candidate tag set; performing time sequence consistency verification on the candidate tag set, and correcting tag time sequence offset in combination with timestamp information; performing cross-modal disambiguation on the corrected label set by using an entity relationship network to eliminate semantic conflicts; and according to a disambiguation result, constructing a scene label knowledge sub-graph containing entity attributes and relation path constraints.
Owner:GUANGZHOU JUNHE INFORMATION TECH CO LTD

Dynamic advertisement placement based on content understanding and user data

Aspects of the disclosed technology provide solutions for dynamically placing an advertisement within media content based on content understanding and / or user data. An example method can include receiving live media content, which captures a live event, analyzing the live media content to identify one or more attributes associated with the live event, and accessing user data associated with a user device displaying the live media content. The example method can further include determining a time at which an advertisement is to be inserted within the live media content based on at least one of the one or more attributes or the user data.
Owner:ROKU INC

Multimedia network-oriented electronic publication automatic typesetting system and method

The invention discloses a multimedia network-oriented electronic publication automatic typesetting system and a multimedia network-oriented electronic publication automatic typesetting method, which relate to the technical field of electronic publication typesetting, and comprise a typesetting module for customizing multiple parameters of an initial layout of an electronic publication reader according to the type of a media terminal, analyzing and identifying the type of the media terminal by adopting equipment UA or API (Application Program Interface), establishing an adaptive font size calculation model to dynamically generate a basic typesetting rule, and outputting the basic typesetting rule to a layout planning module and a typesetting rule library; according to the invention, a basic typesetting rule is dynamically generated by establishing an adaptive font size calculation model, dynamic content data is acquired in a content analysis module by adopting an FFmpeg audio and video processing technology and an ARIMA time sequence analysis model, and format conversion and terminal adaptation are carried out by adopting a dynamic coding optimization algorithm and a multimedia cross-format compatible conversion algorithm. The compatibility of the electronic publication to the multimedia elements on different browsers and devices is improved.
Owner:SUZHOU DIGITAL POWER CULTURE COMM CO LTD

Method and System for Real-Time Collaboration, Task Linking, and Code Design and Maintenance in Software Development

A method for automated document processing and task assignment including receiving an input document from a document source, performing a content analysis on the input document, extracting metadata from the input document, identifying an identified task type to be performed, generating standardized metadata by converting the metadata into a standardized JSON format, storing the standardized metadata in a database, determining a user assignment for the identified task type, the including an assigned user, generating an action item including the identified task type, the user assignment, and the standardized metadata, and adding the action item to a task management system for processing by the assigned user.
Owner:MADISETTI VIJAY

Multi-granularity semantic alignment method for cross-modal image-text retrieval

A multi-granularity semantic alignment method oriented to cross-modal image-text retrieval comprises the steps that multi-layer feature representation of an image and a text is established, and a global vector and a local mark sequence are obtained respectively; capturing fine-grained features with higher information amount by using a mark screening mechanism, and inhibiting interference of redundant and invalid features on representation; under a unified semantic alignment framework, a bidirectional mapping relationship among multi-granularity semantics is constructed, and coarse-granularity, fine-granularity and cross-granularity semantic similarity alignment is comprehensively realized so as to strengthen multilayer semantic consistency constraints; and adaptively integrating the granularity similarities through a weighted fusion module, generating single instance-level similarities, and carrying out descending sorting on candidates. According to the method, the multi-granularity characteristics of the image and the text are comprehensively considered, the image-text retrieval precision in an actual scene can be improved, and the method has good application value in the fields of image-text retrieval, multi-modal content analysis and the like.
Owner:NANJING UNIV OF POSTS & TELECOMM

Language-driven hour-level traffic video analysis method and system

The invention discloses a language-driven hour-level traffic video analysis method and system, relates to the technical field of video content analysis and retrieval, can generate a video-specific open world detector, can accurately identify and filter fine-grained and composite semantic targets, and improves the accuracy of video analysis. And an hour-level video analysis service driven by open vocabularies and natural languages is really realized. In order to achieve the purpose, the technical scheme of the invention comprises the following steps of: 1) receiving an hour-level traffic video, and positioning a video clip related to natural language query in the video; and 2) automatically constructing an open world target detector in the video clip positioned in the step 1). And step 3) finishing cross-time-period target trajectory generation on the video clip positioned in the step 1). The invention also provides a language-driven hour-level traffic video analysis system for executing the method. The system comprises a video clip positioning module, an open world target detector construction module and a target trajectory extraction module.
Owner:BEIJING INST OF TECH

Subjective test question scoring method and device, electronic equipment and storage medium

The embodiment of the invention discloses a subjective test question scoring method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining question stem information and scoring standards of subjective test questions and a first answer content set for the subjective test questions; the first answer content set comprises a part of to-be-scored answer content which is randomly screened out; analyzing context correlation between the first answer content set and the answer key point set, and determining a rough score number of each answer content in the first answer content set according to the context correlation and the score weight; extracting a part of answer content from the first answer content set as a second answer content set according to the rough score number; obtaining an accurate score obtained by manually adjusting the rough score number of the answer content in the second answer content set; and taking the answer key point set and the second answer content set as training data, taking the accurate score as a real score corresponding to the second answer content set, and training a scoring model.
Owner:OPEN UNIVERSITY OF CHINA +1

Digital content analysis

In implementations of systems for digital content analysis, a computing device implements an analysis system to extract a first content component and a second content component from digital content to be analyzed based on content metrics. The analysis system generates first embeddings using a first machine learning model and second embedding using a second machine learning model. The first embeddings and the second embeddings are combined as concatenated embeddings. The analysis system generates an indication of a content metric for display in a user interface using a third machine learning model based on the concatenated embeddings.
Owner:ADOBE INC

Visual image efficient matching and retrieval method based on spatio-temporal characteristics

The invention discloses a visual image efficient matching and retrieval method based on spatio-temporal characteristics, and relates to the field of computer vision. Comprising the following steps: receiving a to-be-authenticated visual data object; starting a physical space source feature analysis channel to extract a physical space feature from an imaging device, and calculating a source credibility score; and starting a physical time content feature analysis channel in parallel to extract time correlation features between the moving target and associated phenomena thereof, and calculating a content authenticity score. And when the content authenticity score meets a preset trigger condition, executing a cross validation step. And comparing local physical space features of the suspicious target region and the background region identified in content analysis to obtain a judgment result about whether implantable tampering exists in the content. And generating a structured data object based on all analysis results. According to the method, the robustness and the reliability of authentication are improved through dual verification and intelligent linkage of the source and the content.
Owner:CHINA UNICOM (SICHUAN) IND INTERNET CO LTD +1

Utilizing intelligent digital content analysis and large language models to resolve transaction disputes

This disclosure describes methods, non-transitory computer readable storage media, and systems that use intelligent digital content analysis and large language models to resolve transaction disputes. The disclosed system selects appropriate digital content analysis tools for analyzing digital documentation including details of a transaction in response to a transaction dispute. The disclosed system utilizes data extracted from the digital documentation to generate a number and type(s) of model prompts to provide to a large language model based on attributes of the transaction dispute. The disclosed system uses responses generated by the large language model for the model prompts to determine a dispute resolution operation (e.g., approval, denial, or agent review). Furthermore, the disclosed system provides a response to the request to one or more computing devices associated with the request in response to execution of the dispute resolution operation.
Owner:MARQETA INC

Video content automatic auditing method and system based on AI drive

The invention discloses a video content automatic auditing method and system based on AI drive, and relates to the technical field of video content analysis, and the method comprises the steps: obtaining video information to be audited, and carrying out the multi-level segmentation of the video information; in the slices, node simulation is carried out according to the content independence of the video information; analyzing associated branches and historical features of each node to generate attribute features of the nodes; content auditing is carried out through an auditing strategy of multiple virtual agents, and illegal space analysis is carried out according to an unstable part in an auditing result; and generating a final video content auditing report according to the auditing result and the analysis result of the illegal space. According to the method, the structured recognition and intelligent auditing efficiency of the multi-modal video content is improved, accurate positioning, label attribution and automatic early warning of complex violation behaviors are realized, and the content security management capability is enhanced.
Owner:JIANGSU BROADCASTING CORPORATION

Retrieval-augmented generation for domain-specific technical documents

Effective Retrieval-Augmented Generation (RAG) pipelines face significant challenges when processing domain-specific technical documents that have diverse content types like text, figures, equations, and tables. To address this challenge, a context-oriented RAG system can be implemented for various domain-specific applications. The RAG system can include a lightweight, two-stage architecture to facilitate contextual understanding: a content analysis and enrichment pipeline for structured metadata extraction and a query processing pipeline for context-aware retrieval. In some cases, tabular data is processed using a dual-stream approach: semantically via text and visually via screenshots. The embedding vectors and the metadata can be stored in a visual data management system. The RAG system, utilizing the visual data management system, can answer questions and precisely retrieve technical information in a way that can preserve structural relationships and semantic connections across different modalities.
Owner:INTEL CORP

Small target detection method for unmanned aerial vehicle data

The invention provides a small target detection method for unmanned aerial vehicle data, and the method can effectively improve the detection precision, achieves better balance in three aspects of light weight, high precision and real-time performance, and can be suitable for more small target detection scenes. The method comprises the following steps: constructing a PLDP-SPP module, calculating a regional significance score S and a texture complexity score T for each pixel point in an input feature map in a feature content analysis module, and calculating to obtain a comprehensive score V by taking S and T as quantitative indexes; different pooling scales are preset according to the comprehensive scores V in different ranges, and it is ensured that more detailed context information extraction can be carried out on an area with large information density; in a size decision mechanism module, feature densities of different areas are matched through dynamic pooling, the scale adaptability is higher, and particularly, the detection precision of small targets and medium targets is remarkably improved.
Owner:AUTOLINK INFORMATION TECHNOLOGY CO LTD

Mine monitoring video key frame extraction method

PendingCN121459254ACharacter and pattern recognitionCluster algorithmVideo content analysis
The invention relates to the technical field of video content analysis, in particular to a mine surveillance video key frame extraction method, which comprises the following steps: extracting multi-dimensional features of each frame of a candidate frame set, obtaining a fusion value, and obtaining a candidate key frame set; based on an Euclidean distance method, calculating an Euclidean distance between continuous frames of the candidate key frame set, and determining a total clustering number according to a relationship between the Euclidean distance and a preset clustering threshold value; and clustering the frames in the candidate key frame set by using a preset mixed GWO-FCM clustering algorithm and the total clustering number to obtain a key frame set. According to the method, a hybrid clustering algorithm GWO-FCM combining grey wolf optimization and fuzzy C-means is adopted, global search and local optimization capabilities are considered, and the accuracy and representativeness of key frame clustering are ensured. The high-quality extraction of the key frames is realized, the number of redundant frames is obviously reduced, and the compactness and readability of the video abstract are improved.
Owner:XJ GRP CORP

Intelligent video monitoring early warning method and system based on multi-modal visual model

The invention provides an intelligent video monitoring early warning method and system based on a multi-mode visual model, and the method comprises the steps: introducing a low-rank decomposition attention mechanism and an improved attention mechanism into an improved YOLOv8 algorithm, so as to obtain a detection model; detecting a specific target in the video based on the detection model; a ByteTrack algorithm is adopted to track a specific target, and an image of the specific target is cut out; performing short-time-sequence action recognition on the image by adopting a SlowFast algorithm; and a Qwen-VL multi-mode reasoning model is adopted to analyze the long-time-sequence content of the video, and the analyzed long-time-sequence content is combined with the RAG knowledge base to realize dynamic anomaly judgment. According to the method, the purpose of effective and continuous tracking can be achieved, short-time-sequence action recognition and long-time-sequence content analysis can be carried out, and dynamic anomaly judgment can be effectively achieved.
Owner:JIANGXI UNIV OF TECH