Marketing system based on intelligent analysis of multi-modal literature data
The multimodal literature data intelligent analysis system solves the problems of structured integration and deep semantic understanding of scientific literature, constructs a multi-dimensional researcher knowledge graph, and realizes efficient and accurate marketing strategy generation, thus solving the problems of data silos in scientific literature and the lag in marketing strategies.
Patent Information
- Application Number
- CN202511437401.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-09
AI Technical Summary
In the fields of scientific literature data analysis and precision marketing, existing technologies are incomplete and inaccurate in extracting structured information from multi-source heterogeneous literature, making it difficult to deeply integrate the information, lacking deep semantic understanding, and resulting in a single user profile that cannot support personalized marketing.
The system employs a multimodal intelligent analysis system for literature data, including a literature parsing and preprocessing module, an AI-enhanced analysis module, a data processing and integration module, a cache management module, and a batch processing module. Through format recognition, specialized parsing, semantic segmentation, multi-model scheduling, and knowledge graph construction, it achieves efficient structured preprocessing and deep semantic extraction of heterogeneous literature, constructs a multi-dimensional researcher knowledge graph, and generates personalized marketing strategies.
It significantly improved the resolution rate and information extraction accuracy of heterogeneous documents, enhanced marketing precision, increased processing speed by 3-4 times, and improved marketing precision from 2.3% to over 12.7%, achieving end-to-end intelligent conversion from original documents to precision marketing strategies.
Smart Images

Figure CN120911441A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of Internet data collection technology and intelligent information analysis and processing, and specifically relates to a marketing system based on multi-modal literature data intelligent analysis. BACKGROUND
[0002] Internet data collection and intelligent analysis technology refers to a technical system that uses automated means to obtain data from the Internet such as websites, APPs, APIs, sensors, etc., and through machine learning, big data processing, etc. to clean, store, analyze and mine valuable information.
[0003] Currently, Internet data collection and intelligent analysis technology has made significant progress, but still faces challenges in data quality, label depth, dynamic modeling, etc.
[0004] In the field of data collection, technology development shows a trend of diversification and efficiency. Focusing on crawler, incremental crawler and other technologies have been widely used in e-commerce, news and other fields, combined with dynamic IP proxy and anti-crawling strategies, greatly improving the accuracy and efficiency of data collection. At the same time, the maturity of distributed architecture such as Scrapy-Redis and real-time stream processing technology such as Kafka+Flink makes it possible to efficiently collect and respond to billions of data in milliseconds. However, there are still two major problems in this field: data quality is uncontrollable such as false information, multi-source data format confusion, and the contradiction between real-time and resource cost such as low latency demand leading to exponential growth in server costs.
[0005] In the field of intelligent analysis, the optimization of machine learning and deep learning algorithms has significantly improved data processing capabilities. The accuracy of CV / NLP tasks has approached or surpassed human level, such as BERT's text classification F1 value reaching 0.92, and lightweight technologies such as MobileBERT enabling models to adapt to edge devices. In addition, the popularity of AutoML tools and low-code visualization platforms such as QuickBI has lowered the threshold for data analysis. However, the field still faces serious challenges: the shallow labeling system leads to insufficient semantic granularity, the weak dynamic modeling capability makes it difficult for traditional algorithms to adapt to user interest drift, such as K-means model accuracy decaying by 15% every week due to changes in data distribution, and multi-modal data fusion difficulties, such as limited improvement in joint analysis of text and video data.
[0006] The application shortcomings of typical scenarios further highlight the limitations of technology. For example, in the scientific research consumables supply chain, due to the lack of fine-grained correlation in the labeling system, the inventory turnover rate is 20% lower than the industry average; in the cross-platform user portrait scenario, coarse-grained labels fail to identify behavioral motivations, resulting in a 12% decrease in ad click-through rate.
[0007] With the rapid expansion of the parameter scale of deep learning models, the decision-making process is increasingly "black-boxed" as the model's representation capacity increases, and feature interpretability decreases significantly. This uninterpretability severely restricts the application of AI in high-risk fields such as medical diagnosis and financial risk control, as it cannot meet the requirements of accurate results, transparent decision logic, and auditability. In recommendation systems, new users or new items have difficulty generating reliable representations due to sparse behavior data, leading to poor initial recommendation results. Frequent model updates to improve results, on the other hand, introduce instability in prediction results, increase business decision risk, and even disrupt the consistency of user experience. These two contradictions essentially reveal the deep tension between "performance improvement" and "reliability guarantee," "data-driven" and "logical transparency" in the development process of AI technology.
[0008] Overall, the current technology is still in the stage of perceptual intelligence, and there are obvious bottlenecks in data authenticity, semantic understanding, and dynamic adaptability.
[0009] Under this technical background, the limitations of existing technologies are particularly prominent when they are applied to vertical and specialized fields. In particular, in the field of scientific literature data analysis and precision marketing, the existing technical architecture cannot effectively solve the following three core problems:
[0010] First, the structured information extraction of multi-source heterogeneous literature is incomplete and inaccurate, making it difficult to integrate deeply. Scientific literature data exists in various formats such as PDF, XML, Word, and HTML, with different internal structures, coding standards, and content presentation methods. Existing general collection and analysis technologies lack the ability to adapt to the structure of academic literature, resulting in insufficient extraction accuracy of complex formats, mathematical formulas, tables, references, and other elements, making it difficult to achieve unified and high-quality structured integration of cross-format and cross-source data, forming "data silos."
[0011] Second, the understanding of literature content is limited to keywords and shallow semantics, lacking insight into deep semantics such as research intent, technology roadmap, and equipment and consumable usage logic. Although existing NLP technologies can achieve high accuracy in text classification and entity recognition in general fields, they struggle to understand complex scientific concepts, logical relationships between experimental methods, and the underlying motivations behind technology selection in highly specialized scientific literature. This results in the inability to accurately extract high-value information such as "the specific experimental step required for precise instrument models," "alternative chemicals for key reagents," or "software tool chains relied upon for data analysis" from literature.
[0012] Third, the user portrait constructed based on shallow labels is single and cannot support truly personalized scientific research product and service recommendation. Existing user portrait technologies mostly rely on shallow features such as behavior clicks and keyword frequencies, and it is difficult to construct a dynamic and multi-dimensional knowledge model that can reflect the professional ability, technical preference, equipment use history and future demand trend of researchers. Therefore, the generated marketing strategies are often generalized and lagging, and cannot achieve accurate matching with the real, deep and evolving scientific research needs of researchers, resulting in low marketing conversion rate. SUMMARY
[0013] The purpose of the present application is to provide a marketing system based on intelligent analysis of multi-modal literature data to solve the problems in the prior art.
[0014] To this end, the above-mentioned purpose of the present application is achieved by the following technical solutions:
[0015] The marketing system based on intelligent analysis of multi-modal literature data comprises,
[0016] The literature analysis and preprocessing module uses a format recognizer to identify the format of the input literature, and uses a format-specific parser for deep content extraction and structure analysis, and then generates a unified document tree structure representation through a document structure parser, a metadata extractor obtains key information, and a content partitioner divides logical parts to achieve efficient and structured preprocessing of heterogeneous scientific research literature, and the information is transmitted to the AI enhanced analysis module through the standardized JSON data structure output;
[0017] The AI enhanced analysis module receives the JSON data, extracts deep semantic information in the literature through a semantic blocker, a domain-specific prompt generator, a multi-LLM API connector and a response parser, and outputs the information as structured data in JSON-LD format;
[0018] The data processing and integration module cleans, verifies and de-duplicates the JSON-LD data, and based on the processed data, constructs a researcher knowledge graph containing multiple dimensions of basic identity, research interest, technology stack, equipment use and demand prediction through a user portrait generator, and generates personalized product combinations and marketing plans using a marketing strategy recommender, achieving complete conversion from raw data to precise marketing strategies, cleaning, verification and structured processing of extraction results;
[0019] The cache management module connects the data processing and integration module and is used for storing and retrieving analysis results and user portraits based on content hash;
[0020] The batch processing module connects the literature analysis and preprocessing module, the AI enhanced analysis module and the cache management module respectively to coordinate parallel processing and cross-correlation analysis of multiple files;
[0021] The literature analysis and preprocessing module, the AI enhanced analysis module, and the data processing and integration module are sequentially connected to form an end-to-end processing pipeline from original literature to marketing strategy.
[0022] In addition to the above technical solutions, the application can also use or combine the following technical solutions:
[0023] As a preferred technical solution of the application, the literature analysis and preprocessing module comprises:
[0024] The format identifier comprises a three-level format identification mechanism based on file extension, file header magic number, and file content features.
[0025] The multiple format-specific parsers comprise an XML file reader, a PDF parser, a JSON parser, an HTML parser, a Word document parser, a TXT parser, and an EPUB parser.
[0026] The document structure parser generates a unified document tree structure representation.
[0027] The metadata extractor extracts the title, author, and DOI information of the literature.
[0028] The content partitioner divides the literature content according to academic structure.
[0029] The literature analysis and preprocessing module realizes the preprocessing of heterogeneous literature standardization processing through three-level format identification, multiple parser adaptation, unified structure reconstruction, metadata extraction, and academic content partitioning.
[0030] As a preferred technical solution of the application, the PDF parser comprises:
[0031] A text, image, and table separation unit based on PyMuPDF.
[0032] A scanned document processing unit integrated with Tesseract-OCR.
[0033] A layout analysis unit using computer vision algorithms to identify the chapter structure of the document.
[0034] A mathematical formula recognition and MathML conversion unit.
[0035] The PDF parser realizes high-precision content and structure extraction of native PDF and scanned PDF through the cooperative work of the above units.
[0036] As a preferred technical solution of the application, the AI enhanced analysis module comprises:
[0037] A text chunker, employing a segmentation algorithm with semantic integrity protection;
[0038] A prompt generator, containing a library of at least 200 field-specific prompt templates, designed specifically for different parts of scientific literature and different extraction targets, to guide large language models to accurately extract experimental instruments, experimental consumables, and software tools information;
[0039] An API connector, supporting dynamic switching between multiple large language model APIs;
[0040] A response parser, for converting unstructured responses into JSON-LD format;
[0041] The AI-enhanced analysis module realizes domain knowledge enhancement and standardized analysis through semantic chunking, specialized prompt templates, multi-model API dynamic scheduling, and JSON-LD structured output.
[0042] As a preferred technical solution of the present application: the data processing and integration module includes:
[0043] A data cleaner, implementing noise filtering based on rules and machine learning;
[0044] A user portrait generator, constructing a five-dimensional portrait containing basic information, research interests, technology stack, equipment usage, and demand prediction;
[0045] A marketing strategy recommender, using a recommendation algorithm combining collaborative filtering and knowledge graph;
[0046] The data processing and integration module realizes precise marketing strategies through data cleaning, five-dimensional user portrait construction, and collaborative filtering and knowledge graph fusion recommendation.
[0047] As a preferred technical solution of the present application: the five-dimensional researcher knowledge graph constructed by the user portrait generator includes:
[0048] The technology stack dimension includes the experimental methods, data analysis software, and programming languages of the researchers;
[0049] The equipment usage dimension includes the equipment models, brands, and usage scenarios mentioned in the literature;
[0050] The demand prediction dimension is based on the researchers' historical technology stack, equipment usage records, and the evolution trend of their research field, and predicts the needed equipment and consumables through a machine learning model.
[0051] As a preferred technical solution of the present application: in the marketing strategy recommender:
[0052] The collaborative filtering algorithm is used to discover the mixed similarity of the current researcher and other researcher nodes in the knowledge graph, the mixed similarity includes linear or nonlinear combination of the similarity based on the graph topology and the similarity based on the multi-dimensional attributes of the researchers, and then the product preferences of other researchers with high mixed similarity to the current researcher are discovered.
[0053] The knowledge graph reasoning is used to deduce potential product demand based on the entity association path between "technology-equipment-supplies".
[0054] As a preferred technical solution of the application, the cache management module comprises:
[0055] A cache key generator based on content hash;
[0056] A data storage supporting LRU and LFU hybrid eviction policy;
[0057] A retrieval engine accelerated by a Bloom filter;
[0058] The cache management module is a high-performance layered cache system that generates unique keys through content hash, optimizes memory using hybrid eviction policy, and prevents cache penetration with the help of a Bloom filter.
[0059] As a preferred technical solution of the application, the batch processing module comprises:
[0060] A task scheduler for dynamic resource allocation;
[0061] A parallel processor based on work-stealing algorithm;
[0062] A cross analyzer supporting multi-dimensional correlation analysis;
[0063] The batch processing module is a computing module that realizes efficient batch task processing through dynamic resource scheduling, work-stealing parallel processing and correlation cross analysis.
[0064] As a preferred technical solution of the application, it further comprises a Web application interactive interface connected to the cache management module to provide a user operation interface.
[0065] Compared with the prior art, the marketing system based on multi-modal literature data intelligent analysis has the following beneficial effects: the application constructs an integrated technical framework of multi-modal analysis, AI enhanced semantic mining, dynamic knowledge graph construction and strategy generation coupling, realizes end-to-end intelligent conversion from raw literature to executable marketing strategy.
[0066] Compared with the prior art, the technical scheme of the present application brings significant effects mainly in the following aspects:
[0067] 1. A multi-modal literature deep fusion preprocessing mechanism based on "three-level judgment-specialized parsing-structure unification" is proposed: through three-level judgment of file extension, magic number and content features, the literature format is accurately identified, and a specialized parser (such as a PDF parser with OCR, layout analysis and formula recognition capabilities) is called to perform deep content extraction. Subsequently, through a document structure parser, different formats of literature are mapped to a unified document tree structure representation, fundamentally solving the structured bottleneck of heterogeneous data sources, providing a high-quality, standardized data foundation for subsequent analysis, and significantly improving the complete parsing rate of mixed format literature.
[0068] 2. An AI-enhanced analysis paradigm of "semantic segmentation-domain hint-multi-model scheduling" is established. The paradigm first uses a semantic integrity protection algorithm to segment the text, ensuring that key information is not fragmented; then uses a hint library containing more than 200 domain-specific hint templates to guide large language models to perform fine-grained semantic mining on different parts of the literature (such as methods, results); finally, through an API connector supporting dynamic switching of multiple LLMs, both analysis performance and cost-effectiveness are considered. This method is particularly suitable for accurately extracting professional information such as experimental instruments, consumables, software tools from academic texts, and its extraction accuracy and recall rate have increased by more than 20% compared to traditional NLP methods.
[0069] 3. A precise marketing decision model of "five-dimensional portrait-knowledge graph-fusion recommendation" is constructed. The present invention breaks through the past single-dimensional user portrait technology and dynamically constructs a researcher knowledge graph containing five dimensions of basic identity, research interest, technology stack, device usage history and potential demand prediction based on extracted structured information. On this basis, the marketing strategy recommender uses a hybrid algorithm combining collaborative filtering and knowledge graph reasoning to not only recommend products directly related to research content, but also predict future demand trends of researchers, thereby realizing forward-looking marketing layout. This model improves the marketing precision from about 2.3% of traditional methods to more than 12.7%.
[0070] 4. A "cache-batch" linked system performance optimization system is designed. Through a content hash-based cache key generator, an LRU / LFU hybrid eviction strategy and a Bloom filter retrieval engine, a high-performance layered cache system is constructed, greatly reducing repeated calculations. At the same time, through a dynamic resource allocation task scheduler and a work-stealing parallel processor, efficient batch literature processing and cross-correlation analysis are realized, with a processing speed 3-4 times that of traditional methods, and the ability to discover cross-document research hotspots and technology associations.
[0071] In summary, the present application successfully solves a series of technical problems such as data fusion of multi-source heterogeneous scientific research literature, deep semantic understanding and accurate marketing strategy generation through the synergistic innovation and deep integration of the above technical links, and realizes the automatic and intelligent conversion from massive and disordered literature data to high-value and actionable marketing insights. BRIEF DESCRIPTION OF DRAWINGS
[0072] Figure 1 Structure diagram of the marketing system based on multi-modal literature data intelligent analysis of the present application Figure 1 ;
[0073] Figure 2 Structure diagram of the marketing system based on multi-modal literature data intelligent analysis of the present application Figure 2 . DETAILED DESCRIPTION
[0074] The present application will be further described in detail with reference to the drawings and specific examples.
[0075] The marketing system based on multi-modal literature data intelligent analysis of the present application comprises,
[0076] The literature analysis and preprocessing module identifies the format of the input literature through a three-level format recognition mechanism, and performs deep content extraction and structure analysis using a format-specific parser, and then unifies the internal representation through a document structure parser, a metadata extractor obtains key information, and a content partitioner divides logical parts, realizing efficient and structured preprocessing of heterogeneous scientific research literature, and passing the information to the AI enhanced analysis module through a standardized JSON data structure;
[0077] The AI enhanced analysis module receives the JSON data, extracts deep semantic information in the literature through a semantic block generator, a domain-specific prompt generator, a multi-LLM API connector and a response parser, and outputs structured data in JSON-LD format;
[0078] The data processing and integration module cleans, verifies and de-duplicates the JSON-LD data, and based on the processed data, constructs a researcher knowledge graph containing basic identity, research interest, technology stack, device usage and demand prediction through a user portrait generator, and generates personalized product combinations and marketing solutions using a marketing strategy recommender, realizing complete conversion from raw data to accurate marketing strategy, cleaning, verification and structured processing of extraction results;
[0079] The cache management module connects the data processing and integration module, and is used for storing and retrieving analysis results and user portraits based on content hash;
[0080] A batch processing module is connected to the literature analysis and preprocessing module, the AI-enhanced analysis module, and the cache management module to coordinate parallel processing and cross-relation analysis of multiple files.
[0081] The literature analysis and preprocessing module, the AI-enhanced analysis module, and the data processing and integration module are connected in series to form an end-to-end processing pipeline from raw literature to marketing strategy.
[0082] The literature analysis and preprocessing module includes:
[0083] A format identifier includes a three-level format identification mechanism based on file extension, secondary verification based on file header magic number, and final decision based on file content features.
[0084] A plurality of format-specific parsers includes an XML file reader, a PDF parser, a JSON parser, an HTML parser, a Word document parser, a TXT parser, and an EPUB parser.
[0085] A document structure parser generates a unified document tree structure representation.
[0086] A metadata extractor extracts the title, author, and DOI information of the literature.
[0087] A content partitioner divides the literature content according to academic structure.
[0088] The literature analysis and preprocessing module realizes the preprocessing of heterogeneous literature standardization through three-level format identification, multi-parser adaptation, unified structure reconstruction, metadata extraction, and academic content partitioning.
[0089] The PDF parser includes:
[0090] A text, image, and table separation unit based on PyMuPDF;
[0091] A scanned document processing unit integrated with Tesseract-OCR;
[0092] A layout analysis unit using computer vision algorithms to identify the chapter structure of the document;
[0093] A mathematical formula recognition and MathML conversion unit;
[0094] The PDF parser realizes high-precision content and structure extraction of native PDF and scanned PDF through the cooperative work of the above units.
[0095] The AI-enhanced analysis module includes:
[0096] A text partitioner using a segmentation algorithm with semantic integrity protection;
[0097] a prompt generator comprising a library of at least 200 domain-specific prompt templates;
[0098] an API connector supporting dynamic switching of multiple large language model APIs;
[0099] a response parser for converting unstructured responses into JSON-LD format;
[0100] The AI-enhanced analysis module implements domain knowledge enhancement and standardized analysis through semantic chunking, domain-specific prompt templates, dynamic scheduling of multiple model APIs, and JSON-LD structured output.
[0101] The data processing and integration module includes:
[0102] a data cleaner implementing rule-based and machine learning-based noise filtering;
[0103] a user portrait generator constructing a five-dimensional portrait including basic information, research interests, technology stack, device usage, and demand prediction;
[0104] a marketing strategy recommender using a hybrid recommendation model combining collaborative filtering algorithms and knowledge graph reasoning;
[0105] The data processing and integration module implements precise marketing strategies through data cleaning, five-dimensional user portrait construction, and collaborative filtering and knowledge graph fusion recommendation.
[0106] The five-dimensional researcher knowledge graph constructed by the user portrait generator includes:
[0107] The technology stack dimension includes the experimental methods, data analysis software, and programming languages commonly used by researchers;
[0108] The device usage dimension includes the equipment models, brands, and usage scenarios mentioned in the literature;
[0109] The demand prediction dimension is based on the researcher's historical technology stack, device usage records, and the evolution trend of their research field, and predicts the equipment and consumables they may need in the future through machine learning models.
[0110] In the marketing strategy recommender:
[0111] The collaborative filtering algorithm is used to calculate the hybrid similarity between the current researcher node and other researcher nodes based on the researcher knowledge graph; the hybrid similarity is a linear or nonlinear combination of a) similarity based on graph topology (such as based on common neighbors) and b) similarity based on researcher multi-dimensional attributes (such as based on research interests, technology stack); and further discover the product preferences of other researchers with high hybrid similarity to the current researcher.
[0112] The knowledge graph reasoning is used to deduce potential product demand based on entity association paths among "technology-equipment-consumables".
[0113] A content hash-based cache key generator is used to generate a globally unique identifier for the analysis result;
[0114] A data storage supporting LRU and LFU hybrid eviction policy is used to optimize memory usage;
[0115] A bloom filter accelerated search engine is used to quickly determine whether the request result does not exist in the cache to prevent cache penetration;
[0116] The cache management module uses a high-performance layered cache system with unique key generation by content hash, memory optimization by hybrid eviction policy, and cache penetration prevention by bloom filter.
[0117] The batch processing module includes:
[0118] A dynamic resource allocation task scheduler dynamically allocates computing resources according to file size, parsing complexity, and current system load;
[0119] A work-stealing algorithm-based parallel processor achieves load balancing of multiple literature parsing and AI analysis tasks;
[0120] A cross analyzer supporting multi-dimensional correlation analysis performs correlation analysis on multiple literature results after batch processing to identify common research teams, technology hotspots, and equipment usage combination patterns;
[0121] The batch processing module uses dynamic resource scheduling, work-stealing parallel processing, and correlation cross-analysis to achieve a high-performance batch task processing computing module.
[0122] It also includes a Web application interaction interface connected to the cache management module, providing user interfaces for uploading literature, configuring parameters, visualizing analysis results, and marketing strategies.
[0123] Compared with the prior art, the present application has the following beneficial effects:
[0124] 1. Unified multi-format literature processing framework: a unified processing framework based on format recognition and special parsers is created, supporting 10 scientific literature formats including XML, PDF, JSON, HTML, Word, EPUB, etc., breaking through the limitations of traditional single format processing, and greatly expanding the comprehensiveness of data sources.
[0125] 2. AI-enhanced deep information extraction technology: Developed an enhanced analysis technology based on large language models, through specially designed prompt engineering strategies, fine-grained analysis of different parts of the literature, significantly improving the extraction accuracy (more than 20%) and recall rate (more than 25%) of professional information such as experimental instruments and consumables.
[0126] 3. Multi-dimensional scientific portrait construction method: Based on the extracted information, a multi-dimensional scientific portrait is constructed, including research direction, equipment used, experimental consumables, software tools, etc., which improves the marketing accuracy from 2.3% to 12.7%, an increase of 5.5 times.
[0127] 4. Efficient parallel batch processing and cross-analysis technology: Designed a dynamic resource allocation parallel processing architecture and multi-document cross-analysis algorithm, processing speed increased by 3-4 times compared to traditional methods, while discovering deep connections between documents, identifying research hotspots and core research teams.
[0128] Example 1
[0129] As Figures 1-2 shown, the marketing system based on multi-modal literature data intelligent analysis of the present application provides a multi-format scientific literature intelligent analysis system based on large language models. The system develops a complete solution to address the format limitations, inaccurate information extraction, and insufficient marketing accuracy in existing scientific literature processing. The system includes a literature analysis and preprocessing module 1, an AI-enhanced analysis module 2, a data processing and integration module 3, a cache management module 4, a Web application interaction interface 5, and a batch processing module 6.
[0130] As Figure 1 shown, the literature analysis and preprocessing module 1 and the AI-enhanced analysis module 2 exchange information through standardized JSON data structures, the AI-enhanced analysis module 2 and the data processing and integration module 3 exchange analysis results through structured object arrays, the data processing and integration module 3 and the cache management module 4 access data through key-value pair mapping, the cache management module 4 and the Web application interaction interface 5 exchange query results through RESTful API interfaces, the Web application interaction interface 5 and the batch processing module 6 manage batch requests through task queue objects, and the batch processing module 6 and the literature analysis and preprocessing module 1 exchange batch processing tasks through file path arrays and configuration objects. The batch processing module 6 also maintains feedback connections with the AI-enhanced analysis module 2 and the cache management module 4, forming a complete data processing closed loop. The data transmission between modules adopts standardized interface definition, ensuring high scalability and modular characteristics of the system. In the architecture diagram, solid arrows represent the main data flow, and dashed arrows represent feedback or control flow.
[0131] The document analysis and preprocessing module 1 includes the following components: a format identifier 10, an XML file reader 11, a PDF parser 11A, a JSON parser 11B, an HTML parser 11C, a Word document parser 11D, a TXT parser 11E, an EPUB parser 11F, a document structure parser 12, a metadata extractor 13, and a content partitioner 14.
[0132] The format identifier 10 is responsible for identifying the format type of the input document. It uses multiple levels of judgment based on file extensions, magic numbers, and content characteristics to accurately identify the file format and call the corresponding parser according to different formats. In specific implementation, the format identification uses a decision tree algorithm. First, it checks the file extension, then verifies the file header magic number, and finally analyzes the content characteristics. The three-level judgment ensures that the format can be correctly identified even if the file extension is modified. The recognition accuracy reaches 99.8%, which is significantly higher than the traditional single feature recognition method of 92.3%;
[0133] The XML file reader 11 is responsible for reading the input XML file. It supports the XML namespaces and tag structures of major publishers such as PubMed, Elsevier, and Springer. It uses both SAX and DOM parsing methods and automatically selects the optimal parsing strategy based on file size.
[0134] The PDF parser 11A is responsible for processing PDF format documents. It includes the following core functions: (1) content extraction based on PyMuPDF, supporting text, image, and table separation; (2) using Tesseract-OCR for scanned document text recognition, supporting multi-language recognition; (3) layout analysis algorithm based on machine learning, identifying the secondary structure of the article; (4) automatic identification of literature references and charts; (5) mathematical formula OCR and structured storage;
[0135] The JSON parser 11B is responsible for parsing JSON format scientific research data. It is suitable for API data provided by arXiv, PubMed Central, etc., including: (1) recursively parsing complex nested structures; (2) processing various Unicode encodings and special characters; (3) automatically identifying and converting different formats of dates, numerical values, and other fields; (4) handling incomplete or incorrect JSON structures;
[0136] The HTML parser 11C processes web page format scientific literature, supporting: (1) using Beautiful Soup and lxml for DOM parsing; (2) extracting structured content based on XPath and CSS selectors; (3) processing JavaScript-rendered dynamic content; (4) automatically identifying and extracting tables, references, and references in HTML;
[0137] The Word document parser 11D processes DOC / DOCX format scientific documents, with functions including: (1) extracting text content and format information using python-docx and antiword; (2) identifying and preserving the style, structure, and hierarchical headings of the document; (3) extracting embedded charts and formulas; (4) processing annotations and revisions in the document;
[0138] The TXT parser 11E processes plain text format scientific materials, with functions including: (1) automatically identifying text encoding; (2) identifying document structure based on natural language processing technology; (3) distinguishing between abstract, method, result, etc. parts using semantic analysis technology; (4) automatically identifying citation and reference format;
[0139] The EPUB parser 11F processes electronic book format scientific materials, with functions including: (1) parsing EPUB container structure; (2) extracting chapter content and metadata; (3) processing embedded HTML / XHTML content; (4) extracting and associating image resources;
[0140] The document structure parser 12 parses the document tree structure based on the characteristics of various documents, adopts a unified internal representation form, and realizes structure mapping between different formats; the metadata extractor 13 extracts article title, DOI, publication date, etc. metadata; the content partitioner 14 divides the article content into different parts such as abstract, method, result, etc.
[0141] The AI-enhanced analysis module 2 includes the following components: text chunker 21, prompt generator 22, API connector 23, response parser 24. The text chunker 21 divides long text into appropriate size processing blocks; the prompt generator 22 generates targeted AI prompts according to different text types; the API connector 23 is responsible for communication with large language models (Gemini API); the response parser 24 converts AI returned text into structured data.
[0142] The data processing and integration module 3 includes the following components: data cleaner 31, verifier 32, deduplication merger 33, user portrait generator 34, and marketing strategy recommender 35. The data cleaner 31 is responsible for removing noise and redundant information in the analysis results; the verifier 32 ensures the accuracy of the extracted information through cross-validation; the deduplication merger 33 merges analysis results from multiple sources into a unified data structure; the user portrait generator 34 constructs a multi-dimensional researcher portrait based on the extracted information; the marketing strategy recommender 35 generates personalized marketing suggestions based on the user portrait.
[0143] The user portrait generator 34 is one of the core components of the system, responsible for converting various types of information extracted from the literature into structured researcher portraits. This component uses a multi-level label system and knowledge graph technology to build a comprehensive portrait containing the following dimensions: (1) Basic identity dimension: including basic information such as researcher name, affiliated institution, title, research field, etc.; (2) Research interest dimension: by analyzing the literature topics, citation patterns and collaboration networks of the researcher, identify their core research interests and development trajectory; (3) Technology application dimension: based on the experimental methods, technical routes and data processing methods mentioned in the literature, build a technology preference model of the researcher; (4) Equipment and consumables dimension: by extracting experimental equipment, reagents and consumables information from the literature, establish the equipment usage portrait of the researcher; (5) Demand prediction dimension: combined with historical research trajectory and field development trend, predict the potential equipment and consumables demand of the researcher. The user portrait generator 34 realizes the conversion from static text to dynamic user model, providing data basis for precision marketing.
[0144] The marketing strategy recommender 35 is based on the researcher portrait built by the user portrait generator 34, combined with the product library and marketing knowledge base, to automatically generate personalized marketing strategy suggestions. This component uses a hybrid recommendation algorithm based on rules and machine learning, according to the researcher's professional background, research stage and equipment demand characteristics, to recommend the most suitable product combination and marketing method. The recommendation system considers multiple factors, including: product and research demand matching degree, researcher's budget level, institutional procurement cycle, past marketing response rate, etc., after comprehensive scoring to generate the optimal recommendation scheme. In addition, this component also provides A / B testing function, which can generate multiple marketing strategy schemes at the same time, and then apply them on a large scale after small-scale testing to verify the effect.
[0145] The cache management module 4 includes the following components: cache key generator 41, data storage 42, retrieval engine 43, cache cleaner 44. The cache key generator 41 generates a unique identifier based on file content and parameters; the data storage 42 saves the analysis results to the local file system; the retrieval engine 43 quickly retrieves existing results according to the cache key; the cache cleaner 44 is responsible for cleaning expired or invalid cache data.
[0146] The Web application interaction interface 5 includes the following components: file uploader 51, parameter configurator 52, progress monitor 53, result visualizer 54, history manager 55. The file uploader 51 supports single or batch uploading of scientific literature files in multiple formats, including XML, PDF, JSON, etc.; the parameter configurator 52 allows users to select the type of information to extract; the progress monitor 53 displays the analysis progress in real time; the result visualizer 54 displays the analysis results in a structured manner; the history manager 55 manages past analysis records and results.
[0147] The batch processing module 6 includes the following components: task scheduler 61, parallel processor 62, result aggregator 63, cross analyzer 64. The task scheduler 61 manages the processing order of multiple files; the parallel processor 62 simultaneously processes multiple scientific literature files in different formats; the result aggregator 63 combines multiple analysis results; the cross analyzer 64 identifies common elements and relationships in multiple files.
[0148] The system initialization process is as follows: when the system starts, first load the configuration file, including module parameter settings, API keys, cache strategies, etc.; then initialize each parser component, preload necessary models and dictionaries; then establish inter-module communication pipelines; finally start the Web service and wait for user requests. The initialization stage also includes environment detection, verifying the availability of all dependent libraries and external services, and automatically downgrading to a backup solution if unavailable. The system uses a lazy loading strategy, loading parsers for specific formats only when needed, to optimize resource usage and startup speed.
[0149] System security and privacy protection measures are as follows: (1) User authentication and authorization: implement role-based access control (RBAC) to ensure that users can only access authorized functions and data; (2) Data encryption: all stored analysis results and user information are encrypted using AES-256, and the transmission process uses TLS 1.3 protocol to ensure data security; (3) Privacy protection: provide data desensitization options, which can partially mask sensitive information such as email addresses when extracting; (4) Audit logs: record all system operations, including user access, file processing, and data export behaviors; (5) Data retention policy: users can set automatic data deletion deadlines, and the system regularly cleans up expired data; (6) Compliance checks: built-in compliance checks for privacy laws such as GDPR and CCPA to ensure that marketing activities comply with legal requirements.
[0150] The workflow of the present application is as follows: first, the user uploads one or more scientific literature files (supports multiple formats such as XML, PDF, JSON, HTML, Word, TXT, EPUB, etc.) through the file uploader 51 of the Web application interaction interface 5, and selects the information type (such as email, experimental instruments, experimental consumables, etc.) to be extracted through the parameter configurator 52.
[0151] Then, if it is single-file analysis, the document parsing and preprocessing module 1 first judges the document type by the format identifier 10. For XML files, the XML file reader 11 is called for processing; for PDF files, the PDF parser 11A is called for OCR processing and structure extraction; for JSON files, the JSON parser 11B is called for parsing; for HTML files, the HTML parser 11C is called for processing; for Word documents, the Word document parser 11D is called for processing; for TXT files, the TXT parser 11E is called for processing; for EPUB files, the EPUB parser 11F is called for processing. Then, the document structure parser 12 parses the document structure based on the respective format characteristics, the metadata extractor 13 extracts the article metadata, and the content partitioner 14 divides the text into different parts. The parsed structured data is uniformly passed to the AI enhanced analysis module 2.
[0152] If it is multi-file analysis, the task scheduler 61 of the batch processing module 6 arranges the processing order, and the parallel processor 62 simultaneously processes multiple files of different formats. Each file is first judged by the format identifier 10, and then the corresponding parser is called for processing. The subsequent process is the same as single-file analysis. This design enables the system to efficiently process batch tasks composed of multiple documents of different formats. In specific implementation, the task scheduler 61 uses a priority queue and a resource-aware algorithm to dynamically adjust the processing order and resource allocation according to file size, complexity, and current system load, maximizing throughput while ensuring processing quality. The parallel processor 62 implements a load balancing strategy based on work-stealing, ensuring that computing resources are fully utilized and avoiding the situation where processors are idle while tasks are stacked.
[0153] In the AI-enhanced analysis module 2, the text chunker 21 divides the text into blocks of appropriate size, the prompt generator 22 generates specialized prompts based on the text type (such as the method section, the results section), the API connector 23 sends the prompts and the text to the large language model, and the response parser 24 parses the returned results into structured data. In technical implementation, the text chunking uses an adaptive chunking algorithm based on semantic integrity, which not only considers the length limit of the text, but also ensures the semantic integrity of each block, avoiding the fragmentation of key information. The prompt generator 22 uses a template library and dynamic combination technology to generate highly specialized prompts for different literature sections and extraction targets. For example, when extracting experimental instruments in the method section, the prompt contains "identify all instruments used in the experimental steps, pay attention to extracting the brand, model, parameter settings and purpose of use" and other professional guidance; when extracting software tools in the results section, the prompt contains "identify the software packages used for data analysis and visualization, including version number and specific parameters" and other targeted content. This specialized prompt engineering improves the extraction accuracy by 18.7% compared to general prompts.
[0154] Subsequently, the data cleaning 31 of the data processing and integration module 3 removes invalid data, the verifier 32 verifies the data format, and the de-duplicator 33 removes duplicates. In actual implementation, the data cleaning 31 uses a combination of rule-based and machine learning methods to identify and handle various abnormal data, including incomplete information, format errors and outliers. The verifier 32 applies specialized verification rules for different types of data (such as email, instrument name, chemical substance) to ensure data validity. The de-duplicator 33 not only identifies identical items based on string matching, but also identifies information expressed differently but referring to the same entity (such as "PCR instrument" and "polymerase chain reaction instrument") through fuzzy matching and semantic similarity analysis, with a merging accuracy of 95.3%.
[0155] Next, the user portrait generator 34 converts the cleaned and verified structured data into a multi-dimensional researcher portrait. First, the system establishes a basic identity tag based on the extracted author information and institutional information; second, it constructs a research interest portrait by analyzing the literature theme, keywords and citation patterns; then, it forms a technology application preference based on the extracted experimental methods and technology routes; next, it establishes a device usage portrait by identifying experimental equipment and consumables mentioned in the literature; finally, it generates a demand prediction model combined with historical data and field trends. The entire portrait construction process uses incremental learning, continuously improving and updating the researcher portrait as more literature is analyzed. When processing multiple papers of the same researcher, the system can automatically merge and update the portrait information, maintaining the timeliness and completeness of the portrait.
[0156] The marketing strategy recommender 35 generates personalized marketing strategies based on user profiles, combined with enterprise product libraries and marketing knowledge bases. The recommendation process is divided into three stages: product matching, channel selection, and content customization. In the product matching stage, the system filters the most suitable product combinations from the product library based on researchers' technical needs and device usage habits; in the channel selection stage, the best contact channel (such as email, academic conferences, professional seminars, etc.) is determined based on researchers' contact method preferences and response history; in the content customization stage, the system generates specialized marketing copy and technical materials based on researchers' professional backgrounds and research stages. The entire recommendation process uses an A / B testing mechanism to continuously optimize the recommendation effect, and the marketing conversion rate is improved by 37.8% compared with traditional methods.
[0157] The processed results are processed by the cache management module 4, the cache key generator 41 generates a unique identifier, and the data storage 42 saves the results to the local cache for future quick retrieval. The cache system adopts a hierarchical storage strategy, with hot data saved in memory and cold data stored on disk, while implementing an adaptive expiration policy that dynamically adjusts cache retention time based on data access frequency and recent usage time. This design shortens the response time by 87% in a large number of repeated query scenarios, significantly improving user experience.
[0158] Finally, the analysis results are returned to the Web application interaction interface 5, which is displayed in a user-friendly manner by the result visualizer 54. Users can view detailed results, download result files, or save them to the history. The interface design uses responsive layout and modular components to support consistent experience on different devices. The result visualizer 54 implements multiple visualization methods, including tables, charts, network relationship graphs, and heat maps, allowing users to understand analysis results from different perspectives. The system also provides interactive filtering and sorting functions, allowing users to adjust the result display as needed.
[0159] If it is batch processing, the result aggregator 63 of the batch processing module 6 combines the results of multiple files, and the cross-analyzer 64 analyzes the common elements between multiple files to generate a comprehensive report. The cross-analyzer 64 implements various advanced analysis algorithms, including graph-based researcher relationship network analysis, topic model-based research hotspot identification, time series-based research trend analysis, and association rule-based device-spare part relationship mining. These algorithms can discover deep associations from the analysis results of multiple articles, such as identifying that "research teams using a specific brand of PCR instrument tend to also use the same brand of DNA extraction reagents," etc. valuable marketing insights. The batch processing module 6 also feeds back analysis results and performance data to the AI-enhanced analysis module 2 and the cache management module 4 through the feedback connection, to optimize the prompt strategy and cache management strategy.
[0160] The system also implements complete error handling and exception handling mechanisms. In the literature analysis stage, the system can detect and handle damaged files, non-standard formats and incomplete content, and extract valid information as much as possible through degradation processing strategies. In the AI analysis stage, the system implements a request retry mechanism and model degradation strategy to ensure that the analysis task can be completed even if the API connection is unstable or the response is timed out. In the data processing stage, the system marks and records all abnormal data and generates a detailed quality evaluation report to enable users to understand the reliability of the results. These mechanisms enable the system to exhibit high robustness in actual application environments, and even if 20% of the input documents have problems, the system can still maintain an overall success rate of more than 85%.
[0161] The core advantage of the system is that it supports the unified processing of scientific literature in multiple formats, including XML, PDF, JSON, etc., greatly expanding the data source range; automatically extracts key information from scientific literature through large language models, including but not limited to author email, experimental instruments, experimental consumables, chemicals, software tools, databases, and statistical methods, greatly reducing the workload of manual extraction; at the same time, modular design makes the system have good scalability, can add new literature format parser and extraction type according to needs; format automatic identification function can intelligently judge the type of literature and call the corresponding processing flow, improve the adaptability and user experience of the system; cache mechanism improves the efficiency of the system, avoids repeated analysis; the design of the Web interface is intuitive and easy to use, providing rich interactive functions and result display methods.
[0162] While achieving the basic functions of the application, other technical paths can also be used to achieve similar results. For example, for the document format processing strategy, in addition to the special parser used by the application to process different formats of documents, a format conversion scheme can also be used, that is, all non-XML format documents (such as PDF, JSON, etc.) are first converted into XML format before subsequent processing. This approach can simplify the subsequent processing process, but may lose some format-specific information during conversion. When processing training data, you can choose to extract important parts of the document for detailed analysis, or directly hand over the complete document to the AI model for overall processing. For literature data sources, in addition to relying on public literature databases, you can also use local search document methods, use key data regularization matching technology to build your own data set, or scrape relevant scientific research information from university websites to build a database. In terms of batch processing mode, the system can use queue processing instead of parallel processing, processing files in priority order. This approach is particularly suitable for resource-limited system environments. In addition, in terms of AI model selection, in addition to using Gemini API, the system can also be configured to use other large language models such as OpenAIGPT series, Claude series or open source models such as LLaMA, etc. to adapt to different scene requirements and cost considerations. These alternative solutions may have different implementation paths, but they can all meet the basic requirements of the application to some extent.
[0163] The marketing system based on multi-modal literature data intelligent analysis of the application successfully solves the data extraction problem of multi-source heterogeneous scientific research literature by building a unified multi-format literature processing framework. The framework uses a modular parser matrix to support XML / PDF / JSON and other 10 formats and a three-layer format recognition system file extension name → magic number verification → content feature analysis, combined with decision tree algorithm to ensure format recognition accuracy of 99.8%. Through the document structure parser, a general document tree model is established to realize structure mapping between different formats and standardized JSON output, and a fault-tolerant processing mechanism is built in. The system has high complete parsing rate when processing mixed format documents, greatly expanding the comprehensiveness of data sources; the application system develops AI model enhanced multi-modal analysis technology, solving the problem of low precision of deep information extraction; the application constructs a five-dimensional portrait modeling system and knowledge graph enhancement technology, solving the problem of insufficient depth of scientific research portrait, realizing synonym merging, technology evolution analysis and demand prediction model.
[0164] The marketing system based on intelligent analysis of multi-modal literature data of the application extracts semantics through analyzing heterogeneous literature and AI enhanced analysis, constructs a researcher knowledge graph and generates a precise marketing strategy, and finally a multi-modal intelligent marketing system is obtained through cache and batch processing to optimize performance, the application realizes deep mining and precise marketing of multi-source scientific research literature through unified multi-format processing framework, AI enhanced information extraction, multi-dimensional scientific research portrait and efficient parallel analysis technology, and significantly improves the information extraction accuracy and marketing conversion efficiency.
[0165] The above specific embodiments are used to explain and illustrate the present application, and are only preferred embodiments of the present application, not to limit the present application, any modification, equivalent replacement, improvement, etc. of the present application within the spirit and protection scope of the claims of the present application, fall into the protection scope of the present application.
Claims
1. A marketing system based on intelligent analysis of multi-modal document data, characterized in that: Comprising, The document analysis and preprocessing module uses a format recognizer to identify the format of the input document and uses a format-specific parser for deep content extraction and structural analysis, and then generates a unified document tree structure representation through a document structure parser. The metadata extractor obtains key information, and the content partitioner divides the logical part to achieve efficient and structured preprocessing of heterogeneous scientific literature. The information is transmitted to the AI-enhanced analysis module through the standardized JSON data structure output; The AI-enhanced analysis module receives the JSON data and extracts deep semantic information from the literature through a semantic chunker, a domain-specific prompt generator, a multi-LLMAPI connector, and a response parser, and outputs the structured data in JSON-LD format; The data processing and integration module cleans, verifies, and de-duplicates the structured data in JSON-LD format, and based on the processed data, constructs a researcher knowledge graph containing basic identity, research interest, technology stack, device usage, and demand prediction through a user portrait generator, and generates personalized product combinations and marketing strategies using a marketing strategy recommender, achieving complete conversion from raw data to precise marketing strategies. Cleaning, verification, and structured processing of extraction results; The cache management module connects the data processing and integration module to store and retrieve analysis results and user portraits based on content hash; The batch processing module connects the document analysis and preprocessing module, AI-enhanced analysis module, and cache management module to coordinate parallel processing and cross-correlation analysis of multiple files; The document analysis and preprocessing module, AI-enhanced analysis module, and data processing and integration module are connected in series to form an end-to-end processing pipeline from raw literature to marketing strategies.
2. The marketing system based on intelligent analysis of multi-modal document data as claimed in claim 1, wherein: The document analysis and preprocessing module includes: The format recognizer includes a three-level format recognition mechanism based on file extension, secondary verification based on file header magic number, and final decision based on file content characteristics; The format-specific parser includes an XML file reader, a PDF parser, a JSON parser, an HTML parser, a Word document parser, a TXT parser, and an EPUB parser; The document structure parser generates a unified document tree structure representation; The metadata extractor extracts the title, author, and DOI information of the document; The content partitioner divides the document content according to the academic structure; The document analysis and preprocessing module implements standardized processing of heterogeneous literature through three-level format recognition, multi-parser adaptation, unified structure reconstruction, metadata extraction, and academic content partitioning. 3.The marketing system based on intelligent analysis of multi-modal document data according to claim 2, wherein, The PDF parser includes: A text, image, and table separation unit based on PyMuPDF; A scanned document processing unit integrated with Tesseract-OCR; A layout analysis unit using computer vision algorithms to identify the chapter structure of the document; A mathematical formula recognition and MathML conversion unit; The PDF parser achieves high-precision content and structure extraction of native PDF and scanned PDF through the coordinated work of the above units. 4.The marketing system based on intelligent analysis of multi-modal document data according to claim 1, wherein, The AI-enhanced analysis module includes: A text chunker with a semantic integrity-protected segmentation algorithm; A prompt generator containing a library of at least 200 domain-specific prompt templates; An API connector supporting dynamic switching of multiple large language model APIs; A response parser for converting unstructured responses into JSON-LD format; The AI-enhanced analysis module implements domain knowledge enhancement and standardized analysis through semantic chunking, specialized prompt templates, dynamic scheduling of multiple model APIs, and JSON-LD structured output. 5.The marketing system based on intelligent analysis of multi-modal document data according to claim 1, wherein, The data processing and integration module includes: A data cleaner that implements rule-based and machine learning-based noise filtering; A user portrait generator that constructs a five-dimensional portrait containing basic information, research interests, technology stack, device usage, and demand prediction; A marketing strategy recommender that uses a hybrid recommendation model combining collaborative filtering algorithms and knowledge graph reasoning; The data processing and integration module achieves precise marketing strategies through data cleaning, five-dimensional user portrait construction, and collaborative filtering and knowledge graph fusion recommendations. 6.The marketing system based on intelligent analysis of multi-modal document data according to claim 5, wherein, In the five-dimensional researcher knowledge graph constructed by the user portrait generator: The technology stack dimension includes the researchers' experimental methods, data analysis software, and programming languages; The device usage dimension includes the equipment models, brands, and usage scenarios mentioned in the literature; The demand prediction dimension is based on the researcher's historical technology stack, device usage records, and the evolution trend of their research field, and predicts the required equipment and consumables through a machine learning model. 7.The marketing system based on intelligent analysis of multi-modal document data according to claim 5, wherein, In the marketing strategy recommender: The collaborative filtering algorithm is used to discover similarities with researchers in the knowledge graph, and to calculate the hybrid similarity between the current researcher and other researcher nodes, which includes linear or nonlinear combinations of similarity based on graph topology and similarity based on researcher multi-dimensional attributes; The knowledge graph reasoning is used to infer potential product demand based on the entity association paths between technology, equipment, and consumables.
8. The marketing system based on intelligent analysis of multi-modal document data as claimed in claim 1 wherein, The cache management module includes: A content hash-based cache key generator for generating globally unique identifiers for analysis results; A data storage with LRU and LFU hybrid eviction policy to optimize memory usage; A bloom filter-accelerated search engine to determine whether the requested results do not exist in the cache to prevent cache penetration; The cache management module uses a high-performance hierarchical cache system that generates unique keys based on content hashes, optimizes memory using a hybrid eviction policy, and prevents cache penetration with the help of a bloom filter. 9.The marketing system based on intelligent analysis of multi-modal document data according to claim 1, wherein, The batch processing module includes: A dynamic resource allocation task scheduler that dynamically allocates computing resources based on file size, parsing complexity, and current system load; A parallel processor based on the work-stealing algorithm to achieve load balancing for multiple literature parsing and AI analysis tasks; A cross-analyzer that supports multi-dimensional correlation analysis to perform correlation analysis on multiple literature results after batch processing to identify common research teams, technology hotspots, and equipment usage combination patterns; The batch processing module implements a high-performance batch task processing computing module through dynamic resource scheduling, work-stealing parallel processing, and cross-correlation analysis.
10. The marketing system based on intelligent analysis of multi-modal document data as claimed in claim 1 wherein: The application further comprises a web application interactive interface connected to the cache management module, which provides an operation interface for users to upload documents, configure parameters, visualize analysis results and marketing strategies.
Citation Information
Patent Citations
Multi-dimensional fusion meta universe and vertical AI model collaborative innovation platform
CN119443116A
AI-driven digital publication content and online derivative resource performance prediction system
CN120579997A
Marketing strategy optimization method based on consumer group behavior analysis
CN120725209A
Text Mining Analysis and Output System
US20130144605A1
Continuously learning and optimizing artificial intelligence (AI) adaptive neural network (ANN) computer modeling methods and systems
US20210034959A1
Cited By
Multi-modal document data processing method and system oriented to large language model training
CN121093293A
Multimodal document data processing methods and systems for training large language models
CN121093293B