A label-based open source data dynamic organization and intelligent retrieval method, device and system
Patent Information
- Application Number
- CN202611255984.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-18
AI Technical Summary
[0003]本发明的目的在于克服现有技术的不足,提供一种基于标签的开源数据动态组织与智能检索方法、装置及系统,可以解决多源异构数据管理中长期存在的存储割裂、标签精度与广度失衡以及检索效率低下等问题
(1)在存储层面:本发明通过精细化的分层解耦设计,能够有效降低数据冗余率,节约存储成本。更重要的是,新增纳管一类开源数据的扩展时间将从小时级压缩至分钟级,实现近实时接入新兴数据源(如突发热点事件的社交平台数据),彻底摆脱了因表结构僵化导致的数据纳管滞后问题。
Smart Images

Figure CN122777589A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, and more specifically, to a tag-based open-source data dynamic organization and intelligent retrieval method, apparatus and system. Background Technology
[0002] The global open-source data scale is growing exponentially, with the daily volume of multi-source heterogeneous data reaching several petabytes (covering all scenarios such as social media, news media, and forums). Existing open-source data management systems are facing structural bottlenecks in data storage, tag generation, data organization, and data retrieval: In terms of data storage, they adhere to the traditional monolithic model, and the hard-coded table structure leads to delays in the management of new data sources that are fixed at the hour level, significantly deviating from industry best practices. This results in high latency rates for data ingestion during sudden hot events, directly causing the loss of real-time analysis windows; Tag construction relies on pure rule engines, resulting in high conflict rates and low coverage, and an inability to accurately distinguish semantic overlap. Stacked entities (such as a "large model" simultaneously triggering "AI" and "cloud computing" tags) force manual intervention, leading to decreased processing efficiency. The static directory path strategy used in data organization causes directory depth to expand exponentially as open-source data grows dynamically, requiring data administrators to invest significant time in maintaining directory paths, drastically reducing data organization efficiency. Furthermore, the lack of dynamic view adaptation for diverse and heterogeneous open-source data results in a degraded user experience. Regarding data querying, the fragmented storage of multi-source heterogeneous data and the real-time parsing of complex query conditions during multi-dimensional tag combination queries prolong query response time and reduce QPS throughput. These problems are not isolated but interdependent and cumulative. Storage lag exacerbates tag errors, tag errors worsen organizational efficiency, and organizational rigidity negatively impacts query performance, ultimately causing the system to reach a bottleneck in all aspects of elastic scalability, timeliness, and availability, severely restricting the management efficiency and usability of open-source data. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a tag-based open-source data dynamic organization and intelligent retrieval method, device, and system. This addresses long-standing problems in multi-source heterogeneous data management, such as storage fragmentation, imbalance between tag accuracy and breadth, and low retrieval efficiency. By constructing a dynamic hierarchical data storage architecture, it achieves data decoupling and flexible expansion. It innovatively integrates a rule-driven and deep learning-based dual-track tag generation mechanism to overcome the accuracy-breadth paradox. Based on dynamic directory trees and intelligent index pre-computation technology, it achieves the core goals of highly customizable data organization, comprehensive and accurate tag coverage, and rapid user retrieval response. This completely eliminates data integration barriers under traditional static directory and flat storage models, comprehensively improving the efficiency and effectiveness of open-source data management.
[0004] The objective of this invention is achieved through the following solution: A tag-based open-source method for dynamic data organization and intelligent retrieval includes: Step 1, Data Storage: Decouple data storage using a layered architecture; Step 2, Data Processing: Deploy an end-to-end real-time intelligent processing pipeline for data topic extraction, translation, and keyword generation; Step 3, Label Construction: A dual-track fusion mechanism of rules and models is adopted. Rule labeling is based on a dynamic rule base and generates accurate labels by matching keywords through a regular expression engine to ensure high confidence output. Model labeling relies on a classifier fine-tuned based on historical labeled data to output labels with probability weights. By fusing the results of the two, a label set with confidence is constructed. Among them, when the confidence of rule matching is greater than a set value, rule labels are used first; otherwise, model labels are used. Step 4, Data Organization: Adopt a dynamic rule-driven data organization method to dynamically organize personalized data views and data labels; Step 5, Data Retrieval: Through the parsing of multi-dimensional tag query conditions and index pre-calculation, intelligent and efficient interactive data query is achieved.
[0005] Furthermore, the adoption of a layered architecture to decouple data storage specifically includes: A hierarchical storage architecture is set up, comprising a master data table, a text table, a translation table, an attachment table, and a personalization table. The master data table serves as the global hub, storing common data attributes to ensure the consistency of basic metadata. The text table independently carries the original text content and natively saves multilingual input using UTF-8 encoding. The translation table triggers a multilingual translation engine through an automatic language detection module, translating the title and text separately and storing them as independent fields. The attachment table establishes a hash value index mechanism to associate with binary attachment files. The personalization table is dynamically created according to data type, storing personalized fields for different data, so that when adding new data types, only the personalization table needs to be expanded without modifying the master table.
[0006] Furthermore, the deployment of the end-to-end real-time intelligent processing pipeline, used for data topic term extraction, translation, and keyword generation, specifically includes the following sub-steps: 1) Topic word extraction and noise suppression: In the Apache Flink real-time computing engine, a semantic constraint matrix is constructed by fusing the domain dictionary, and topic scores are calculated by combining the TF-IDF weighted algorithm to output high-frequency topic words. The noise suppression mechanism is used to reduce the interference rate and eliminate irrelevant words from interfering with topic recognition. 2) Independent translation processing for titles and body text: The title translation module calls the locally deployed mBART-50 model to output a semantically complete title verified by BLEU-4 scoring; the body text translation calls a finely tuned mT5-large model, and after translation, entity recognition is used to extract personal names and place names to ensure consistency of technical terms and contextual coherence. 3) Keyword structure generation: High-frequency entities are extracted from the topic word set, and target knowledge is associated through entity linking algorithms to generate a semantically related structured keyword library, which provides input for the subsequent tag system, while recording error samples for model iteration and optimization.
[0007] Furthermore, the adoption of a dynamic rule-driven data organization method for the dynamic organization of personalized data views and data tags specifically includes the following sub-steps: First, a non-fixed directory tree structure that can evolve in real time is generated by configuring tag combination logic, so that directory nodes can dynamically expand with data. At the same time, a one-to-many relationship model between tags and personalized data views is established with the unique data ID as the hub. After row and column transposition, the scattered tag relationships are aggregated into a wide table of tag data. Finally, the wide table is imported into the data search engine in real time through an incremental synchronization mechanism.
[0008] Furthermore, the intelligent and efficient interaction of data query through the parsing of multi-dimensional label query conditions and index pre-calculation specifically includes the following sub-steps: By identifying the category relationships of the tags selected by the user, the query conditions are transformed into a standardized data search engine index structure and pre-cached in the in-memory database Redis; at the same time, the search results are aggregated and contextualized to provide users with visualization and data services.
[0009] Furthermore, the semantically complete titles verified by the BLEU-4 score include semantically complete titles verified by a BLEU-4 score ≥ 0.92.
[0010] Furthermore, the classifier fine-tuned from the historical labeled data includes a BERT classifier fine-tuned from the historical labeled data.
[0011] Furthermore, the data search engine includes the Elasticsearch engine.
[0012] A tag-based open-source data dynamic organization and intelligent retrieval device includes a processor and a memory, wherein the memory stores a computer program that, when loaded by the processor, executes the method described in any of the preceding claims.
[0013] A tag-based open-source data dynamic organization and intelligent retrieval system includes the tag-based open-source data dynamic organization and intelligent retrieval device described above.
[0014] The beneficial effects of this invention include: (1) At the storage level: This invention can effectively reduce data redundancy and save storage costs through a refined layered decoupling design. More importantly, the expansion time for adding a type of open source data will be compressed from hours to minutes, enabling near real-time access to emerging data sources (such as social media platform data of sudden hot events), and completely getting rid of the data management lag problem caused by rigid table structure.
[0015] (2) In terms of tag construction: This invention greatly improves the efficiency, accuracy and coverage of tagging by using a dual-track tagging scheme of rules and models. For example, when identifying topics related to "large models", it can accurately distinguish whether they belong to the category of "AI" or "cloud computing", avoiding the conflict rate of traditional rule engines and improving tag coverage.
[0016] (3) In terms of data organization: First, dynamic cataloging configuration automatically generates a non-fixed directory tree structure based on tag relationships, replacing the traditional static path, so that the directory can evolve in real time with data changes, effectively improving the timeliness of data retrieval and system scalability; Second, personalized dynamic data view configuration customizes exclusive display formats and data services for open source data sources such as social media and news, optimizing the flexibility of data presentation and user experience; Finally, tag data organization constructs a wide table of tag data through row and column transposition method and achieves incremental synchronization with Elasticsearch, greatly enhancing data aggregation efficiency and query performance, thereby ensuring the dynamic adaptability, efficiency and maintainability of data organization in complex business scenarios.
[0017] (4) In terms of data query: It automatically maps user query conditions (same category tag generation OR logic, cross category tag generation AND logic) to an optimized Elasticsearch query DSL, thereby completing index pre-computation and caching without real-time parsing, significantly improving the response speed and accuracy of multi-dimensional tag retrieval, and completely solving the query performance bottleneck caused by the complexity of tag logic under traditional flat storage. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the tag-based open-source data dynamic organization and intelligent retrieval workflow of an embodiment of the present invention. Figure 2 This is a flowchart illustrating the data storage information of an embodiment of the present invention. Figure 3 This is a flowchart illustrating the data processing information of an embodiment of the present invention. Figure 4 This is a flowchart illustrating the construction of the label system according to an embodiment of the present invention. Figure 5 This is a flowchart illustrating the dynamic organization of information data in an embodiment of the present invention. Figure 6 This is a flowchart illustrating the data retrieval and service information process according to an embodiment of the present invention. Detailed Implementation
[0020] All features disclosed in all embodiments of this specification, or steps in all methods or processes implied in the disclosure, may be combined and / or extended or replaced in any way, except for mutually exclusive features and / or steps.
[0021] In a preferred embodiment, a method for optimizing storage structure, intelligent data processing, automated tag construction, dynamic data organization, and intelligent retrieval of open-source data (including social media, news, multilingual text, etc.) is provided, which is suitable for efficient management and intelligent retrieval of large-scale unstructured open-source text data.
[0022] More specifically, such as Figure 1 As shown, this invention proposes a tag-based open-source data dynamic organization and intelligent retrieval method, including the following steps: Step 1: Open Source Data Storage The invention employs a refined layered architecture to decouple data: the master data table serves as the global hub, storing common attributes such as unique data identifiers, source channels, keywords, and collection timestamps to ensure the consistency of basic metadata; the text table independently carries the original text content, using UTF-8 encoding to natively store multilingual input; the translation table triggers a multilingual translation engine through an automatic language detection module, translating the title and body text separately and storing them as independent fields, effectively avoiding the loss of key information caused by overall translation; the attachment table establishes a hash value index mechanism to efficiently associate binary attachment files such as images and videos; and the personalized tables are dynamically created according to open-source data types, such as social media personalized tables and news personalized tables. The personalized tables store personalized fields for different open-source data, so that when adding new data types, only the personalized tables need to be extended without modifying the master table.
[0023] More specifically, such as Figure 2 As shown, in order to achieve efficient decoupled storage and elastic scaling of massive, heterogeneous, and multi-source public information, and to avoid data redundancy and access delays caused by rigid table structures in traditional methods, this invention constructs a refined hierarchical storage architecture, which specifically includes the following four steps: 1) Standardized definition of master table metadata. Create a global metadata table in the relational database (such as MySQL, Oracle, Kingbase, etc.), containing only data_id (unique identifier in UUID format), source_channel (enumeration type: data source channel identifier), keywords, subject_terms, publish_time (publication time), collect_time (collection time), update_time (modification time), etc., to ensure the uniformity of basic metadata and minimize storage overhead, providing a stable central hub for subsequent data association.
[0024] 2) Native storage of text tables in multiple languages. Create text tables in relational databases (such as MySQL, Oracle, Kingbase, etc.), including text fields, and natively store the original text content encoded in UTF-8 using TEXT or BLOB type fields. This ensures the integrity and traceability of multilingual (Chinese / English / mixed languages) input and avoids information distortion caused by encoding conversion.
[0025] 3) Intelligent translation result storage in the translation table. Create a translation table in a relational database (such as MySQL, Oracle, Kingbase, etc.), including the title_translated (translated title) and content_translated (translated body text) fields. The translated title is limited to 128 characters in length, and the translated body text field uses the TEXT or BLOB type.
[0026] 4) Dynamic configuration of attachments and personalized tables. The attachment table uses a SHA-256 hash index to efficiently associate images, videos, and other files (filename encoding is {data_id}_{file_type}) through MinIO object storage service. The personalized table adopts a dynamic partitioning strategy, automatically identifying new data source types and dynamically creating dedicated tables (such as social_data_metrics) in the relational database (such as MySQL, Oracle, Kingbase, etc.) containing only characteristic fields (such as retweet_count, quote_count), avoiding full table structure modification.
[0027] Step 2: Intelligent Data Processing In the inventive concept, end-to-end automatic processing is constructed through an intelligent processing pipeline to realize data value extraction, which includes subject term extraction, translation, keyword generation and other steps. The improved LDA model is adopted in the subject term extraction step, which fuses domain dictionaries to construct a semantic constraint matrix, and accurately identifies high-frequency subject terms in combination with the TF-IDF weighting algorithm, effectively suppressing noise interference; in the title and body text translation step, the capability of large models is utilized to realize independent translation processing of the title and the body text. Title translation focuses on semantic completeness, while body text translation maintains the consistency of technical terms, and the outputs of both are verified by an NLP processing engine and then stored in a translation table; in the keyword generation step, based on the subject term set output from the subject term extraction step, high-frequency entities (such as person names, place names and event names) are extracted therefrom to generate keywords.
[0028] More specifically, as Figure 3 shown, in order to realize automatic data value extraction and quality improvement, the present invention deploys an end-to-end real-time intelligent processing pipeline, and the specific operation includes the following 3 steps: 1) Subject term extraction and noise suppression. In the Apache Flink real-time computing engine, the improved LDA model is adopted to fuse domain dictionaries to construct a semantic constraint matrix, and the TF-IDF weighting algorithm is combined to accurately calculate subject scores and output high-frequency subject terms (such as "AI model", "computing power). The interference rate is reduced through a noise suppression mechanism, which effectively eliminates the interference of irrelevant words (such as "de", "le") to subject identification.
[0029] 2) Independent translation processing of title and body text. The title translation module calls the locally deployed mBART-50 model (GPU-accelerated, with an input length limit of 128 characters), and outputs semantically complete titles verified by BLEU-4 score greater than or equal to 0.92; for body text translation, the fine-tuned mT5-large model is called (the domain knowledge base is injected during inference, forcing "Transformer" to be translated as "Transformer"), after translation, entities such as person names and place names are extracted through entity recognition, which ensures the consistency of technical terms and the coherence of context, and improves translation accuracy.
[0030] 3) Structured generation of keywords. High-frequency entities are extracted from the subject term set, and target knowledge is associated through an entity linking algorithm to generate a structured keyword library with semantic associations, which provides accurate input for the subsequent label system, and automatically records error samples for model iterative optimization.
[0031] Step 3: Construction of label system The invention employs a dual-track fusion mechanism of rules and models to completely break the precision-breadth paradox of tag generation. Rule-based tagging is based on a dynamic rule base (e.g., "#technology# matches the keyword 'AI'"), using a regular expression engine to match keywords in real time and generate accurate tags (e.g., industry: technology), ensuring high-confidence output. Model-based tagging relies on a BERT classifier finely tuned from historical labeled data to output tags with probability weights (e.g., sentiment: positive (0.85)), significantly improving the coverage of emerging hot topics. By intelligently fusing the results of the two, a tag set with confidence is constructed. When the rule matching confidence is greater than 0.7, rule tags are used first; otherwise, model tags are used, thus achieving both improved tag coverage and maintained accuracy.
[0032] More specifically, such as Figure 4 As shown, to resolve the precision-breadth paradox in label generation (rules have broad coverage but high conflict rate, while models have high precision but narrow coverage), this invention implements a dual-track fusion mechanism of rules and models, which includes the following three steps: 1) Dynamic execution of rule tagging. Based on the Redis Hash structure to store the dynamic rule base, the regular expression engine matches keywords in real time to generate high-confidence tags (such as industry: AI, with a default confidence of 0.9), ensuring a rapid response to mature hot topics (such as "AI").
[0033] 2) Intelligent prediction by model labeling. The BERT-base model (fine-tuned based on historical labeled data, with 50 categories) is called. The input topic word sequence outputs a probability distribution (e.g., topic: AI (0.85), topic: cloud computing (0.12)). Emerging hot topics are identified through probability weights, improving the coverage of unknown areas.
[0034] 3) Intelligent Tag Fusion Decision. Built-in fusion decision tree: When the rule matching confidence is greater than 0.7, the rule label is used first; otherwise, the model label is enabled, and the overlap of the label set is calculated by cosine similarity to automatically remove conflicting items (such as avoiding "large models" from being labeled as AI and cloud computing at the same time), so as to improve the label coverage and reduce the conflict rate, and achieve a balance between accuracy and breadth.
[0035] Step 4: Dynamically Organize Data The invention employs a dynamic rule-driven data organization approach to achieve dynamic organization of personalized data views and data tags. First, by configuring tag combination logic (such as "tag A + tag B" triggering grouping or "tag C excluding tag D" filtering rules), a non-fixed directory tree structure that can evolve in real time is automatically generated, allowing directory nodes to dynamically expand with the data. Simultaneously, a one-to-many relationship model between tags and personalized data views is established using unique data IDs as the hub. Through row and column transposition technology, the scattered tag relationships are aggregated into a wide tag data table. Finally, an incremental synchronization mechanism imports this wide table into Elasticsearch in real time. This approach maintains a high degree of customizability in data organization while completely eliminating the storage fragmentation problem of multi-source heterogeneous data, improving multi-dimensional tag retrieval efficiency to the level of a single-table scan, and significantly solving the technical bottlenecks of rigid data organization and low retrieval performance under traditional static directory and flat storage models.
[0036] More specifically, such as Figure 5 As shown, to support data retrieval and services, this invention develops a dynamic data organization engine, which specifically includes the following three steps: 1) Dynamic cataloging configuration. First, by selecting tag categories or configuring tag combinations (setting association rules between tags, such as defining a specific grouping logic to trigger when "tag A + tag B" is combined, or setting a filtering condition to exclude tag D from tag C), accurate adaptation to complex business scenarios is achieved. Then, based on the above configuration, the tag relationships are automatically parsed and a non-fixed directory tree structure (rather than a preset static path) is generated, ensuring that the directory can evolve in real time as the data changes. Finally, a dynamic directory is generated based on the configuration, outputting a highly customizable interactive directory interface. This directory supports real-time updates (such as automatically expanding child nodes when a new tag is added), thereby improving the timeliness and scalability of data directory retrieval.
[0037] 2) Personalized Dynamic Data View Configuration. Based on a hierarchical data storage structure, personalized data views are created for various open-source data such as social media and news. The display format and personalized data services are configured according to the characteristics of each type of data. The display format supports personalized configuration of titles, attributes, and body content. The data services provide personalized services by adding third-party API interfaces and setting their input and output parameters.
[0038] 3) Tag Data Organization. First, data tags and data views are associated using unique data IDs to form a one-to-many tag association. Then, these tag association data are aggregated using row and column transposition based on the unique data IDs to form a large wide table of tag data. Finally, the data collection service periodically and incrementally synchronizes the data to Elasticsearch.
[0039] Step 5: Intelligent Data Retrieval In the inventive concept, intelligent efficient interaction of data query is implemented through an intelligent analysis of multi-dimensional tag query conditions and an index precomputation mechanism. By automatically identifying the classification relationship of the tags selected by the user (such as OR logical expressions generated for tags within the same classification, AND logical expressions generated for cross-classification tags), the query conditions are efficiently converted into a standardized Elasticsearch index structure and precached in Redis, which completely eliminates real-time analysis overhead to significantly reduce query response time; meanwhile, intelligent aggregation and scenario-based processing are performed on retrieval results, for example, timestamps are automatically extracted for person and event tags to generate timeline data in ISO 8601 format, providing users with seamlessly connected visual display and data services, thereby greatly improving data exploration efficiency and significantly lowering the user operation threshold.
[0040] More specifically, as Figure 6 shown, in order to improve user data exploration efficiency and reduce operation threshold, the present invention constructs an intelligent data retrieval interaction and display module, and specific operations comprise the following 4 steps: 1) Intelligent analysis of query conditions. When a user selects a dynamic directory and tags for data retrieval, the query conditions are automatically analyzed to generate a query logical expression: wherein tags within the same classification (such as "AI" and "blockchain" under the theme classification) generate an "OR" logical expression (tags.topic: AI OR tags.topic: blockchain), while cross-classification tags (theme and industry) generate an "AND" logical expression (tags.topic: AI AND tags.industry: technology).
[0041] 2) Index precomputation and cache optimization. After analysis is completed, ES index precomputation is triggered, the obtained query logical expression is converted into an index structure based on Query DSL (such as {"bool": {"should": [{"term": {"tags.topic.keyword": "AI"}}], "must": [{"term": {"tags.industry.keyword": "technology"}}]}}), syntax tree verification is performed, and then it is cached into Redis, so as to avoid real-time analysis overhead and reduce response time of complex queries.
[0042] 3) Data query. First, data stored in Elasticsearch is queried based on the generated index structure, then the obtained query results are comprehensively sorted according to configured sorting fields, hit open source data is aggregated by type, and finally query result data is generated.
[0043] 4) Visualization and Services. The results of data queries are assembled and visualized according to personalized data view configurations, and personalized data services are provided, such as data download, intelligent recommendations, and intelligent Q&A. Furthermore, for "People" and "Events" tags, event sequences are generated by scanning the timestamps of the result data, sorting them according to ISO 8601 format, and outputting JSON-formatted timeline data for front-end rendering, enabling data display along a timeline.
[0044] It should be noted that, within the scope of protection defined in the claims of this invention, the following embodiments can be combined and / or extended or replaced in any logical manner from the above specific embodiments, such as the disclosed technical principles, disclosed technical features or implicitly disclosed technical features.
[0045] Example 1 A tag-based open-source method for dynamic data organization and intelligent retrieval includes: Step 1, Data Storage: Decouple data storage using a layered architecture; Step 2, Data Processing: Deploy an end-to-end real-time intelligent processing pipeline for data topic extraction, translation, and keyword generation; Step 3, Label Construction: A dual-track fusion mechanism of rules and models is adopted. Rule labeling is based on a dynamic rule base and generates accurate labels by matching keywords through a regular expression engine to ensure high confidence output. Model labeling relies on a classifier fine-tuned based on historical labeled data to output labels with probability weights. By fusing the results of the two, a label set with confidence is constructed. Among them, when the confidence of rule matching is greater than a set value, rule labels are used first; otherwise, model labels are used. Step 4, Data Organization: Adopt a dynamic rule-driven data organization method to dynamically organize personalized data views and data labels; Step 5, Data Retrieval: Through the parsing of multi-dimensional tag query conditions and index pre-calculation, intelligent and efficient interactive data query is achieved.
[0046] Example 2 Based on Example 1, the decoupling of data storage using a layered architecture specifically includes: A hierarchical storage architecture is set up, comprising a master data table, a text table, a translation table, an attachment table, and a personalization table. The master data table serves as the global hub, storing common data attributes to ensure the consistency of basic metadata. The text table independently carries the original text content and natively saves multilingual input using UTF-8 encoding. The translation table triggers a multilingual translation engine through an automatic language detection module, translating the title and text separately and storing them as independent fields. The attachment table establishes a hash value index mechanism to associate with binary attachment files. The personalization table is dynamically created according to data type, storing personalized fields for different data, so that when adding new data types, only the personalization table needs to be expanded without modifying the master table.
[0047] Example 3 Based on Example 1, the deployment of the end-to-end real-time intelligent processing pipeline, used for data topic term extraction, translation, and keyword generation, specifically includes the following sub-steps: 1) Topic word extraction and noise suppression: In the Apache Flink real-time computing engine, a semantic constraint matrix is constructed by fusing the domain dictionary, and topic scores are calculated by combining the TF-IDF weighted algorithm to output high-frequency topic words. The noise suppression mechanism is used to reduce the interference rate and eliminate irrelevant words from interfering with topic recognition. 2) Independent translation processing for titles and body text: The title translation module calls the locally deployed mBART-50 model to output a semantically complete title verified by BLEU-4 scoring; the body text translation calls a finely tuned mT5-large model, and after translation, entity recognition is used to extract personal names and place names to ensure consistency of technical terms and contextual coherence. 3) Keyword structure generation: High-frequency entities are extracted from the topic word set, and target knowledge is associated through entity linking algorithms to generate a semantically related structured keyword library, which provides input for the subsequent tag system, while recording error samples for model iteration and optimization.
[0048] Example 4 Based on Example 1, the method of dynamically organizing personalized data views and data tags using a dynamic rule-driven data organization approach specifically includes the following sub-steps: First, a non-fixed directory tree structure that can evolve in real time is generated by configuring tag combination logic, so that directory nodes can dynamically expand with data. At the same time, a one-to-many relationship model between tags and personalized data views is established with the unique data ID as the hub. After row and column transposition, the scattered tag relationships are aggregated into a wide table of tag data. Finally, the wide table is imported into the data search engine in real time through an incremental synchronization mechanism.
[0049] Example 5 Based on Example 1, the intelligent and efficient interaction of data query through parsing of multi-dimensional tag query conditions and index pre-calculation specifically includes the following sub-steps: By identifying the category relationships of the tags selected by the user, the query conditions are transformed into a standardized data search engine index structure and pre-cached in the in-memory database Redis; at the same time, the search results are aggregated and contextualized to provide users with visualization and data services.
[0050] Example 6 Based on Example 3, the semantically complete titles verified by BLEU-4 scoring include semantically complete titles verified by BLEU-4 scoring ≥ 0.92.
[0051] Example 7 Based on Example 1, the classifier fine-tuned by the historical labeled data includes a BERT classifier fine-tuned by the historical labeled data.
[0052] Example 8 Based on Example 4, the data search engine includes the Elasticsearch engine.
[0053] Example 9 A tag-based open-source data dynamic organization and intelligent retrieval device includes a processor and a memory, wherein the memory stores a computer program that, when loaded by the processor, executes the method described in any one of Embodiments 1 to 8.
[0054] Example 10 A tag-based open-source data dynamic organization and intelligent retrieval system includes the tag-based open-source data dynamic organization and intelligent retrieval device described in Example 9.
[0055] The units described in the embodiments of the present invention can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0056] According to one aspect of the present invention, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.
[0057] In another aspect, embodiments of the present invention also provide a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.
Claims
1. A tag-based open-source data dynamic organization and intelligent retrieval method, characterized in that, include: Step 1, Data Storage: Decouple data storage using a layered architecture; Step 2, Data Processing: Deploy an end-to-end real-time intelligent processing pipeline for data topic extraction, translation, and keyword generation; Step 3, Label Construction: A dual-track fusion mechanism of rules and models is adopted. Rule labeling is based on a dynamic rule base and generates accurate labels by matching keywords through a regular expression engine to ensure high confidence output. Model labeling relies on a classifier fine-tuned based on historical labeled data to output labels with probability weights. By fusing the results of the two, a label set with confidence is constructed. Among them, when the confidence of rule matching is greater than a set value, rule labels are used first; otherwise, model labels are used. Step 4, Data Organization: Adopt a dynamic rule-driven data organization method to dynamically organize personalized data views and data labels; Step 5, Data Retrieval: Through the parsing of multi-dimensional tag query conditions and index pre-calculation, intelligent and efficient interactive data query is achieved.
2. The tag-based open-source data dynamic organization and intelligent retrieval method according to claim 1, characterized in that, The layered architecture for decoupling data storage specifically includes: A hierarchical storage architecture is set up, comprising a master data table, a text table, a translation table, an attachment table, and a personalization table. The master data table serves as the global hub, storing common data attributes to ensure the consistency of basic metadata. The text table independently carries the original text content and natively saves multilingual input using UTF-8 encoding. The translation table triggers a multilingual translation engine through an automatic language detection module, translating the title and text separately and storing them as independent fields. The attachment table establishes a hash value index mechanism to associate with binary attachment files. The personalization table is dynamically created according to data type, storing personalized fields for different data, so that when adding new data types, only the personalization table needs to be expanded without modifying the master table.
3. The tag-based open-source data dynamic organization and intelligent retrieval method according to claim 1, characterized in that, The deployed end-to-end real-time intelligent processing pipeline is used for data topic term extraction, translation, and keyword generation, and specifically includes the following sub-steps: 1) Topic word extraction and noise suppression: In the Apache Flink real-time computing engine, a semantic constraint matrix is constructed by fusing the domain dictionary, and topic scores are calculated by combining the TF-IDF weighted algorithm to output high-frequency topic words. The noise suppression mechanism is used to reduce the interference rate and eliminate irrelevant words from interfering with topic recognition. 2) Independent translation processing for titles and body text: The title translation module calls the locally deployed mBART-50 model to output a semantically complete title verified by BLEU-4 scoring; the body text translation calls a finely tuned mT5-large model, and after translation, entity recognition is used to extract personal names and place names to ensure consistency of technical terms and contextual coherence. 3) Keyword structure generation: High-frequency entities are extracted from the topic word set, and target knowledge is associated through entity linking algorithms to generate a semantically related structured keyword library, which provides input for the subsequent tag system, while recording error samples for model iteration and optimization.
4. The tag-based open-source data dynamic organization and intelligent retrieval method according to claim 1, characterized in that, The method of dynamically organizing personalized data views and data tags using a dynamic rule-driven data organization approach includes the following sub-steps: First, a non-fixed directory tree structure that can evolve in real time is generated by configuring tag combination logic, so that directory nodes can dynamically expand with data. At the same time, a one-to-many relationship model between tags and personalized data views is established with the unique data ID as the hub. After row and column transposition, the scattered tag relationships are aggregated into a wide table of tag data. Finally, the wide table is imported into the data search engine in real time through an incremental synchronization mechanism.
5. The tag-based open-source data dynamic organization and intelligent retrieval method according to claim 1, characterized in that, The intelligent and efficient interaction of data query through parsing of multi-dimensional label query conditions and index pre-calculation specifically includes the following sub-steps: By identifying the category relationships of the tags selected by the user, the query conditions are transformed into a standardized data search engine index structure and pre-cached in the in-memory database Redis; at the same time, the search results are aggregated and contextualized to provide users with visualization and data services.
6. The tag-based open-source data dynamic organization and intelligent retrieval method according to claim 3, characterized in that, The semantically complete titles verified by the BLEU-4 score include those with a BLEU-4 score ≥ 0.
92.
7. The tag-based open-source data dynamic organization and intelligent retrieval method according to claim 1, characterized in that, The classifier fine-tuned using historical labeled data includes the BERT classifier fine-tuned using historical labeled data.
8. The tag-based open-source data dynamic organization and intelligent retrieval method according to claim 4, characterized in that, The data search engine includes the Elasticsearch engine.
9. A tag-based open-source data dynamic organization and intelligent retrieval device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when loaded by the processor, executes the method as described in any one of claims 1 to 8.
10. A tag-based open-source data dynamic organization and intelligent retrieval system, characterized in that, It includes the tag-based open-source data dynamic organization and intelligent retrieval device as described in claim 9.