Intelligent organization information retrieval and recommendation system based on data cloud computing
By using a data cloud computing-based intelligent retrieval and recommendation system for organizational information, and employing a three-dimensional classification algorithm and a multi-layered knowledge architecture, dynamic knowledge fusion of structured and unstructured data is achieved. This improves decision-making efficiency and the timeliness of information association, solves the problems of data disconnect and low resource utilization in existing systems, and meets data security and compliance requirements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-10
AI Technical Summary
Existing retrieval and recommendation systems cannot achieve dynamic knowledge fusion of structured data, unstructured data, and real-time streaming data. They cannot respond promptly to project changes and supply chain fluctuations, and the retrieval results are out of touch with actual needs. They cannot transform retrieval results into decision support capabilities, resulting in low decision-making efficiency.
This intelligent retrieval and recommendation system for organizational information, based on data cloud computing, employs access, storage, fusion, retrieval, recommendation, scheduling, and management modules. It utilizes a three-dimensional classification algorithm based on modality, time series, and business value to construct a knowledge architecture consisting of a static foundation layer and a real-time dynamic layer. The system supports keyword, semantic, and multi-hop reasoning retrieval, provides personalized recommendations based on user profiles, and dynamically allocates cloud computing resources through load prediction.
It achieves the continuity and timeliness of information association, improves decision-making efficiency and scientificity, solves the problems of rigid resource allocation and low utilization rate in traditional systems, and meets the compliance requirements of full-process data security management.
Smart Images

Figure CN121636552A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data cloud computing information retrieval and recommendation technology, specifically an intelligent retrieval and recommendation system for organizational information based on data cloud computing. Background Technology
[0002] As organizational digitalization enters a closed-loop phase encompassing business, data, and decision-making, information systems need to meet the core requirements of real-time knowledge association and decision support. However, existing retrieval and recommendation systems suffer from the following technical problems: The traditional model of offline preprocessing and static storage makes it difficult to achieve dynamic knowledge fusion of structured data, unstructured data and real-time streaming data. Structured data includes real-time data from business systems and dynamic customer information, while unstructured data covers R&D documents and meeting minutes. Real-time streaming data includes business operation logs and equipment sensor data. Knowledge graph updates are significantly delayed and cannot respond in a timely manner to real-time business dynamics such as project changes and supply chain fluctuations. This results in a disconnect between search results and actual needs, and a lack of coherent knowledge association links.
[0003] It only has basic query and recall functions, and is not deeply integrated with actual decision-making scenarios. It cannot transform search results into decision support capabilities. For the need to analyze the technical barriers of competitors before the launch of new products, it only returns raw information such as competitor patent documents and market reports. It is difficult to automatically extract the key indicators needed for decision-making, such as technical parameters, patent protection boundaries, and cost structure. It also cannot perform structured comparison with its own product data. For querying solutions to process anomalies, it only pushes historical documents and does not combine real-time anomaly-related data to screen and adapt solutions, resulting in low decision-making efficiency. Summary of the Invention
[0004] The purpose of this invention is to provide an intelligent retrieval and recommendation system for organizational information based on cloud computing, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an intelligent retrieval and recommendation system for organizational information based on data cloud computing, comprising an access module, a storage module, a fusion module, a retrieval module, a recommendation module, a scheduling module, and a management and control module; Preferably, the access module interfaces with business platforms including ERP systems, production line sensors, and OA office systems to achieve the collection of static files and real-time data; A three-dimensional classification algorithm based on modality, time series, and business value is adopted. Data is classified according to modality (structured / unstructured), update time series (high frequency / low frequency), and business value (core / important / ordinary). It is synchronously associated with business scenario tags (R&D, supply chain, finance, etc.) and business urgency (high / medium / low). High urgency scenarios include supply chain interruption warnings, order delays exceeding preset time, and process abnormal errors. Core data includes core patents, customer transaction bottom prices, and project key node progress. The tags are automatically completed through the business department's preset rule table. For unstructured data, extract core information including project number, date and responsible person, form a table that records the correspondence between classification labels and data entities, and output it to the corresponding storage area of the storage module.
[0006] Preferably, the storage module receives the classification tag and data entity correspondence table output by the access module, and establishes a three-layer hybrid storage architecture based on a distributed file system and object storage, consisting of a real-time data cache area, a high-access-frequency storage area, and a low-access-frequency storage area; the real-time data cache area stores high-frequency streaming data marked by the access module; the high-access-frequency storage area stores information marked by the access module that has a high retrieval frequency and high business relevance; and the low-access-frequency storage area stores information marked by the access module that has a low retrieval frequency. Equipped with an intelligent access frequency adaptation migration engine, the system uses the business urgency label of the access module (high-urgency data is migrated to the real-time data cache area first) and retrieval frequency (high-frequency retrieval data is migrated to the high-access frequency storage area) as core indicators to migrate data to the adaptation storage area and generate a metadata table that records the storage location, data label and access permissions, which is then synchronized to the fusion module and the management module.
[0007] Preferably, the fusion module reads the storage location, data tags, and access permission metadata table of the storage module, prioritizes calling data from the high-access-frequency storage area and the real-time data cache area, and establishes a two-layer knowledge architecture consisting of a static basic layer (mainly relying on non-real-time data from the high-access-frequency storage area) and a real-time dynamic layer (relying on real-time data from the real-time data cache area); the static basic layer combines the core entity dictionary of each business scenario to extract entity and relation triples, and eliminates ambiguity and filters redundant data through entity linking algorithms; The real-time dynamic layer calls high / medium / low urgency data from the real-time data cache of the storage module, and allocates update resources according to the urgency tag of the access module through the business priority update scheduler; high urgency data completes the knowledge graph update and occupies 80% of the real-time computing power, medium urgency data updates and occupies 15% of the real-time computing power, and low urgency data updates and occupies 5% of the real-time computing power. After resolving data conflicts using a business rules-first, confidence-based fallback algorithm, a database table is generated that records the correspondence between the knowledge graph and business scenarios in real time, and then pushed to the retrieval module.
[0008] Preferably, the retrieval module accesses the real-time updated knowledge graph and business scenario correspondence database table output by the fusion module, supporting keyword retrieval, semantic retrieval, and multi-hop reasoning retrieval, forming a closed loop of retrieval, analysis, risk labeling, and decision support; it parses the query intent through the keyword and scenario correspondence table; fact queries match entities in the static base layer of the knowledge graph and label basic risks; association queries perform reasoning within two hops based on the static base layer and label association risks; and decision queries call the real-time dynamic layer of the knowledge graph to supplement association data and label decision risks. For decision-making queries, key information is first extracted from the candidate document library through regular expressions and dictionary matching. Then, multi-hop reasoning is performed based on the knowledge graph. Finally, corresponding decision indicators and structured comparison results are output according to roles and business scenarios. The search results and decision indicators are synchronized to the recommendation module.
[0009] Preferably, the recommendation module receives the search results and decision indicators from the retrieval module, and combines static and dynamic dual-dimensional user profiles and a decision scenario tag library to push information according to the user's current business scenario and profile characteristics; the static dimension includes user role and permission level; the dynamic dimension includes search history, document type preferences for stay time exceeding a preset time, and current business task; For decision-making users, push industry trend data, internal indicator comparison reports output by the search module, and risk warning suggestions; for execution users, push adapted solutions, resource allocation lists, and operation procedure guides. A recommendation performance analysis table is generated based on user behavior feedback, and resource usage is simultaneously recorded. The results are then combined to form a comprehensive analysis table that records both recommendation performance and resource usage, which is then synchronized to the scheduling module.
[0010] Preferably, the scheduling module is based on the comprehensive analysis table of recommendation effect and resource consumption of the recommendation module, the knowledge update frequency of the fusion module, and the decision query concurrency of the retrieval module (all of the above indicators are related to business requirements) to achieve dynamic optimization allocation of cloud computing resources including CPU, memory and bandwidth; the load prediction input features include historical business cycle data and real-time load, and the load is predicted by the algorithm. The microservice architecture is broken down into functions including data access, knowledge fusion, retrieval and recommendation, and security management. Each module occupies resources independently and a dedicated resource pool is preset for each business. The resource pool is automatically activated according to the corresponding business cycle, and the CPU or memory configuration is adjusted to support the retrieval concurrency requirements of the retrieval module. When high-urgency data is accessed, some resources from the ordinary resource pool are temporarily scheduled to the real-time dynamic layer of the fusion module and released after processing. The resource allocation strategy is synchronized to the management module.
[0011] Preferably, the control module establishes a five-dimensional permission model based on roles, scenarios, data sensitivity, decision scenario sensitivity, and business lifecycle; receives the storage location, data tags, and access permission metadata table from the storage module; encrypts the data transmitted by the access module; and restricts the resource access permissions of unauthorized modules based on the resource allocation strategy of the scheduling module. For the search results of the retrieval module, corresponding anonymization is implemented according to role, business scenario, data sensitivity, decision scenario sensitivity, and business lifecycle; sensitive data access logs of each module are recorded, and tamper-proof audit records are generated. At the same time, permission exception logs are pushed to the scheduling module.
[0012] The beneficial effects of this invention are as follows: 1. This invention connects to multiple business platforms such as ERP and sensors through an access module, and uses a three-dimensional classification algorithm based on modality, time series, and business value to complete data annotation, laying a standardized foundation for subsequent processing; the static base layer built by the fusion module, combined with a business entity dictionary, eliminates data ambiguity and filters redundancy; the real-time dynamic layer allocates computing power according to business urgency to complete core data updates and knowledge graph iterations, effectively resolving data conflicts; it solves the problem of traditional system retrieval results being disconnected from real-time business dynamics such as supply chain fluctuations and project changes, improves the coherence and timeliness of information association, and is more in line with actual business needs.
[0013] 2. The retrieval module of this invention supports multiple retrieval methods such as keywords, semantics, and multi-hop reasoning. It accurately analyzes query intent through a keyword-scenario correspondence table, automatically extracts key indicators for decision-making queries, and generates structured comparison results. The recommendation module combines static and dynamic user profiles to push industry trend and risk warning reports to the decision-making level and adapted solutions and operation guides to the execution level, forming a complete closed loop of retrieval-analysis-decision-recommendation. It eliminates the need for manual screening and integration of raw information, directly outputting decision support content adapted to different business scenarios, improving decision-making efficiency and scientific rigor.
[0014] 3. The scheduling module of this invention dynamically allocates cloud computing resources such as CPU and memory based on core indicators such as recommendation effect, knowledge update frequency, and decision query concurrency through a load prediction algorithm. It splits the microservice architecture and presets dedicated resource pools for business operations such as R&D review and supply chain emergency response. In high-urgency scenarios, it can temporarily schedule computing power from ordinary resource pools, which are released and recycled after the task is completed, effectively solving the problems of rigid allocation and low utilization rate of traditional resources. At the same time, the management module constructs a five-dimensional permission model including roles, scenarios, and data sensitivity. It uses AES-256 encryption for accessed data, implements customized desensitization according to business scenarios and roles, and records sensitive data access logs through blockchain to generate tamper-proof audit records. While improving resource utilization efficiency, it achieves full-process data security management and control, effectively meeting compliance requirements. Attached Figure Description
[0015] Figure 1 This is a flowchart of the intelligent retrieval and recommendation system for organizational information based on cloud computing, as described in this invention. Figure 2 This is a flowchart illustrating the entire process of five-dimensional access control in this invention. Figure 3 This is a flowchart of the dynamic scheduling process for cloud computing resources in this invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] like Figures 1 to 3 As shown, this embodiment of the invention provides an intelligent retrieval and recommendation system for organizational information based on cloud computing, including an access module, a storage module, a fusion module, a retrieval module, a recommendation module, a scheduling module, and a management and control module. The specific implementation of each module is as follows: The access module supports Kafka / Flume protocol integration with business platforms such as ERP systems, production line sensors, and OA office systems, enabling the collection of static files and real-time data. Static files include PDFs, Word development documents, and scanned copies of handwritten notes; real-time data includes order logs, real-time sensor data, and business system operation records. A three-dimensional classification algorithm based on modality, time series, and business value is adopted to classify data according to modality, update time series, and business value, and synchronously associate business scenario tags with business urgency. High-urgency scenarios include supply chain interruption warnings, order delays exceeding 24 hours, and process abnormal errors. Core data includes core patents, customer transaction bottom prices, and project key node progress. The tags are automatically completed through the business department's preset rule table. For example, supply chain system data is tagged as high urgency and important data, and historical archived reports are tagged as low urgency and ordinary data. Core rules for determining business value: Key data: Must meet the criteria of "impacting ≥5% of enterprise revenue" or "involving core intellectual property / core customer information", such as core patent technical parameters, customer transaction bottom price, etc. Key data: Must meet the requirement of "impacting ≥30% of departmental business", such as quotations from ordinary suppliers, monthly expense reports of departments, etc. General data: Basic information that does not fall into the core or important category, such as historical supplier archives, backups of R&D meeting minutes, etc.
[0018] Data modalities are divided into structured data and unstructured data; update time series are divided into high frequency and low frequency; business value is divided into core, important, and ordinary; business scenario tags include R&D, supply chain, finance, etc.; and business urgency is divided into high urgency, medium urgency, and low urgency. Configure urgency level mechanism: Support business departments to customize and adjust the urgency level threshold range through the system backend. For example, the high urgency delay threshold in the supply chain scenario can be modified within the range of 12-48h according to business needs. For unstructured data, keyword dictionary matching is used to quickly extract core information such as project number, date, and responsible person, forming a table that records the correspondence between classification tags and data entities, and outputting it to the corresponding storage area of the storage module, providing standardized data input for subsequent storage stratification and knowledge processing.
[0019] The keyword dictionary is constructed using a dual-dimensional approach, consisting of core fields and business tags. The core fields are derived from industry data meta standards, while the business tags are generated from high-frequency terms submitted by various departments and then deduplicated using NLP tools. The dictionary is automatically synchronized monthly with newly added terms from the ERP / OA system, triggering updates to unstructured data extraction rules.
[0020] Business scenario, value, and urgency mapping rules: Supply chain scenario: Core data is delivery data of core suppliers. The high urgency threshold is "delivery delay ≥ 24h" or "supply chain interruption warning", the medium urgency threshold is "delivery delay 12-24h", and the low urgency threshold is "delivery delay < 12h". Important data is quotations from ordinary suppliers, and ordinary data is historical supplier information. Research and development scenario: Core data consists of core patent technical parameters; high urgency threshold is "patent infringement risk warning"; medium urgency threshold is "technical parameter optimization needs"; and low urgency threshold is "document version update"; important data consists of non-core technical documents; and ordinary data consists of backups of research and development meeting minutes. Financial scenario: Core data is the customer's transaction price or annual budget core data. The high urgency threshold is "budget overrun ≥ 10%", the medium urgency threshold is "budget overrun 5%-10%", and the low urgency threshold is "budget overrun < 5%". Important data is the department's monthly expense report, and ordinary data is scanned copies of historical financial vouchers.
[0021] The storage module receives the classification tag and data entity correspondence table output by the access module and establishes a three-layer hybrid storage architecture based on the Distributed File System (HDFS) and object storage: a real-time data cache area, a high-access-frequency storage area, and a low-access-frequency storage area. The real-time data cache area is deployed using a Redis cluster and stores high-frequency streaming data marked by the access module for the past hour, such as real-time order delay logs and production line anomaly signals. The high-access-frequency storage area uses SSD media and stores information marked by the access module that has a high retrieval frequency and high business relevance for the past three months, such as current project documents and active customer data. The low-access-frequency storage area uses low-cost cloud storage media and stores information marked by the access module that has a low retrieval frequency for more than one year, such as historical archived reports and expired project materials. Equipped with an intelligent access frequency adaptation migration engine, the system uses the business urgency label of the access module (high-urgency data is migrated to the real-time data cache area first) and the retrieval frequency in the past 7 days (high-frequency retrieval data is migrated to the high-access frequency storage area) as core indicators to migrate data to the adaptation storage area and generate a metadata table that records the storage location, data label and access permissions, which is then synchronized to the fusion module and the management module.
[0022] The fusion module reads the storage location, data tags, and access permission metadata table of the storage module, prioritizes calling data from the high-access-frequency storage area and the real-time data cache area, and establishes a two-layer knowledge architecture consisting of a static foundation layer (mainly relying on non-real-time data from the high-access-frequency storage area) and a real-time dynamic layer (relying on real-time data from the real-time data cache area). The static foundation layer uses GraphRAG technology, combined with the core entity dictionary of each business scenario, to extract entity and relationship triples. Entities include personnel, projects, supply chain nodes, patents, etc., and relationships include participation, association, reference, supply, etc. Ambiguity is eliminated through entity linking algorithms, for example, "Zhang XX" and "Zhang Mou" are normalized into a unified entity, and redundant data such as duplicate documents and invalid information are filtered out. The core entity dictionary for business scenarios includes supplier ID, material code, and delivery cycle for supply chain scenarios; and patent number, technical parameters, and project stage for R&D scenarios. The real-time dynamic layer calls high / medium / low urgency data from the real-time data cache of the storage module, and allocates update resources according to the urgency tag of the access module through the business priority update scheduler; for example, high urgency data such as supply chain interruption signals completes knowledge graph updates and occupies 80% of real-time computing power within 10 seconds, medium urgency data such as project progress fine-tuning is updated within 1 minute and occupies 15% of real-time computing power, and low urgency data such as ordinary document uploads is updated within 5 minutes and occupies 5% of real-time computing power. Core business rule types: Data source priority rules: The data credibility ranking is clearly defined as "ERP system data (credibility 0.9) > sensor data (credibility 0.7) > OA document data (credibility 0.6)", and in case of conflict, the data with higher credibility shall be used first; Timestamp priority rule: Newly generated data automatically overwrites historical duplicate data without manual intervention; Business logic constraint rules: Based on business formula verification and correction, such as "order amount = unit price × quantity", if there is a data conflict, the outlier value will be recalculated and corrected according to this formula; Confidence level calculation method: The confidence level is calculated by "data source credibility × data consistency score × business relevance score": the data consistency score is the degree of matching between the data and historical similar data (value 0-1), and the business relevance score is the degree of matching between the data and the current business scenario (value 0-1). Conflict resolution process: The first step is to prioritize matching the above business rules. If there are explicit rules, the conflict will be resolved directly according to those rules. The second step is to retain data with a confidence level of ≥0.7 if no matching business rule exists. Third, if the confidence level is less than 0.7, the system will automatically trigger a manual review process and push a review notification to the corresponding business manager. The confidence threshold (≥0.7) is set based on the conventional design idea of "balancing processing efficiency and accuracy": if the threshold is lower than 0.7, low-confidence data will enter the system, increasing the risk of subsequent decision-making; if the threshold is higher than 0.7, a large amount of data will trigger manual review, reducing processing efficiency.
[0023] Thresholds can be adjusted to suit different business scenarios. For example, in supply chain emergency scenarios where efficiency is a priority, the threshold can be lowered to 0.6; in R&D patent scenarios where accuracy is a priority, the threshold can be raised to 0.8. The system supports modifying the threshold through configuration files without changing the core algorithm, thus adapting to the priority requirements of different scenarios.
[0024] After resolving data conflicts using a business rules-first, confidence-based fallback algorithm, a database table is generated that records the correspondence between the knowledge graph and business scenarios in real time, and then pushed to the retrieval module.
[0025] Real-time dynamic layer computing power allocation formula: ; In the formula: This represents the total real-time computing power, which is all the real-time computing resources that the real-time dynamic layer of the fusion module can call upon, measured with 100% computing power as the benchmark. This indicates the proportion of computing power allocated to high-urgency tasks, which is fixed at 80% and is specifically allocated to data updates for high-urgency tasks, such as supply chain disruption warnings, order delays exceeding 24 hours, and process error reports. This indicates the proportion of computing power allocated to medium-urgency data updates, which is fixed at 15% and distributed to medium-urgency data updates, such as data updates in scenarios like minor adjustments to project progress. This represents the proportion of computing power allocated to low-urgency data updates, which is fixed at 5% and distributed to low-urgency data updates, such as data updates in scenarios like ordinary document uploads.
[0026] High-urgency data concurrent processing rules: When multiple high-urgency data update requests exist simultaneously, they are sorted in the order of "business scenario priority + data generation time" (business scenario priority: supply chain emergency > R&D patent early warning > financial budget overrun); if the business scenario priorities are the same, computing power is allocated in reverse order of data generation time, and the minimum computing power ratio of a single request is not less than 20%, and the remaining computing power is allocated to other high-urgency requests proportionally. Triggering conditions for elastic adjustment of computing power: When the total computing power of the real-time dynamic layer is ≥95%, the system automatically sends a temporary computing power expansion request to the scheduling module. When the total computing power of the real-time dynamic layer is ≤30%, the system automatically releases 50% of the redundant computing power to the ordinary resource pool to avoid resource waste.
[0027] The retrieval module connects to the real-time updated knowledge graph and business scenario correspondence database output by the fusion module, supporting keyword retrieval, semantic retrieval, and multi-hop reasoning retrieval, forming a closed loop of retrieval, analysis, risk labeling, and decision support. It parses query intent through the keyword-scenario correspondence table; fact queries (e.g., "current progress of Project A") directly match entities in the static base layer of the knowledge graph and label basic risks; association queries (e.g., "supply chain nodes involved in Project A") perform reasoning within two hops based on the static base layer and label associated risks; and decision queries (e.g., "competitor technical risks before Project A goes live") call the real-time dynamic layer of the knowledge graph to supplement associated data and label decision risks. Inference jump limit rules: Fact lookup (e.g., "current progress of project A"): Only supports 1-hop reasoning, which directly matches the basic information of the target entity; Related queries (e.g., "supply chain nodes involved in project A"): support no more than 2 hops of reasoning, that is, extending from the target entity to related entities at one level; Decision query (e.g., "competitive technical risks before Project A goes live"): Supports no more than 3 hops of reasoning. The maximum number of hops can be adjusted through the system configuration interface. The default is ≤3 hops.
[0028] Inference path pruning rules: Only inference paths with a relevance score ≥ 0.6 are retained, while paths with a score below 0.6 are filtered out. The relevance score is calculated as "entity relationship importance × data timeliness score". Importance of entity relationships: Core business relationships (such as "supply-dependency") are valued at 0.9, ordinary relationships (such as "reference-reference") are valued at 0.6, and weak relationships (such as "mention-association") are valued at 0.3; Data timeliness score: calculated as "1 - (current time - data generation time) / data validity period". Core data is valid for 30 days, important data is valid for 90 days, and ordinary data is valid for 180 days.
[0029] Example of decision query reasoning: Taking the query "competitive product technical risks before the launch of Project A" as an example, the reasoning process is as follows: 1 jump (extract the core technical parameters of Project A) → 2 jump (match the scope of protection of competitor patents) → 3 jump (associate the historical infringement cases of competitors), and finally output specific risk conclusions, such as "Technical parameter X falls within the scope of protection of competitor patent Y, and the risk of infringement is high".
[0030] For decision-making queries, key information is first extracted from the candidate document library through regular expressions and dictionary matching. Then, multi-hop reasoning is performed based on the knowledge graph. Finally, corresponding decision indicators and structured comparison results are output according to roles and business scenarios. When supply chain managers deal with order delays, indicators such as delay duration, supplier historical delay rate, and alternative suppliers are automatically extracted, and a comparison table is generated with "delay rate < 5%" as the qualified threshold. When R&D engineers analyze competitor technologies, core patent claims, technical parameters, and cost differences are extracted to determine whether our parameters fall within the patent scope. The search results and decision indicators are synchronized to the recommendation module to provide decision-dimensional data support for personalized recommendations.
[0031] Multi-hop reasoning entity relationship extraction adopts a hybrid extraction scheme of rule base and BERT fine-tuning model. The rule base covers core business relationships such as "supply-dependency" and "reference-reference". Based on historical business graph summarization, the BERT model is fine-tuned using industry corpus (100,000 business documents), and the entity relationship recognition F1 score reaches 0.89. During reasoning, the rule base is matched first. If no match is found, the model prediction is called. Relationship scores below 0.6 are automatically filtered.
[0032] The recommendation module receives the search results and decision indicators from the retrieval module, and combines static and dynamic dual-dimensional user profiles and a decision scenario tag library to push personalized information according to the user's current business scenario and profile characteristics. The static dimension includes user roles (CEO, department manager, executive employee) and permission levels; the dynamic dimension includes the search history of the past 30 days, document type preferences with a dwell time of ≥3 minutes, and current business tasks, such as new product launch preparation and production process anomaly investigation, which are related to the decision scenario of the retrieval module. The decision-making scenario tag library includes scenarios such as product launch decisions, process exception handling, quarterly budget planning, and R&D project management; the permission levels include core data access rights and ordinary data access rights, which are associated with the permission configuration of the management module. For decision-making users, push industry trend data, internal indicator comparison reports output by the retrieval module, and risk warning suggestions; for execution users, push adaptation solutions, resource allocation lists, and operation step guides, such as the solution for process exception code E08: 1. Restart the server; 2. Check the data interface, which requires calling IT resource B, and the relevant resource idle status data is taken from the scheduling module; Dynamic dimension weight calculation formula: Dynamic dimension score = 0.4 × search history matching degree + 0.3 × document dwell time score + 0.3 × current business task matching degree; The document dwell time score is determined according to the following criteria: dwell time < 3 minutes, 0.3 points; 3-5 minutes, 0.6 points; 5-10 minutes, 0.8 points; > 10 minutes, 1 point.
[0033] The weight values are set based on the priority of business logic: the historical matching degree of the search (weight 0.4) has the highest weight, because the user's historical search behavior directly reflects the relevance of the current needs and is the core factor affecting the accuracy of the recommendation; the document dwell time score (weight 0.3) is the second highest, which can help determine the user's preference for document type; the current business task matching degree (weight 0.3) serves as a supplement for scenario adaptation to ensure that the recommended content fits the real-time business needs.
[0034] - Supports dynamic weight adjustment: The system backend provides a weight configuration interface, and business departments can manually adjust the proportion of each dimension according to the actual scenario (such as increasing the weight of "current business task matching degree" in the R&D scenario and increasing the weight of "search history matching degree" in the financial scenario), without relying on fixed values, to adapt to different business needs.
[0035] Recommendation priority rules: For decision-making users: The recommended content priority is "risk warning suggestions (weight 0.4) > internal indicator comparison report (weight 0.3) > industry trend data (weight 0.3)", prioritizing the push of information directly related to decision-making risks; For execution layer users: The recommended content priority is "adaptation solution (weight 0.5) > operation step guide (weight 0.3) > resource allocation list (weight 0.2)", and priority is given to pushing execution solutions that can be directly implemented.
[0036] Recommended performance evaluation criteria: The core evaluation metrics include click-through rate (target value ≥ 30%), average dwell time (target value ≥ 5 minutes), and feedback satisfaction (target value ≥ 4 / 5 points). If a metric falls below the target value, the system will automatically adjust the dimensional weights. For example, if the click-through rate is low, the weight of "current business task matching degree" will be increased.
[0037] A recommendation performance analysis table is generated based on user behavior feedback, and resource usage is simultaneously recorded. The results are then combined to form a comprehensive analysis table that records both recommendation performance and resource usage, which is synchronized to the scheduling module to optimize subsequent resource allocation strategies.
[0038] The scheduling module, based on the comprehensive analysis table of recommendation effect and resource consumption of the recommendation module, the knowledge update frequency of the fusion module, and the decision query concurrency of the retrieval module (all of the above indicators are related to business requirements), realizes dynamic optimization allocation of cloud computing resources such as CPU, memory, and bandwidth; the load prediction input features include historical business cycle data (e.g., high search volume on the 5th of each month during the "project review period" and high data access volume on the 20th of each month during the "financial settlement period", which are related to the business scenario tags of the access module) and the real-time load of the past hour, and predicts the load through algorithms; The microservice architecture is broken down into functions such as data access, knowledge fusion, retrieval and recommendation, and security management. Each module occupies resources independently, and dedicated resource pools such as "R&D review resource pool" and "supply chain emergency resource pool" are preset. The resource pool is automatically activated according to the corresponding business cycle, and the CPU or memory configuration is adjusted to support the retrieval concurrency requirements of the retrieval module. When high-urgency data is accessed, some resources from the ordinary resource pool are temporarily scheduled to the real-time dynamic layer of the fusion module and released after processing. The resource allocation strategy is synchronized to the management module to ensure that resource access permissions and data security management are coordinated.
[0039] Load prediction weighting rules: Typical scenario: Business cycle weight ( The value is 0.6, representing the real-time load weight. The value is set to 0.4, taking into account historical business cycle patterns. High-urgency scenarios (e.g., supply chain disruptions, patent infringement warnings): Business cycle weight ( The real-time load weight was adjusted to 0.3. The value has been adjusted to 0.7 to prioritize adapting to current real-time business needs.
[0040] Weight values are set based on the priority of scenario requirements: business cycle weight in normal scenarios ( =0.6) is higher than the real-time load weight ( =0.4), because historical business patterns can predict resource demand in advance, reducing the frequency of temporary capacity expansion; in high-urgency scenarios, it is adjusted to real-time load weight ( The higher value (=0.7) is to prioritize responding to sudden business needs and avoid data processing delays.
[0041] Customizable weights: The system has two preset weight templates, one for regular and one for high urgency. It also allows IT departments to modify the weight values according to the scale of the business (for example, small organizations can reduce the business cycle weight to 0.5) to adapt to the business characteristics of different organizations.
[0042] Business-specific resource pool configuration standards: R&D review resource pool: CPU configuration 32 cores, memory configuration 64GB, bandwidth configuration 100Mbps, automatic activation cycle is 5-10 days per month, mainly adapted to R&D related scenarios such as project review and patent analysis; Supply chain emergency resource pool: CPU configuration of 64 cores, memory configuration of 128GB, bandwidth configuration of 200Mbps, activated only when triggered in high-urgency scenarios (such as supply chain disruptions and order delays) to meet emergency data processing needs; Financial settlement resource pool: CPU configuration of 16 cores, memory configuration of 32GB, bandwidth configuration of 50Mbps, automatic activation cycle of the 20th-25th of each month, adapted to settlement-related scenarios such as financial statement generation and budget verification.
[0043] Temporary resource scheduling rules: When high-urgency data is received, no more than 50% of the idle resources are scheduled from the ordinary resource pool and temporarily allocated to the real-time dynamic layer of the fusion module; after the high-urgency data is processed, the temporary resources are released within 10 minutes and returned to the ordinary resource pool for recycling.
[0044] Load forecasting formula: ; In the formula: This represents the predicted load, which is the target load value output by the scheduling module and is used to guide the dynamic allocation decisions of cloud computing resources such as CPU, memory, and bandwidth. This indicates the business cycle weight, which is associated with the business scenario label of the access module and is preset by the system based on historical business patterns. This indicates historical load, which is the historical average load data corresponding to the business cycle. For example, the historical search volume during the project review period on the 5th of each month, or the load value corresponding to the data access volume during the financial closing period on the 20th of each month. This represents the real-time load weight, which is a fixed coefficient preset by the system to balance historical business patterns with the current real-time state of the system. This indicates real-time load, referring to the actual load data of the system within the past hour, which directly reflects the current real-time resource occupancy.
[0045] The control module runs through the entire process of multimodal heterogeneous data access, cloud computing elastic storage, multi-source dynamic knowledge fusion, decision-oriented intelligent retrieval engine, scenario-based recommendation, and intelligent scheduling of cloud computing resources. It establishes a five-dimensional permission model based on roles, scenarios, data sensitivity, decision scenario sensitivity, and business lifecycle. It receives the storage location, data tags, and access permission metadata table from the storage module and encrypts the data transmitted by the access module using the AES-256 algorithm. Based on the resource allocation strategy of the scheduling module, it restricts the resource access permissions of unauthorized modules. For the search results of the retrieval module, customized anonymization is implemented according to role, business scenario, data sensitivity, decision-making scenario sensitivity, and business life cycle: During the R&D phase (core stage), the decision-making level can see the specific value, and the department manager can see the "±5% range", such as 1 million to 1.05 million yuan; after the launch (non-core stage), the department manager can only see the "±10% range". Five-dimensional desensitization ratio core configuration rules: Decision-making level - R&D scenario: When data sensitivity is the core and decision-making scenario sensitivity is high, if the business lifecycle is in the R&D phase, the anonymization rate is 0%, and the explicit value can be viewed; Department Manager - R&D Scenario: When data sensitivity is the core and decision-making scenario sensitivity is high, if the business lifecycle is in the R&D phase, the anonymization ratio is 5%, and only the ±5% range is viewed; if the business lifecycle is after launch, the anonymization ratio is adjusted to 10%, and only the ±10% range is viewed. Execution layer - financial scenario: When data sensitivity is the core and decision-making scenario sensitivity is medium, regardless of the stage of the business lifecycle, the anonymization ratio is 15%, and only the ±15% range value is viewed; Execution layer - supply chain scenario: When the data sensitivity is critical and the decision-making scenario sensitivity is medium, the anonymization ratio is 10% regardless of the stage of the business lifecycle, and only the ±10% range value is viewed.
[0046] Implementation details of blockchain audit logs: Architecture Design: A consortium blockchain architecture is adopted, with nodes including IT management, compliance and supervision departments, and core business departments, to ensure multi-party supervision of log recording; Consensus mechanism: The PBFT (Practical Byzantine Fault Tolerance) algorithm is adopted to ensure that trusted logs can still be generated normally when some nodes are abnormal; The core log fields include visitor ID, access timestamp, operation type (query / download / modify), unique data identifier, data sensitivity level, user role, business scenario, business lifecycle stage, and blockchain block height, covering full-process traceability information. Security mechanism: Logs are encrypted using the SHA-256 hash algorithm to ensure they cannot be tampered with; compliance departments can query logs in real time through the system backend to complete source tracing and auditing.
[0047] By using blockchain technology to record access logs for sensitive data in each module, including the visitor, time, operation content, and associated module, an immutable audit record is generated. At the same time, abnormal permission logs are pushed to the scheduling module to trigger the resource access restriction mechanism.
[0048] Desensitization interval calculation formula: ; This represents the range of values after desensitization. It is the output result of the control module after desensitizing sensitive data and is presented in the form of a range. The specific range is dynamically adjusted according to user roles and business lifecycles. This refers to the original values, i.e., the specific values of the core data that have not been anonymized, including the true values of sensitive data such as the value of core patents, the bottom price of customer transactions, and project budgets; The desensitization ratio is determined by the five-dimensional permission model, corresponding to the department manager during the R&D phase. =5%, corresponding to the department manager after launch. =10%, the corresponding percentage of the decision-making level at the core stage. =0, you can directly view the specific value; Indicates the lower limit of the interval. This indicates the upper limit of the interval.
[0049] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0050] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An organization information intelligent retrieval and recommendation system based on data cloud computing, characterized in that, The access module, the storage module, the fusion module, the retrieval module, the recommendation module, the scheduling module and the control module are included. The access module collects static files and real-time data, uses a three-dimensional classification algorithm to label the data, and forms a corresponding relationship table. The storage module establishes a three-layer storage architecture, uses an intelligent engine to schedule data, and generates a metadata table. The fusion module preferentially calls data in the high access frequency storage area and the real-time data cache area to establish a double-layer knowledge architecture, the static layer handles entity relationships and eliminates ambiguity, the real-time layer allocates resources to update the graph according to the business urgency, and generates a library table after resolving conflicts. The retrieval module supports multiple retrieval methods to form a retrieval, analysis, risk labeling decision support closed loop, analyzes the query intent, extracts key information for decision queries, executes reasoning, and outputs indicators and comparison results according to roles. The recommendation module receives the results and indicators of the retrieval module, combines user portraits and scene tag libraries to push information, and generates a comprehensive analysis table recording the recommended effect and resource occupation through user feedback. The scheduling module dynamically allocates cloud computing resources, predicts load, splits microservices and presets resource pools, schedules resources according to business needs, and outputs resource allocation strategies. The control module builds a five-dimensional permission model including roles, scenes, data sensitivity, decision scene sensitivity and business life cycle, and records sensitive data access logs. 2.The data cloud computing based organization information intelligent retrieval and recommendation system according to claim 1, characterized in that, The access module interfaces with business platforms including ERP systems, production line sensors and OA office systems to realize the collection of static files and real-time data. A three-dimensional classification algorithm of mode, time sequence and business value is used to classify data by mode, update time sequence and business value, synchronously associate business scene tags and business urgency, and automatically complete label labeling through preset rules table of business departments. For unstructured data, core information is extracted to form a table of classification tags and data entity correspondence, and output to the corresponding storage area of the storage module. 3.The data cloud computing based organization information intelligent retrieval and recommendation system according to claim 2, characterized in that, The storage module receives the corresponding relationship table output by the access module, establishes a three-layer hybrid storage architecture of real-time data cache area, high access frequency storage area and low access frequency storage area based on distributed file system and object storage; the real-time data cache area stores high-frequency stream data; the high access frequency storage area stores information with high retrieval frequency and high business correlation; the low access frequency storage area stores information with low retrieval frequency; An intelligent access frequency adaptive migration engine is loaded to migrate data to the adaptive storage area based on business urgency label and retrieval frequency, and generate a metadata table recording storage location, data label and access permission, which is synchronized to the fusion module and the control module. 4.The organization information intelligent retrieval and recommendation system based on data cloud computing according to claim 3, characterized in that, The fusion module reads the metadata table of the storage module, preferentially calls data in the high access frequency storage area and the real-time data cache area, and establishes a double-layer knowledge architecture of static basic layer and real-time dynamic layer; the static basic layer combines the core entity dictionary of each business scene to extract entity and relationship triples, eliminate ambiguity through entity linking algorithm, and filter redundant data; The real-time dynamic layer calls data in the real-time data cache area and allocates update resources according to the urgency label through a business priority update scheduler. After resolving data conflicts by using business rules priority and confidence bottom-up algorithm, a library table corresponding to the relationship between real-time updated knowledge graph and business scenarios is generated and pushed to the retrieval module. 5.The organization information intelligent retrieval and recommendation system based on data cloud computing according to claim 4, characterized in that, The retrieval module accesses the library table output by the fusion module, supports keyword retrieval, semantic retrieval, and multi-hop reasoning retrieval, forms a closed loop of retrieval, analysis, and risk labeling decision support, analyzes the query intent through the keyword and scenario correspondence table, and labels the corresponding risks for different types of queries; For decision queries, key information is extracted from the candidate document library, multi-hop reasoning is performed based on the knowledge graph, and finally the corresponding decision indicators and structured comparison results are output according to the role and business scenario; the retrieval results and decision indicators are synchronized to the recommendation module. 6.The organization information intelligent retrieval and recommendation system based on data cloud computing according to claim 5, characterized in that, The recommendation module receives the retrieval results and decision indicators from the retrieval module, combines static and dynamic two-dimensional user portraits and decision scenario label library, and pushes information according to the current business scenario and portrait characteristics of the user; The static dimension includes user roles and permission levels; the dynamic dimension includes search history, document type preference, and current business tasks; Through user behavior feedback and statistical resource occupation, a comprehensive analysis table of recommendation effect and resource occupation is formed and synchronized to the scheduling module. 7.The organization information intelligent retrieval and recommendation system based on data cloud computing according to claim 6, characterized in that, The scheduling module dynamically optimizes the allocation of cloud computing resources based on the comprehensive analysis table of the recommendation module, the knowledge update frequency of the fusion module, and the decision query concurrency of the retrieval module; load prediction is based on historical business cycle data and real-time load; Split microservice architecture and preset business-specific resource pool, schedule resources according to business cycle or urgency requirements; Synchronize resource allocation strategies to the control module. 8.The organization information intelligent retrieval and recommendation system based on data cloud computing according to claim 7, characterized in that, The five-dimensional permission model of the control module includes role, scene, data sensitivity, decision scene sensitivity, and business life cycle, receives the metadata table from the storage module, and encrypts the data transmitted by the access module; based on the resource allocation strategy of the scheduling module, limit the resource access rights of unauthorized modules; Desensitize the retrieval results according to the five-dimensional model; Record sensitive data access logs and generate audit records, and push permission exception logs to the scheduling module.
Citation Information
Patent Citations
Product decision optimization method and device based on mapping knowledge domain, equipment and medium
CN120707188A
Cooperative processing and intelligent conversion method for multi-mode service data
CN121029411A