Intelligent retrieval and archiving platform for business administration electronic archives supported by cloud computing
The five-layer architecture supported by cloud computing solves the problems of data dispersion, inefficient retrieval, and delayed risk warning in the management of electronic business registration records. It realizes the integration of data and precise binding of entities, and improves the standardization level and supervision efficiency of business registration record management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-10
AI Technical Summary
Traditional electronic record management for business administration suffers from problems such as data dispersion, heterogeneous data formats, inconsistent fields, difficulties in entity identification and association, limited search functions, and delayed risk warnings, making it difficult to meet the needs of efficient modern business administration.
A five-layer core architecture supported by cloud computing is constructed, including a multi-source data collection standardization layer, a regulatory knowledge graph construction layer, an electronic archive intelligent archiving layer, an archive panoramic retrieval and visualization layer, and an operational risk early warning push layer, to achieve data integration, entity disambiguation, accurate archive binding, multi-dimensional retrieval, and risk early warning.
It has achieved standardization of business registration file management and improved data utilization efficiency, provided efficient search functions and timely risk warnings, and helped regulatory authorities to comprehensively and accurately grasp enterprise information, thereby enhancing the pertinence and foresight of supervision.
Smart Images

Figure CN121636788A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the field of information technology of business administration, in particular to a cloud computing supported intelligent retrieval and archiving platform for electronic archives of business administration. BACKGROUND
[0002] With the continuous deepening of digital transformation in the field of government affairs, the demand for intelligent management of electronic archives of business administration is increasingly urgent. The mature application of cloud computing, big data, knowledge graph and other technologies provides a solid support for the integration, analysis and utilization of multi-source business data. The business activities of enterprises in the whole life cycle will generate a large amount of archive data, covering multiple business links such as establishment registration, qualification license, change record, administrative punishment, annual report and cancellation registration. These data sources are scattered and diverse, including both structured data and offline paper material digital data. Supervision departments need to comprehensively, accurately and timely grasp the relevant information of enterprises to realize precise supervision and efficient service, and promote the continuous optimization of market environment. Under this background, the management of electronic archives of business administration gradually transforms from the traditional scattered storage and manual processing mode to the integrated, intelligent and visual direction. It is urgent to build a comprehensive platform that can realize multi-source data integration, intelligent archiving, accurate retrieval and risk early warning to meet the efficient development needs of modern business administration.
[0003] The traditional electronic archive management technology of business administration has many obvious shortcomings and cannot adapt to the development requirements of digital supervision. There is a lack of unified standardized processing procedures in the data collection link. The heterogeneous data formats obtained from different channels are different and the fields are not unified. The problems of data duplication, deletion and conflict are prominent, forming an "information island" and hindering the effective integration and utilization of data. In terms of entity recognition and association, there is a lack of scientific disambiguation algorithm and association analysis mechanism, which is prone to repeated records of the same subject and errors in entity relationship matching, resulting in insufficient data accuracy. The triggering mechanism of the archiving process is single and relies on manual operation or single business linkage. Some special sources or forms of archives are difficult to archive in time. The binding of archives and enterprise entities lacks multi-dimensional verification and the matching accuracy is low. The retrieval function is limited to simple keyword matching or single condition query, which cannot meet the multi-dimensional retrieval needs of complex supervision scenarios. The retrieval results lack integrated associated information and cannot fully present the overall picture of enterprise archives. Risk early warning relies on manual inspection and cannot systematically extract risk characteristics, lacks quantitative analysis and grading mechanism, leading to incomplete risk identification and untimely early warning, lagging supervision response and difficulty in effectively preventing and resolving business risks. SUMMARY
[0004] The cloud computing supported business management electronic archive intelligent retrieval and archiving platform provided by the present application realizes intelligent management of the whole process of business archives through a five-layer core architecture, integrates heterogeneous data through a multi-source data acquisition standardization layer, and completes standard processing, defines entities and relationships, and eliminates entity ambiguity through a regulatory knowledge graph construction layer to form an enterprise full life cycle regulatory knowledge graph; relying on the multi-scene triggering mechanism and multi-dimensional matching technology of the electronic archive intelligent archiving layer, accurate archive binding and compliance archiving are realized; with the help of the multi-element retrieval function and special templates of the archive panoramic retrieval visualization layer, the archive retrieval efficiency is improved; through the risk feature extraction and risk level division of the business risk early warning pushing layer, closed-loop early warning is realized, and the platform comprehensively solves the problems of data dispersion, inaccurate matching, inefficient retrieval, and lagging risk early warning in business archive management, and provides efficient support for regulatory work.
[0005] The present application provides the following technical solutions to solve the above technical problems: a cloud computing supported business management electronic archive intelligent retrieval and archiving platform, which comprises: A multi-source data acquisition standardization layer is used to connect through official interfaces, comply with cloud crawlers, and obtain multi-source business data through offline data digitization interfaces, and output business supervision data sets after format unification, field calibration, and data cleaning processing; A regulatory knowledge graph construction layer is used to define core entities and regulatory related relationships, use an entity correlation algorithm for entity disambiguation, extract correlation information between entities and store it through a distributed graph database to form an enterprise full life cycle regulatory knowledge graph; the entity correlation algorithm is used to calculate the identity of entities in different data sources based on attribute matching degree, scene weight, relationship level, and time sequence matching degree; An electronic archive intelligent archiving layer is used to divide archive categories and complete initial collection, start the archiving process through a multi-scene triggering mechanism, and complete archive and enterprise entity binding through multi-dimensional feature matching fusion technology, synchronously update archive associated attributes, and carry out integrity verification; An archive panoramic retrieval visualization layer is used to provide accurate retrieval, combined retrieval, and fuzzy retrieval functions, preset high-frequency regulatory demand retrieval templates, optimize retrieval results through the entity correlation algorithm, and display archives and associated information through a three-region linkage interface; A business risk early warning pushing layer is used to extract risk features from the knowledge graph and archived archives, calculate risk confidence and divide risk levels through a risk confidence algorithm, push early warning information, and track processing results to form a closed-loop management; the risk confidence algorithm is used to calculate the overall risk level of an enterprise based on the weighted sum of multiple risk features.
[0006] Further, the specific way of acquiring multi-source business data in the multi-source data acquisition standardization layer is: through official interface docking, structured business data is retrieved from enterprise credit information public system and market supervision integrated platform by using encryption transmission protocol; through compliance cloud crawler, enterprise equity change and qualification license public information are regularly captured from public websites in compliance with crawler protocol and authorization, and source marks are made; through offline data digitization interface, paper material scans are received, and OCR identification technology is used to convert them into structured text data, which is compared with standard field templates.
[0007] Further, the specific steps of format unification, field calibration and data cleaning processing in the multi-source data acquisition standardization layer are: format unification, converting the heterogeneous data obtained by interface docking, crawler capture and offline recognition into the same data format; field calibration, dividing the fields according to priority, the first priority is unified social credit code and enterprise name, the second priority is legal representative and registered address, the third priority is business scope and registered capital, and the format and content of each field are calibrated; data cleaning, removing duplicate data with unified social credit code as the unique identifier, calling supplementary interface to complete missing fields, and recording missing reasons when interface calling fails; after processing, outputting business supervision data set and storing it to cloud storage node.
[0008] Further, the core entities in the supervision knowledge graph construction layer include: enterprise, legal representative, shareholder, registration authority, license and approval agency, administrative punishment agency and archive management personnel; the defined supervision related relationships include the employment relationship between enterprise and legal representative, the shareholding relationship between enterprise and shareholder, the registration relationship between enterprise and registration authority, the license qualification granting relationship between enterprise and license and approval agency, the punishment relationship between enterprise and administrative punishment agency, and the management relationship between archive management personnel and business archives, each relationship contains two mandatory attributes of relationship effective time and relationship source.
[0009] Further, the specific way of extracting the correlation information between entities in the supervision knowledge graph construction layer is: extracting the employment and change correlation between enterprise and legal representative from enterprise registration archives and change record materials; extracting the shareholding and equity change correlation between enterprise and shareholder from company charter and equity agreement; extracting the registration correlation between enterprise and registration authority from registration application file and approval record; extracting the qualification granting and extension correlation between enterprise and license and approval agency from qualification application materials and approval file; extracting the punishment and execution correlation between enterprise and administrative punishment agency from penalty decision and execution receipt; extracting the archiving and consulting correlation between archive management personnel and business archives from archive management record.
[0010] Further, the specific way of adopting the business file entity correlation degree algorithm for entity disambiguation in the supervision knowledge graph construction layer is: extracting attribute information and associated business data of the entity to be disambiguated, and substituting them into the business file entity correlation degree algorithm to calculate the correlation strength The mathematical expression of the business file entity correlation degree algorithm is: Wherein is the degree of fit of the entity information to be disambiguated and the existing entity information in the entity library, is the weight of the supervision scene to which the entity to be disambiguated belongs, is the relationship level coefficient of the entity to be disambiguated and the existing entity in the entity library, is the time sequence matching degree of the associated business time of the entity to be disambiguated and the related business time of the existing entity in the entity library; the preset entity disambiguation correlation strength threshold is 0.75, when the calculated correlation strength , it is determined that the entity to be disambiguated and the corresponding entity in the entity library are the same subject, and the entity record with complete attribute information is retained; when the calculated correlation strength , the key attributes of the entity to be disambiguated and the existing entity in the entity library are compared, and if the key attributes are completely consistent, they are merged into the same entity, and if there is a conflict in the key attributes, the conflict information is marked as an abnormal entity and retained, so as to complete entity disambiguation; the entity library contains complete attribute information of enterprises, legal representatives, shareholders, registration authorities, license and approval agencies, administrative punishment agencies and file management personnel, wherein the enterprise attributes include unified social credit code, enterprise name, establishment date, registered address, the legal representative attributes include name, ID number, service status, the shareholder attributes include name or name, holding ratio, investment method, the registration authority attributes include agency name, administrative division code, business scope, the license and approval agency attributes include agency name, approval scope, approval level, the administrative punishment agency attributes include agency name, punishment type, jurisdiction, and the file management personnel attributes include name, department, operation permission.
[0011] Further, the enterprise full life cycle supervision knowledge graph in the supervision knowledge graph construction layer comprises: entity information, entity attribute information and entity supervision correlation information; wherein the entity supervision correlation information covers the establishment registration, change registration and cancellation registration correlation between enterprises and registration authorities, the qualification application, grant, extension and cancellation correlation between enterprises and license and approval agencies, the illegal record and penalty decision execution correlation between enterprises and administrative punishment agencies, the service and change correlation between enterprises and legal representatives, the holding and equity change correlation between enterprises and shareholders, and the filing, viewing and modification correlation between file management personnel and business archives.
[0012] Further, the specific way of dividing the archive category in the electronic archive intelligent archiving layer is: according to the business type of industrial and commercial supervision, it is divided into establishment registration type, change registration type, license qualification type, annual report type, administrative punishment type and cancellation registration type; wherein, the establishment registration type contains establishment registration application, company charter, capital verification report and shareholder identification; the change registration type contains change registration application table, relevant change proof materials and approval decision; the license qualification type contains qualification application table, approval file and qualification certificate scan; the annual report type contains annual report declaration table and relevant supporting materials; the administrative punishment type contains administrative punishment decision, penalty notice and execution receipt; the cancellation registration type contains cancellation registration application, liquidation report and cancellation approval file.
[0013] Further, the multi-scene triggering mechanism in the electronic archive intelligent archiving layer specifically includes: business system linkage triggering, after the enterprise handles establishment registration, change registration, license qualification application, administrative punishment record, annual report submission and cancellation registration business, the corresponding business system automatically sends a trigger signal to the archiving layer; timing scanning triggering, scanning the unarchived industrial and commercial data at fixed time every day, and starting the archiving process when the data meeting the archiving conditions are identified; offline data digitization completion triggering, after the offline paper materials are recognized by OCR, field extraction and format verification, the archiving process is automatically triggered; manual triggering, after the archive management personnel upload the industrial and commercial archive materials which cannot be processed through the above automatic process due to special source or form, the archiving trigger instruction is initiated through the system specified function.
[0014] Further, the specific way of using multi-dimensional feature matching fusion technology to complete the binding of archives and enterprise entities in the electronic archive intelligent archiving layer is: extracting the enterprise name, registered address, legal representative name, business scope recorded in the to-be-archived archives, as well as the business type and generation time information of the archives itself, and calling the enterprise entity data in the supervision knowledge graph; through multi-dimensional feature matching fusion technology, the enterprise-related information of the archives and the enterprise entity data are matched in attribute dimension, business dimension and time sequence dimension in multiple levels, the preliminary matching of enterprise name, registered address, legal representative name, business scope and enterprise entity data is carried out first, then the secondary matching is carried out in combination with the relevance of archive business type and enterprise historical business, and finally the final matching is completed according to the degree of coincidence of archive generation time and enterprise-related business occurrence time; when the three-layer matching results all meet the preset matching standard, the archives and the corresponding enterprise entity are automatically bound.
[0015] Further, the specific implementation of the precise search, combined search and fuzzy search functions in the archive panoramic search visualization layer is as follows: the precise search function takes unique identification information as the search basis, supports user input of enterprise unified social credit codes, archive unique numbers and legal representative ID numbers, and the system locates corresponding archives and associated information through accurate matching; the combined search function supports user selection of at least two conditions from enterprise name, archive category, archiving time range, associated business type and generating agency for combined query, and the system performs screening according to the logical relationship of the conditions set by the user; the fuzzy search function supports user input of partial enterprise name, partial legal representative name and archive content containing text, and the system searches for all relevant archive information through keyword matching, with the search range covering enterprise basic information, archive text content and associated business note fields.
[0016] Further, the high-frequency regulatory demand search template in the archive panoramic search visualization layer specifically includes: an enterprise qualification check template, which internally stores unified social credit codes, enterprise names, qualification types, qualification validity periods and approval agency search condition items; an administrative penalty record query template, which contains enterprise identification information, penalty authorities, penalty time ranges and penalty type search condition items; an annual report traceability template, which sets enterprise names, unified social credit codes, report years and report submission state search condition items; a stock right change tracking template, which covers enterprise information, change time intervals, original shareholder information and new shareholder information search condition items; and an archive completeness check template, which contains enterprise identification information, archive categories, archiving time ranges and completeness check state search condition items.
[0017] Further, the way of extracting risk features from the knowledge graph and the archived archives in the operating risk early warning push layer is as follows: the risk features extracted from the knowledge graph and the archived archives include: records of expired or about-to-expire qualifications in the association information between the enterprise and the license approval agency; records of 3 or more administrative penalties within the last 24 months in the association information between the enterprise and the administrative penalty authority; records of frequent stock right changes without required recordation in the association information between the enterprise and the shareholders; conflict records between enterprise registration information and actual operation information; and risk features extracted from the archived archives include: records of untimely submission or false information in the annual report archived archives; records of unapproved changes in specific registration matters in the change registration archived archives; abnormal signs of unliquidated debts in the cancellation registration archived archives; and records of repeated defaults without performance of compensation obligations in the operating contract archived archives.
[0018] Further, the business risk early warning pushing layer adopts an enterprise management risk confidence algorithm to calculate the risk confidence and divide the risk level in the following specific manner: extracting enterprise risk features from the knowledge graph and the archived archives, including license qualification expiration status, multiple administrative penalty records, annual report non-submission status, abnormal information of equity change, and records of inconsistent registration information and actual operation information; putting each risk feature into the enterprise management risk confidence algorithm to calculate the overall risk confidence of the enterprise The mathematical expression of the enterprise management risk confidence algorithm is as follows: Wherein is the matching degree of a single risk feature, with a value range of 0-1; is the weight of the corresponding risk feature, which is set according to the importance of supervision, and the total weight is 1; three risk level thresholds are preset, which are high risk threshold 0.8, medium risk threshold 0.5, and low risk threshold 0.2; when the risk confidence , it is determined as a high risk level, and the corresponding enterprise has a major business risk hidden danger; when , it is determined as a medium risk level, and the corresponding enterprise has a relatively obvious business risk; when , it is determined as a low risk level, and the corresponding enterprise has a potential business risk; when , it is determined as a no-risk level, thereby completing the risk confidence calculation and risk level division, and providing a basis for subsequent early warning information pushing.
[0019] Compared with the prior art, the cloud computing supported business management electronic archive intelligent retrieval and archiving platform has the following beneficial effects: 1. The present application realizes the comprehensive integration and standard management of business data through the cooperative operation of the standardized processing of multi-source data collection and the construction of the supervision knowledge graph, and forms a standardized data set through format unification, field calibration and data cleaning of the heterogeneous data obtained through multiple channels, and completes entity disambiguation in combination with an entity correlation degree algorithm, constructs a supervision knowledge graph covering the whole life cycle of an enterprise, clearly presents various entities and associated relationships, and realizes the accurate binding of archives and enterprise entities through the multi-scene triggering mechanism and multi-dimensional feature matching fusion technology of the electronic archive intelligent archiving layer, synchronously completes integrity verification, effectively solves the problems of data dispersion, weak correlation and untimely archiving in traditional archive management, significantly improves the standardization level and data utilization efficiency of business archive management, provides a solid data foundation for supervision work, and helps the supervision department to comprehensively and accurately grasp the enterprise operation related information.
[0020] Secondly, the application builds an efficient business archive application and supervision system through rich search functions and intelligent risk early warning mechanism, the archive panoramic search visualization layer provides accurate, combined and fuzzy search functions, matches high-frequency supervision demand search templates, combines entity correlation degree algorithm to optimize search results, enables users to quickly locate the required archives and associated information, reduces search difficulty and time cost, the business risk early warning push layer extracts multiple risk characteristics from the knowledge graph and archived archives, divides risk levels through a scientific risk confidence algorithm, timely pushes early warning information and tracks processing results, forms a closed-loop management, this mechanism helps the supervision department to perceive enterprise business risks in advance, enhances the pertinence and foresight of supervision, reduces the waste of supervision resources, effectively maintains the stability of market operation order, and provides protection for the healthy development of the market.
[0021] Other advantages, objects, and features of the application will be set forth in part in the following specification, and in part will become apparent to those skilled in the art upon examination of the following specification, or can be learned by practice of the application. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0023] Figure 1 Flow chart of the intelligent search and archiving platform for cloud computing supported business management electronic archives; Figure 2 Data transmission relationship diagram of each level of the business management electronic archive platform; Figure 3 Enterprise life cycle supervision knowledge graph construction and application association diagram. DETAILED DESCRIPTION
[0024] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined application purpose, the following will combine the drawings and the preferred embodiments to specifically describe the specific implementation, structure, features and effects of the present application.
[0025] Embodiment one: Market supervision department enterprise qualification compliance sampling and risk early warning scene.
[0026] The market supervision department plans to carry out qualification compliance spot check on key industry enterprises in the jurisdiction, and synchronously investigate the operation risk, carry out related work through the platform. Firstly, the platform multi-source data collection standardization layer is connected with the enterprise credit information public system and the market supervision integrated platform through official interface, adopts an encryption transmission protocol to call enterprise structured business data, and guarantees the safety in the data transmission process; the compliance cloud crawler follows the crawler agreement and authorization, regularly crawls the qualification license public information of enterprises on the public website and marks the source, ensures the compliance and traceability of data acquisition; meanwhile, the enterprise qualification paper materials scanned by offline submission are converted into structured text data by using OCR identification technology, compared with the standard field template, the digital transformation of offline unstructured data is realized, and the data acquisition channel is widened. Then, the above-mentioned heterogeneous data is uniformly processed in format, the data of different sources and different formats are converted into the same standard format, the processing obstacles caused by data format difference are eliminated; the core fields such as unified social credit code and enterprise name are calibrated according to priority, the consistency and accuracy of key information are ensured; the repeated data is removed with the unified social credit code as the unique identifier, the storage resources are avoided from being occupied by data redundancy, the missing fields are completed by calling the supplementary interface, the missing reasons are recorded when the interface calling fails, the data integrity is guaranteed, and finally the business supervision data set is stored in the cloud storage node, the safe storage and convenient calling of data are realized, and high-quality data support is provided for subsequent processing, such as shown in Figure 1
[0027] The supervision knowledge graph construction layer defines the core entities such as enterprises, license approval agencies and legal representatives, and the supervision related relationships such as the qualification granting relationship between enterprises and license approval agencies and the employment relationship between enterprises and legal representatives, and each relationship contains the relationship validity time and source attribute, clearly defines the core dimensions and key information of supervision association, and lays a foundation for subsequent association analysis. The qualification granting and continuation association information between enterprises and license approval agencies is extracted from the enterprise qualification application materials and approval files, the employment association information between enterprises and legal representatives is extracted from the enterprise registration archives, the supervision association details between entities are fully captured, and various business associations are clear and traceable. The business file entity association degree algorithm is used for disambiguation of enterprise entities in different data sources, and the mathematical expression of the business file entity association degree algorithm is as follows: wherein, is the degree of fit between the entity information to be disambiguated and the existing entity information in the entity library, is the weight of the supervision scene to which the entity to be disambiguated belongs, is the relationship level coefficient of the entity to be disambiguated and the existing entity in the entity library, is the time sequence matching degree of the business time of the entity to be disambiguated and the related business time of the existing entity in the entity library; the preset entity disambiguation association strength threshold is 0.75, and when the calculated association strength When the calculated correlation strength is greater than the preset threshold, it is determined that the to-be-resolved entity and the corresponding entity in the entity library are the same subject, and the entity record with complete attribute information is retained. When the calculated correlation strength is greater than the preset threshold, it is determined that the to-be-resolved entity and the corresponding entity in the entity library are the same subject, and the entity record with complete attribute information is retained. Figure 3 When the calculated correlation strength is greater than the preset threshold, it is determined that the to-be-resolved entity and the corresponding entity in the entity library are the same subject, and the entity record with complete attribute information is retained.
[0028] When the enterprise qualification archives complete data collection and knowledge graph association, the electronic archives intelligent archiving layer triggers the business system linkage trigger mechanism due to the qualification extension business handled by the enterprise in the past, automatically starts the archiving process, ensures seamless connection between business handling and archiving, and improves the timeliness of archiving. The platform divides the archives into licensed qualification categories, clearly containing qualification application forms, approval documents, and scanned copies of qualification certificates, so that the archives are clearly classified and clearly attributed, facilitating subsequent retrieval and management. Then, the information recorded in the archives, such as the enterprise name, legal representative name, and qualification type, is extracted, the enterprise entity data in the supervision knowledge graph is retrieved, and multi-dimensional feature matching and fusion technology is used for three-level matching of attributes, businesses, and time sequences - first, the core information of the enterprise is compared with the entity data to preliminarily lock the matching target; then, the archive business type is associated with the historical qualification business of the enterprise to narrow the matching range; finally, the consistency of the archive generation time and the qualification approval time is verified to ensure the accuracy of the matching result and avoid incorrect or missed binding of the archives. After the three-level matching meets the standard, the automatic binding of the archives and the enterprise entity is completed, the associated attributes of the archives are updated synchronously to keep the archive information consistent with the enterprise dynamics, and the integrity verification ensures that the archived archive materials are complete and no key files are missing, providing complete archive basis for subsequent retrieval and supervision.
[0029] Regulatory personnel can use the enterprise qualification verification template in the panoramic file retrieval visualization layer to combine search criteria such as enterprise name, qualification type, and approval agency for combined searches. The template includes built-in search conditions required for high-frequency supervision, reducing manual input and improving search convenience. The platform optimizes search results using an entity correlation algorithm, filtering irrelevant information for more accurate matching and avoiding interference from invalid information. A three-area linked interface clearly displays enterprise qualification files, approval agency information, and qualification granting relationships, intuitively presenting core files, related entities, and business relationships, helping regulatory personnel quickly obtain comprehensive information and significantly improving search efficiency and information completeness. Simultaneously, the operational risk warning push layer extracts risk characteristics of an enterprise's impending qualification expiration from the knowledge graph, combining this with related information such as one administrative penalty record within the past 24 months, comprehensively covering both qualification compliance and administrative penalty risk points. The overall risk level is calculated using an enterprise management risk confidence algorithm, the mathematical expression of which is: ,in, The matching degree for a single risk feature, with a value ranging from 0 to 1. The weights corresponding to the risk characteristics are set according to regulatory importance, with a total weight of 1; three risk level thresholds are preset: high risk threshold 0.8, medium risk threshold 0.5, and low risk threshold 0.2; when the risk confidence level... When a risk level is determined to be high, the corresponding enterprise faces significant operational risks; when When the risk level is determined to be medium, the corresponding enterprise faces significant operational risks; when When it is determined to be at a low risk level, the corresponding enterprise has potential operational risks; when When the risk level is determined to be no risk, the risk status of the enterprise is accurately quantified. The risk confidence level is determined to be 0.65, which is considered a medium risk level, allowing regulators to quickly clarify the degree of risk. The platform immediately pushes early warning information to the corresponding regulators to ensure that the risk is detected in a timely manner. It tracks and records the process of regulators conducting on-site inspections and urging enterprises to renew their qualifications, forming a closed-loop management system to ensure that the risk is effectively handled and to prevent compliance risks caused by expired qualifications.
[0030] In summary, in scenarios involving market supervision departments' random checks and risk warnings on enterprise qualification compliance, the platform integrates multi-channel industrial and commercial data through a multi-source data collection standardization layer, providing high-quality data support after standardized processing; the regulatory knowledge graph construction layer accurately identifies entity relationships and disambiguates them, forming a complete regulatory knowledge graph; the electronic file intelligent archiving layer achieves efficient classification and accurate binding of qualification files; the file panoramic retrieval visualization layer improves retrieval efficiency through dedicated templates and optimized algorithms; and the operational risk warning push layer comprehensively extracts risk characteristics, quantifies the level using an enterprise management risk confidence algorithm, and pushes warnings, forming a closed-loop management system. This collaborative effort across the entire process significantly improves the accuracy and timeliness of supervision, effectively preventing qualification compliance risks.
[0031] Example 2: Archival management personnel handle various types of enterprise archives filing and historical business retrieval scenarios.
[0032] The document management personnel need to handle the filing of three types of documents for a certain enterprise: establishment registration, annual reports, and administrative penalties. They also need to assist business departments in querying the enterprise's historical equity change records. This invention platform efficiently completes the relevant tasks. First, the platform's multi-source data collection standardization layer connects to the integrated market supervision platform through an official interface to obtain structured data from the enterprise's establishment registration and annual reports, ensuring the authority and accuracy of the data sources. It then uses a compliance cloud crawler to capture publicly available information on the enterprise's administrative penalties and marks the sources, ensuring that data acquisition is compliant and traceable. Finally, it receives offline submitted paper administrative penalty decisions and execution receipts, converts them into structured text through OCR recognition, and compares them with standard fields, achieving the digital transformation of offline paper documents and opening up online and offline data channels. This level then performs format standardization on the data, eliminating format differences between data from different sources and creating conditions for subsequent standardized processing. First-priority fields such as the unified social credit code and company name are calibrated according to priority to ensure the accuracy of core identification information. In the data cleaning stage, duplicate data is removed using the unified social credit code as the unique identifier to avoid redundant archival information. Missing penalty execution date fields are supplemented to ensure data integrity. The final output industrial and commercial supervision dataset provides a standardized, accurate, and complete data foundation for archiving, correlation matching, and subsequent retrieval. Figure 2 As shown.
[0033] The regulatory knowledge graph construction layer extracts employment-related information from the company's registration file, shareholding-related information from equity agreements, and penalty-related information from penalty decisions, comprehensively capturing the company's relationships across different business scenarios and making its business trajectory clearly traceable. It defines core entities such as the company, shareholders, administrative penalty authorities, and file management personnel, along with their corresponding shareholding, penalty, and management relationships, clarifying the regulatory and management relationship dimensions between these entities and providing a clear logical framework for entity matching and relationship queries. The business registration-file-entity correlation algorithm is used to complete entity disambiguation; the mathematical expression of the business registration-file-entity correlation algorithm is: ,in, To determine the degree of fit between the entity information to be disambiguated and the existing entity information in the entity database, The weight of the regulatory scenario to which the entity to be disambiguated belongs. The relationship hierarchy coefficient between the entity to be disambiguated and existing entities in the entity database. This refers to the temporal matching degree between the business time associated with the entity to be disambiguated and the related business time of existing entities in the entity database; the preset threshold for entity disambiguation association strength is 0.75. When the calculated association strength... When the entity to be disambiguated is determined to be the same subject as the corresponding entity in the entity database, the entity record with complete attribute information is retained; when the calculated association strength is obtained... During the process, the key attributes of the entity to be disambiguated are compared with those of existing entities in the entity database. If the key attributes are completely identical, they are merged into the same entity; if the key attributes conflict, they are marked as abnormal entities and the conflict information is retained. This completes entity disambiguation. By comprehensively analyzing attribute information, related business data, and other factors, the identity of the same entity in different data sources is accurately determined, avoiding confusion or duplication of entity information and ensuring the uniqueness and accuracy of entity information. Entity information, attribute information, and full lifecycle regulatory association information are stored in a distributed graph database, achieving efficient data storage and rapid retrieval. The resulting complete regulatory knowledge graph provides solid data support for the accurate binding of archives and enterprise entities and the retrieval of historical business associations.
[0034] In terms of the intelligent archiving layer for electronic records, after a company completes its establishment registration, the business system automatically sends a trigger signal to automatically connect the business processing and archiving processes, ensuring timely archiving of establishment registration documents. The platform categorizes relevant documents into establishment registration categories, including establishment registration applications, company articles of association, capital verification reports, and other materials, making document classification clear and facilitating subsequent retrieval and management. Multi-dimensional feature matching and fusion technology is used to complete the binding and integrity verification with the enterprise entity, ensuring accurate document attribution and complete materials. Through a timed scanning trigger mechanism, the platform scans for unarchived annual report materials of the enterprise at fixed times each day. Once the materials meet the archiving conditions, the process is initiated, avoiding delays in archiving annual report documents due to omissions. This also completes classification, binding, and verification archiving. Offline administrative penalty paper materials are automatically triggered for archiving after OCR recognition and format verification, achieving rapid digitization of offline documents and categorizing them as administrative penalties for binding and archiving. Additionally, a special format of supplementary penalty explanation material is manually uploaded by the document management personnel and then triggers an archiving command through a designated function in the system. This compensates for the limitations of the automatic archiving process, ensuring that special types of documents can also be successfully archived, comprehensively covering the archiving needs of all types of documents.
[0035] Business departments need to query the historical equity change information of a company. Records management personnel can use the equity change tracking template in the panoramic archive retrieval visualization layer to search by entering the company name and the target change time range. The template specifically covers the core conditions required for equity change queries, eliminating the need for manual filtering of numerous search items and improving the convenience of the search operation. The platform optimizes search results using an entity correlation algorithm, filtering out information irrelevant to the query conditions to make the search results more accurate and avoid interference from invalid information. A three-area linked interface displays the company's equity change archives, information on original and new shareholders, and related change filing materials, intuitively presenting the core archives and related entity information of equity changes, helping business departments quickly obtain complete historical equity change information, significantly improving the efficiency of archive queries and the completeness of information acquisition. Simultaneously, the operational risk warning push layer extracts the risk characteristic of the company's failure to submit an annual report on time from the archived archives. Combined with the records in the knowledge graph that the company has two equity change failures to be filed in a timely manner, it comprehensively captures the potential risks of the company in report submission, equity change filing, and other aspects. The enterprise management risk confidence algorithm is used to calculate the risk confidence level. The mathematical expression of the enterprise management risk confidence algorithm is: ,in The matching degree for a single risk feature, with a value ranging from 0 to 1. The weights corresponding to the risk characteristics are set according to regulatory importance, with a total weight of 1; three risk level thresholds are preset: high risk threshold 0.8, medium risk threshold 0.5, and low risk threshold 0.2; when the risk confidence level... When a risk level is determined to be high, the corresponding enterprise faces significant operational risks; when When the risk level is determined to be medium, the corresponding enterprise faces significant operational risks; when When it is determined to be at a low risk level, the corresponding enterprise has potential operational risks; when When the risk level is determined to be no risk, the result is 0.35, which accurately quantifies the enterprise's risk level and classifies it as low risk, allowing relevant personnel to clearly understand the degree of risk. The platform pushes early warning information to ensure that the risk is noticed in a timely manner. The file management personnel track and record the processing results of the business departments reminding enterprises to supplement the filing materials, forming a closed-loop management system, helping enterprises to standardize relevant business operations and reduce the probability of subsequent risks.
[0036] In summary, in scenarios involving enterprise archiving and historical business retrieval by archivists, the platform's multi-source data collection standardization layer connects online and offline data channels, outputting standardized datasets; the regulatory knowledge graph construction layer captures multi-dimensional related information of enterprises, ensuring the accuracy and uniqueness of entity information; the intelligent electronic archive archiving layer achieves comprehensive and timely archiving of various types of archives through multi-scenario triggering mechanisms; the panoramic archive retrieval visualization layer, leveraging equity change tracking templates and optimized algorithms, quickly outputs complete search results; and the operational risk warning push layer accurately identifies potential risks and pushes reminders, helping to standardize business operations. The platform simplifies the archiving process and improves retrieval efficiency throughout, providing efficient support for archive management and business operations.
[0037] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A cloud computing supported intelligent search and archiving platform for business electronic files, characterized in that, The platform comprises: A multi-source data acquisition standardization layer: used for acquiring multi-source business data through official interface docking, compliant cloud crawler, offline data digitization interface, outputting business supervision data set after format unification, field calibration and data cleaning processing; A supervision knowledge graph construction layer: used for defining core entities and supervision related relationships, performing entity disambiguation using an entity correlation degree algorithm, extracting correlation information between entities and storing them in a distributed graph database to form an enterprise life cycle supervision knowledge graph; the entity correlation degree algorithm is used to calculate the identity of entities in different data sources based on attribute matching degree, scene weight, relationship level and time sequence matching degree; An electronic archive intelligent archiving layer: used for dividing archive categories and completing initial collection, starting the archiving process through a multi-scene triggering mechanism, binding archives and enterprise entities using multi-dimensional feature matching fusion technology, synchronously updating archive associated attributes and carrying out integrity verification; An archive panoramic search visualization layer: used for providing accurate search, combined search and fuzzy search functions, presetting high-frequency supervision demand search templates, optimizing search results using the entity correlation degree algorithm, and displaying archives and associated information through a three-region linkage interface; An operating risk early warning pushing layer: used for extracting risk features from the knowledge graph and archived archives, calculating risk confidence and dividing risk levels using a risk confidence algorithm, pushing early warning information and tracking processing results to form a closed-loop management; the risk confidence algorithm is used to calculate the overall risk level of an enterprise based on the weighted sum of multiple risk features.
2. The cloud computing supported intelligent retrieval and archiving platform for e-filing of business management according to claim 1, wherein, The specific way of acquiring multi-source business data in the multi-source data acquisition standardization layer is: through official interface docking, structured business data is retrieved from the enterprise credit information public system and the market supervision integrated platform using an encrypted transmission protocol; through compliant cloud crawler, enterprise equity changes and qualification license public information are regularly scraped from public websites in accordance with the crawler protocol and authorization, and source marks are added; through offline data digitization interface, paper materials scans are received, and OCR recognition technology is used to convert them into structured text data, which is compared with a standard field template.
3. The cloud computing supported intelligent retrieval and archiving platform for e-filing of business management according to claim 1, wherein, The specific steps of format unification, field calibration and data cleaning processing in the multi-source data acquisition standardization layer are: format unification: converting heterogeneous data obtained through interface docking, crawler scraping and offline recognition into the same data format; field calibration: dividing fields according to priority, the first priority being unified social credit code and enterprise name, the second priority being legal representative and registered address, and the third priority being business scope and registered capital, and calibrating the format and content of each field; data cleaning: removing duplicate data using unified social credit code as the unique identifier, calling a supplementary interface to complete missing fields, and recording missing reasons when interface calling fails; after processing, outputting business supervision data set and storing it in a cloud storage node.
4. The cloud computing supported e-filing intelligent retrieval and archiving platform for business management according to claim 1, characterized in that, The core entities in the supervision knowledge graph construction layer include: enterprises, legal representatives, shareholders, registration authorities, license and approval agencies, administrative punishment agencies, and archive management personnel; the defined supervision-related relationships include the employment relationship between an enterprise and a legal representative, the shareholding relationship between an enterprise and a shareholder, the registration relationship between an enterprise and a registration authority, the license and qualification granting relationship between an enterprise and a license and approval agency, the punishment relationship between an enterprise and an administrative punishment agency, and the management relationship between an archive management personnel and a business archive, each relationship includes two mandatory attributes: relationship effective time and relationship source.
5. The cloud computing supported e-filing intelligent retrieval and archiving platform for business management according to claim 1, characterized in that, The specific way of extracting the correlation information between entities in the supervision knowledge graph construction layer is: extracting the employment and change correlation between an enterprise and a legal representative from the enterprise registration archives and change record materials; extracting the shareholding and equity change correlation between an enterprise and a shareholder from the company charter and equity agreement; extracting the registration correlation between an enterprise and a registration authority from the registration application file and approval record; extracting the qualification granting and extension correlation between an enterprise and a license and approval agency from the qualification application material and approval file; extracting the punishment and execution correlation between an enterprise and an administrative punishment agency from the punishment decision and execution receipt; and extracting the archiving and consulting correlation between an archive management personnel and a business archive from the archive management record.
6. The cloud computing supported e-filing intelligent retrieval and archiving platform for business management according to claim 1, characterized in that, The enterprise full-life-cycle supervision knowledge graph in the supervision knowledge graph construction layer includes: entity information, entity attribute information, and entity-to-entity supervision correlation information; wherein the entity-to-entity supervision correlation information covers the establishment registration, change registration, and cancellation registration correlation between an enterprise and a registration authority, the qualification application, granting, extension, and cancellation correlation between an enterprise and a license and approval agency, the illegal record and punishment decision execution correlation between an enterprise and an administrative punishment agency, the employment and change correlation between an enterprise and a legal representative, the shareholding and equity change correlation between an enterprise and a shareholder, and the archiving, consulting, and modification correlation between an archive management personnel and a business archive.
7. The cloud computing supported e-filing intelligent retrieval and archiving platform for business management according to claim 1, characterized in that, The specific way of dividing the archive categories in the electronic archive intelligent archiving layer is: dividing into establishment registration category, change registration category, license and qualification category, annual report category, administrative punishment category, and cancellation registration category according to the business types of business supervision; wherein the establishment registration category includes establishment registration application, company charter, capital verification report, and shareholder identification; the change registration category includes change registration application table, relevant change proofing materials, and approval decision; the license and qualification category includes qualification application table, approval file, and qualification certificate scan; the annual report category includes annual report declaration table and relevant supporting materials; the administrative punishment category includes administrative punishment decision, punishment notice, and execution receipt; and the cancellation registration category includes cancellation registration application, liquidation report, and cancellation approval file.
8. The cloud computing supported e-filing intelligent retrieval and archiving platform for business management according to claim 1, characterized in that, The specific way of binding the archives and the enterprise entity in the electronic archives intelligent archiving layer by using the multi-dimensional feature matching fusion technology is: extracting the enterprise name, registered address, legal representative name, business scope recorded in the to-be-archived archives, and the business type and generation time information of the archives itself, and calling the enterprise entity data in the supervision knowledge graph; performing multi-level matching of the attribute dimension, business dimension and time sequence dimension of the enterprise-related information of the archives and the enterprise entity data by using the multi-dimensional feature matching fusion technology, first performing preliminary matching of the enterprise name, registered address, legal representative name, business scope and the enterprise entity data, then performing secondary matching in combination with the relevance of the business type of the archives and the historical business of the enterprise, and finally performing final matching according to the degree of fit between the generation time of the archives and the occurrence time of the business of the enterprise; when the three-layer matching results all meet the preset matching standard, the archives and the corresponding enterprise entity are automatically bound.
9. The cloud computing supported e-filing intelligent retrieval and archiving platform for business management according to claim 1, wherein, The specific implementation of the precise search, combined search and fuzzy search functions in the archives panoramic search visualization layer is: the precise search function takes unique identification information as the search basis, supports the user to input the enterprise unified social credit code, archive unique number and legal representative ID number, and the system locates the corresponding archives and associated information through accurate matching; the combined search function supports the user to select at least two conditions from the enterprise name, archive category, archiving time range, associated business type and generating agency for combined query, and the system filters according to the logical relationship of the conditions set by the user; the fuzzy search function supports the user to input part of the enterprise name, part of the legal representative name and text contained in the archive content, and the system searches all related archive information through keyword matching, and the search range covers enterprise basic information, archive text content and associated business note field.
10. The cloud computing supported e-filing intelligent retrieval and archiving platform for business management according to claim 1, wherein, The high-frequency regulatory demand search template in the archives panoramic search visualization layer specifically includes: an enterprise qualification checking template, which internally stores the unified social credit code, enterprise name, qualification type, qualification validity period and approval agency search condition items; an administrative penalty record query template, which contains the enterprise identification information, penalty organ, penalty time range, penalty type search condition items; an annual report tracing template, which sets the enterprise name, unified social credit code, report year, report submission state search condition items; a stock right change tracking template, which covers the enterprise information, change time interval, original shareholder information and new shareholder information search condition items; and an archive integrity checking template, which contains the enterprise identification information, archive category, archiving time range and integrity checking state search condition items.