Big data-based intelligent management and control site operation method, device, equipment and medium

By constructing a big data site panorama database and machine learning models, the problem of insufficient data integration and analysis capabilities in traditional site management systems has been solved. This has enabled the automated identification of abnormal rents and fake sites, improved information integrity and risk prediction capabilities, and supported refined operational decision-making.

CN120598641BActive Publication Date: 2025-12-12SHENZHEN LEAPFROG NEW TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511103584.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-12-12
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Traditional site management systems suffer from a lack of data integration capabilities, a deficiency in intelligent analysis capabilities, and outdated decision support methods during the process of enterprise scaling up. They are unable to achieve automatic collection of multi-source data, identification of abnormal rents, and real-time monitoring of fake sites, resulting in long information acquisition cycles, low completeness, insufficient analytical capabilities, and a lack of decision support.

Method used

By constructing a site panoramic database based on big data, and combining geographic location analysis and machine learning algorithms, we can achieve the cleaning, standardization, and unification of multi-source heterogeneous data. We can also build a rental anomaly identification rule engine and a fake site identification model, dynamically configure anomaly judgment logic, generate anomaly profiles and risk assessments, and provide support for site operation decisions.

Benefits of technology

It improves information integrity and analysis accuracy, enables automated identification of abnormal rents and fake venues, reduces manual investigation costs, enhances risk prediction capabilities and decision support efficiency, and meets the refined operation needs of enterprises.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598641B_ABST
    Figure CN120598641B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of resource management, and discloses a method, device, equipment and medium for intelligently managing and controlling site operation based on big data, which comprises the following steps: acquiring multi-source heterogeneous site related information in a preset data source, and generating a site panoramic database; collecting site leasing data based on the site panoramic database, constructing a rent abnormality identification rule engine, and calculating site information of each site; constructing a false site identification model, and continuously analyzing newly entered site information and site information in an operation process in a mode of fusing a machine learning algorithm and the rule engine to identify a risk site; and generating structured data containing site identification, an abnormal item list, an index score, a recommendation level and a strategy label according to the false site identification model and the site panoramic database. The method is suitable for whole life cycle management of site location, evaluation, monitoring and strategy optimization of enterprises such as logistics, e-commerce and manufacturing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of resource management, and in particular to a method and device for intelligent management and control of site operation based on big data, equipment and medium. BACKGROUND

[0002] In the process of enterprise scale expansion, site operation management faces multi-dimensional challenges. Traditional site management relies on manual screening, Excel recording and experience judgment, which has the following significant defects:

[0003] 1. Lack of data integration capability: existing systems can only manage single site basic information, and cannot automatically collect heterogeneous data such as rental platform, map service, enterprise internal system, etc. through multi-source interface, resulting in long information acquisition period and low completeness of nationwide site information, and it is difficult to form a global data view.

[0004] 2. Lack of intelligent analysis capability: site evaluation relies on manual experience, lacks dynamically configurable rent anomaly recognition rule engine, and cannot automatically analyze complex scenarios such as excessive deposit, overlapping rental period, and abnormal rent price increase; at the same time, risk identification is limited to manual audit level, without building a double-layer model combining machine learning and rule engine, making it difficult to identify false sites or potential operational risks in real time.

[0005] 3. Decision support means is backward: data visualization only realizes basic dashboard display, cannot integrate multi-dimensional information such as contract data, expense details, and operation status to form strategic output, and lacks intelligent recommendation mechanism based on site panoramic database, making it difficult to meet the enterprise's fine needs for site selection efficiency, cost control and risk prediction.

[0006] Therefore, a method is needed to solve at least one of the above problems. SUMMARY

[0007] The present application provides a method and device for intelligent management and control of site operation based on big data, equipment and medium, which solves the problem that the existing technology in the process of enterprise scale expansion, site operation management faces multi-dimensional challenges. Traditional site management relies on manual screening, Excel recording and experience judgment.

[0008] In a first aspect, the present application provides a method for intelligent management and control of site operation based on big data, which comprises:

[0009] Obtaining multi-source heterogeneous site related information in a preset data source, performing cleaning and standardization processing operation on structured data and unstructured data corresponding to the site related information, and combining geographic position analysis technology to unify the coordinate system and label system corresponding to the site related information, to generate a site panoramic database;

[0010] Based on the site panoramic database, site leasing data is collected, and a rent abnormality identification rule engine is constructed, which is used to dynamically configure abnormality judgment logic to process the collected leasing data, calculate site information of each site, and the site information includes abnormality portrait and key indicators.

[0011] A false site identification model is constructed, and a machine learning algorithm is combined with a rule engine to continuously analyze newly entered site information and site information during operation to identify risk sites and generate processing information corresponding to the risk sites and send the processing information to corresponding user terminals.

[0012] According to the false site identification model and the site panoramic database, structured data including site identification, abnormal item list, indicator score, recommended level, and strategy label is generated to support site operation decision-making.

[0013] In some embodiments, the structured data and unstructured data corresponding to the site-related information are subjected to cleaning and standardization processing operations, including removing duplicate data, filling or filtering missing values, unifying data formats, and semantic conversion processing, so that the same name different meanings or different name same meanings fields of different data sources form a unified semantic definition.

[0014] In some embodiments, the geographic location analysis technology is combined to unify the coordinate system and label system corresponding to the site-related information, including: based on site address or coordinate data, geographic coding analysis is performed to convert the coordinate system of different data sources into a unified geographic coordinate system or projection coordinate system; a standardized label system is established according to site use, regional attributes, traffic conditions, and other dimensions, so that each site corresponds to a unique geographic coordinate identifier and multi-dimensional classification label.

[0015] In some embodiments, the site panoramic database is generated, including: integrating and storing the site information after cleaning, standardization, and geographic labeling to form site panoramic data including site basic attribute information, rent cost data, physical space parameters, address coordinates, surrounding supporting facilities, real-time operation status, and historical leasing records, the historical leasing records including lease term, rent variation, contract performance, and lease termination records.

[0016] In some embodiments, the site panoramic database is used to collect site leasing data, and a rent anomaly identification rule engine is constructed, including: obtaining rent, deposit, property fee, water and electricity fee, decoration fee and operation data through a combination of offline batch collection and real-time streaming collection; constructing dynamically configurable abnormality judgment logic based on business rules, the abnormality judgment logic including at least rules of deposit being higher than a preset rent multiple, public area proportion exceeding a limit, new and old site lease period overlapping, long-term non-operation after payment, rent unit price anomaly rising and rent loss exceeding a limit, and the preset rent multiple and proportion threshold supporting dynamic adjustment through a business configuration interface.

[0017] In some embodiments, the collected leasing data is processed to calculate site information of each site, including abnormal portraits and key indicators, including: according to the judgment logic of the rent anomaly identification rule engine, grouping, aggregating and correlating the leasing data of a single site to identify abnormal rule items triggered by the site and generate an abnormal portrait; calculating key indicators based on abnormal rule triggering frequency, involved fee amount, influence degree on operation efficiency, etc., synchronizing the abnormal portrait and key indicators to a data analysis platform, and supporting manual notes and feedback corrections on abnormal identification results.

[0018] In some embodiments, a false site identification model is constructed, and a machine learning algorithm is used in combination with a rule engine to continuously analyze newly entered site information and site information during operation to identify risk sites, including: based on deposit and rent proportion, site operation rate, replacement site unit price change amplitude, discount implementation and historical blowout records, a double-layer identification model including a rule matching layer and a machine learning prediction layer is constructed; the rule engine is used to preliminarily filter information that obviously deviates from business logic, and a machine learning algorithm is used to perform pattern recognition on the filtered information to identify potential false information or high-risk operation states, and the identified risk sites are automatically marked and alarm information including risk level, abnormal type and influence range is generated.

[0019] In a second aspect, the present application provides a site operation intelligent management and control device based on big data, including:

[0020] An information acquisition unit is configured to acquire multi-source heterogeneous site-related information in a preset data source, perform cleaning and standardization processing operations on structured data and unstructured data corresponding to the site-related information, and unify coordinate systems and label systems corresponding to the site-related information by using a geographic position analysis technique, to generate a site panoramic database.

[0021] The data collection unit is configured to collect site leasing data based on the site panoramic database, and construct a rent abnormality identification rule engine configured to dynamically configure abnormality judgment logic to process the collected leasing data and calculate site information of each site, the site information including an abnormality portrait and key indicators.

[0022] The model construction unit is configured to construct a fake site identification model, and continuously analyze newly entered site information and site information in an operation process in a manner of combining a machine learning algorithm and a rule engine to identify a risk site and generate processing information corresponding to the risk site and send the processing information to a corresponding user terminal.

[0023] The decision support unit is configured to generate structured data including site identification, an abnormality item list, indicator scores, recommendation levels, and strategy labels based on the fake site identification model and the site panoramic database, and provide support for site operation decision-making.

[0024] In a third aspect, the present application provides a computer device, which includes a memory and a processor.

[0025] The memory is configured to store a computer program.

[0026] The processor is configured to execute the computer program and implement any of the site operation intelligent management and control methods provided in the embodiments of the present application when executing the computer program.

[0027] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to make the processor implement any of the site operation intelligent management and control methods provided in the embodiments of the present application.

[0028] The application discloses a big data intelligent management and control site operation method and device, equipment and medium. The provided method automatically acquires heterogeneous data through a multi-source interface and processes the data in a standardized manner, constructs a panoramic database covering site lifecycle information, improves information integrity, and solves the problems of hysteresis and one-sidedness of traditional manual collection. Based on a configurable rule engine and multi-dimensional index analysis, the method realizes automatic identification of complex abnormal scenarios such as excessive deposit and overlapping rental period, improves identification accuracy, and significantly reduces manual investigation costs. The method fuses machine learning and rule engine to construct a false site identification model, monitors site operation status in real time, and automatically generates alarms, forms a "recognition-processing-closed loop" mechanism, effectively improves risk prediction ability, and reduces user loss and operation risk. The method integrates multi-dimensional data through a unified visual board, outputs structured results including site score, recommended level and strategy label, provides real-time and intelligent support for site selection, lease renewal and lease termination decisions, improves abnormal processing efficiency, and meets the needs of fine operation of enterprises.

[0029] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the application. BRIEF DESCRIPTION OF DRAWINGS

[0030] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0031] Figure 1 is a step schematic flow chart of a big data intelligent management and control site operation method provided by the embodiments of the application;

[0032] Figure 2 is a schematic block diagram of a big data intelligent management and control site operation device provided by the embodiments of the application;

[0033] Figure 3 is a structural schematic block diagram of a computer device provided by the embodiments of the application.

[0034] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the application. DETAILED DESCRIPTION

[0035] With reference to the drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts are within the scope of the present application.

[0036] The flowcharts shown in the drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be further decomposed, combined or partially merged, so that the actual execution order can be changed according to the actual situation.

[0037] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and the appended claims of the present application, unless otherwise clear from the context, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0038] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0039] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The embodiments described below and the features in the embodiments can be combined with each other without conflict.

[0040] In the process of enterprise expansion, site operation management faces multi-dimensional challenges. Traditional site management relies on manual screening, Excel recording and experience judgment, which has the following significant defects:

[0041] Lack of data integration capability: the existing system can only manage single site basic information, and cannot automatically collect heterogeneous data such as rental platform, map service, enterprise internal system, etc. through multi-source interface, resulting in long information acquisition period, low completeness, and difficulty in forming a global data view.

[0042] Lack of intelligent analysis capability: site evaluation relies on manual experience, lacks dynamically configurable rent anomaly recognition rule engine, and cannot automatically analyze complex scenarios such as excessive deposit, overlapping rental period, abnormal rent price increase, etc. At the same time, risk identification is limited to manual review, without building a double-layer model combining machine learning and rule engine, making it difficult to identify false sites or potential operational risks in real time.

[0043] The decision support means is backward: data visualization only realizes basic board display, cannot integrate multi-dimensional information such as contract data, cost details, operation status to form strategic output, and lacks intelligent recommendation mechanism based on site panoramic database, and it is difficult to meet the fine needs of enterprises for site selection efficiency, cost control and risk prediction.

[0044] Although there are site information libraries or BI tools in the prior art, none of them systematically combines automatic multi-source data integration, dynamic rule configuration, machine learning risk identification and visualized strategy output, and cannot solve the chain problem of 'incomplete data collection - insufficient analysis capability - lack of decision support'. The present application fills the gap in the intelligent management and control of the whole link of site operation by constructing a complete process of multi-source information collection and integration, intelligent identification of rent anomalies, dynamic risk warning and strategic data output.

[0045] Please refer to Figure 1 , Figure 1 is a schematic flowchart of a site operation intelligent management and control method based on big data provided by the present application. The method is applied to a computer device, which can be deployed on a single server or a server cluster. It can also be deployed on a handheld terminal, a notebook computer, a wearable device or a robot, etc.

[0046] It should be noted that the acquisition of any information mentioned in the provided method is in accordance with relevant regulations and with the consent of the user, and does not infringe on the user's privacy or violate relevant laws and regulations.

[0047] As shown in Figure 1 , the specific steps of the site operation intelligent management and control method based on big data include steps S101 to S104.

[0048] S101, acquire multi-source heterogeneous site related information in a preset data source, perform cleaning and standardization processing operations on structured data and unstructured data corresponding to the site related information, and unify the coordinate system and label system corresponding to the site related information by combining geographic location analysis technology, to generate a site panoramic database.

[0049] Specifically, multi-dimensional site related information is acquired through a preset data source (including a lease platform, a map service system, an enterprise internal management system, etc.), structured data (such as tabular rent and area data) and unstructured data (such as supporting facilities described in text and picture type site panorama) are cleaned, standardized and geotagged, and finally a comprehensive database containing site full life cycle information is generated.

[0050] Multi-source data collection is achieved by network crawler technology (such as HTTP protocol interface call) to grab site basic information (rent, area, leasing status) from the leasing platform, through the map service API (get address coordinates, surrounding traffic and supporting facility data, through enterprise internal system interface (such as ERP, OA system) to synchronize historical leasing records, operation status and financial data. Support dynamic expansion of data source, through configuration file to define new data source access protocol (RESTful, SOAP, etc.) and data mapping rules, realize flexible access of heterogeneous data source.

[0051] Data cleaning and standardization includes: structured data processing: remove duplicate records (based on site unique identifier such as address + area combination), fill in missing values (such as filling in missing rent values by averaging similar site data in the same area), unify data format (such as unifying "rent unit" of different platforms to "yuan / square meter / month"). Unstructured data processing: through natural language processing (NLP) technology to analyze text description, extract keywords (such as "direct subway" "with fire safety qualification") to generate standardized labels; feature extraction is performed on picture data to generate structured facility list (such as parking space number, floor height parameter).

[0052] Geographic coordinates and label system are unified through geocoding to convert text addresses such as "XX City, XX District, XX Road" into a unified geographic coordinate system (such as WGS84, GCJ-02), and to convert the coordinate system of different data sources to ensure that spatial data can be overlaid and analyzed.

[0053] Label system construction is achieved by establishing a multi-level label system, including site purpose (warehouse, office, production), regional attribute (core business district, industrial park), traffic level (subway coverage, highway entrance within 1 km); secondary label is refined into supporting facilities (such as whether to contain loading and unloading platform, whether to support layered leasing), forming a unique geographic coordinate identifier and multi-dimensional classification label for each site.

[0054] Panorama database generation is achieved by integrating cleaned structured data, geographic coordinates and label data, using relational database (such as MySQL) or distributed database (such as HBase) storage, forming a database table containing the following core fields: basic information: site ID, name, ownership attribute, building age; spatial data: area, floor height, public proportion, coordinate system coordinate; leasing data: current rent, historical rent fluctuation record, deposit standard, property fee; operation data: current tenant information, historical leasing period, reason for termination label, vacancy duration; geographic label: regional economic level (generated according to surrounding GDP data), traffic convenience score (based on subway station distance calculation).

[0055] Through multi-source automatic collection instead of manual input, the problem of field information fragmentation in traditional management is solved, the information coverage is improved, and enterprise-level field data assets are formed. The unified coordinate system and label system enable field data from different sources to have spatial analysis and cross-comparison capabilities (such as automatic clustering according to regional rent levels), providing a unified data foundation for subsequent rent assessment and risk identification, and avoiding analysis errors caused by "data silos".

[0056] S102, based on the field panoramic database, collecting field leasing data, and constructing a rent anomaly identification rule engine, the rent anomaly identification rule engine is used to dynamically configure abnormal judgment logic to process the collected leasing data, calculate the field information of each field, and the field information includes abnormal portrait and key indicators.

[0057] Specifically, based on the field panoramic database generated in step S101, detailed leasing business data (rent, deposit, property fee, etc.) is collected, abnormal scenarios are identified through a dynamically configurable rule engine, and an abnormal portrait (triggered abnormal rule set) and key indicators (such as abnormal severity score) of a single field are output.

[0058] Leasing data collection includes: offline batch collection: historical leasing data (such as rent adjustment records in the past 3 years, deposit payment vouchers) is extracted from the financial system and contract management system regularly through ETL tools (such as Apache Sqoop), and stored in a data warehouse (such as Hive). Real-time streaming collection: new generated leasing data (such as deposit terms of newly signed contracts, monthly property fee bills) is captured in real time through a message queue (such as Kafka), and real-time cleaning and format conversion are performed through a real-time computing engine (such as Flink), to ensure that the data is synchronized to the rule engine in seconds.

[0059] Rule engine construction: dynamic rule configuration: provide a visual rule editing interface, support business personnel to customize abnormal judgment logic, rule types include: numerical type rule: such as "deposit ≥ 3 times monthly rent" (trigger "high deposit" exception), "public area ratio > 30% and not mentioned in advance" (trigger "public area exception"); logical type rule: such as "new lease period and historical lease period overlap for more than 30 days" (trigger "lease period overlap"), "6 consecutive months of property fee payment records are 0 and the field state shows 'operating'" (trigger "unoperated field exception"); trend type rule: such as "rent unit price same period increase > regional average increase 150%" (trigger "rent abnormal rise"). Rule priority management: support setting priority (high / medium / low) for different rules, such as "lease period overlap" is set as high priority, which triggers the exception and blocks the contract approval process immediately.

[0060] Abnormal image and index calculation: By grouping and aggregating the lease data of a single site (e.g., aggregating rent data by year, quarter), the rule engine matches each rule, records the triggered rule items and trigger time, and generates an "abnormal image" containing the abnormal type, involved amount, and duration.

[0061] Key indicator calculation: According to the priority and impact of abnormal rules, a quantitative score is generated (e.g., "excessive deposit" scores 10 points, "rental period overlap" scores 20 points), and an "abnormal index" is calculated based on the frequency of abnormal occurrences, which is used for subsequent risk ranking and visualization.

[0062] Instead of the traditional way of manually screening each table, the system realizes 7x24 hours of full data monitoring, significantly improves the efficiency of abnormal identification, and significantly reduces labor costs. Through dynamic rule configuration, it supports customizing abnormal standards for different industries (logistics and warehousing, commercial real estate, industrial plants), such as configuring the tolerance threshold of "public area" in the logistics industry, solving the problem of rule fixation in traditional systems, and expanding the system's adaptability from a single scenario to multi-format management.

[0063] S103, construct a fake site identification model, use machine learning algorithm and rule engine to analyze new input site information and site information in operation, identify risk sites, and generate processing information corresponding to risk sites and send to corresponding user terminal.

[0064] Specifically, by constructing a double-layer risk identification model of "rule engine preliminary screening + machine learning deep analysis", the new input site information and the operating site are continuously monitored, and the fake site or high-risk operating state is identified in real time, and the processing information containing the risk level is generated and pushed to the management terminal.

[0065] Multi-dimensional index system: Collect the following core indicators as model input: financial indicators: deposit / rent ratio (threshold ≥2 triggers rule warning), rent collection rate (continuous 3 months <80%); operation indicators: site operation rate (actual use area / lease area <60%), replacement unit price change (new rent is more than 30% lower than the average price of surrounding areas, which may indicate a low-price trap); historical records: historical warehouse explosion times (warehouse site accidents due to load overload), rent termination rate (12 months within 3 times); related data: site owner credit record (obtained through enterprise credit API), surrounding similar site vacancy rate (grabbed through map service API).

[0066] Double-layer model architecture: Rule engine layer: First filter obvious abnormal information through preset business rules, such as "address resolution result is invalid coordinates" and "owner has court credit record", directly marked as "high risk" and blocked the input process. Machine learning layer: Gradient Boosting Decision Tree (GBDT), Random Forest and other algorithms are used to train the model based on historical risk site samples (such as confirmed false sites and sites with rent disputes leading to litigation), learn risk feature patterns (such as "low deposit + high rent increase + frequent ownership changes"). The model is updated regularly (weekly automatic synchronization of the latest sample data), supporting the identification of potential risks (such as complex contract clause vulnerabilities) that are not covered by the rule engine.

[0067] When the model identifies a risk site, it automatically generates alert information containing risk types (such as "false address" and "operational abnormalities") and evidence chains (triggered rule items / model prediction probability), which are pushed to regional managers through SMS and system pop-up windows. After the managers handle it (such as verifying and terminating the leasing process), they will record the handling results in the system, mark the risk as "closed loop", and synchronize the handling records to the panoramic database, forming a "identification-warning-handling-record" whole-process management and control.

[0068] Unlike the traditional "after-the-fact" mode, the machine learning method can identify potential risk features and provide 3-6 month early warning for unexploded risks (such as site value plummeting due to surrounding planning adjustments), improving the risk identification accuracy compared to single rule engine.

[0069] S104、According to the false site identification model and the site panoramic database, structured data containing site identification, abnormal item list, index score, recommended level and strategy label are generated to support site operation decision-making.

[0070] Specifically, integrate site panoramic data, abnormal identification results and risk assessment model output, build visual operation dashboard, and output structured data containing strategy suggestions through standardized interface, realize seamless connection with existing enterprise systems (office, recruitment, contract system), provide quantitative support for site selection, lease renewal, reconstruction and other decisions.

[0071] Visual dashboard construction: Multi-dimensional dashboard modules are designed: Global monitoring: National site distribution heatmap (colored by rental level and risk level), core indicator dashboard (vacancy rate, percentage of abnormal sites, rental compliance rate); Individual site details: Basic site information, historical anomaly records, current risk level, comparison with surrounding competitors (rent, area, supporting facilities); Anomaly handling tracking: List of pending anomalies, countdown timer for handling, historical handling effect analysis (e.g., changes in renewal rate after handling a certain type of anomaly). Drill-down analysis is supported: Clicking on a high-risk area in the heatmap allows drilling down to a list of anomaly details for all sites in that area, supporting export to Excel or PDF reports.

[0072] Structured data output: Generates a standardized data structure containing the following fields: { "Site Identifier": "CN_SH_20250729_001", "Anomaly List": ["Excessive Deposit", "Overlapping Lease Terms"], "Indicator Score": { "Rental Reasonableness": 65 points, "Risk Level": "Medium Risk", "Operational Efficiency": 72 points}, "Recommendation Level": "Grade B (Recommendation to renegotiate rent)", "Strategy Tags": ["Rental Reduction Negotiation", "Vacation Period Optimization", "Legal Clause Verification"]}; where "Strategy Tags" automatically matches preset strategies based on the anomaly type (e.g., "Excessive Deposit" matches the "Initiate Deposit Installment Payment Negotiation" strategy).

[0073] The system integration architecture adopts a microservice architecture (Spring Cloud) deployment, with each module (data collection, rules engine, and visualization service) deployed independently, and providing a unified interface (RESTful format) to the outside world through an API gateway. It supports containerized deployment (Docker + Kubernetes) to achieve dynamic expansion of computing resources and meet the needs of concurrent analysis of tens of thousands of sites. Standardized interface adaptation: Data integration documentation is provided, supporting integration with enterprise OA systems for alarm push notifications, with procurement systems for site screening condition synchronization, and with data dashboard systems for real-time data visualization.

[0074] By replacing traditional manual recommendations with a "recommendation level + strategy tag" approach, the site selection decision-making cycle is shortened while the accuracy of strategy matching is improved, addressing the subjectivity issues of experience-based decision-making. The microservice architecture supports rapid iteration (such as the addition of a "carbon neutral site assessment" module), and containerized deployment allows the system's capacity to scale linearly with enterprise expansion, avoiding the performance bottlenecks of traditional monolithic architectures and adapting to the evolving management needs of enterprises from regional to national operations.

[0075] In some embodiments, the structured data and unstructured data corresponding to the site-related information are subjected to cleaning and standardization processing operations, including: removing duplicate data, filling or filtering missing values, unifying data formats, and semantic conversion processing, so that the same name different meanings or different name same meanings fields of different data sources form a unified semantic definition.

[0076] For multi-source heterogeneous site-related information, structured data (such as table data) and unstructured data (such as text, pictures) are subjected to cleaning and standardization processing, including removing duplicate data, filling or filtering missing values, unifying data formats, and solving the semantic differences of fields in different data sources through semantic conversion, such as "same name different meanings" (such as A system "area" refers to building area, B system refers to use area) or "different name same meanings" (such as "rental price per unit area" and "rental price per unit area"), to form a unified data definition.

[0077] The data cleaning operation includes: de-duplication processing: based on the unique identification of the site (such as "address + property number" combination) or feature field (area ± 5% error is considered as the same site), the duplicate data is removed through database uniqueness constraint or ETL tool (such as ApacheNiFi). Missing value processing: mark the missing data of key fields (such as rent, address) as "to be supplemented" and trigger the manual review process; for non-key fields (such as decoration period), use interpolation method (mean / median filling in the same area) or machine learning model (such as random forest) to predict and fill.

[0078] Standardization and semantic conversion include: format unification: unify the date format of different data sources to "YYYY-MM-DD", unify the amount unit, and unify the area unit to "square meters". Semantic mapping: establish a data dictionary (MetadataDictionary) to define the standard name of the field and the business meaning (such as unify "rental period" to "contractual rental duration, unit: month"), and through the mapping table, the fields of each data source are associated to the standard semantics (for example: the "rental period" of A system and "rental period" of B system are both mapped to the standard field "rental period"). Unstructured data analysis: through NLP technology, the entity recognition is performed on the text description (such as extracting the "property fee inclusion status" label from "rental fee includes property fee"), and the domain dictionary (such as real estate industry term library) is combined to ensure the accuracy of semantic analysis.

[0079] Through de-duplication, completion and semantic unification, the data accuracy is improved, and analysis errors caused by data ambiguity (such as mistakenly taking "building area" as "use area" to calculate rental cost) are avoided. The unified semantic definition eliminates the understanding barriers between data sources, provides a standardized data basis for subsequent multi-dimensional correlation analysis (such as comparison of rental data from different platforms), and reduces the cost of manual coordination in data integration.

[0080] In some embodiments, the geographic location analysis technology unifies the coordinate system and label system corresponding to the site-related information, including: based on the site address or coordinate data, the geographic coding analysis is performed to convert the coordinate system of different data sources into a unified geographic coordinate system or a projection coordinate system; and a standardized label system is established according to the dimensions of site use, regional attribute, traffic condition and the like, so that each site corresponds to a unique geographic coordinate identifier and a multi-dimensional classification label.

[0081] By using the geographic location analysis technology, the coordinate system of different data sources is converted into a unified coordinate system (such as the 2000 Geodetic Coordinate System), and a standardized label system is established based on the dimensions of site use, regional attribute, traffic condition and the like, so that each site has a unique geographic coordinate identifier and a multi-dimensional classification label.

[0082] The geographic coordinate unification includes: address analysis: through a geographic coding service (such as Geocoding API), a text address such as “XX City, XX District, XX Number” is converted into longitude and latitude coordinates, and for an address that fails to be analyzed (such as a fuzzy address), fuzzy matching is performed in combination with surrounding POI (point of interest) data. Coordinate system conversion: using a coordinate conversion algorithm (such as a seven-parameter conversion model) or an open source tool (such as Proj.4), coordinate systems of different data sources are unified into a reference coordinate system (such as a unified plane coordinate system) specified by an enterprise, so that spatial data can be superimposed and displayed on the same map.

[0083] Label system construction: multi-level label definition: first-level label: site use (office / warehouse / production), regional attribute (city core area / development zone / rural area), traffic condition (subway coverage / highway direct access / main road side); second-level label: refined attribute (such as “Class A office building” and “co-working space” under the “office” use, and “subway station within 500 meters” and “public transportation hub within 1 kilometer” under the “traffic condition”). Automatic label generation: through a spatial analysis tool (such as ArcGIS), the straight-line distance from the site to the subway station is calculated, and the “subway coverage” label is automatically matched; based on a public regional planning file (such as “XX District is a key industrial park”), a “regional attribute” label is generated through text matching.

[0084] The unified coordinate system makes it possible to analyze the site distribution and density (such as drawing a national warehouse site heat map), and supports regional management based on geographic fences (such as automatically calculating the average rent price of all sites in a certain business circle).

[0085] In some embodiments, the generating the site panoramic database comprises: integrating and storing the cleaned, standardized and geotagged site information to form site panoramic data containing site basic attribute information, rent cost data, physical space parameters, address coordinates, surrounding supporting facilities, real-time operation status and historical leasing records, the historical leasing records including lease term, rent variation, contract performance and lease termination records.

[0086] The cleaned and standardized site information (including geotag) is integrated and stored to form a "panoramic database" covering the whole life cycle of the site, and the data content includes basic attributes, rent cost, physical space parameters, address coordinates, surrounding supporting facilities, real-time operation status and historical leasing records (such as lease term, rent variation, contract performance, lease termination reason).

[0087] Data integration range: basic attributes: site ID, owner, building structure, property right term, fire safety qualification; rent cost: current rent, historical rent adjustment record, deposit, property management fee, water and electricity fee allocation method; physical space: building area, usable area, floor height, bearing load, number of parking spaces; operation data: current tenant name, lease status (in use / empty / lease termination), empty start time, historical lease termination times and reason label (such as "excessive rent" "insufficient area").

[0088] The storage architecture design adopts a "main table + dimension table" structure: the main table stores the site core identifier (site ID) and associated foreign keys, and the dimension table stores basic attributes, rent data, operation records, etc., and the data decoupling is realized through foreign key association. For more than one million site data, Hadoop distributed file system (HDFS) is used to store unstructured data (pictures, documents), HBase is used to store structured data, and Spark is used for distributed query calculation.

[0089] The panoramic database covers the whole process data of the site from site selection, leasing to lease termination, supports tracing the historical performance of the site (such as the rent increase trend of a site in the past three years), provides data basis for lease renewal negotiation and asset valuation. The unified database breaks down the barriers between business systems (such as data interconnection between financial system and operation system), so that the management can real-time view the correlation between "rent cost - operation efficiency - risk status" of the site (such as whether the high rent site corresponds to high lease termination rate), and promotes data-driven cross-departmental collaborative decision-making.

[0090] In some embodiments, the site panoramic database collects site rental data, and constructs a rent anomaly identification rule engine, comprising: obtaining rent, deposit, property fee, water and electricity fee, decoration fee and operation data through a combination of offline batch collection and real-time streaming collection; based on business rules, a dynamically configurable abnormality judgment logic is constructed, which at least includes rules such as deposit higher than a preset rent multiple, public area proportion exceeding limit, new and old site rental period overlap, long-term non-operation after payment, rent unit price abnormal rise, and rent loss exceeding limit, and the preset rent multiple and proportion threshold can be dynamically adjusted through a business configuration interface.

[0091] Through a combination of offline batch collection (historical data) and real-time streaming collection (newly generated data), rental-related data (rent, deposit, property fee, etc.) is obtained, a dynamically configurable rent anomaly identification rule engine is constructed, and self-defined abnormality judgment logic (such as deposit higher than 3 times the monthly rent, public area exceeding 30%) is supported, and the threshold in the rule (such as "preset rent multiple" and "proportion threshold") can be flexibly adjusted through a business interface.

[0092] Data collection method: offline collection: historical data (such as rent adjustment records for the past 5 years, deposit payment vouchers) is extracted from the financial system and contract management system through ETL tools every morning and stored in the data lake (Data Lake). Real-time collection: new contract creation and fee bill generation events in the contract management system are monitored in real time through API interface, data is transmitted at millisecond level using Kafka message queue, and after cleaning by Flink real-time computing engine, it is written into the rule engine database.

[0093] Rule engine configuration: visual rule editing: a graphical interface is provided, and business personnel configure rules through "drag-and-drop" operation (such as selecting the "deposit" field, setting "greater than" "3 times the monthly rent", and triggering "high deposit" anomaly), supporting logical operator combination (AND / OR / NOT). Threshold dynamic adjustment: different thresholds are set for different industries / regions (such as 25% for "public area proportion exceeding limit" in first-tier city warehouse sites, and 30% for second-tier city), and the threshold configuration result is stored in a configuration center (such as Apollo) that can be hot updated, and can take effect without restarting the system.

[0094] Through offline + real-time collection, "historical backtracking + real-time monitoring" of rental data is realized, the rule engine covers high-frequency abnormal scenarios such as deposit, public area, rental period, rent increase, etc., and the coverage of traditional manual review is improved. Support dynamic adjustment of rule thresholds for different business scenarios (such as temporarily relaxing the time threshold of "non-operation after payment" during e-commerce promotions), avoid "misjudgment" or "omission" caused by rule fixation, and make the system adaptability expand from a single enterprise to multi-format group management.

[0095] In some embodiments, the collected lease data is processed to calculate site information for each site, including an abnormal profile and key indicators, including: according to the judgment logic of the rent abnormality identification rule engine, the lease data of a single site is grouped, aggregated and associated, the abnormal rule items triggered by the site are identified and an abnormal profile is generated; based on the dimensions of abnormal rule triggering frequency, involved fee amount, influence on operating efficiency, key indicators are calculated, the abnormal profile and key indicators are synchronized to the data analysis platform, and manual notes and feedback corrections of abnormal identification results are supported.

[0096] Based on the judgment logic of the rent abnormality identification rule engine, the lease data of a single site is grouped, aggregated (such as aggregating rent by year / quarter) and associated (such as the association between rent and property fees), the triggered abnormal rule items are identified and an abnormal profile (including abnormal type, triggering time, involved amount) is generated, and key indicators (such as abnormal index) are calculated from the dimensions of triggering frequency, fee impact, and operating efficiency, supporting manual notes and corrections of the identification results.

[0097] Abnormal profile generation includes: grouping and aggregation analysis: grouping lease data by “site ID + time period”, calculating indicators such as deposit / rent ratio and public area proportion in each period, comparing with the threshold value in the rule engine, recording the triggered rule items and times (for example: a site triggered the “rent unit price abnormal rise” rule twice in Q2 of 2025). Through SQL association query or graph database (such as Neo4j) analysis of the association between abnormal rules (such as whether “high deposit” is often accompanied by “rent period overlap”), an abnormal rule association graph is generated to assist in identifying systematic risks.

[0098] Key indicator calculation: abnormal index model: abnormal index = Σ (rule priority weight x triggering times) + involved amount coefficient (for example: high priority rule triggers 20 points each time, medium priority 10 points, and additional 15 points for involved amount over 100,000 yuan). An abnormal result editing portal is provided on the data analysis platform, administrators can mark “misjudgment” or add notes, and feedback information is used as training data for rule optimization.

[0099] Through abnormal profile and key indicators, abstract risks are converted into quantifiable scores (such as abnormal index ≥ 80 points marked as red alert), which facilitates management to quickly locate high-risk sites and improves decision-making efficiency. Artificial feedback data reversely optimizes the rule engine (such as correcting the threshold of misjudgment rules), forming a closed loop of “identification-verification-optimization”, so that the accuracy of abnormal identification gradually improves over time, and the misjudgment rate can be reduced within 3 months.

[0100] In some embodiments, the construction of the false site identification model, using a machine learning algorithm combined with a rule engine, continuously analyzes newly entered site information and site information during operation to identify risk sites, including: based on the deposit and rent ratio, site operation rate, replacement site unit price change range, preferential implementation and historical blowout records, etc. Multidimensional indicators, a double-layer identification model containing a rule matching layer and a machine learning prediction layer is constructed; through the rule engine, the information obviously contrary to the business logic is preliminarily screened, and then the machine learning algorithm is used to identify the pattern of the screened information, identify potential false information or high-risk operation state, and automatically mark the identified risk sites and generate alarm information containing risk level, abnormal type and impact range.

[0101] A double-layer false site identification model containing a "rule matching layer" and a "machine learning prediction layer" is constructed, based on the deposit and rent ratio, site operation rate, replacement unit price change, historical blowout records, etc. First, the rule engine preliminarily screens the obviously abnormal information (such as invalid address, owner's bad faith), and then uses machine learning algorithms (such as random forest, XGBoost) to mine potential risk patterns (such as low deposit + high rent rate combination features), and automatically marks the identified risk sites and generates alarm information containing risk level, abnormal type, and impact range.

[0102] The double-layer model architecture includes: rule matching layer: preset rigid rules (such as "address resolution failure" "owner is listed in the blacklist of bad faith"), directly marked as "high risk" and blocked from entering the process; flexible rules (such as "deposit / rent <0.5 and replacement unit price decreased by more than 20%") are marked as "medium risk" and trigger manual review. Machine learning layer: feature engineering: extract 50+ risk features (such as "number of rent returns in the past 12 months" "vacancy rate of similar sites within 3 kilometers" "frequency of ownership changes"), select core features (retain 20-30 high-correlation features) through feature selection algorithms (such as mutual information method). Model training uses historical risk site samples (positive samples) and normal sites (negative samples) to train the model, uses cross-validation (Cross-Validation) to optimize hyperparameters, and regularly (every week) automatically synchronizes the latest data for incremental training to avoid the decline in recognition ability caused by "concept drift" (Concept Drift).

[0103] Alarm mechanism: risk level is divided into "high / medium / low" three levels, high-risk sites trigger instant SMS alarm, medium-risk sites are displayed in red on the operation board, and low-risk sites are recorded in the observation list. Alarm information contains evidence chain (such as "trigger rule R003: deposit / rent = 4.2 > 3 times threshold" "model prediction probability = 92%"), which facilitates managers to quickly determine the priority.

[0104] In some embodiments, based on natural language processing (NLP) and knowledge graph technology, a field-adaptive semantic alignment model is constructed, semantic ambiguity fields in multi-source data are automatically identified, field semantic mapping and standardization across data sources are realized through a deep learning model, and the problem that traditional rule matching is difficult to cover long-tail semantic differences is solved.

[0105] Domain knowledge graph construction: Collect real estate industry terminology library, enterprise internal data dictionary and historical data mapping cases, and construct a knowledge graph containing "field name-business meaning-data type-associated field", with nodes as field entities (such as "building area" and "usable area") and edges as semantic association relationships (such as "synonym", "contain" and "conversion formula").

[0106] Semantic alignment model training: Use BERT pre-training model, input field name and context description in multi-source data (such as table title and data description document where the field is located), and output semantic vector representation of the field; design a comparison learning task: pull the vectors of synonymous fields (such as "rental price" and "yuan / m 2 / month") closer, and push the vectors of different fields (such as "area" in different systems) further apart, and through Triplet Loss optimization model, make fields with similar semantics closer in vector space.

[0107] Intelligent mapping engine: For newly accessed data sources, predict the standard semantic category of the field through the model (such as input "site size" of A system, and match to standard field "building area"), and automatically generate data conversion logic combining conversion rules in the knowledge graph (such as "usable area = building area x 0.75"), supporting low-code configuration.

[0108] In some embodiments, by fusing computer vision (CV) and graph neural network (GNN), through joint analysis of site image data and surrounding POI data, automatic generation and dynamic update of geographic labels are realized, and the problem that traditional rules rely on manual annotation and labels lag behind regional development is solved.

[0109] Multi-modal data input: visual data: crawl site aerial photos / street view photos of A map software and B map software, detect building appearance features (such as "glass curtain wall" and "warehouse roof") through YOLO model, and judge site purpose (office / warehouse); POI data: get POI types (subway, mall, factory) and density within 1 km around the site from OpenStreetMap, and construct a "POI adjacency graph" centered on the site.

[0110] Graph neural network modeling: node features: site coordinates, building area, visually recognized purpose labels; POI node features: type, distance, rating; graph convolution layer (GCN): learn regional attribute labels (such as “commercial core area” and “industrial park”) of sites through spatial distance and functional correlation between nodes (such as strong correlation between “subway node” and “office site node”); dynamic updating mechanism: set quarterly tasks, re-crawl POI data and satellite images, detect newly added facilities (such as newly built subway stations), and automatically trigger label updates (such as adding a “subway coverage” label to a site).

[0111] Geographic coordinate error correction: for sites with failed address resolution, use the house number OCR recognition result in the visual data, combined with the “house number reverse geocoding” function of Amap software API, to improve the accuracy from 70% to more than 90%.

[0112] In some embodiments, by constructing a data quality intelligent monitoring system based on self-supervised learning, potential data quality problems (such as outliers, logical contradictions) in the panoramic database are automatically identified, and data repair strategies are generated through reinforcement learning, replacing the fixed verification logic of traditional rule engines.

[0113] Self-supervised quality detection model: unsupervised pre-training: use the time series (such as historical rent changes) and spatial correlation (area distribution of sites in the same region) of site data to reconstruct the data through Autoencoder, and mark samples with reconstruction error exceeding the threshold as “potential anomalies” (such as a site rent increases by 50% and surrounding site rent does not fluctuate).

[0114] Cross-field logical verification: build a Bayesian network model to learn the logical relationships between fields such as “rent = unit price x area” and “deposit ≤ 3 times monthly rent”, and trigger an early warning for records that violate the probability model (such as rent calculation result deviates from the input value by >5%).

[0115] Reinforcement learning repair strategy: state space: data quality problem type (missing value / outlier / logical contradiction), field importance level, historical repair record; action space: fill (mean / model prediction), correction (trigger manual review), delete; reward function: dynamically adjust the strategy according to the impact of the repaired field on subsequent analysis (such as the rent data after repair improves the accuracy of anomaly identification), forming a “detection-repair-evaluation” closed loop.

[0116] Dynamic quality report: generate a data quality dashboard every day to display the completeness, accuracy, and consistency scores of data from each business line, and automatically trigger a data governance ticket for data sources that continuously fall below the threshold (such as a sub-company data completeness rate <80%).

[0117] In some embodiments, a "time series prediction-correlation risk" dual-module system is constructed based on a long short-term memory network (LSTM) and a causal inference model. This system not only identifies current abnormalities, but also predicts future rent fluctuation risks and excavates causal relationships between abnormal events (such as whether the rent surge triggers the rent-out tide).

[0118] Time series prediction module: input features: historical rent, CPI index, regional vacancy rate, average rent of similar sites, etc. time series data, learn the long-term dependence of rent changes (such as the regular decline of Q4 rent every year due to the leasing off-season) through LSTM model; abnormal detection: the deviation between the predicted value and the actual value exceeding 3 times the standard deviation is considered as "abnormal fluctuation", and "seasonal fluctuation" and "real abnormality" are distinguished (such as a site rent surges by 20% in the off-season, triggering an early warning).

[0119] Causal inference module builds a causal graph: taking "rent abnormality", "rent-out record", "property fee increase" and other events as nodes, analyzing the causal effect between variables through Do-Calculus algorithm (such as verifying whether "excessive deposit" will lead to "rent-out rate increase"), counterfactual reasoning: simulating "if rent does not rise by 15%, will the rent-out rate decrease", providing decision basis for risk intervention (such as suggesting to start rent negotiation mechanism for sites with high rent increase).

[0120] Risk transmission analysis: using graph network to analyze the correlation between sites (such as multiple sites of the same owner, competitive sites in the same area), when a site triggers "abnormal rent increase", automatically assesses the rent linkage risk to surrounding sites (such as surrounding sites may cause tenant loss due to comparison effect).

[0121] In some embodiments, by building a real estate industry federated learning platform, under the premise of protecting the data privacy of each enterprise, joint multi-party data is used to train a false site identification model, solving the problem of weak model generalization ability caused by insufficient data volume of a single enterprise, and realizing cross-enterprise risk information sharing.

[0122] Federated learning architecture design: participants: real estate enterprises, property management companies, credit investigation agencies, each party retains the original data locally and only uploads model parameters (gradient, weight);

[0123] Hierarchical training: bottom feature layer: each enterprise trains a local feature extractor (such as extracting "ownership change frequency", "historical contract compliance rate" and other features); high-level decision layer: the central server aggregates parameters from each party to update the global risk identification model, avoiding data leakage (complying with GDPR / PIPL compliance requirements).

[0124] Privacy protection mechanism: homomorphic encryption: encrypt the uploaded model gradient, the central server can only aggregate in the ciphertext state; differential privacy: add Gaussian noise in the parameters to prevent the original data features from being inferred through the model. Defense mechanism: when a company marks "fake site", automatically synchronize the risk features to the federal model through the smart contract, and the identification model of other companies updates the risk mode in real time, forming an industry-level risk blacklist.

[0125] In some embodiments, by constructing a site-level digital twin model, the operation effect under different leasing strategies (such as the impact of rent adjustment on the occupancy rate) is simulated through three-dimensional modeling and discrete event simulation, and the optimal risk control strategy is found through reinforcement learning, which assists in rent anomaly identification and fake site decision-making.

[0126] Digital twin construction: three-dimensional modeling: generate a three-dimensional space model of the site through unmanned aerial vehicle aerial photography and point cloud data processing, and label physical parameters such as building area, floor layout, and parking space; business process modeling: abstract the leasing process into discrete events (contract signing, rent payment, and rent application), and construct a simulation model based on AnyLogic, input historical operation data (such as rent rate and rent increase) as initial parameters.

[0127] Reinforcement learning optimization: state space: real-time operation state of the site (vacancy rate, rent level, and number of surrounding competitive sites); action space: rent adjustment range (±5%, ±10%), deposit scheme (whether to accept installment payment), and preferential policy (extension of rent-free period); reward function: taking "rent income-risk loss" as the goal, the long-term income of different actions is simulated through simulation, and the optimal risk control strategy is output (such as when the surrounding vacancy rate is >40%, the rent is recommended to be reduced by 10% to reduce the risk of rent).

[0128] Abnormal scenario deduction: preset extreme scenarios, simulate site operation data changes, and verify the effectiveness of existing abnormal identification rules (such as whether the time threshold of "long-term non-operation after payment" needs to be dynamically adjusted), and provide data support for rule engine optimization.

[0129] See Figure 2 , Figure 2 Embodiments of the present application also provide a schematic block diagram of a big data-based intelligent site operation management device. The big data-based intelligent site operation management device 200 is used to execute the aforementioned big data-based intelligent site operation management method. The big data-based intelligent site operation management device can be configured in a server or a terminal.

[0130] The server can be a standalone server, a server cluster, a cloud server providing cloud services, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content delivery network (CDN), and basic cloud computing services such as big data and artificial intelligence platform. The terminal can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, a user digital assistant, and a wearable device, and the like.

[0131] As shown in Figure 2 The site operation device based on big data intelligent management and control 200 includes:

[0132] The information acquisition unit 201 is configured to acquire multi-source heterogeneous site related information in a preset data source, perform cleaning and standardization processing operations on structured data and unstructured data corresponding to the site related information, and unify coordinate systems and label systems corresponding to the site related information by combining geographic position analysis technology, to generate a site panoramic database.

[0133] The data acquisition unit 202 is configured to acquire site rental data based on the site panoramic database, and construct a rent abnormality identification rule engine. The rent abnormality identification rule engine is configured to dynamically configure abnormality judgment logic to process the acquired rental data, calculate site information of each site, and the site information includes abnormal portraits and key indicators.

[0134] The model construction unit 203 is configured to construct a false site identification model, and use a machine learning algorithm combined with a rule engine to continuously analyze newly entered site information and site information in the operation process, to identify risk sites, and generate processing information corresponding to the risk sites and send the processing information to corresponding user terminals.

[0135] The decision support unit 204 is configured to generate structured data including site identification, abnormal item list, indicator score, recommendation level, and strategy label based on the false site identification model and the site panoramic database, to provide support for site operation decision.

[0136] In some embodiments, the cleaning and standardization processing operations on the structured data and unstructured data corresponding to the site related information include: removing duplicate data, filling or filtering missing values, unifying data formats, and performing semantic conversion processing on the structured data and unstructured data, so that the same name different meaning or different name same meaning fields of different data sources form a unified semantic definition.

[0137] In some embodiments, the combination of the geographic location resolution technology unifies the coordinate system and label system of the site-related information, including: based on the site address or coordinate data, the geographic coding resolution is performed to convert the coordinate system of different data sources into a unified geographic coordinate system or a projection coordinate system; a standardized label system is established according to the site purpose, regional attribute, traffic condition and other dimensions, so that each site corresponds to a unique geographic coordinate identifier and multi-dimensional classification label.

[0138] In some embodiments, the generation of the site panoramic database includes: integrating and storing the site information after cleaning, standardization and geographic labeling to form a site panoramic data including site basic attribute information, rent cost data, physical space parameters, address coordinates, surrounding supporting facilities, real-time operation status and historical leasing records, and the historical leasing records include lease period, rent change, contract performance and lease termination records.

[0139] In some embodiments, the site leasing data is collected based on the site panoramic database, and a rent abnormality identification rule engine is constructed, including: acquiring rent, deposit, property fee, water and electricity fee, decoration fee and operation data through a combination of offline batch acquisition and real-time stream acquisition; based on business rules, a dynamically configurable abnormality judgment logic is constructed, which at least includes rules of deposit higher than a preset rent multiple, public area proportion exceeding a limit, new and old site lease period overlapping, long-term non-operation after payment, rent unit price abnormal rise and lease termination loss exceeding a limit, and the preset rent multiple and proportion threshold support dynamic adjustment through a business configuration interface.

[0140] In some embodiments, the collected leasing data is processed to calculate the site information of each site, including abnormal profile and key indicators, including: according to the judgment logic of the rent abnormality identification rule engine, the leasing data of a single site is grouped, aggregated and associated analyzed to identify the abnormal rule items triggered by the site and generate an abnormal profile; based on the dimensions of abnormal rule triggering frequency, involved fee amount, influence degree on operation efficiency, key indicators are calculated, the abnormal profile and key indicators are synchronized to a data analysis platform, and manual notes and feedback corrections of abnormal identification results are supported.

[0141] In some embodiments, the constructing a false site identification model, using a machine learning algorithm combined with a rule engine, continuously analyzes newly entered site information and site information during operation to identify risk sites, including: based on the deposit and rent ratio, site operation rate, replacement site unit price change range, preferential implementation, and historical blowout records, etc. Multidimensional indicators, a double-layer identification model including a rule matching layer and a machine learning prediction layer is constructed; the rule engine is used to preliminarily screen information that obviously violates business logic, and the machine learning algorithm is used to identify patterns of the screened information to identify potential false information or high-risk operating status. The identified risk sites are automatically marked and alarm information containing risk level, abnormal type and impact range is generated.

[0142] It should be noted that, for the convenience and brevity of description, the specific working process of the model training device and each module described above can refer to the corresponding process in the foregoing embodiment of the method for intelligently controlling site operation based on big data, which will not be described here.

[0143] The above-mentioned device for intelligently controlling site operation based on big data can be realized in the form of a computer program, which can run on a computer device as shown in Figure 3 .

[0144] Please refer to Figure 3 , Figure 3 is a structural schematic block diagram of a computer device provided by an embodiment of the present application. The computer device can be a server or a terminal.

[0145] Please refer to Figure 3 , the computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory can include a storage medium and an internal memory.

[0146] The storage medium can store an operating system and a computer program. The computer program includes program instructions which, when executed, can cause the processor to execute any one of the methods for intelligently controlling site operation based on big data provided by the embodiments of the present application.

[0147] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0148] The internal memory provides an environment for the running of the computer program in the storage medium, and the computer program, when executed by the processor, can cause the processor to execute any one of the methods for intelligently controlling site operation based on big data. The storage medium can be non-volatile or volatile.

[0149] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that,Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0150] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0151] For example, in one embodiment, the processor is configured to run a computer program stored in the memory to implement the following steps:

[0152] Obtain multi-source heterogeneous site-related information in a preset data source, perform cleaning and standardization processing operations on structured data and unstructured data corresponding to the site-related information, and unify coordinate systems and label systems corresponding to the site-related information by combining geographic position analysis technology, to generate a site panoramic database;

[0153] Collect site rental data based on the site panoramic database, and construct a rent abnormality identification rule engine, which is configured to dynamically configure abnormality judgment logic to process the collected rental data, calculate site information of each site, and the site information includes an abnormal profile and key indicators;

[0154] Construct a fake site identification model, and use a machine learning algorithm combined with a rule engine to continuously analyze newly entered site information and site information in the operation process to identify risk sites, and generate processing information corresponding to the risk sites and send the processing information to corresponding user terminals;

[0155] Generate structured data including site identification, abnormal item list, indicator score, recommended level and strategy label based on the fake site identification model and the site panoramic database, to provide support for site operation decision-making.

[0156] In some embodiments, the structured data and unstructured data corresponding to the site-related information are subjected to cleaning and standardization processing operations, including: removing duplicate data, filling or filtering missing values, unifying data formats, and performing semantic conversion processing, so that the same name different meanings or different name same meanings fields of different data sources form a unified semantic definition.

[0157] In some embodiments, the geographic position resolution technology is combined to unify the coordinate system and label system corresponding to the site-related information, including: based on site address or coordinate data, performing geographic coding resolution to convert the coordinate systems of different data sources into a unified geographic coordinate system or projection coordinate system; and according to dimensions such as site purpose, regional attribute, and traffic condition, establishing a standardized label system, so that each site corresponds to a unique geographic coordinate identifier and multi-dimensional classification label.

[0158] In some embodiments, the site panoramic database is generated, including: integrating and storing the site information after cleaning, standardization, and geographic labeling, forming a site panoramic data containing site basic attribute information, rent cost data, physical space parameters, address coordinates, surrounding supporting facilities, real-time operation status, and historical leasing records, the historical leasing records including lease period, rent variation, contract performance, and lease termination record.

[0159] In some embodiments, the site leasing data is collected based on the site panoramic database, and a rent abnormality identification rule engine is constructed, including: acquiring rent, deposit, property fee, water and electricity fee, decoration fee, and operation data through a combination of offline batch acquisition and real-time stream acquisition; based on business rules, a dynamically configurable abnormality judgment logic is constructed, the abnormality judgment logic at least including rules of deposit being higher than a preset rent multiple, public area proportion exceeding a limit, new and old site lease period overlapping, long-term non-operation after payment, rent unit price abnormal rise, and lease termination loss exceeding a limit, and the preset rent multiple and proportion threshold support dynamic adjustment through a business configuration interface.

[0160] In some embodiments, the collected leasing data is processed to calculate site information of each site, the site information including an abnormal profile and key indicators, including: based on the judgment logic of the rent abnormality identification rule engine, the leasing data of a single site is subjected to grouping aggregation and correlation analysis to identify abnormal rule items triggered by the site and generate an abnormal profile; based on dimensions such as abnormal rule triggering frequency, involved fee amount, and influence degree on operation efficiency, key indicators are calculated, the abnormal profile and key indicators are synchronized to a data analysis platform, and manual notes and feedback corrections to abnormality identification results are supported.

[0161] In some embodiments, the method for constructing a false site identification model, using a machine learning algorithm combined with a rule engine, continuously analyzes newly entered site information and site information during operation to identify risk sites, including: based on the proportion of deposit and rent, site operation rate, replacement site unit price change range, preferential implementation and historical blowout records, etc. Multi-dimensional indicators, a double-layer identification model including a rule matching layer and a machine learning prediction layer is constructed; the rule engine is used to preliminarily screen information that obviously violates business logic, and then a machine learning algorithm is used to identify patterns in the screened information to identify potential false information or high-risk operating states. The identified risk sites are automatically marked and alarm information containing risk levels, abnormal types and impact ranges is generated.

[0162] The application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements the steps of the risk early warning method according to the first aspect.

[0163] The computer-readable storage medium can be an internal storage unit of the computer device, such as a hard disk or a memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0164] The above merely illustrates the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A big data-based intelligent management and control site operation method, characterized in that, The method comprises the following steps: Obtaining multi-source heterogeneous site-related information in a preset data source, performing cleaning and standardization processing operations on structured data and unstructured data corresponding to the site-related information, and unifying the coordinate system and label system corresponding to the site-related information by combining geographic position analysis technology to generate a site panoramic database; Collecting site rental data based on the site panoramic database and constructing a rent abnormality identification rule engine, which is used to dynamically configure abnormality judgment logic to process the collected rental data and calculate site information of each site, including abnormal portraits and key indicators, including: grouping, aggregating and correlating the rental data of a single site according to the judgment logic of the rent abnormality identification rule engine, identifying the abnormal rule items triggered by the site and generating an abnormal portrait; calculating the key indicators based on the dimensions corresponding to the abnormal rule triggering frequency, the involved fee amount and the influence degree on the operation efficiency, synchronizing the abnormal portrait and the key indicators to a data analysis platform, and supporting manual notes and feedback corrections on the abnormality identification results; Constructing a false site identification model, using a machine learning algorithm combined with a rule engine to continuously analyze newly entered site information and site information in the operation process to identify risk sites, including: constructing a double-layer identification model including a rule matching layer and a machine learning prediction layer based on multi-dimensional indicators corresponding to the deposit and rent ratio, site operation rate, single price change range of site replacement, preferential implementation and historical blowout records; the rule engine is used to preliminarily screen information that obviously deviates from the business logic, and the machine learning algorithm is used to identify potential false information or high-risk operation state, automatically mark the identified risk sites and generate alarm information including risk level, abnormal type and influence range, and generate processing information corresponding to the risk site and send it to the corresponding user terminal; According to the false site identification model and the site panoramic database, structured data including site identification, abnormal item list, indicator score, recommended level and strategy label are generated to support site operation decision-making.

2. The method of claim 1, wherein, The cleaning and standardization processing operations performed on the structured data and unstructured data corresponding to the site-related information include: Removing duplicate data, filling or filtering missing values, unifying data format and semantic conversion processing are performed on the structured data and unstructured data, so that the same name different meaning or different name same meaning fields of different data sources form a unified semantic definition.

3. The method of claim 1, wherein, The unification of the coordinate system and the label system corresponding to the site-related information by combining the geographic position analysis technology includes: Based on the site address or coordinate data, the geographic coding analysis is performed to convert the coordinate system of different data sources into a unified geographic coordinate system or a projection coordinate system; A standardized label system is established according to the dimensions corresponding to the site purpose, regional attribute and traffic conditions, so that each site corresponds to a unique geographic coordinate identifier and a multi-dimensional classification label.

4. The method of claim 1, wherein, The site panoramic database is generated, including: The cleaned, standardized and geotagged site information is integrated and stored to form site panoramic data including site basic attribute information, rent cost data, physical space parameters, address coordinates, surrounding supporting facilities, real-time operation status and historical leasing records, and the historical leasing records include lease period, rent change, contract performance and lease termination records.

5. The method of claim 1, wherein, The site leasing data is collected based on the site panoramic database, and a rent abnormality identification rule engine is constructed, which includes: Rent, deposit, property management fee, water and electricity fee, decoration fee and operation data are acquired through a combination of offline batch collection and real-time stream collection; A dynamically configurable abnormality judgment logic is constructed based on business rules, and the abnormality judgment logic includes at least rules of deposit being higher than a preset rent multiple, public area proportion exceeding a limit, new and old site lease periods overlapping, long-term non-operation after payment, rent unit price abnormally rising and lease termination loss exceeding a limit, and the preset rent multiple and proportion threshold values are dynamically adjustable through a business configuration interface.

6. A device for intelligently managing and controlling site operation based on big data, characterized in that, It includes: An information acquisition unit is configured to acquire multi-source and heterogeneous site-related information from a preset data source, perform cleaning and standardization processing operations on structured data and unstructured data corresponding to the site-related information, and unify coordinate systems and label systems corresponding to the site-related information by using a geographic position analysis technique to generate a site panoramic database; A data acquisition unit is configured to collect site leasing data based on the site panoramic database and construct a rent abnormality identification rule engine, which is used to dynamically configure abnormality judgment logic to process the collected leasing data and calculate site information of each site, including an abnormal profile and key indicators, including: grouping, aggregating and correlating the leasing data of a single site according to the judgment logic of the rent abnormality identification rule engine to identify abnormal rule items triggered by the site and generate an abnormal profile; calculating key indicators based on dimensions corresponding to abnormal rule triggering frequencies, involved fee amounts and influence degrees on operation efficiency, synchronizing the abnormal profile and the key indicators to a data analysis platform, and supporting manual notes and feedback corrections on abnormality identification results; A model construction unit is configured to construct a false site identification model, which uses a machine learning algorithm in combination with a rule engine to continuously analyze newly entered site information and site information in the operation process to identify risk sites, including: constructing a double-layer identification model including a rule matching layer and a machine learning prediction layer based on multi-dimensional indicators corresponding to a deposit to rent ratio, a site operation rate, a replacement site unit price change amplitude, a discount implementation situation and a historical burst record; using the rule engine to preliminarily filter information that obviously deviates from business logic, and using a machine learning algorithm to perform pattern recognition on the filtered information to identify potential false information or high-risk operation status, automatically marking the identified risk sites and generating alarm information including a risk level, an abnormal type and an influence range, and generating processing information corresponding to the risk sites and sending it to corresponding user terminals. A decision support unit is configured to generate structured data including site identification, anomaly item list, index score, recommendation level and strategy label based on the false site identification model and the site panorama database, and provide support for site operation decision.

7. A computer device, comprising: The computer device comprises a memory and a processor; The memory is configured to store a computer program; The processor is configured to execute the computer program and implement the method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is configured to enable the processor to implement the method according to any one of claims 1 to 5 when executed by the processor.

Citation Information

Patent Citations

  • House false information identification method and device thereof, electronic equipment and storage medium

    CN112699659A

  • Field big data risk screening method and device

    CN114997571A