Intelligent management and control site operation method and device based on big data, equipment and medium

By building a big data site panoramic database and machine learning model, the data integration and risk identification problems of traditional site management systems have been solved, the intelligent and refined management of site operations has been realized, and the accuracy and efficiency of rental anomaly identification and risk prediction have been improved.

CN120598641AActive Publication Date: 2025-09-05SHENZHEN LEAPFROG NEW TECH CO LTD

Patent Information

Application Number
CN202511103584.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-09-05
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

During the scale expansion of enterprises, traditional venue management systems suffer from a lack of data integration capabilities, a lack of intelligent analysis capabilities, and backward decision-making support methods. They are unable to effectively manage multi-source heterogeneous data, automatically identify rental anomalies and fake venues, and lack a panoramic database and intelligent recommendation mechanism.

Method used

By building a panoramic site database based on big data, combining geographic location analysis and machine learning algorithms, we can achieve automatic integration and cleaning of multi-source data, dynamically configure the rent anomaly identification rule engine, build a fake site identification model, generate processing information for risky sites, and provide decision support.

Benefits of technology

It realizes a global data view of site information, improves the accuracy of rental anomaly identification and risk prediction capabilities, reduces manual investigation costs, improves the efficiency of exception handling, and supports the company's refined operational decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598641A_ABST
    Figure CN120598641A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of resource management, and discloses an intelligent management and control site operation method and device based on big data, equipment and a medium, and the method comprises the steps: obtaining multi-source heterogeneous site related information from a preset data source, and generating a site panoramic database; acquiring site lease data based on the site panoramic database, constructing a rent anomaly recognition rule engine, and calculating site information of each site; constructing a false site identification model, and continuously analyzing newly input site information and site information in an operation process by using a mode of fusing a machine learning algorithm and a rule engine so as to identify a risk site; and according to the false site identification model and the site panorama database, generating structured data comprising a site identifier, an abnormal item list, an index score, a recommendation level and a strategy label. The method is suitable for full life cycle management of site selection, evaluation, monitoring and strategy optimization of enterprises such as logistics, e-commerce, manufacturing and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of resource management technology, and in particular to a method, device, equipment and medium for intelligent site management based on big data. Background Art

[0002] During the scale expansion of enterprises, site operation and management face multi-dimensional challenges. Traditional site management relies on manual screening, Excel records and empirical judgment, which has the following significant shortcomings:

[0003] 1. Lack of data integration capabilities: The existing system can only manage the basic information of a single venue and is unable to automatically collect heterogeneous data from leasing platforms, map services, internal enterprise systems, and other sources through multi-source interfaces. This results in a long acquisition cycle and low integrity of nationwide venue information, making it difficult to form a global data view.

[0004] 2. Lack of intelligent analysis capabilities: Site assessment relies on manual experience and lacks a dynamically configurable rule engine for identifying rent anomalies. This makes it impossible to perform automated analysis for complex scenarios such as excessive deposits, overlapping lease periods, and abnormal rent price increases. Furthermore, risk identification remains at the manual review level, and a two-tier model combining machine learning and a rule engine has not been built, making it difficult to identify fake sites or potential operational risks in real time.

[0005] 3. Outdated decision-making support methods: Data visualization only provides basic dashboard displays and is unable to integrate multi-dimensional information such as contract data, cost details, and operational status to form strategic outputs. Furthermore, there is a lack of an intelligent recommendation mechanism based on a panoramic site database, making it difficult to meet companies' refined needs for site selection efficiency, cost control, and risk prediction.

[0006] Therefore, a method is urgently needed to solve at least one of the above problems. Summary of the Invention

[0007] This application provides a method, device, equipment, and medium for intelligently managing site operations based on big data, designed to address the multi-dimensional challenges faced by existing technologies in site operations management during enterprise scale expansion. Traditional site management relies on manual screening, Excel records, and empirical judgment.

[0008] In a first aspect, the present application provides a method for intelligently controlling site operations based on big data, the method comprising:

[0009] Acquire multi-source heterogeneous site-related information from pre-set data sources, cleanse and standardize the structured and unstructured data corresponding to the site-related information, and use geographic location analysis technology to unify the coordinate system and label system corresponding to the site-related information to generate a panoramic site database;

[0010] Based on the site panorama database, site rental data is collected and a rent anomaly identification rule engine is built. The rent anomaly identification rule engine is used to dynamically configure anomaly judgment logic to process the collected rental data and calculate site information for each site, including anomaly profiles and key indicators;

[0011] Build a fake site identification model, using a combination of machine learning algorithms and rule engines to continuously analyze newly entered site information and site information during operation to identify risky sites, and generate corresponding processing information for risky sites and send it to the corresponding user terminals;

[0012] Based on the fake site identification model and the site panoramic database, structured data including site identification, anomaly list, indicator score, recommendation level and strategy label is generated to provide support for site operation decision-making.

[0013] In some embodiments, the structured data and unstructured data corresponding to the venue-related information are cleaned and standardized, including: removing duplicate data, filling or filtering missing values, unifying data formats and semantic conversion processing on the structured data and unstructured data, so that the fields with the same name but different meanings or synonyms with different names from different data sources form a unified semantic definition.

[0014] In some embodiments, the combination of geographic location resolution technology unifies the coordinate system and labeling system corresponding to site-related information, including: performing geocoding resolution based on the site address or coordinate data, converting the coordinate systems of different data sources into a unified geographic coordinate system or projection coordinate system; establishing a standardized labeling system based on dimensions such as site use, regional attributes, and traffic conditions, so that each site corresponds to a unique geographic coordinate identifier and multi-dimensional classification label.

[0015] In some embodiments, generating a panoramic site database includes integrating and storing cleaned, standardized, and geo-tagged site information to form panoramic site data including basic site attribute information, rental fee data, physical space parameters, address coordinates, surrounding supporting facilities, real-time operating status, and historical leasing records. The historical leasing records include lease terms, rent changes, contract performance, and lease termination records.

[0016] In some embodiments, the site rental data is collected based on the site panoramic database, and a rent anomaly identification rule engine is constructed, including: obtaining rent, deposit, property fees, water and electricity fees, decoration fees and operating data through a combination of offline batch collection and real-time streaming collection; constructing dynamically configurable anomaly judgment logic based on business rules, and the anomaly judgment logic includes at least rules such as the deposit is higher than the preset rent multiple, the proportion of common area exceeds the limit, the lease period of the new and old sites overlaps, the site has not been operated for a long time after payment, the unit rent price increases abnormally, and the loss of termination exceeds the limit, and the preset rent multiple and proportion thresholds support dynamic adjustment through the business configuration interface.

[0017] In some embodiments, the collected rental data is processed to calculate the venue information of each venue, and the venue information includes abnormal portraits and key indicators, including: according to the judgment logic of the rent abnormality identification rule engine, the rental data of a single venue is grouped, aggregated and associated analyzed, the abnormal rule items triggered by the venue are identified and abnormal portraits are generated; key indicators are calculated based on dimensions such as the frequency of abnormal rule triggering, the amount of expenses involved, and the degree of impact on operational efficiency, the abnormal portraits and key indicators are synchronized to the data analysis platform, and manual annotations and feedback corrections to the abnormal identification results are supported.

[0018] In some embodiments, the construction of a false venue identification model uses a fusion of machine learning algorithms and rule engines to continuously analyze newly entered venue information and venue information during operation to identify risky venues, including: based on multi-dimensional indicators such as the deposit to rent ratio, venue operation rate, change in unit price of replacement venues, preferential implementation status and historical liquidation records, a two-layer identification model consisting of a rule matching layer and a machine learning prediction layer; the rule engine is used to preliminarily screen information that clearly violates business logic, and then the machine learning algorithm is used to perform pattern recognition on the screened information to identify potential false information or high-risk operating conditions, automatically mark the identified risky venues, and generate alarm information including risk level, abnormality type and impact range.

[0019] In a second aspect, the present application provides a device for intelligently controlling site operations based on big data, the device comprising:

[0020] An information acquisition unit is used to obtain multi-source heterogeneous site-related information from a preset data source, perform cleaning and standardization operations on the structured and unstructured data corresponding to the site-related information, and unify the coordinate system and label system corresponding to the site-related information in combination with geographic location analysis technology to generate a site panoramic database;

[0021] A data collection unit is used to collect site rental data based on a site panorama database and build a rent anomaly identification rule engine. The rent anomaly identification rule engine is used to dynamically configure anomaly judgment logic to process the collected rental data and calculate site information for each site, including anomaly profiles and key indicators;

[0022] The model building unit is used to build a fake site identification model. It uses a machine learning algorithm and a rule engine to continuously analyze newly entered site information and site information during operation to identify risky sites, and generates processing information corresponding to risky sites and sends it to the corresponding user terminal;

[0023] The decision support unit is used to generate structured data including site identification, abnormal item list, indicator score, recommendation level and strategy label based on the false site identification model and the site panoramic database to provide support for site operation decision-making.

[0024] In a third aspect, the present application provides a computer device, the computer device comprising a memory and a processor;

[0025] The memory is used to store computer programs;

[0026] The processor is used to execute the computer program and implement any one of the methods for intelligently controlling site operations based on big data provided in the embodiments of the present application when executing the computer program.

[0027] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor enables the processor to implement any one of the methods for intelligently controlling site operations based on big data provided in the embodiments of the present application.

[0028] The present application discloses a method, device, equipment and medium for intelligent site management and control based on big data. The provided method automatically acquires heterogeneous data and processes it in a standardized manner through multi-source interfaces, constructs a panoramic database covering the entire life cycle of the site, improves the integrity of the information, and solves the lag and one-sidedness of traditional manual collection. Based on a configurable rule engine and multi-dimensional indicator analysis, it realizes the automatic identification of complex abnormal scenarios such as excessive deposits and overlapping leases, improves the recognition accuracy, and significantly reduces the cost of manual investigation. It integrates machine learning and rule engines to build a false site identification model, monitors the site operation status in real time and automatically generates alarms, forming an "identification-processing-closed loop" mechanism, effectively improving risk prediction capabilities and reducing user losses and operational risks. Through the unified visual dashboard, multi-dimensional data is integrated to output structured results including site scores, recommendation levels, and policy tags, providing real-time and intelligent support for decisions such as site selection, renewal, and termination, thereby improving the efficiency of exception handling and meeting the company's refined operation needs.

[0029] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0031] Figure 1 This is a flowchart showing the steps of a method for intelligently controlling site operations based on big data provided by an embodiment of the present application;

[0032] Figure 2 This is a schematic block diagram of a device for intelligently controlling site operations based on big data, provided in an embodiment of the present application;

[0033] Figure 3 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application.

[0034] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0036] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0037] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0038] It will also be understood that the term "and / or" as used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0039] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0040] During the scale expansion of enterprises, site operation and management face multi-dimensional challenges. Traditional site management relies on manual screening, Excel records and empirical judgment, which has the following significant shortcomings:

[0041] Lack of data integration capabilities: The existing system can only manage the basic information of a single venue and is unable to automatically collect heterogeneous data such as leasing platforms, map services, and internal corporate systems through multi-source interfaces. This results in a long acquisition cycle and low integrity of nationwide venue information, making it difficult to form a global data view.

[0042] Lack of intelligent analysis capabilities: Site assessment relies on manual experience and lacks a dynamically configurable rule engine for identifying rent anomalies. This makes it impossible to perform automated analysis for complex scenarios such as excessive deposits, overlapping lease periods, and abnormal increases in rent prices. Furthermore, risk identification remains at the manual review level, and a two-layer model combining machine learning and a rule engine has not been built, making it difficult to identify fake sites or potential operational risks in real time.

[0043] Outdated decision-making support methods: Data visualization only enables basic dashboard display and is unable to integrate multi-dimensional information such as contract data, cost details, and operational status to form strategic outputs. It also lacks an intelligent recommendation mechanism based on a panoramic site database, making it difficult to meet companies' refined needs for site selection efficiency, cost control, and risk prediction.

[0044] While existing venue information repositories or BI tools exist, none systematically integrate automated multi-source data integration, dynamic rule configuration, machine learning risk identification, and visualized strategy output, failing to address the interconnected problem of "incomplete data collection, insufficient analytical capabilities, and lack of decision support." This invention fills the gap in existing technologies for intelligent management and control of the entire site operation chain by establishing a complete process for multi-source information collection and integration, intelligent identification of rental anomalies, dynamic risk warnings, and strategic data output.

[0045] See also Figure 1 , Figure 1 This is a schematic flow chart of a method for intelligently managing and controlling site operations based on big data, as provided in an embodiment of this application. The method is applied to a computer device, which can be deployed on a single server or a server cluster. Alternatively, the method can be deployed on handheld devices, laptops, wearable devices, or robots.

[0046] It should be noted that the acquisition of any information mentioned in the provided method complies with relevant regulations and is carried out with the user's consent, and will not infringe on the user's privacy or violate relevant laws and regulations.

[0047] like Figure 1 As shown, the specific steps of the method for intelligently controlling site operations based on big data include: steps S101 to S104.

[0048] S101. Acquire multi-source heterogeneous site-related information from a preset data source, perform cleaning and standardization operations on the structured data and unstructured data corresponding to the site-related information, and unify the coordinate system and label system corresponding to the site-related information in combination with geographic location analysis technology to generate a site panoramic database.

[0049] Specifically, multi-dimensional venue-related information is obtained through preset data sources (including leasing platforms, map service systems, internal enterprise management systems, etc.), and structured data (such as tabular rent and area data) and unstructured data (such as text-described supporting facilities and picture-based venue panoramas) are cleaned, standardized, and geo-tagged, ultimately generating a comprehensive database containing information on the entire life cycle of the venue.

[0050] Multi-source data collection uses web crawler technology (such as HTTP protocol interface calls) to capture basic venue information (rent, area, and lease status) from the leasing platform. Address coordinates, surrounding transportation and supporting facilities data are obtained through map service APIs. Historical leasing records, operational status, and financial data are synchronized through internal enterprise system interfaces (such as ERP and OA systems). Dynamic data source expansion is supported, and access protocols (such as RESTful and SOAP) and data mapping rules for new data sources are defined in configuration files, enabling flexible access to heterogeneous data sources.

[0051] Data cleaning and standardization include: Structured data processing: removing duplicate records (based on unique site identifiers such as address + area), interpolating missing values ​​(e.g., filling missing rent values ​​with the mean of similar sites in the same area), and standardizing data formats (e.g., standardizing "rent units" across different platforms to "yuan / square meter / month"). Unstructured data processing: parsing text descriptions using natural language processing (NLP) technology, extracting keywords (e.g., "subway access" and "firefighting certification") to generate standardized tags; and extracting features from image data to generate structured facility lists (e.g., number of parking spaces, floor height parameters).

[0052] The geographic coordinate and labeling system is unified through geocoding parsing and address conversion (Geocoding) technology to convert text addresses such as "XX City XX District XX Road" into a unified geographic coordinate system (such as WGS84, GCJ-02), and coordinate conversion is performed on the coordinate systems of different data sources to ensure that spatial data can be superimposed for analysis.

[0053] The labeling system is constructed by establishing a multi-level labeling system. The first-level label includes the site purpose (warehousing, office, production), regional attributes (core business district, industrial park), and transportation level (subway coverage, within 1 km of the highway entrance); the second-level label is refined into supporting facilities (such as whether it contains a loading and unloading platform, whether it supports tiered leasing), forming a unique geographic coordinate identification and multi-dimensional classification label for each site.

[0054] The panoramic database is generated by integrating and cleaning structured data, geographic coordinates and label data, and using a relational database (such as MySQL) or a distributed database (such as HBase) for storage, forming a database table containing the following core fields: basic information: site ID, name, ownership attributes, construction year; spatial data: area, floor height, common area ratio, coordinate system coordinates; rental data: current rent, historical rent fluctuation records, deposit standards, property fees; operational data: current tenant information, historical rental cycles, termination reason labels, vacancy duration; geographic tags: regional economic level (generated based on surrounding GDP data), transportation convenience score (calculated based on the distance to subway stations).

[0055] By replacing manual data entry with multi-source automated data collection, we address the fragmentation of site information in traditional management, improve information coverage, and form enterprise-level site data assets. A unified coordinate system and labeling system enables spatial analysis and cross-comparison of site data from different sources (e.g., automatic clustering by regional rental level), providing a unified data foundation for subsequent rental assessments and risk identification, and avoiding analytical errors caused by "data silos."

[0056] S102. Collect site rental data based on a site panorama database and build a rent anomaly identification rule engine. The rent anomaly identification rule engine is used to dynamically configure anomaly judgment logic to process the collected rental data and calculate site information for each site. The site information includes anomaly profiles and key indicators.

[0057] Specifically, based on the site panoramic database generated in step S101, detailed leasing business data (rent, deposit, property fees, etc.) is collected, abnormal scenarios are identified through a dynamically configurable rule engine, and an abnormal profile (a set of triggered abnormal rules) and key indicators (such as an abnormality severity score) of a single site are output.

[0058] Lease data collection includes: Offline batch collection: Using ETL tools (such as Apache Sqoop) to regularly extract historical lease data (such as rent adjustment records and deposit payment receipts from the past three years) from financial systems and contract management systems, and storing it in a data warehouse (such as Hive). Real-time streaming collection: Using message queues (such as Kafka) to capture newly generated lease data (such as deposit terms in newly signed contracts and monthly property management bills), and using real-time computing engines (such as Flink) for real-time cleaning and format conversion, ensuring data synchronization to the rule engine within seconds.

[0059] Rule Engine Construction: Dynamic rule configuration provides a visual rule editing interface, allowing business personnel to customize exception detection logic. Rule types include: numerical rules, such as "Deposit ≥ 3 times the monthly rent" (triggering the "Excessive Deposit" exception) and "Common Area Percentage > 30% without prior notice" (triggering the "Common Area Exception"); logical rules, such as "The new lease overlaps with the previous lease by more than 30 days" (triggering the "Lease Overlap" exception) and "No property fee payment records for six consecutive months and the site status is displayed as 'Operating'" (triggering the "Non-Operating Site Exception"); trend rules, such as "Year-on-year rent price increase > 150% of the regional average" (triggering the "Abnormal Rent Increase"). Rule priority management supports setting priorities (high, medium, or low) for different rules. For example, setting "Lease Overlap" to high priority will immediately mark the exception and block the contract approval process.

[0060] Abnormal profiles and indicator calculations are performed by grouping and aggregating the rental data of a single site (such as aggregating rental data by year or quarter), matching abnormal rules one by one through the rule engine, recording the triggered rule items and triggering time, and generating an "abnormal profile" that includes the abnormal type, amount involved, and duration.

[0061] Key indicator calculation generates a quantitative score based on the priority and impact of the exception rules (e.g., "excessive deposit" is scored 10 points, "overlapping lease period" is scored 20 points), and the "anomaly index" is calculated based on the frequency of exceptions for subsequent risk ranking and visualization.

[0062] This replaces the traditional manual table-by-table screening method, enabling 24 / 7 full data monitoring, improving anomaly identification efficiency and significantly reducing labor costs. Through dynamic rule configuration, it supports customized anomaly standards for different industries (logistics and warehousing, commercial real estate, and industrial plants) (for example, the tolerance threshold for "common area" in the logistics industry can be configured separately). This addresses the rigidity of traditional system rules and expands the system's adaptability from single-scenario management to multi-business management.

[0063] S103: Build a fake site identification model, and use a combination of machine learning algorithms and rule engines to continuously analyze newly entered site information and site information during operation to identify risky sites, and generate processing information corresponding to risky sites and send it to the corresponding user terminal.

[0064] Specifically, by building a two-layer risk identification model of "rule engine initial screening + machine learning in-depth analysis", newly entered site information and operating sites are continuously monitored, false sites or high-risk operating conditions are identified in real time, and processing information including risk levels is generated and pushed to the management terminal.

[0065] Multi-dimensional indicator system: The following core indicators are collected as model inputs: Financial indicators: Deposit / rent ratio (rule warning is triggered when the threshold is ≥2), rent collection rate (<80% for three consecutive months); Operational indicators: Site operation rate (actual use area / leased area <60%), replacement unit price change (newly signed rent is more than 30% lower than the surrounding average price, and there may be a low-price trap); Historical records: Historical number of warehouse explosions (accidents caused by excessive load in storage sites), Lease termination rate (number of lease terminations ≥3 times within 12 months); Related data: Credit records of site owners (obtained through the enterprise credit API), vacancy rates of similar sites in the surrounding area (captured through the map service API).

[0066] The two-tier model architecture: The rule engine layer first filters out obvious anomalies using pre-set business rules, such as "invalid address resolution results" or "owner has a court record of default." These anomalies are then directly labeled "high risk" and the entry process is blocked. The machine learning layer uses algorithms such as gradient boosted tree (GBDT) and random forests to train the model based on historical risk site samples (such as confirmed fraudulent sites and sites involved in litigation due to rental disputes). The model learns risk signature patterns (e.g., a combination of "low deposit, high rent increases, and frequent ownership changes"). The model is regularly updated (automatically syncing with the latest sample data weekly) to identify potential risks not covered by the rule engine, such as loopholes in complex contract terms.

[0067] When the model identifies a risky site, it automatically generates an alert message containing the risk type (e.g., "fake address," "abnormal operations") and the evidence chain (triggered rule items / model prediction probabilities). This alert is then pushed to regional managers via SMS and system pop-up windows. After the manager takes action (e.g., conducting an on-site inspection or terminating the lease), the results are entered into the system, which marks the risk as "closed-loop" and synchronizes the action record to the Panorama database, establishing a comprehensive management process of "identification-warning-action-recording."

[0068] Different from the traditional "after-the-fact processing" model, this system uses machine learning to explore potential risk characteristics and provide early warnings 3-6 months in advance for risks that have not yet erupted (such as a sudden drop in site value due to adjustments to surrounding planning). The accuracy of risk identification is improved compared to a single rule engine.

[0069] S104. Generate structured data including site identification, abnormal item list, indicator score, recommendation level and strategy label based on the false site identification model and the site panorama database to provide support for site operation decision-making.

[0070] Specifically, it integrates panoramic site data, anomaly identification results, and risk assessment model outputs to build a visual operation dashboard. It also outputs structured data containing strategic recommendations through standardized interfaces, enabling seamless integration with the company's existing systems (office, procurement, and contract systems) and providing quantitative support for decisions such as site selection, lease renewal, and renovation.

[0071] Visual Dashboard Construction: Design multi-dimensional dashboard modules: Global Monitoring: Nationwide site distribution heat map (colored by rent level and risk level), core indicator dashboard (vacancy rate, percentage of abnormal sites, rent compliance rate); Single site details: basic site information, historical exception records, current risk level, and comparison with surrounding competitors (rent, area, and supporting facilities); Exception Handling Tracking: A list of pending exceptions, a countdown to handling timelines, and historical handling effect analysis (such as changes in renewal rates after handling a certain type of exception). Drill-down analysis is supported: Click on a high-risk area in the heat map to drill down to a detailed list of exceptions for all sites in that area, with support for exporting to Excel or generating PDF reports.

[0072] Structured data output: Generates a standardized data structure containing the following fields: {"Site ID": "CN_SH_20250729_001","Exception Item List": ["Deposit Excessive", "Lease Overlap"],"Indicator Score": {"Rent Reasonableness": 65 points,"Risk Level": "Medium Risk","Operational Efficiency": 72 points},"Recommendation Level": "Grade B (Recommend Rent Renegotiation)","Strategy Label": ["Rent Reduction Negotiation", "Vacancy Period Optimization", "Legal Clause Verification"]}; "Strategy Label" automatically matches the preset strategy based on the exception type (for example, "Deposit Excessive" matches the "Initiate Deposit Installment Negotiation" strategy).

[0073] The system integration architecture utilizes a microservices framework (Spring Cloud). Each module (data collection, rule engine, and visualization service) is independently deployed, with a unified external interface (RESTful format) provided through an API gateway. Containerized deployment (Docker + Kubernetes) is supported, enabling dynamic scaling of computing resources to meet the concurrent analysis needs of tens of thousands of sites. Standardized interface adaptation provides data integration documentation, supports integration with enterprise OA systems for alert push, integration with procurement systems for site screening synchronization, and integration with data dashboards for real-time data visualization.

[0074] By replacing traditional manual recommendations with "recommendation levels + strategic tags," the site selection decision cycle is shortened, strategy matching accuracy is improved, and the subjectivity of empirical decision-making is addressed. The microservices architecture supports rapid iteration (e.g., the addition of the "Carbon Neutral Site Assessment" module). Containerized deployment enables linear scaling of system capacity with enterprise expansion, avoiding the performance bottlenecks of traditional monolithic architectures and adapting to the evolving management needs of enterprises from regional to national operations.

[0075] In some embodiments, the structured data and unstructured data corresponding to the venue-related information are cleaned and standardized, including: removing duplicate data, filling or filtering missing values, unifying data formats and semantic conversion processing on the structured data and unstructured data, so that the fields with the same name but different meanings or synonyms with different names from different data sources form a unified semantic definition.

[0076] For multi-source and heterogeneous site-related information, structured data (such as tabular data) and unstructured data (such as text and images) are cleaned and standardized. Specifically, this includes removing duplicate data, filling or filtering missing values, unifying data formats, and resolving semantic differences in fields with "same names but different meanings" (such as "area" in system A refers to building area, while "area" in system B refers to usable area) or "different names but same meanings" (such as "unit rental price" and "rent per unit area") in different data sources through semantic conversion, thus forming a unified data definition.

[0077] Data cleaning operations include: Deduplication: Based on the unique identifier of the site (such as the "address + property number" combination) or characteristic fields (area within ±5% is considered the same site), duplicate data is eliminated through database uniqueness constraints or ETL tools (such as Apache NiFi). Missing value processing: Missing data in key fields (such as rent and address) is marked as "pending" and triggers a manual verification process; non-key fields (such as the number of years of renovation) are filled using interpolation (interpolation using the mean / median of the same area) or prediction using machine learning models (such as random forest).

[0078] Standardization and semantic conversion include: format unification: unify the date format of different data sources to "YYYY-MM-DD", unify the amount unit, and unify the area unit to "square meters". Semantic mapping: establish a data dictionary (MetadataDictionary), define the standard field names and business meanings (such as unifying "lease term" to "contractual lease duration, unit: month"), and associate the fields of each data source with standard semantics through a mapping table (for example, map the "lease period" of system A and the "lease period" of system B to the standard field "lease term"). Unstructured data analysis: use NLP technology to perform entity recognition on text descriptions (such as extracting the "property fee includes status" label from "rent includes property fee"), and combine with domain dictionaries (such as real estate industry terminology library) to ensure the accuracy of semantic analysis.

[0079] Through deduplication, completion, and semantic unification, data accuracy is improved, avoiding analytical errors caused by data ambiguity (such as mistaking "building area" for "usable area" in calculating rental costs). Unified semantic definitions eliminate barriers to understanding between data sources, providing a standardized data foundation for subsequent multi-dimensional correlation analysis (such as comparing rental data from different platforms), and reducing the manual coordination costs of data integration.

[0080] In some embodiments, the combination of geographic location resolution technology unifies the coordinate system and labeling system corresponding to site-related information, including: performing geocoding resolution based on the site address or coordinate data, converting the coordinate systems of different data sources into a unified geographic coordinate system or projection coordinate system; establishing a standardized labeling system based on dimensions such as site use, regional attributes, and traffic conditions, so that each site corresponds to a unique geographic coordinate identifier and multi-dimensional classification label.

[0081] Utilize geographic location analysis technology to convert the coordinate systems of different data sources into a unified coordinate system (such as the 2000 geodetic coordinate system), and establish a standardized labeling system based on dimensions such as site use, regional attributes, and traffic conditions to ensure that each site has a unique geographic coordinate identifier and multi-dimensional classification labels.

[0082] Geographic coordinate unification includes: Address resolution: Using geocoding services (such as the Geocoding API), text addresses such as "XX City, XX District, XX" are converted into longitude and latitude coordinates. Addresses that fail resolution (such as ambiguous addresses) are then fuzzy matched with surrounding POI (Point of Interest) data. Coordinate system conversion: Using coordinate conversion algorithms (such as the seven-parameter conversion model) or open source tools (such as Proj.4), coordinate systems from different data sources are unified into a company-specified base coordinate system (such as a unified planar coordinate system), ensuring that spatial data can be displayed overlaid on the same map.

[0083] Label system construction: Multi-level label definition: First-level labels: site use (office / warehousing / production), regional attributes (urban core area / development zone / suburbs), transportation conditions (subway coverage / direct highway access / near main road); second-level labels: detailed attributes (such as "Grade A office building" and "co-working space" under "office" use, and "within 500 meters of a subway station" and "within 1 kilometer of a bus hub" under "transportation conditions"). Automated label generation: Use spatial analysis tools (such as ArcGIS) to calculate the straight-line distance from the site to the subway station and automatically match the "subway coverage" label; based on publicly available regional planning documents (such as "XX District is a key industrial park"), generate "regional attribute" labels through text matching.

[0084] A unified coordinate system enables site distribution and density analysis (e.g., creating a nationwide heat map of warehouse space) and supports regional management based on geofencing (e.g., automatically calculating the average rental price for all sites within a specific business district). A standardized tagging system enables fast, multi-dimensional searches, improving search efficiency compared to traditional keyword searches and meeting the precise screening needs of large-scale enterprise site selection.

[0085] In some embodiments, generating a panoramic site database includes integrating and storing cleaned, standardized, and geo-tagged site information to form panoramic site data including basic site attribute information, rental fee data, physical space parameters, address coordinates, surrounding supporting facilities, real-time operating status, and historical leasing records. The historical leasing records include lease terms, rent changes, contract performance, and lease termination records.

[0086] The cleaned and standardized site information (including geographic tags) is integrated and stored to form a "panoramic database" covering the entire life cycle of the site. The data content includes basic attributes, rental fees, physical space parameters, address coordinates, surrounding facilities, real-time operation status and historical leasing records (such as lease term, rent changes, contract performance, and reasons for termination).

[0087] Data integration scope: Basic attributes: site ID, owner, building structure, property ownership period, fire protection qualification; rental expenses: current rent, historical rent adjustment records, deposit, property management fees, water and electricity fee sharing method; physical space: building area, usable area, floor height, load-bearing capacity, number of parking spaces; operational data: current tenant name, lease status (in use / vacant / in the process of vacancy), vacancy start time, historical vacancy number and reason label (such as "rent is too high" or "insufficient area").

[0088] The storage architecture adopts a "master table + dimension table" structure: the master table stores the core site identifier (site ID) and associated foreign keys, while the dimension tables store basic attributes, rental data, operational records, and other information. Foreign key associations achieve data decoupling. For site data exceeding one million, the Hadoop Distributed File System (HDFS) is used to store unstructured data (images, documents), while HBase is used to store structured data, and Spark is used for distributed query and calculation.

[0089] This comprehensive database covers the entire site process, from site selection to leasing and lease termination. It supports tracing a site's historical performance (e.g., rent increase trends for a particular site over the past three years), providing data support for renewal negotiations and asset valuation. The unified database breaks down silos (e.g., data interoperability between financial and operational systems), allowing management to view the real-time correlation between site rental costs, operational efficiency, and risk status (e.g., whether high-rent sites correlate with high lease termination rates), fostering data-driven, cross-departmental collaborative decision-making.

[0090] In some embodiments, the site rental data is collected based on the site panoramic database, and a rent anomaly identification rule engine is constructed, including: obtaining rent, deposit, property fees, water and electricity fees, decoration fees and operating data through a combination of offline batch collection and real-time streaming collection; constructing dynamically configurable anomaly judgment logic based on business rules, and the anomaly judgment logic includes at least rules such as the deposit is higher than the preset rent multiple, the proportion of common area exceeds the limit, the lease period of the new and old sites overlaps, the site has not been operated for a long time after payment, the unit rent price increases abnormally, and the loss of termination exceeds the limit, and the preset rent multiple and proportion thresholds support dynamic adjustment through the business configuration interface.

[0091] By combining offline batch collection (historical data) with real-time streaming collection (newly generated data), we obtain rental-related data (rent, deposit, property fees, etc.), build a dynamically configurable rent anomaly identification rule engine, support customized anomaly judgment logic (such as the deposit is higher than 3 times the monthly rent, the common area exceeds 30%), and the thresholds in the rules (such as "preset rent multiples" and "proportion thresholds") can be flexibly adjusted through the business interface.

[0092] Data Collection Method: Offline Collection: Historical data (such as rent adjustment records and deposit payment receipts from the past five years) is extracted from the financial and contract management systems using ETL tools at dawn each day and stored in the Data Lake. Real-time Collection: New contract creation and expense bill generation events from the contract management system are monitored in real time through APIs. Data is transmitted in milliseconds using Kafka message queues. Data is cleaned by the Flink real-time computing engine and written to the rule engine database.

[0093] Rule Engine Configuration: Visual Rule Editing: A graphical interface is provided, allowing business personnel to configure rules through drag-and-drop operations (e.g., selecting the "Deposit" field and setting "greater than" or "three times the monthly rent" to trigger a "Deposit Excessive" exception). Logical operator combinations (AND / OR / NOT) are supported. Dynamic Threshold Adjustment: Differentiated thresholds can be set for different industries and regions (e.g., the threshold for "excessive common area ratio" for warehouse sites in first-tier cities is set at 25% and 30% in second-tier cities). Threshold configuration results are stored in a hot-update configuration center (such as Apollo) and take effect without a system restart.

[0094] Through offline and real-time data collection, we achieve both historical backtracking and real-time monitoring of rental data. Our rule engine covers high-frequency anomalies such as deposits, common areas, lease terms, and rent increases, offering improved coverage compared to traditional manual audits. We also support dynamic adjustment of rule thresholds for different business scenarios (e.g., temporarily relaxing the "paid but not in operation" time threshold during e-commerce promotions), avoiding misjudgments or missed judgments caused by rigid rules. This expands the system's adaptability from single enterprises to multi-format group management.

[0095] In some embodiments, the collected rental data is processed to calculate the venue information of each venue, and the venue information includes abnormal portraits and key indicators, including: according to the judgment logic of the rent abnormality identification rule engine, the rental data of a single venue is grouped, aggregated and associated analyzed, the abnormal rule items triggered by the venue are identified and abnormal portraits are generated; key indicators are calculated based on dimensions such as the frequency of abnormal rule triggering, the amount of expenses involved, and the degree of impact on operational efficiency, the abnormal portraits and key indicators are synchronized to the data analysis platform, and manual annotations and feedback corrections to the abnormal identification results are supported.

[0096] Based on the judgment logic of the rent anomaly identification rule engine, the rental data of a single venue is grouped and aggregated (such as rent aggregation by year / quarter) and correlation analysis (such as the correlation between rent and property fees) is performed. The triggered abnormal rule items are identified and an "anomaly profile" (including abnormality type, triggering time, and amount involved) is generated. At the same time, key indicators (such as anomaly index) are calculated from dimensions such as triggering frequency, cost impact, and operational efficiency. Manual annotation and correction of identification results are supported.

[0097] Anomaly profile generation includes: Grouping and aggregation analysis: Leasing data is grouped by "venue ID + time period," and metrics such as the deposit / rent ratio and the proportion of shared area within each period are calculated. These metrics are then compared with thresholds in the rule engine, and the triggered rules and their frequency are recorded (for example, a venue triggers the "abnormal rent price increase" rule twice consecutively in Q2 2025). The correlation between anomaly rules is analyzed through SQL correlation queries or graph databases (such as Neo4j) (for example, whether "excessive deposit" is often accompanied by "overlapping leases") to generate anomaly rule correlation maps to assist in identifying systemic risks.

[0098] Key indicator calculation: Anomaly Index Model: Anomaly Index = Σ(Rule Priority Weight × Number of Triggers) + Amount Involved Coefficient (e.g., each trigger of a high-priority rule is worth 20 points, a medium-priority rule is worth 10 points, and an additional 15 points are added for amounts exceeding 100,000 yuan). The data analysis platform provides an editing portal for anomaly results, allowing administrators to mark them as "false positives" or provide additional explanations. This feedback serves as training data for rule optimization.

[0099] By using anomaly profiling and key indicators, abstract risks are converted into quantifiable scores (e.g., an anomaly index ≥80 marks a red alert), enabling management to quickly identify high-risk sites and improve decision-making efficiency. Human feedback data is used to optimize the rule engine (e.g., correcting thresholds for misjudgment rules), forming a closed loop of "identification-verification-optimization." This allows anomaly identification accuracy to gradually improve over time, with the misjudgment rate reduced within three months.

[0100] In some embodiments, the construction of a false venue identification model uses a fusion of machine learning algorithms and rule engines to continuously analyze newly entered venue information and venue information during operation to identify risky venues, including: based on multi-dimensional indicators such as the deposit to rent ratio, venue operation rate, change in unit price of replacement venues, preferential implementation status and historical liquidation records, a two-layer identification model consisting of a rule matching layer and a machine learning prediction layer; the rule engine is used to preliminarily screen information that clearly violates business logic, and then the machine learning algorithm is used to perform pattern recognition on the screened information to identify potential false information or high-risk operating conditions, automatically mark the identified risky venues, and generate alarm information including risk level, abnormality type and impact range.

[0101] A two-layer fake venue identification model consisting of a "rule matching layer" and a "machine learning prediction layer" is constructed. Based on multi-dimensional indicators such as the deposit-to-rent ratio, venue operation rate, changes in replacement unit price, historical liquidation records, etc., the rule engine is first used to initially screen out obvious abnormal information (such as invalid address, owner's dishonesty), and then machine learning algorithms (such as random forest, XGBoost) are used to explore potential risk patterns (such as the combination of low deposit + high rental termination rate). The identified risky venues are automatically marked and alarm information containing risk level, abnormality type, and impact range is generated.

[0102] The two-layer model architecture includes the following: A rule-matching layer: Pre-set rigid rules (such as "address resolution failure" or "owner listed as dishonest") directly mark a site as "high risk" and block the entry process; Flexible rules (such as "deposit / rent < 0.5 and replacement price drops by more than 20%) mark a site as "medium risk" and trigger manual review. The machine learning layer: Feature engineering extracts over 50 risk features (such as "number of lease terminations in the past 12 months," "vacancy rate of similar sites within 3 kilometers," and "frequency of equity changes in ownership") and uses feature selection algorithms (such as the mutual information method) to select core features (retaining 20-30 highly relevant features). Model training utilizes historical samples of risky sites (positive samples) and healthy sites (negative samples). Cross-validation is used to optimize hyperparameters. The model automatically synchronizes with the latest data weekly for incremental training to prevent "concept drift" and other degradations in recognition capabilities.

[0103] Alert Mechanism: Risk levels are categorized as "high / medium / low." High-risk sites trigger instant SMS alerts, medium-risk sites are highlighted in red on the operations dashboard, and low-risk sites are placed on a watch list. Alert information includes a chain of evidence (e.g., "Triggering Rule R003: Deposit / Rent = 4.2 > 3x Threshold," "Model Prediction Probability = 92%)," enabling managers to quickly prioritize actions.

[0104] In some embodiments, based on natural language processing (NLP) and knowledge graph technology, a domain-adaptive semantic alignment model is constructed to automatically identify semantic ambiguity fields in multi-source data, and field semantic mapping and standardization across data sources are achieved through a deep learning model, solving the problem that traditional rule matching is difficult to cover long-tail semantic differences.

[0105] Domain knowledge graph construction: Collect the real estate industry term library, enterprise internal data dictionary, and historical data mapping cases, and construct a knowledge graph containing "field name - business meaning - data type - associated fields". The nodes are field entities (such as "floor area", "usable area"), and the edges are semantic association relationships (such as "synonym", "contains", "conversion formula").

[0106] Semantic alignment model training: Use the BERT pre-trained model to input the field names and context descriptions in multi-source data (such as the table title where the field is located, the data description document), and output the semantic vector representation of the field; Design a contrastive learning task: Pull the vectors of synonymous fields (such as "rent unit price" and "yuan / m 2 / month") closer, and push the vectors of non-synonymous fields (such as the ambiguity of "area" in different systems) farther away. Optimize the model through Triplet Loss to make fields with similar semantics closer in the vector space.

[0107] Intelligent mapping engine: For newly connected data sources, predict the standard semantic category of the field through the model (such as inputting "site size" in System A and matching it to the standard field "floor area"), and automatically generate data conversion logic in combination with the conversion rules in the knowledge graph (such as "usable area = floor area × 0.75"), supporting low-code configuration.

[0108] In some embodiments, by integrating computer vision (CV) and graph neural network (GNN), through the joint analysis of site image data and surrounding POI data, automatic generation and dynamic update of geographic labels are achieved, solving the problems that traditional rules rely on manual annotation and labels lag behind regional development.

[0109] Multi-modal data input: Visual data: Crawl the aerial images / street view images of the site from Map A software and Map B software, detect the building appearance features (such as "glass curtain wall", "warehouse ceiling") through the YOLO model, and judge the site use (office / warehouse); POI data: Obtain the POI types (subway, shopping mall, factory) and density within 1 kilometer around the site from OpenStreetMap, and construct a "POI adjacency graph" centered on the site.

[0110] Graph neural network modeling: Node features: site coordinates, building area, and visual recognition usage labels; POI node features: type, distance, and score; Graph Convolutional Layer (GCN): Learn the regional attribute labels of the site (such as "commercial core area" and "industrial park") through the spatial distance and functional correlation between nodes (such as the strong correlation between "subway node" and "office space node"); Dynamic update mechanism: Set quarterly tasks, re-crawl POI data and satellite images, detect new facilities in the surrounding area (such as new subway stations), and automatically trigger label updates (such as adding a "subway coverage" label to a certain site).

[0111] Geographic coordinate error correction: For venues where address resolution fails, the accuracy rate is increased from 70% to over 90% by utilizing the house number OCR recognition results in the visual data and combining it with the "house number reverse geocoding" function of the A map software API.

[0112] In some embodiments, by building an intelligent data quality monitoring system based on self-supervised learning, potential data quality issues (such as outliers and logical contradictions) in the panoramic database are automatically identified, and data repair strategies are generated through reinforcement learning to replace the fixed verification logic of the traditional rule engine.

[0113] Self-supervised quality inspection model: Unsupervised pre-training: Utilizes the time series of site data (e.g., historical rent changes) and spatial associations (area distribution of sites in the same region) to reconstruct the data using an autoencoder. Samples with reconstruction errors exceeding a threshold are marked as "potential anomalies" (e.g., a sudden 50% rent increase at a site with no fluctuations in surrounding site rents).

[0114] Cross-field logic verification: Build a Bayesian network model to learn the logical relationships between fields such as "rent = unit price × area" and "deposit ≤ 3 times monthly rent", and trigger warnings for records that violate the probability model (for example, the deviation between the calculated rent and the input value is greater than 5%).

[0115] Reinforcement learning repair strategy: State space: data quality issue type (missing values / outliers / logical contradictions), field importance level, historical repair records; action space: filling (mean / model prediction), correction (triggering manual verification), and deletion; reward function: dynamically adjust the strategy based on the impact of the repaired field on subsequent analysis (for example, repaired rental data improves the accuracy of anomaly identification), forming a "detection-repair-evaluation" closed loop.

[0116] Dynamic quality report: Generates a data quality dashboard daily, displaying the completeness, accuracy, and consistency scores of each business line's data. Automatically triggers data governance work orders for data sources that consistently fall below a threshold (e.g., a subsidiary's data completeness rate <80%).

[0117] In some embodiments, a dual-module system of "time series prediction-correlation risk" for rent anomalies is constructed based on the long short-term memory network (LSTM) and causal inference model. This system can not only identify current anomalies, but also predict future rent fluctuation risks and explore the causal relationship between abnormal events (such as whether the wave of rent terminations is caused by skyrocketing rents).

[0118] Time Series Forecasting Module: Input features include time series data such as historical rents, CPI index, regional vacancy rate, and average rent of similar venues. The LSTM model is used to learn long-term dependencies on rent changes (for example, rents in Q4 of each year regularly decrease due to the off-season). Anomaly Detection: Deviations between the predicted and actual values ​​exceeding three standard deviations are considered "abnormal fluctuations," distinguishing between "seasonal fluctuations" and "true anomalies" (for example, a sudden 20% rent increase at a venue during the off-season triggers an alert).

[0119] The causal inference module constructs a causal graph: using events such as "abnormal rent", "rental termination records", and "property fee increases" as nodes, it analyzes the causal effects between variables through the Do-Calculus algorithm (for example, verifying whether "excessively high deposits" will lead to "increased rent termination rates"). Counterfactual reasoning: simulating "whether the rent termination rate would decrease if the rent had not increased by 15%), providing a decision-making basis for risk intervention (such as recommending the initiation of a rent negotiation mechanism for venues with high rent increases).

[0120] Risk transmission analysis: Utilizes graph networks to analyze the relationships between venues (e.g., multiple venues owned by the same owner, or competing venues in the same region). When a venue triggers an "abnormal rent increase," the risk of rent linkage to surrounding venues is automatically assessed (e.g., the potential for tenant loss due to the contrast effect of surrounding venues).

[0121] In some embodiments, by building a federated learning platform for the real estate industry, while protecting the data privacy of each enterprise, the fake site identification model is trained by combining data from multiple parties to solve the problem of weak model generalization ability caused by insufficient data from a single enterprise, and realize cross-enterprise risk information sharing.

[0122] Federated learning architecture design: Participants: real estate companies, property management companies, and credit reporting agencies. Each party retains the original data locally and only uploads model parameters (gradients, weights);

[0123] Layered training: Bottom-level feature layer: Each enterprise trains a local feature extractor (such as extracting features such as "ownership change frequency" and "historical contract fulfillment rate"); High-level decision-making layer: The central server aggregates parameters from all parties, updates the global risk identification model, and avoids data leakage (in compliance with GDPR / PIPL compliance requirements).

[0124] Privacy Protection Mechanisms: Homomorphic encryption: Uploaded model gradients are encrypted, allowing the central server to aggregate only in ciphertext form. Differential privacy: Gaussian noise is added to parameters to prevent the model from inferring original data features. Joint Defense Mechanism: When a company is labeled a "fake site," risk characteristics are automatically synchronized to the federated model through a smart contract. The identification models of other companies are then updated with the risk pattern in real time, forming an industry-wide risk blacklist.

[0125] In some embodiments, by constructing a site-level digital twin model, through three-dimensional modeling and discrete event simulation, the operational effects under different leasing strategies (such as the impact of rent adjustments on occupancy rates) are simulated, and reinforcement learning is combined to find the optimal risk control strategy to assist in the identification of rent anomalies and false site decisions.

[0126] Digital Twin Construction: 3D Modeling: Using drone aerial photography and point cloud data processing, a 3D spatial model of the site is generated, annotating physical parameters such as building area, floor layout, and parking spaces. Business Process Modeling: The leasing process is abstracted into discrete events (contract signing, rent payment, and lease termination application). An AnyLogic-based simulation model is constructed, inputting historical operational data (such as lease termination rate and rent increase) as initial parameters.

[0127] Reinforcement learning optimization: State space: real-time operational status of the venue (vacancy rate, rental level, number of surrounding competing venues); Action space: rent adjustment range (±5%, ±10%), deposit plan (whether to accept installment payments), preferential policies (extending the rent-free period); Reward function: With "rental income minus risk loss" as the goal, simulate the long-term benefits of different actions and output the optimal risk control strategy (for example, when the surrounding vacancy rate is >40%, it is recommended to reduce the rent by 10% to reduce the risk of rent termination).

[0128] Abnormal scenario simulation: Preset extreme scenarios, simulate changes in site operation data, verify the effectiveness of existing abnormality identification rules (such as whether the time threshold of "long-term non-operation after payment" needs to be dynamically adjusted), and provide data support for rule engine optimization.

[0129] See also Figure 2 , Figure 2 The embodiment of the present application further provides a schematic block diagram of a device for intelligently controlling a venue's operations based on big data. The device 200 is configured to execute the aforementioned method for intelligently controlling a venue's operations based on big data. The device can be configured in a server or terminal.

[0130] The server can be a standalone server or a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal can be an electronic device such as a mobile phone, tablet computer, laptop computer, desktop computer, user digital assistant, and wearable device.

[0131] like Figure 2 As shown, the big data-based intelligent site operation management device 200 includes:

[0132] The information acquisition unit 201 is configured to acquire multi-source heterogeneous site-related information from a preset data source, perform cleaning and standardization operations on the structured and unstructured data corresponding to the site-related information, and unify the coordinate system and label system corresponding to the site-related information using geographic location analysis technology to generate a site-wide database.

[0133] Data collection unit 202, for collecting site rental data based on the site panorama database and building a rent anomaly identification rule engine, which is used to dynamically configure anomaly judgment logic to process the collected rental data and calculate site information for each site, including anomaly profiles and key indicators;

[0134] The model building unit 203 is used to build a fake site identification model. It uses a machine learning algorithm and a rule engine to continuously analyze newly entered site information and site information during operation to identify risky sites, and generates processing information corresponding to risky sites and sends it to the corresponding user terminal.

[0135] The decision support unit 204 is used to generate structured data including site identification, abnormal item list, indicator score, recommendation level and strategy label based on the false site identification model and the site panorama database to provide support for site operation decision-making.

[0136] In some embodiments, the structured data and unstructured data corresponding to the venue-related information are cleaned and standardized, including: removing duplicate data, filling or filtering missing values, unifying data formats and semantic conversion processing on the structured data and unstructured data, so that the fields with the same name but different meanings or synonyms with different names from different data sources form a unified semantic definition.

[0137] In some embodiments, the combination of geographic location resolution technology unifies the coordinate system and labeling system corresponding to site-related information, including: performing geocoding resolution based on the site address or coordinate data, converting the coordinate systems of different data sources into a unified geographic coordinate system or projection coordinate system; establishing a standardized labeling system based on dimensions such as site use, regional attributes, and traffic conditions, so that each site corresponds to a unique geographic coordinate identifier and multi-dimensional classification label.

[0138] In some embodiments, generating a panoramic site database includes integrating and storing cleaned, standardized, and geo-tagged site information to form panoramic site data including basic site attribute information, rental fee data, physical space parameters, address coordinates, surrounding supporting facilities, real-time operating status, and historical leasing records. The historical leasing records include lease terms, rent changes, contract performance, and lease termination records.

[0139] In some embodiments, the site rental data is collected based on the site panoramic database, and a rent anomaly identification rule engine is constructed, including: obtaining rent, deposit, property fees, water and electricity fees, decoration fees and operating data through a combination of offline batch collection and real-time streaming collection; constructing dynamically configurable anomaly judgment logic based on business rules, and the anomaly judgment logic includes at least rules such as the deposit is higher than the preset rent multiple, the proportion of common area exceeds the limit, the lease period of the new and old sites overlaps, the site has not been operated for a long time after payment, the unit rent price increases abnormally, and the loss of termination exceeds the limit, and the preset rent multiple and proportion thresholds support dynamic adjustment through the business configuration interface.

[0140] In some embodiments, the collected rental data is processed to calculate the venue information of each venue, and the venue information includes abnormal portraits and key indicators, including: according to the judgment logic of the rent abnormality identification rule engine, the rental data of a single venue is grouped, aggregated and associated analyzed, the abnormal rule items triggered by the venue are identified and abnormal portraits are generated; key indicators are calculated based on dimensions such as the frequency of abnormal rule triggering, the amount of expenses involved, and the degree of impact on operational efficiency, the abnormal portraits and key indicators are synchronized to the data analysis platform, and manual annotations and feedback corrections to the abnormal identification results are supported.

[0141] In some embodiments, the construction of a false venue identification model uses a fusion of machine learning algorithms and rule engines to continuously analyze newly entered venue information and venue information during operation to identify risky venues, including: based on multi-dimensional indicators such as the deposit to rent ratio, venue operation rate, change in unit price of replacement venues, preferential implementation status and historical liquidation records, a two-layer identification model consisting of a rule matching layer and a machine learning prediction layer; the rule engine is used to preliminarily screen information that clearly violates business logic, and then the machine learning algorithm is used to perform pattern recognition on the screened information to identify potential false information or high-risk operating conditions, automatically mark the identified risky venues, and generate alarm information including risk level, abnormality type and impact range.

[0142] It should be noted that, those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the model training device and each module described above can refer to the corresponding processes in the aforementioned embodiment of the method for intelligent control of site operations based on big data, and will not be repeated here.

[0143] The above-mentioned big data-based intelligent control site operation device can be implemented in the form of a computer program. The computer program can be used in Figure 3 Runs on the computer equipment shown.

[0144] See also Figure 3 , Figure 3 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device may be a server or a terminal.

[0145] See Figure 3 The computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory may include a storage medium and an internal memory.

[0146] The storage medium can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can cause the processor to execute any one of the methods for intelligently controlling site operations based on big data provided in the embodiments of the present application.

[0147] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0148] Internal memory provides an environment for the execution of computer programs stored in storage media. When executed by a processor, this computer program enables the processor to execute any of the methods for intelligently managing and controlling site operations based on big data. The storage medium can be either non-volatile or volatile.

[0149] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0150] It should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0151] Exemplarily, in one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0152] Acquire multi-source heterogeneous site-related information from pre-set data sources, cleanse and standardize the structured and unstructured data corresponding to the site-related information, and use geographic location analysis technology to unify the coordinate system and label system corresponding to the site-related information to generate a panoramic site database;

[0153] Based on the site panorama database, site rental data is collected and a rent anomaly identification rule engine is built. The rent anomaly identification rule engine is used to dynamically configure anomaly judgment logic to process the collected rental data and calculate site information for each site, including anomaly profiles and key indicators;

[0154] Build a fake site identification model, using a combination of machine learning algorithms and rule engines to continuously analyze newly entered site information and site information during operation to identify risky sites, and generate corresponding processing information for risky sites and send it to the corresponding user terminals;

[0155] Based on the fake site identification model and the site panoramic database, structured data including site identification, anomaly list, indicator score, recommendation level and strategy label is generated to provide support for site operation decision-making.

[0156] In some embodiments, the structured data and unstructured data corresponding to the venue-related information are cleaned and standardized, including: removing duplicate data, filling or filtering missing values, unifying data formats and semantic conversion processing on the structured data and unstructured data, so that the fields with the same name but different meanings or synonyms with different names from different data sources form a unified semantic definition.

[0157] In some embodiments, the combination of geographic location resolution technology unifies the coordinate system and labeling system corresponding to site-related information, including: performing geocoding resolution based on the site address or coordinate data, converting the coordinate systems of different data sources into a unified geographic coordinate system or projection coordinate system; establishing a standardized labeling system based on dimensions such as site use, regional attributes, and traffic conditions, so that each site corresponds to a unique geographic coordinate identifier and multi-dimensional classification label.

[0158] In some embodiments, generating a panoramic site database includes integrating and storing cleaned, standardized, and geo-tagged site information to form panoramic site data including basic site attribute information, rental fee data, physical space parameters, address coordinates, surrounding supporting facilities, real-time operating status, and historical leasing records. The historical leasing records include lease terms, rent changes, contract performance, and lease termination records.

[0159] In some embodiments, the site rental data is collected based on the site panoramic database, and a rent anomaly identification rule engine is constructed, including: obtaining rent, deposit, property fees, water and electricity fees, decoration fees and operating data through a combination of offline batch collection and real-time streaming collection; constructing dynamically configurable anomaly judgment logic based on business rules, and the anomaly judgment logic includes at least rules such as the deposit is higher than the preset rent multiple, the proportion of common area exceeds the limit, the lease period of the new and old sites overlaps, the site has not been operated for a long time after payment, the unit rent price increases abnormally, and the loss of termination exceeds the limit, and the preset rent multiple and proportion thresholds support dynamic adjustment through the business configuration interface.

[0160] In some embodiments, the collected rental data is processed to calculate the venue information of each venue, and the venue information includes abnormal portraits and key indicators, including: according to the judgment logic of the rent abnormality identification rule engine, the rental data of a single venue is grouped, aggregated and associated analyzed, the abnormal rule items triggered by the venue are identified and abnormal portraits are generated; key indicators are calculated based on dimensions such as the frequency of abnormal rule triggering, the amount of expenses involved, and the degree of impact on operational efficiency, the abnormal portraits and key indicators are synchronized to the data analysis platform, and manual annotations and feedback corrections to the abnormal identification results are supported.

[0161] In some embodiments, the construction of a false venue identification model uses a fusion of machine learning algorithms and rule engines to continuously analyze newly entered venue information and venue information during operation to identify risky venues, including: based on multi-dimensional indicators such as the deposit to rent ratio, venue operation rate, change in unit price of replacement venues, preferential implementation status and historical liquidation records, a two-layer identification model consisting of a rule matching layer and a machine learning prediction layer; the rule engine is used to preliminarily screen information that clearly violates business logic, and then the machine learning algorithm is used to perform pattern recognition on the screened information to identify potential false information or high-risk operating conditions, automatically mark the identified risky venues, and generate alarm information including risk level, abnormality type and impact range.

[0162] The present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements the steps of the risk warning method described in the first aspect above.

[0163] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the computer device.

[0164] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for intelligently controlling site operations based on big data, characterized in that: include: Acquire multi-source heterogeneous site-related information from pre-set data sources, cleanse and standardize the structured and unstructured data corresponding to the site-related information, and use geographic location analysis technology to unify the coordinate system and label system corresponding to the site-related information to generate a panoramic site database; Based on the site panorama database, site rental data is collected and a rent anomaly identification rule engine is built. The rent anomaly identification rule engine is used to dynamically configure anomaly judgment logic to process the collected rental data and calculate site information for each site, including anomaly profiles and key indicators; Build a fake site identification model, using a combination of machine learning algorithms and rule engines to continuously analyze newly entered site information and site information during operation to identify risky sites, and generate corresponding processing information for risky sites and send it to the corresponding user terminals; Based on the fake site identification model and the site panoramic database, structured data including site identification, anomaly list, indicator score, recommendation level and strategy label is generated to provide support for site operation decision-making.

2. The method according to claim 1, characterized in that The cleaning and standardization operations for the structured data and unstructured data corresponding to the site-related information include: The structured data and unstructured data are processed by removing duplicate data, filling or filtering missing values, unifying data formats and semantic conversion, so that fields with the same name but different meanings or synonyms with different names from different data sources form a unified semantic definition.

3. The method according to claim 1, characterized in that The above-mentioned geographic location analysis technology is used to unify the coordinate system and label system corresponding to the site-related information, including: Perform geocoding analysis based on site addresses or coordinate data to convert the coordinate systems of different data sources into a unified geographic coordinate system or projected coordinate system; A standardized labeling system is established based on dimensions such as site use, regional attributes, and traffic conditions, so that each site corresponds to a unique geographic coordinate identifier and multi-dimensional classification label.

4. The method according to claim 1, wherein The generating of the site panoramic database includes: The cleaned, standardized and geo-tagged site information is integrated and stored to form a panoramic site data that includes basic site attribute information, rental cost data, physical space parameters, address coordinates, surrounding supporting facilities, real-time operating status and historical rental records. The historical rental records include lease term, rent changes, contract performance and lease termination records.

5. The method according to claim 1, wherein The method of collecting site rental data based on the site panoramic database and building a rent anomaly identification rule engine includes: Acquire rent, deposit, property management fee, utility fee, renovation fee and operation data through a combination of offline batch collection and real-time streaming collection; Based on business rules, dynamically configurable exception judgment logic is constructed. The exception judgment logic at least includes rules such as the deposit is higher than the preset rent multiple, the common area ratio exceeds the limit, the lease period of the new and old venues overlaps, the site has not been operated for a long time after payment, the unit rent price increases abnormally, and the loss of termination exceeds the limit. The preset rent multiple and ratio thresholds support dynamic adjustment through the business configuration interface.

6. The method according to claim 1, characterized in that The collected rental data is processed to calculate the site information of each site. The site information includes abnormal profiles and key indicators, including: According to the judgment logic of the rent anomaly identification rule engine, the rental data of a single venue is grouped, aggregated, and analyzed for association, the anomaly rule items triggered by the venue are identified, and an anomaly profile is generated; Key indicators are calculated based on dimensions such as the frequency of abnormal rule triggering, the amount of expenses involved, and the degree of impact on operational efficiency. The abnormal profiles and key indicators are synchronized to the data analysis platform, and manual comments and feedback corrections to the abnormal identification results are supported.

7. The method according to claim 1, characterized in that The aforementioned fake site identification model uses a machine learning algorithm integrated with a rule engine to continuously analyze newly entered site information and site information during operation to identify risky sites, including: Based on multi-dimensional indicators such as the deposit-to-rent ratio, site operation rate, price fluctuation of replacement sites, preferential implementation, and historical liquidation records, a two-layer recognition model consisting of a rule matching layer and a machine learning prediction layer was constructed; The rule engine is used to preliminarily screen information that clearly violates business logic, and then a machine learning algorithm is used to perform pattern recognition on the screened information to identify potential false information or high-risk operating conditions. The identified risk sites are automatically marked and alarm information containing risk level, anomaly type and impact range is generated.

8. A device for intelligently controlling site operations based on big data, characterized in that: include: An information acquisition unit is used to obtain multi-source heterogeneous site-related information from a preset data source, perform cleaning and standardization operations on the structured and unstructured data corresponding to the site-related information, and unify the coordinate system and label system corresponding to the site-related information in combination with geographic location analysis technology to generate a site panoramic database; A data collection unit is used to collect site rental data based on a site panorama database and build a rent anomaly identification rule engine. The rent anomaly identification rule engine is used to dynamically configure anomaly judgment logic to process the collected rental data and calculate site information for each site, including anomaly profiles and key indicators; The model building unit is used to build a fake site identification model. It uses a machine learning algorithm and a rule engine to continuously analyze newly entered site information and site information during operation to identify risky sites, and generates processing information corresponding to risky sites and sends it to the corresponding user terminal; The decision support unit is used to generate structured data including site identification, abnormal item list, indicator score, recommendation level and strategy label based on the false site identification model and the site panoramic database to provide support for site operation decision-making.

9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • A housing rent forecasting method based on FFM algorithm

    CN109389530A

  • Method for identifying abnormal house resources information

    CN112148945A

  • House false information identification method and device thereof, electronic equipment and storage medium

    CN112699659A

  • Field big data risk screening method and device

    CN114997571A

  • Dynamic lease risk assessment method and system based on index data

    CN119722277A

Cited By

  • Enterprise data link treatment and value management method and system

    CN120832348A

  • Intelligent lease operation system based on multi-dimensional data analysis

    CN121213208A

  • User interest matching network marketing system fusing knowledge graph

    CN121303291A